AbstractReader
AbstractReader is the base class every instrument reader subclasses. It owns
the pipeline — file discovery, native-grid detection, caching, the L3
presentation steps, reporting and df.attrs stamping — and leaves three hooks
to the subclass: _raw_reader, _QC and (optionally) _process.
- What each stage may add or destroy, and what is cached: Data Levels (L0–L3).
- The QC machinery and status decoding the hooks plug into: RawDataReader Reference.
- Writing a new subclass, step by step, with the registry entry, the tests and the docs page it needs: Contributing a reader.
You do not instantiate it directly — the
RawDataReader factory picks the subclass from the
instrument name.
API
AeroViz.rawDataReader.core.AbstractReader
Bases: ABC
Abstract class for reading raw data from different instruments.
This class serves as a base class for reading raw data from various instruments. Each instrument
should have a separate class that inherits from this class and implements the abstract methods.
The abstract methods are _raw_reader and _QC.
The class handles file management, including reading from and writing to pickle files, and implements quality control measures. It can process data in both batch and streaming modes.
Attributes:
| Name | Type | Description |
|---|---|---|
nam |
str
|
Name identifier for the reader class |
path |
Path
|
Path to the raw data files |
meta |
dict
|
Metadata configuration for the instrument |
logger |
ReaderLogger
|
Custom logger instance for the reader |
reset |
bool
|
Flag to indicate whether to reset existing processed data |
append |
bool
|
Flag to indicate whether to append new data to existing processed data |
qc |
bool or str
|
Quality control settings |
qc_freq |
str or None
|
Frequency for quality control calculations |
Initialize the AbstractReader.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
path
|
Path or str
|
Path to the directory containing raw data files |
required |
reset
|
bool or str
|
If True, forces re-reading of raw data If 'append', appends new data to existing processed data |
False
|
qc
|
bool or str
|
If True, performs quality control If str, specifies the frequency for QC calculations |
True
|
**kwargs
|
dict
|
Additional keyword arguments: raw_freq : str Override raw data frequency (e.g., '6min', '1h'). If not set, frequency is auto-inferred from the data. drop_outlier_dates : bool, default=False Stray timestamps far outside the data's bulk (e.g. a year-2000 row in 2023 data) are always detected and warned about, since they balloon the native grid. By default they are kept (the warning explains how to fix the source); set True to drop them automatically before the grid is built. log_level : str Logging level for the log file quiet : bool If True, suppresses all console output |
{}
|
Notes
Creates necessary output directories and initializes logging system. Sets up paths for pickle files, CSV files, and report outputs.
Attributes
_output_prefix
instance-attribute
logger
instance-attribute
logger = ReaderLogger(self.nam, output_folder, kwargs.get('log_level', 'INFO').upper(), quiet=self.quiet)
qc_severity_overrides
instance-attribute
save_intermediate_csv
instance-attribute
Methods:
_QC
abstractmethod
Abstract method for quality control processing.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
df
|
DataFrame
|
Input DataFrame containing raw data |
required |
Returns:
| Type | Description |
|---|---|
DataFrame
|
Quality controlled data with QC_Flag column |
Notes
Must be implemented by child classes to handle instrument-specific QC. This method should only check raw data quality (status, range, completeness). Derived parameter validation should be done in _process().
__call__
Process data for a specified time range.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
start
|
datetime
|
Start time for data processing; defaults to the data's first timestamp |
None
|
end
|
datetime
|
End time for data processing; defaults to the data's last timestamp |
None
|
mean_freq
|
str
|
Frequency for resampling the output; if None, no resampling is done and the data is returned at its native resolution |
None
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
Processed and resampled data for the specified time range |
Notes
The processed data is also saved to a CSV file.
_cache_is_current
True if the cached frame was written by the current cache format.
_flag_outlier_dates
Detect and warn about stray timestamps; drop them only if asked.
A single bad row — e.g. a 2000-01-01 stamp in otherwise-2023 data —
stretches the canonical native grid (built over the data's own min->max
in _read_raw_files, before any requested range applies) across the
whole bogus span, inflating the cached frame to millions of NaN rows
even when the caller only asked for 2023. Such stamps are almost always
a source-data error, so by default we warn and tell the user how to fix
it rather than silently changing their data; pass
drop_outlier_dates=True to have them excluded automatically.
_generate_report
Calculate and log data quality rates for different time periods.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
raw_data
|
DataFrame
|
Raw data before quality control |
required |
qc_data
|
DataFrame
|
Data after quality control |
required |
qc_flag
|
Series
|
QC flag series indicating validity of each row |
None
|
Notes
Calculates rates for specified QC frequency if set. Updates the quality report with calculated rates.
_load_or_parse
Return the canonical parsed (raw, qc) frames, using the pkl cache when it exists and is current.
Canonical = snapped to the native grid over the files' own coverage,
NOT padded to any requested range. Parse provenance (n_files,
raw_freq, freq_mixed) is persisted in df.attrs so a cache
hit restores it onto self. The requested range / fill_missing
is applied later, in _run.
_outlier_process
Process outliers in the data.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
_df
|
DataFrame
|
Input DataFrame containing potential outliers |
required |
Returns:
| Type | Description |
|---|---|
DataFrame
|
DataFrame with outliers processed |
Notes
Implementation depends on specific instrument requirements.
_partition_compatible_scans
Drop frames whose scan schema differs from the dominant group.
Default is a no-op — overridden by readers (currently SMPS) where the
same instrument can export at different size-bin grids depending on
the host software version (AIM 10.3 .TXT vs AIM 11.x .CSV). The outer
join inside pd.concat happily concatenates frames with disjoint
columns, but the NaN holes break per-bin completeness QC. The
well-defined repair is to treat each grid as its own scan: keep the
majority group, drop the minority and tell the user which files were
skipped so they can re-run them in isolation if they want both.
df_list and files are aligned and contain only successfully
parsed entries.
_process
Process data to calculate derived parameters.
This method is called after _QC() to calculate instrument-specific derived parameters (e.g., absorption coefficients, AAE, SAE).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
df
|
DataFrame
|
Quality-controlled DataFrame carrying |
required |
Returns:
| Type | Description |
|---|---|
DataFrame
|
DataFrame with derived parameters added and the QC columns updated |
Notes
Default implementation returns the input unchanged. Override in child classes to implement instrument-specific processing.
The method should:
- Calculate derived parameters for every row, flagged or not.
- Judge the derived parameters and record the verdict with
update_qc_flag(df, mask, name, severity=...). - Log the combined summary via
extend_qc_summary+log_qc_summarywhen it added a rule of its own.
Do not skip rows that _QC already flagged. An earlier version of
this docstring offered that as an optimisation; it is wrong twice over.
L2's contract (rule R2 in docs/guide/data-levels.md) is to judge
without destroying, so a derived value belongs in _read_*_qc.csv
whatever the verdict — that file is where a user goes to see why a row
was rejected, and a NaN there answers nothing. And since severity
arrived, a flagged row is not necessarily an invalid one: skipping
"flagged" rows would silently drop derived values for rows that are kept.
Rules are independent and deliberately overlap — one row can be both
Status Error and Invalid AAE — so the per-rule counts in a QC
summary do not sum to the total. Valid and Usable are the totals
that mean something.
_qc_summary_rows
The final QC summary as JSON-friendly rows, for df.attrs.
acquisition_rate/yield_rate say how much data was lost;
these rows say which rule lost it. A downstream monitor holding only
the rates can report an outage but not explain one — "all values are
null" is a symptom, "Status Error 12%" is a cause.
Percentages come back as floats: the summary table formats them as
'12.3%' for the log, which is the wrong type to do arithmetic or
thresholding on.
_raw_reader
abstractmethod
Abstract method to read raw data files.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
file
|
Path or str
|
Path to the raw data file |
required |
Returns:
| Type | Description |
|---|---|
DataFrame
|
Raw data read from the file |
Notes
Must be implemented by child classes to handle specific file formats.
_read_raw_files
Read and process raw data files.
Returns:
| Type | Description |
|---|---|
tuple[DataFrame | None, DataFrame | None]
|
Tuple containing: - Raw data DataFrame or None - Quality controlled DataFrame or None |
Notes
Handles file reading and initial processing.
_resample
Average df onto mean_freq; a no-op when none was requested.
mean() silently drops non-numeric columns, which is how text metadata
(a status string, an instrument ID) vanishes between the native-resolution
and resampled outputs. That is the right behaviour — there is no sensible
mean of a status string — but it should not be silent, so the dropped
columns are named in the log once.
_restore_parse_meta
Pull parse provenance off a cached frame back onto self (cache hit).
_run
Main execution method for data processing.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
user_start
|
datetime
|
Start time for processing |
required |
user_end
|
datetime
|
End time for processing |
required |
Returns:
| Type | Description |
|---|---|
tuple[DataFrame, DataFrame]
|
Raw and quality-controlled frames for the requested range. |
Notes
Two layers. _load_or_parse returns the canonical parsed frames
(from the pkl cache when valid, else by reading the raw files). The
presentation step below — grid placement to the requested range, with
fill_missing — runs on every call, so a cache hit honours the
current call's range/fill_missing instead of replaying whatever was
stored. Parse provenance restored by _load_or_parse feeds the
df.attrs stamp in __call__.
_save_data
Save processed data to files.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
raw_data
|
DataFrame
|
Raw data to save |
required |
qc_data
|
DataFrame
|
Quality controlled data to save |
required |
Notes
Saves data in both pickle and CSV formats.
_stamp
Attach reader metadata to df.attrs just before returning.
Always records provenance (instrument, station, coverage, requested
range, native frequency). When with_qc is True it additionally
records the output frequency and the overall QC rates; the plain raw
path (qc=False) gets provenance only.
See core.metadata for why this is the single, final stamping point.
_stamp_parse_meta
Persist parse provenance into df.attrs so it survives the pkl cache.
_status_condition_rows
Which status conditions actually fired, and how often.
The Status Error rule can only say that a non-whitelisted bit was
set — the QC verdict is a boolean, so the identity of the bit is lost
the moment it is computed. Decoding the register here is the difference
between "Status Error 8.3%" and "Ambient RH & Temp sensor 8.3%", which
is the difference between knowing something is wrong and knowing what
to go and fix.
Counted over every row, flagged or not: a condition that fired on rows that survived QC is still worth seeing.
percentage shares its denominator with qc_rules — the whole
frame, including the empty rows that placing a sparse instrument on a
time grid creates. That makes the two directly comparable (a consumer
can put "Status Error 0.1%" and the condition that caused it in the same
sentence), at the cost of looking small for a sparse reader. count
is the absolute, and is the number to trust when the grid is mostly
padding.
_status_register
The status column as integers, whichever way it was written down.
_timeIndex_process
Process time index of the DataFrame.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
_df
|
DataFrame
|
Input DataFrame to process |
required |
user_start
|
datetime
|
User-specified start time |
None
|
user_end
|
datetime
|
User-specified end time |
None
|
append_df
|
DataFrame
|
DataFrame to append to |
None
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
DataFrame with processed time index |
Notes
Frequency is resolved once per run in _read_raw_files (per-file
detection, see self._resolved_freq); this method only places the
data on that grid via to_grid — snapping off-grid timestamps to
their nearest bin without the duplicate-fill of method='nearest'.
check_status_columns
Which of candidates are present, warning loudly when none are.
filter_error_status returns all-False for a column it cannot find, so
a renamed status column degrades to "this instrument reported no errors,
ever" — indistinguishable from a healthy instrument. Vendors do rename
it between host-software versions (SMPS AIM 10.3 vs 11.x split the same
information across differently-named columns), so the absence has to be
visible.
Returns the present names so a caller can OR their masks together.
extend_qc_summary
extend_qc_summary(summary: DataFrame, df: DataFrame, rule: str, mask: Series, description: str = '', severity: str = ERROR) -> DataFrame
Add a _process-stage rule to a _QC summary and refresh the totals.
_QC builds the summary before derived quantities exist, so a rule like
Invalid AAE can only be counted later. The new row is inserted above
the trailing Valid / Usable totals, and both totals are then
recomputed from df's QC columns so they account for the late rule.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
summary
|
DataFrame
|
The table returned by |
required |
df
|
DataFrame
|
The frame after |
required |
rule
|
str
|
The late rule's name, boolean mask, description and severity. As in
|
required |
mask
|
str
|
The late rule's name, boolean mask, description and severity. As in
|
required |
description
|
str
|
The late rule's name, boolean mask, description and severity. As in
|
required |
severity
|
str
|
The late rule's name, boolean mask, description and severity. As in
|
required |
log_below_mdl
Report, per column, how much of it sits below its detection limit.
A value below the MDL is a valid measurement of a low concentration (or
a non-detect), not a broken row — so this is a diagnostic, not a QC rule.
Deliberately so: any non-Valid flag NaNs the whole row in
__call__, and with tens of species/elements per row "any one below
MDL" is true almost always, which would delete the dataset. Use this to
see which species are near their limits, then decide per analysis.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
df
|
DataFrame
|
The frame to inspect (columns not in |
required |
mdl
|
dict
|
|
required |
top
|
int
|
How many of the worst-affected columns to log. |
10
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
One row per column with |
log_qc_summary
Log a QCFlagBuilder.get_summary table.
Advisory rules are marked so it is obvious which flags kept their data.
Valid (passed everything) and Usable (nothing invalidating) are
both reported — they differ by the rows carrying only advisory flags.
progress_reading
Context manager for tracking file reading progress.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
files
|
list
|
List of files to process |
required |
Yields:
| Type | Description |
|---|---|
Progress
|
Progress bar object for tracking |
Notes
Uses rich library for progress display.
qc_builder
A QCFlagBuilder carrying this run's severity overrides.
Readers should use this instead of QCFlagBuilder() directly so that
flag_severity={'Insufficient': 'warning'} reaches their rules, so a
rule that raises is reported through the reader's log, and so an override
naming no rule of this instrument is rejected instead of ignored.
qc_columns
staticmethod
The QC bookkeeping columns present in df, in a stable order.
Readers that narrow their output to a fixed column list must carry both
of them through — QC_Flag (the record) and QC_Invalid (the
verdict the presentation layer masks on). Slicing with a hard-coded
+ ['QC_Flag'] silently drops the verdict, which would make every flag
fatal again.
qc_severity
Effective severity of flag_name under this run's flag_severity.
QCFlagBuilder resolves its own rules; this is the same lookup for the
flags a reader raises later in _process, which never reach the
builder. Route every late flag through here — hard-coding a severity is
what made Invalid AAE impossible to reclassify.
reorder_dataframe_columns
staticmethod
Reorder DataFrame columns according to specified lists.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
df
|
DataFrame
|
Input DataFrame |
required |
order_lists
|
list[list]
|
Lists specifying column order |
required |
keep_others
|
bool
|
If True, keeps unspecified columns at the end |
False
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
DataFrame with reordered columns |
update_qc_flag
Add a flag to QC_Flag for rows matching the mask, after _QC ran.
Used by _process to flag something that can only be judged once
derived quantities exist (e.g. Invalid AAE).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
df
|
DataFrame
|
DataFrame with QC_Flag column |
required |
mask
|
Series
|
Boolean mask indicating rows to flag |
required |
flag_name
|
str
|
Name of the flag to add. Must appear in the reader's
|
required |
severity
|
(error, warning)
|
|
'error'
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
DataFrame with updated |