Configuration
ImpulseConfig configures everything about a report: the silver-layer input
tables, the gold-layer output location, container-level filters, the
query-engine solver, incremental processing, and which container columns
get surfaced into the gold-layer measurement dimension. Configuration is
defined as JSON (or an equivalent Python dictionary) and validated using
Pydantic models. The canonical schema lives in
src/impulse_reporting/config/config_parser.py.
Quick example
{
"source": {
"container_metrics_table": "my_catalog.silver.container_metrics",
"channel_metrics_table": "my_catalog.silver.channel_metrics",
"channels_uri": "my_catalog.silver.channels",
"container_tags_table": "my_catalog.silver.container_tags",
"channel_tags_table": "my_catalog.silver.channel_tags"
},
"unity_sink": {
"catalog": "my_catalog",
"schema": "gold",
"table_prefix": "my_report"
},
"query_engine": {
"solver": "DefaultSolver",
"data_type": "RAW"
},
"container_filters": {
"tag_filters": [
[
{ "tag_name": "uut_id", "comparator": "==", "value": "ABC123", "cast_type": "string" }
]
],
"metric_filters": [
[
{ "column_name": "start_dt", "comparator": ">=", "value": "2025-04-27T05:20:54.000Z", "value_type": "timestamp" }
]
]
},
"measurement_dimensions": ["container_id", "vehicle_key", "start_ts", "stop_ts"]
}
A configuration is passed to Report either as a Python dict
(config=...) or as a JSON file path (config_path=...). Sinkless
mode is also supported — see Sinkless reports.
source
Maps the silver-layer input tables.
| Field | Type | Required | Description |
|---|---|---|---|
container_metrics_table | str | Yes | Full Unity Catalog path. Container metadata (timestamps, duration). |
channel_metrics_table | str | Yes | Full Unity Catalog path. Channel-level statistics. |
channels_uri | str | Yes | Full Unity Catalog path. Time-series sample data. |
poi_channels_uri | str | No | Full Unity Catalog path to the Points-in-Time (POI) channel data table. Required only when selecting POI channels via poi_channel(); omit for sample-only data models. |
container_tags_table | str | No | Full Unity Catalog path. Container EAV tags. |
channel_tags_table | str | No | Full Unity Catalog path. Channel EAV tags. |
channel_mapping_table | str | No | Full Unity Catalog path. Logical-to-physical channel alias table. Required when using QueryBuilder.channel_with_alias(). In reporting mode the resolved alias-to-physical-channel mapping is materialized to the gold-layer channel_mapping_resolution_dimension. |
unit_conversion_table | str | No | Full Unity Catalog path. Per-unit-family conversion factors. When configured together with a channel_mapping_table whose rows carry source_unit / target_unit columns, aliased selectors auto-convert values from source to target unit during solve(). |
A container_tags_table is required to use tag-based container filters;
a channel_tags_table is required to select channels by tag rather than
by columns on channel_metrics.
unity_sink
Defines where gold-layer tables are written.
| Field | Type | Required | Description |
|---|---|---|---|
catalog | str | Yes | Target catalog name. |
schema | str | Yes | Target schema name. |
table_prefix | str | Yes | Prefix for all generated table names. |
cleanup_temp_tables | bool | No | Drop the batch-solving __impulse_temp_* tables from this schema after persist_results() completes successfully. Defaults to false. Overridable per call via persist_results(cleanup_temp_tables=...). |
Output tables are named {table_prefix}_{entity} (e.g.
my_report_histogram_fact).
Sinkless reports
unity_sink is optional. When omitted, the report runs in sinkless
mode: determine_report() still computes events, aggregations, and
container dimensions and exposes them on the report object, but
persist_results() becomes a no-op. Useful for ad-hoc analysis,
notebooks, and tests where writing to Unity Catalog is not desired.
container_filters (optional)
Restricts the set of processed containers. Filters are expressed in disjunctive normal form (OR of ANDs): each inner list is AND-combined, the outer list is OR-combined.
Two independent filter families:
tag_filters— applied oncontainer_tags_table(EAV key/value model).metric_filters— applied oncontainer_metrics_table(columnar model).
| Field | Type | Default | Description |
|---|---|---|---|
tag_filters | list[list[TagFilter]] | [] | Tag-based filter groups (DNF). |
metric_filters | list[list[MetricFilter]] | [] | Metric-based filter groups (DNF). |
TagFilter
| Field | Type | Required | Description |
|---|---|---|---|
tag_name | str | Yes | Tag key to filter on. |
comparator | str | Yes | One of ==, !=, >, >=, <, <=. |
value | any | Yes | Expected value. Must match cast_type. |
cast_type | str | No | string (default), int, double, or timestamp (ISO-format string). |
MetricFilter
| Field | Type | Required | Description |
|---|---|---|---|
column_name | str | Yes | Column on container_metrics_table to filter on (e.g. start_dt, stop_dt). When solver_config.container_metrics.column_name_mapping is set, this refers to the internal name (after renaming). |
comparator | str | Yes | One of ==, !=, >, >=, <, <=. |
value | any | Yes | Expected value. Must match value_type when provided. |
value_type | str | No | When provided, validates/converts the value: string, int, double, timestamp. |
query_engine (optional)
| Field | Type | Default | Description |
|---|---|---|---|
solver | str | "DefaultSolver" | "DefaultSolver" adapts to the silver layer: it selects channels from a narrow EAV channel_tags table when source.channel_tags_table is set and otherwise from columns on channel_metrics; it filters containers via a narrow EAV container_tags table or, when source.container_tags_table is omitted, a wide-only container_metrics. "DeltaSolver" and "KeyValueStoreSolver" are deprecated aliases that resolve to DefaultSolver. |
data_type | str | "RLE" | "RLE" (intervals [tstart, tend)) or "RAW" (raw timestamps; converted to intervals at query time by the configured raw_encoder). |
raw_encoder | str | null (resolves to "RLE" when data_type = "RAW") | How RAW point data is converted to intervals. "RLE" (the default) run-length encodes on the fly to reduce memory consumption: consecutive samples with the same value collapse into a single [tstart, tend) interval per run. "INTERVAL" keeps every original sample, only deriving tend from the next sample's timestamp and dropping exact duplicate points — use it when downstream analysis needs all original timestamps. Only takes effect with data_type = "RAW"; ignored for "RLE" input. See How Impulse interprets intervals for validity semantics and what counts as a duplicate point. |
drop_implausible_data | bool | false | When true, drops channels rows where is_plausible = false. Requires data_type = "RAW"; combining with "RLE" raises a validation error. |
max_channels_per_batch | int | 500 | Caps how many distinct channels are read in one solve batch, to bound per-solve memory. Counts channel selections (each distinct channel(...) / poi_channel(...) in your expressions). Applies to events, aggregations, and calculated channels; for calculated channels it counts the input channels they read, not the number of calculated channels. The cap is a packing target, not a hard limit: an expression that on its own reads more channels than the cap still runs, in a batch of its own. Must be >= 1. The old name batch_size is a deprecated alias (still accepted, warns, removed in a future release); setting both raises a validation error. |
max_containers_per_batch | int | null | Caps how many containers are solved at once, to bound per-solve memory. null disables the cap; when set it must be >= 1. On an incremental run, Report.run() commits at most this many upserted containers per iteration and loops until the population is drained (see incremental). On a full run, and when a changed definition recomputes over all containers, the population is split into chunks of about this size and solved chunk by chunk. Chunk sizes are approximate (the cap is a memory heuristic, not an exact bound). |
solver_config | SolverConfig | null | Per-table column mappings, per-table equality filters, and project scoping. Set project_id to scope reads by project — it is applied to container_tags (if configured), container_metrics, and channel_mapping (if configured), so it works in both narrow EAV and wide-only data models. Omit it when you don't need project scoping. See Solver column mappings and filters. |
If query_engine is omitted, the default is DefaultSolver with
data_type = "RLE".
Solver column mappings and filters
The framework references columns by a fixed set of internal names (e.g. container_id, channel_id,
tstart, tend, value). When your silver-layer tables use different physical names, declare the mapping
in solver_config so the solver renames each table's columns at read time.
SolverConfig has one section per silver table. Each section is a TableConfig with two fields:
column_name_mapping(dict[str, str]):{ "physical_column": "internal_column" }. The mapping is applied once, when the table is read. All downstream processing (filters, joins, aggregations) uses the internal names.filters(dict[str, str]): equality filters applied after renaming. Keys are internal column names; values are literals to match. Useful for project/toolbox scoping where a single value should always be enforced. Values are strings but are coerced to the target column's type, so a boolean column takes"true"/"false"(an unparseable value errors at read time under ANSI mode).
Top-level fields on SolverConfig:
project_id(str, optional): Project scope. When set, the solver applies an equality filter on theproject_idcolumn (after column-name mapping) of every table it reads that carries one —container_tags(if configured),container_metrics, andchannel_mapping(if configured). Omit it if you don't need project-level scoping; the solver does not require it.
Per-table sections (each a TableConfig):
| Section | When it applies | Typical mappings |
|---|---|---|
container_tags | when container_tags_table is configured | entity_id → container_id, custom EAV key/value columns |
container_metrics | always | Custom container_id column, custom timestamp columns |
channel_tags | when channel_tags_table is configured | Tag key/value column renames |
channel_metrics | always | Custom channel_id column, custom value/timestamp columns |
channel_mapping | when channel_mapping_table is configured | Alias-table column renames; priority column; optional join_keys for non-default alias-resolution composite keys |
channels | always | RLE column renames (tstart/tend/value); filters supported (applied at read time, before non-solve columns are dropped) |
unit_conversion | when unit_conversion_table is configured | Unit-conversion table column renames (unit, group_id, conversion_factor) |
Internal column names that mappings can target:
| Internal name | Description |
|---|---|
container_id | Container identifier |
channel_id | Channel identifier |
tstart, tend | Sample interval start/end on the channels table (RLE) |
timestamp | Raw sample timestamp on the channels table (RAW mode; encoded into tstart/tend) |
is_plausible | Boolean plausibility flag on the channels table (RAW mode); consumed by drop_implausible_data |
start_ts, stop_ts | Measurement start/stop epoch timestamps on the container_metrics table — referenced by ContainerEvent to derive event-fact start/end |
value | Sample value (or attribute value on the EAV tag table) |
key | Attribute key on the EAV container_tags table |
priority | Tie-breaker column on the channel_mapping table |
project_id | Project scoping column |
parent_id | Parent/scope identifier |
source_channel | Source-channel identifier on the channel_mapping table |
data_key | Data-key identifier (default present on both channel_mapping and channel_metrics) |
channel_alias | Alias identifier on the channel_mapping table |
channel_name | Channel-name identifier on the channel_metrics table |
source_unit, target_unit | Source/target unit columns on the channel_mapping table |
unit | Unit name column on the unit_conversion table |
group_id | Unit-family identifier on the unit_conversion table |
conversion_factor | Per-unit factor on unit_conversion; also the per-channel factor name carried into the solve UDF |
DefaultSolver consumes every section's column_name_mapping, plus the
top-level project_id and the channel_mapping / unit_conversion sections.
Per-table filters are applied for container_tags, container_metrics,
channel_mapping, and channels. Filters on channel_tags, channel_metrics,
and poi_channels are accepted for forward compatibility but not yet applied.
Sections for tables you do not configure (e.g. channel_tags, channel_mapping)
are simply unused.
channels.filters are applied before raw encoding. A filter that removes
samples from the middle of a channel therefore bridges the surrounding interval
(the last good value is held across the gap) rather than splitting it.
Intended scope: use channels.filters in RAW mode only for whole-channel
scoping (a value constant across all of a channel's samples, e.g. a single
source/stream or project scoping), never for per-sample cleaning (value ranges,
quality flags, NaN drops). Per-sample cleaning leaves interior gaps that get bridged,
distorting the signal. To drop implausible samples with correct boundaries, use
drop_implausible_data.
When data_type = RAW, the validator rejects a channels.filters entry on
is_plausible (unambiguously per-sample) and warns on any other entry, since it
cannot tell scoping from cleaning statically.
In RLE mode there is no such restriction: the channels table is already encoded, so
a filter simply drops the matching [tstart, tend) interval rows without bridging, and
may target any column.
Example: DefaultSolver with renamed columns and per-table filters
"query_engine": {
"solver": "DefaultSolver",
"solver_config": {
"project_id": "my_project",
"container_tags": {
"column_name_mapping": {"entity_id": "container_id"},
"filters": {"parent_id": "my_parent_id"}
},
"container_metrics": {
"column_name_mapping": {"start_dt": "tstart", "stop_dt": "tend"}
},
"channel_metrics": {
"column_name_mapping": {}
},
"channel_mapping": {
"column_name_mapping": {},
"filters": {"toolbox_id": "my_toolbox"}
},
"channels": {
"column_name_mapping": {},
"filters": {"source": "live"}
}
}
}
Sections you don't customize can be omitted; defaults are an empty mapping and no filters.
Unit conversion (optional)
Set source.unit_conversion_table and extend channel_mapping with source_unit / target_unit columns
to have aliased selectors auto-convert values from source to target unit during solve(). Direct selectors
via query.channel(...) always return raw values, even on a channel that an aliased sibling converts —
conversion is a property of the alias, not of the channel. See
unit_conversion for the table schema.
"source": {
"container_metrics_table": "my_catalog.silver.container_metrics",
"channel_metrics_table": "my_catalog.silver.channel_metrics",
"channels_uri": "my_catalog.silver.channels",
"channel_mapping_table": "my_catalog.silver.channel_mapping",
"unit_conversion_table": "my_catalog.silver.unit_conversion"
},
"query_engine": {
"solver": "DefaultSolver",
"solver_config": {
"unit_conversion": {
"column_name_mapping": {}
}
}
}
Alias-resolution join keys (optional)
DefaultSolver.filter_aliased_channel_metrics joins channel_mapping
to channel_metrics to resolve aliased selectors. The default composite key
is (source_channel, channel_name) + (data_key, data_key). Override
channel_mapping.join_keys to change the arity or column choice — for
example, a single-column join when data_key is not part of the channel
identity in your silver layout:
"solver_config": {
"channel_mapping": {
"join_keys": [
{"mapping_col": "source_channel", "metrics_col": "channel_name"}
]
}
}
Each mapping_col / metrics_col is an internal name (the name as the
solver sees the column after column_name_mapping has been applied on
the respective table). The two sides of a pair are independent, so the same
column can carry different names on the two tables. For instance, a layout
where the data-key column has different physical names on the two tables
has two equivalent paths:
# Path 1 — rename both physical columns to the same internal name; the
# default join_keys then works unchanged.
"solver_config": {
"channel_mapping": {
"column_name_mapping": {"mapping_data_key": "data_key"}
},
"channel_metrics": {
"column_name_mapping": {"metrics_data_key": "data_key"}
}
}
# Path 2 — leave the physical names as-is and reference them directly.
"solver_config": {
"channel_mapping": {
"join_keys": [
{"mapping_col": "source_channel", "metrics_col": "channel_name"},
{"mapping_col": "mapping_data_key", "metrics_col": "metrics_data_key"}
]
}
}
query.channel(...) and query.channel_with_alias(...) kwargs are column
references against the post-column_name_mapping schema. If you
override join_keys (or skip renames) so that the solver sees a column
under a non-default name, the same name must be used as the kwarg. Example:
if join_keys references metrics_col: "my_chan_name" and the column is
not renamed via column_name_mapping, call
query.channel(my_chan_name=...). The internal-name properties on
SolverConfig exist primarily to remove magic strings from the solver
code; the user-facing contract is "kwarg name == column name as the solver
sees it".
When to use what
solver_config.<table>.column_name_mapping— your silver-layer column is named differently from the framework's internal name (e.g.entity_idinstead ofcontainer_id).container_filters.tag_filters/metric_filters— choose which containers participate in this particular run (supports comparators, OR/AND combinations, and type casting). Refer to internal column names whensolver_configrewrites them.
incremental (optional)
Incremental processing reuses results from prior runs for unchanged definitions and reprocesses only containers that are new or have been updated in silver. See the Report reference for mode-resolution rules and what counts as a definition change.
| Field | Type | Default | Description |
|---|---|---|---|
enabled | bool | false | Turns incremental processing on. |
silver_last_modified_column | str | "timestamp" | Silver column used to detect container updates. Name after column_name_mapping is applied (the physical name unless you remap that column). |
gold_last_modified_column | str | "_created_at" | Gold-side column used to detect prior-run freshness. |
When query_engine.max_containers_per_batch is set, a single
Report.run() commits at most that many upserted containers per iteration and loops until the
population is drained, so no one solve holds the whole population in memory. If a batch is detected
again unchanged after being persisted (no forward progress, e.g. a misconfigured
gold_last_modified_column), run() logs a warning and stops with partial completion rather than
looping forever.
full_recalculation (optional)
Forces a full recalculation of specific entities on an incremental run: the listed
aggregations, events, and/or calculated channels recompute over all matching containers and
have their gold rows fully replaced, regardless of whether their definition_hash changed. This is
the same treatment a definition change already receives, applied on demand — useful to backfill
after fixing a persistence bug or correcting upstream data, without rerunning everything in full
mode. In full mode it is a no-op, since every entity already recomputes over all containers.
Entities are identified by their human-readable name. A name that matches no registered entity
fails the run fast with a ValueError, so typos surface immediately rather than silently doing
nothing.
| Field | Type | Default | Description |
|---|---|---|---|
aggregations | list[str] | [] | Names of aggregations to fully recompute. |
events | list[str] | [] | Names of events to fully recompute. |
calculated_channels | list[str] | [] | Names of calculated channels to fully recompute. |
{
"full_recalculation": {
"aggregations": ["rpm_hist_p1"],
"events": ["overspeed"],
"calculated_channels": ["power_kw"]
}
}
full_recalculation only has an effect when the run is incremental. A first run, or any full run,
already recomputes every entity over all containers, so the scope is ignored.
calculated_channels (optional)
Controls the optional calculated_channel_metrics output. By default a report
writes calculated channels to calculated_channel_fact (the derived signal) and
calculated_channel_dimension (the definitions). Setting emit_channel_metrics
adds a third table, calculated_channel_metrics, shaped like the silver
channel_metrics table so the fact + metrics pair can serve as an Impulse silver
source. See the Channels reference for the
output schema.
| Field | Type | Default | Description |
|---|---|---|---|
emit_channel_metrics | bool | false | Turns on the calculated_channel_metrics table. |
attribute_columns | list[str] | [] | Calculated-channel attributes keys to surface as columns on the metrics table (e.g. ["unit"]). |
kpis | list[str] | ["duration", "min", "max", "mean"] | KPIs computed per (container_id, channel_id), one column each. Must be registered KPI names. |
When enabled, each row of calculated_channel_metrics is one
(container_id, channel_id) pair, carrying the selected kpis plus dynamic
identity columns (the union of identity keys across the report's channels) and
the configured attribute_columns. A channel that omits an identity or attribute
key gets null for that column; an identity key wins over an attribute key of the
same name.
The available kpis are duration, min, max, and mean (all
duration-weighted, matching the silver ingestion semantics). An unknown KPI name is
rejected at config validation with a ValueError naming the valid KPIs.
When emit_channel_metrics is false (the default), no metrics table is written
and attribute_columns / kpis have no effect.
measurement_dimensions (optional)
List of container_metrics column names to surface into the gold-layer
measurement_dimension table. Names are matched after
solver_config.container_metrics.column_name_mapping
has been applied — i.e. these are the internal (post-mapping) column
names, not the physical silver column names. Each name passes through
to gold verbatim, so the configured name is also the gold column name.
Any column present in your post-mapping container_metrics DataFrame
is a valid entry — there is no closed allow-list. Typical choices
include container_id, uut_id, project, vehicle_key, file_name,
file_path, start_ts, stop_ts, and environment, but any column
your silver schema carries (under its internal name) is fair game.
container_id is part of the default list and is recommended for any
real-world config: it is the upsert key used by incremental processing
and the join key between the measurement dimension and the event-fact
tables. If you override measurement_dimensions you take full
ownership of what ends up in gold — the framework does not inject
container_id for you. Omit it only if you know the consequences for
downstream joins and incremental runs.
Default:
[
"container_id",
"start_ts",
"stop_ts"
]
If any listed column is not present in the post-mapping
container_metrics DataFrame when the report runs, the run fails fast
with a ValueError naming the missing columns.
Worked example: physical name differs from internal name
Suppose your silver container_metrics table has a physical column
my_measurement_id (no container_id). Map it to the internal name
in solver_config, then reference the internal name in
measurement_dimensions:
{
"query_engine": {
"solver": "DefaultSolver",
"solver_config": {
"container_metrics": {
"column_name_mapping": { "my_measurement_id": "container_id" }
}
}
},
"measurement_dimensions": ["container_id", "start_ts", "stop_ts"]
}
The gold measurement_dimension table will have columns
container_id, start_ts, stop_ts. Listing "my_measurement_id" in
measurement_dimensions here would fail — by the time the framework
selects the dimensions, the column has already been renamed to
container_id.
Migration note (pre-0.1): earlier versions exposed a fixed enum
that renamed two silver columns on the way to gold (project →
project_id, file_path → source_file_path). The rename has been
removed; if you previously listed "project_id" or "source_file_path",
list "project" and "file_path" instead. The default list also
shrank — if you relied on the old eight-column default, add the
columns you want explicitly.