Skip to main content

Configuration

ImpulseConfig configures everything about a report: the silver-layer input tables, the gold-layer output location, container-level filters, the query-engine solver, incremental processing, and which container columns get surfaced into the gold-layer measurement dimension. Configuration is defined as JSON (or an equivalent Python dictionary) and validated using Pydantic models. The canonical schema lives in src/impulse_reporting/config/config_parser.py.

Quick example​

{
"source": {
"container_metrics_table": "my_catalog.silver.container_metrics",
"channel_metrics_table": "my_catalog.silver.channel_metrics",
"channels_uri": "my_catalog.silver.channels",
"container_tags_table": "my_catalog.silver.container_tags",
"channel_tags_table": "my_catalog.silver.channel_tags"
},
"unity_sink": {
"catalog": "my_catalog",
"schema": "gold",
"table_prefix": "my_report"
},
"query_engine": {
"solver": "DefaultSolver",
"data_type": "RAW"
},
"container_filters": {
"tag_filters": [
[
{ "tag_name": "uut_id", "comparator": "==", "value": "ABC123", "cast_type": "string" }
]
],
"metric_filters": [
[
{ "column_name": "start_dt", "comparator": ">=", "value": "2025-04-27T05:20:54.000Z", "value_type": "timestamp" }
]
]
},
"measurement_dimensions": ["container_id", "vehicle_key", "start_ts", "stop_ts"]
}

A configuration is passed to Report either as a Python dict (config=...) or as a JSON file path (config_path=...). Sinkless mode is also supported — see Sinkless reports.


source​

Maps the silver-layer input tables.

FieldTypeRequiredDescription
container_metrics_tablestrYesFull Unity Catalog path. Container metadata (timestamps, duration).
channel_metrics_tablestrYesFull Unity Catalog path. Channel-level statistics.
channels_uristrYesFull Unity Catalog path. Time-series sample data.
poi_channels_uristrNoFull Unity Catalog path to the Points-in-Time (POI) channel data table. Required only when selecting POI channels via poi_channel(); omit for sample-only data models.
container_tags_tablestrNoFull Unity Catalog path. Container EAV tags.
channel_tags_tablestrNoFull Unity Catalog path. Channel EAV tags.
channel_mapping_tablestrNoFull Unity Catalog path. Logical-to-physical channel alias table. Required when using QueryBuilder.channel_with_alias(). In reporting mode the resolved alias-to-physical-channel mapping is materialized to the gold-layer channel_mapping_resolution_dimension.
unit_conversion_tablestrNoFull Unity Catalog path. Per-unit-family conversion factors. When configured together with a channel_mapping_table whose rows carry source_unit / target_unit columns, aliased selectors auto-convert values from source to target unit during solve().

A container_tags_table is required to use tag-based container filters; a channel_tags_table is required to select channels by tag rather than by columns on channel_metrics.


unity_sink​

Defines where gold-layer tables are written.

FieldTypeRequiredDescription
catalogstrYesTarget catalog name.
schemastrYesTarget schema name.
table_prefixstrYesPrefix for all generated table names.
cleanup_temp_tablesboolNoDrop the batch-solving __impulse_temp_* tables from this schema after persist_results() completes successfully. Defaults to false. Overridable per call via persist_results(cleanup_temp_tables=...).

Output tables are named {table_prefix}_{entity} (e.g. my_report_histogram_fact).

Sinkless reports​

unity_sink is optional. When omitted, the report runs in sinkless mode: determine_report() still computes events, aggregations, and container dimensions and exposes them on the report object, but persist_results() becomes a no-op. Useful for ad-hoc analysis, notebooks, and tests where writing to Unity Catalog is not desired.


container_filters (optional)​

Restricts the set of processed containers. Filters are expressed in disjunctive normal form (OR of ANDs): each inner list is AND-combined, the outer list is OR-combined.

Two independent filter families:

  • tag_filters — applied on container_tags_table (EAV key/value model).
  • metric_filters — applied on container_metrics_table (columnar model).
FieldTypeDefaultDescription
tag_filterslist[list[TagFilter]][]Tag-based filter groups (DNF).
metric_filterslist[list[MetricFilter]][]Metric-based filter groups (DNF).

TagFilter​

FieldTypeRequiredDescription
tag_namestrYesTag key to filter on.
comparatorstrYesOne of ==, !=, >, >=, <, <=.
valueanyYesExpected value. Must match cast_type.
cast_typestrNostring (default), int, double, or timestamp (ISO-format string).

MetricFilter​

FieldTypeRequiredDescription
column_namestrYesColumn on container_metrics_table to filter on (e.g. start_dt, stop_dt). When solver_config.container_metrics.column_name_mapping is set, this refers to the internal name (after renaming).
comparatorstrYesOne of ==, !=, >, >=, <, <=.
valueanyYesExpected value. Must match value_type when provided.
value_typestrNoWhen provided, validates/converts the value: string, int, double, timestamp.

query_engine (optional)​

FieldTypeDefaultDescription
solverstr"DefaultSolver""DefaultSolver" adapts to the silver layer: it selects channels from a narrow EAV channel_tags table when source.channel_tags_table is set and otherwise from columns on channel_metrics; it filters containers via a narrow EAV container_tags table or, when source.container_tags_table is omitted, a wide-only container_metrics. "DeltaSolver" and "KeyValueStoreSolver" are deprecated aliases that resolve to DefaultSolver.
data_typestr"RLE""RLE" (intervals [tstart, tend)) or "RAW" (raw timestamps; converted to intervals at query time by the configured raw_encoder).
raw_encoderstrnull (resolves to "RLE" when data_type = "RAW")How RAW point data is converted to intervals. "RLE" (the default) run-length encodes on the fly to reduce memory consumption: consecutive samples with the same value collapse into a single [tstart, tend) interval per run. "INTERVAL" keeps every original sample, only deriving tend from the next sample's timestamp and dropping exact duplicate points — use it when downstream analysis needs all original timestamps. Only takes effect with data_type = "RAW"; ignored for "RLE" input. See How Impulse interprets intervals for validity semantics and what counts as a duplicate point.
drop_implausible_databoolfalseWhen true, drops channels rows where is_plausible = false. Requires data_type = "RAW"; combining with "RLE" raises a validation error.
max_channels_per_batchint500Caps how many distinct channels are read in one solve batch, to bound per-solve memory. Counts channel selections (each distinct channel(...) / poi_channel(...) in your expressions). Applies to events, aggregations, and calculated channels; for calculated channels it counts the input channels they read, not the number of calculated channels. The cap is a packing target, not a hard limit: an expression that on its own reads more channels than the cap still runs, in a batch of its own. Must be >= 1. The old name batch_size is a deprecated alias (still accepted, warns, removed in a future release); setting both raises a validation error.
max_containers_per_batchintnullCaps how many containers are solved at once, to bound per-solve memory. null disables the cap; when set it must be >= 1. On an incremental run, Report.run() commits at most this many upserted containers per iteration and loops until the population is drained (see incremental). On a full run, and when a changed definition recomputes over all containers, the population is split into chunks of about this size and solved chunk by chunk. Chunk sizes are approximate (the cap is a memory heuristic, not an exact bound).
solver_configSolverConfignullPer-table column mappings, per-table equality filters, and project scoping. Set project_id to scope reads by project — it is applied to container_tags (if configured), container_metrics, and channel_mapping (if configured), so it works in both narrow EAV and wide-only data models. Omit it when you don't need project scoping. See Solver column mappings and filters.

If query_engine is omitted, the default is DefaultSolver with data_type = "RLE".


Solver column mappings and filters​

The framework references columns by a fixed set of internal names (e.g. container_id, channel_id, tstart, tend, value). When your silver-layer tables use different physical names, declare the mapping in solver_config so the solver renames each table's columns at read time.

SolverConfig has one section per silver table. Each section is a TableConfig with two fields:

  • column_name_mapping (dict[str, str]): { "physical_column": "internal_column" }. The mapping is applied once, when the table is read. All downstream processing (filters, joins, aggregations) uses the internal names.
  • filters (dict[str, str]): equality filters applied after renaming. Keys are internal column names; values are literals to match. Useful for project/toolbox scoping where a single value should always be enforced. Values are strings but are coerced to the target column's type, so a boolean column takes "true" / "false" (an unparseable value errors at read time under ANSI mode).

Top-level fields on SolverConfig:

  • project_id (str, optional): Project scope. When set, the solver applies an equality filter on the project_id column (after column-name mapping) of every table it reads that carries one — container_tags (if configured), container_metrics, and channel_mapping (if configured). Omit it if you don't need project-level scoping; the solver does not require it.

Per-table sections (each a TableConfig):

SectionWhen it appliesTypical mappings
container_tagswhen container_tags_table is configuredentity_id → container_id, custom EAV key/value columns
container_metricsalwaysCustom container_id column, custom timestamp columns
channel_tagswhen channel_tags_table is configuredTag key/value column renames
channel_metricsalwaysCustom channel_id column, custom value/timestamp columns
channel_mappingwhen channel_mapping_table is configuredAlias-table column renames; priority column; optional join_keys for non-default alias-resolution composite keys
channelsalwaysRLE column renames (tstart/tend/value); filters supported (applied at read time, before non-solve columns are dropped)
unit_conversionwhen unit_conversion_table is configuredUnit-conversion table column renames (unit, group_id, conversion_factor)

Internal column names that mappings can target:

Internal nameDescription
container_idContainer identifier
channel_idChannel identifier
tstart, tendSample interval start/end on the channels table (RLE)
timestampRaw sample timestamp on the channels table (RAW mode; encoded into tstart/tend)
is_plausibleBoolean plausibility flag on the channels table (RAW mode); consumed by drop_implausible_data
start_ts, stop_tsMeasurement start/stop epoch timestamps on the container_metrics table — referenced by ContainerEvent to derive event-fact start/end
valueSample value (or attribute value on the EAV tag table)
keyAttribute key on the EAV container_tags table
priorityTie-breaker column on the channel_mapping table
project_idProject scoping column
parent_idParent/scope identifier
source_channelSource-channel identifier on the channel_mapping table
data_keyData-key identifier (default present on both channel_mapping and channel_metrics)
channel_aliasAlias identifier on the channel_mapping table
channel_nameChannel-name identifier on the channel_metrics table
source_unit, target_unitSource/target unit columns on the channel_mapping table
unitUnit name column on the unit_conversion table
group_idUnit-family identifier on the unit_conversion table
conversion_factorPer-unit factor on unit_conversion; also the per-channel factor name carried into the solve UDF
Feature support

DefaultSolver consumes every section's column_name_mapping, plus the top-level project_id and the channel_mapping / unit_conversion sections. Per-table filters are applied for container_tags, container_metrics, channel_mapping, and channels. Filters on channel_tags, channel_metrics, and poi_channels are accepted for forward compatibility but not yet applied. Sections for tables you do not configure (e.g. channel_tags, channel_mapping) are simply unused.

Channels filters in RAW mode

channels.filters are applied before raw encoding. A filter that removes samples from the middle of a channel therefore bridges the surrounding interval (the last good value is held across the gap) rather than splitting it.

Intended scope: use channels.filters in RAW mode only for whole-channel scoping (a value constant across all of a channel's samples, e.g. a single source/stream or project scoping), never for per-sample cleaning (value ranges, quality flags, NaN drops). Per-sample cleaning leaves interior gaps that get bridged, distorting the signal. To drop implausible samples with correct boundaries, use drop_implausible_data.

When data_type = RAW, the validator rejects a channels.filters entry on is_plausible (unambiguously per-sample) and warns on any other entry, since it cannot tell scoping from cleaning statically.

In RLE mode there is no such restriction: the channels table is already encoded, so a filter simply drops the matching [tstart, tend) interval rows without bridging, and may target any column.

Example: DefaultSolver with renamed columns and per-table filters​

"query_engine": {
"solver": "DefaultSolver",
"solver_config": {
"project_id": "my_project",
"container_tags": {
"column_name_mapping": {"entity_id": "container_id"},
"filters": {"parent_id": "my_parent_id"}
},
"container_metrics": {
"column_name_mapping": {"start_dt": "tstart", "stop_dt": "tend"}
},
"channel_metrics": {
"column_name_mapping": {}
},
"channel_mapping": {
"column_name_mapping": {},
"filters": {"toolbox_id": "my_toolbox"}
},
"channels": {
"column_name_mapping": {},
"filters": {"source": "live"}
}
}
}

Sections you don't customize can be omitted; defaults are an empty mapping and no filters.

Unit conversion (optional)​

Set source.unit_conversion_table and extend channel_mapping with source_unit / target_unit columns to have aliased selectors auto-convert values from source to target unit during solve(). Direct selectors via query.channel(...) always return raw values, even on a channel that an aliased sibling converts — conversion is a property of the alias, not of the channel. See unit_conversion for the table schema.

"source": {
"container_metrics_table": "my_catalog.silver.container_metrics",
"channel_metrics_table": "my_catalog.silver.channel_metrics",
"channels_uri": "my_catalog.silver.channels",
"channel_mapping_table": "my_catalog.silver.channel_mapping",
"unit_conversion_table": "my_catalog.silver.unit_conversion"
},
"query_engine": {
"solver": "DefaultSolver",
"solver_config": {
"unit_conversion": {
"column_name_mapping": {}
}
}
}

Alias-resolution join keys (optional)​

DefaultSolver.filter_aliased_channel_metrics joins channel_mapping to channel_metrics to resolve aliased selectors. The default composite key is (source_channel, channel_name) + (data_key, data_key). Override channel_mapping.join_keys to change the arity or column choice — for example, a single-column join when data_key is not part of the channel identity in your silver layout:

"solver_config": {
"channel_mapping": {
"join_keys": [
{"mapping_col": "source_channel", "metrics_col": "channel_name"}
]
}
}

Each mapping_col / metrics_col is an internal name (the name as the solver sees the column after column_name_mapping has been applied on the respective table). The two sides of a pair are independent, so the same column can carry different names on the two tables. For instance, a layout where the data-key column has different physical names on the two tables has two equivalent paths:

# Path 1 — rename both physical columns to the same internal name; the
# default join_keys then works unchanged.
"solver_config": {
"channel_mapping": {
"column_name_mapping": {"mapping_data_key": "data_key"}
},
"channel_metrics": {
"column_name_mapping": {"metrics_data_key": "data_key"}
}
}

# Path 2 — leave the physical names as-is and reference them directly.
"solver_config": {
"channel_mapping": {
"join_keys": [
{"mapping_col": "source_channel", "metrics_col": "channel_name"},
{"mapping_col": "mapping_data_key", "metrics_col": "metrics_data_key"}
]
}
}

query.channel(...) and query.channel_with_alias(...) kwargs are column references against the post-column_name_mapping schema. If you override join_keys (or skip renames) so that the solver sees a column under a non-default name, the same name must be used as the kwarg. Example: if join_keys references metrics_col: "my_chan_name" and the column is not renamed via column_name_mapping, call query.channel(my_chan_name=...). The internal-name properties on SolverConfig exist primarily to remove magic strings from the solver code; the user-facing contract is "kwarg name == column name as the solver sees it".

When to use what​

  • solver_config.<table>.column_name_mapping — your silver-layer column is named differently from the framework's internal name (e.g. entity_id instead of container_id).
  • container_filters.tag_filters / metric_filters — choose which containers participate in this particular run (supports comparators, OR/AND combinations, and type casting). Refer to internal column names when solver_config rewrites them.

incremental (optional)​

Incremental processing reuses results from prior runs for unchanged definitions and reprocesses only containers that are new or have been updated in silver. See the Report reference for mode-resolution rules and what counts as a definition change.

FieldTypeDefaultDescription
enabledboolfalseTurns incremental processing on.
silver_last_modified_columnstr"timestamp"Silver column used to detect container updates. Name after column_name_mapping is applied (the physical name unless you remap that column).
gold_last_modified_columnstr"_created_at"Gold-side column used to detect prior-run freshness.
Draining a capped incremental run

When query_engine.max_containers_per_batch is set, a single Report.run() commits at most that many upserted containers per iteration and loops until the population is drained, so no one solve holds the whole population in memory. If a batch is detected again unchanged after being persisted (no forward progress, e.g. a misconfigured gold_last_modified_column), run() logs a warning and stops with partial completion rather than looping forever.


full_recalculation (optional)​

Forces a full recalculation of specific entities on an incremental run: the listed aggregations, events, and/or calculated channels recompute over all matching containers and have their gold rows fully replaced, regardless of whether their definition_hash changed. This is the same treatment a definition change already receives, applied on demand — useful to backfill after fixing a persistence bug or correcting upstream data, without rerunning everything in full mode. In full mode it is a no-op, since every entity already recomputes over all containers.

Entities are identified by their human-readable name. A name that matches no registered entity fails the run fast with a ValueError, so typos surface immediately rather than silently doing nothing.

FieldTypeDefaultDescription
aggregationslist[str][]Names of aggregations to fully recompute.
eventslist[str][]Names of events to fully recompute.
calculated_channelslist[str][]Names of calculated channels to fully recompute.
{
"full_recalculation": {
"aggregations": ["rpm_hist_p1"],
"events": ["overspeed"],
"calculated_channels": ["power_kw"]
}
}
Incremental only

full_recalculation only has an effect when the run is incremental. A first run, or any full run, already recomputes every entity over all containers, so the scope is ignored.


calculated_channels (optional)​

Controls the optional calculated_channel_metrics output. By default a report writes calculated channels to calculated_channel_fact (the derived signal) and calculated_channel_dimension (the definitions). Setting emit_channel_metrics adds a third table, calculated_channel_metrics, shaped like the silver channel_metrics table so the fact + metrics pair can serve as an Impulse silver source. See the Channels reference for the output schema.

FieldTypeDefaultDescription
emit_channel_metricsboolfalseTurns on the calculated_channel_metrics table.
attribute_columnslist[str][]Calculated-channel attributes keys to surface as columns on the metrics table (e.g. ["unit"]).
kpislist[str]["duration", "min", "max", "mean"]KPIs computed per (container_id, channel_id), one column each. Must be registered KPI names.

When enabled, each row of calculated_channel_metrics is one (container_id, channel_id) pair, carrying the selected kpis plus dynamic identity columns (the union of identity keys across the report's channels) and the configured attribute_columns. A channel that omits an identity or attribute key gets null for that column; an identity key wins over an attribute key of the same name.

The available kpis are duration, min, max, and mean (all duration-weighted, matching the silver ingestion semantics). An unknown KPI name is rejected at config validation with a ValueError naming the valid KPIs.

Off by default

When emit_channel_metrics is false (the default), no metrics table is written and attribute_columns / kpis have no effect.


measurement_dimensions (optional)​

List of container_metrics column names to surface into the gold-layer measurement_dimension table. Names are matched after solver_config.container_metrics.column_name_mapping has been applied — i.e. these are the internal (post-mapping) column names, not the physical silver column names. Each name passes through to gold verbatim, so the configured name is also the gold column name.

Any column present in your post-mapping container_metrics DataFrame is a valid entry — there is no closed allow-list. Typical choices include container_id, uut_id, project, vehicle_key, file_name, file_path, start_ts, stop_ts, and environment, but any column your silver schema carries (under its internal name) is fair game.

container_id is part of the default list and is recommended for any real-world config: it is the upsert key used by incremental processing and the join key between the measurement dimension and the event-fact tables. If you override measurement_dimensions you take full ownership of what ends up in gold — the framework does not inject container_id for you. Omit it only if you know the consequences for downstream joins and incremental runs.

Default:

[
"container_id",
"start_ts",
"stop_ts"
]

If any listed column is not present in the post-mapping container_metrics DataFrame when the report runs, the run fails fast with a ValueError naming the missing columns.

Worked example: physical name differs from internal name​

Suppose your silver container_metrics table has a physical column my_measurement_id (no container_id). Map it to the internal name in solver_config, then reference the internal name in measurement_dimensions:

{
"query_engine": {
"solver": "DefaultSolver",
"solver_config": {
"container_metrics": {
"column_name_mapping": { "my_measurement_id": "container_id" }
}
}
},
"measurement_dimensions": ["container_id", "start_ts", "stop_ts"]
}

The gold measurement_dimension table will have columns container_id, start_ts, stop_ts. Listing "my_measurement_id" in measurement_dimensions here would fail — by the time the framework selects the dimensions, the column has already been renamed to container_id.

Migration note (pre-0.1): earlier versions exposed a fixed enum that renamed two silver columns on the way to gold (project → project_id, file_path → source_file_path). The rename has been removed; if you previously listed "project_id" or "source_file_path", list "project" and "file_path" instead. The default list also shrank — if you relied on the old eight-column default, add the columns you want explicitly.