Home / Products / Cogito
Data Cleaning & ValidationCogito

Translating raw utility records to analysis-ready customer data

A walkthrough of EcoMetricx's general cleaning, reconciliation, and cohort readiness — turning billing, AMI, and metadata feeds into clean, defensible output.

See the Operating Model
BillingAMIMetadataClean Output
Why Cleaning Matters

The primary challenge is disagreement across sources

Monthly billing, interval, and customer data describe the same service differently. Our process scales to any dataset size, addressing instances when even rare data issues occur.

01

Time

Bill cycles, service days, and AMI intervals do not share a calendar.

02

Customer + Meter Matching

Accounts, premises, meters, and customer IDs change at different rates.

03

Usage Quality

Rebills, estimated usage, grid/metering differences, and missing intervals create apparent variance — flagging usage anomalies (EV, Solar, etc.).

Proven at scale: millions of customers with monthly billing + AMI processed through this cleaning system.

Operating Model

A six-stage cleaning system turns raw feeds into defensible outputs

Each stage has explicit checks, decisions, and evidence — from raw feeds to analysis-ready history.

1 · Ingest

Prove completeness

Completeness is established before any cleaning rule runs. Electric and gas bill batches are combined with an outer join so customers are retained even when one fuel is absent, making billing gaps visible. Ingest controls include file count and arrival logging, schema and type validation, duplicate file detection, and row counts by batch.

2 · Standardize

Align schema + time

Service periods vary in length and cross calendar boundaries, so billing cycles are normalized to a common service-day basis. Calendar month usage is reconstructed by prorating average daily usage across the days each bill overlaps a month.

3 · Detect + Repair

Find gaps + outliers

Anomalies are classified and handled by type: gaps, spikes, flatlines, duplicates, and meter resets. Each defect is flagged with the original value retained, repaired under an explicit rule, or quarantined and excluded from release.

4 · Reconcile

Identify and resolve source differences

AMI interval usage summed over the bill service period is compared with billed usage for the same account, meter, and period. Differences are explained with reason codes, then accepted, repaired, reviewed, or excluded.

5 · Classify

Assign context + cohort

Customers are assigned to cohorts, with particular attention to solar and EV identification. Signal fusion combines utility flags, usage signals, and disaggregation evidence for stronger flags.

6 · Validate

Release with evidence

Before release, every change is validated and traceable: record-level evidence, customer evidence, cohort evidence, and release evidence (version, rule set, exceptions).

Validation questions at the final gate: Did every source file land? Can every repair be explained? Are cohorts mutually exclusive? Can the release be reproduced?

Anomaly Detection

A toolkit for flagging unusual customers

Pairing statistical and ML detectors with AMI usage and customer-attribute extracts to detect anomalous usage.

01 · Statistical

Univariate flags on account fields

IQR, z-score, MAD, jackknife

02 · ML-Based

Multivariate customer-profile outliers

Isolation Forest, LOF, OCSVM, MCD

03 · Time Series

Anomalies within one meter's history

Windowed stats, DARTS, TCN autoencoder

04 · Load-Shape

Group customers by usage pattern

PSD + PCA + clustering

05 · Applied Checks

Solar usage validation

Multiple signals produce a stronger solar flag: utility + usage + disaggregation evidence.

EV Detection

AMI spikes and ML models strengthen EV detection

Detect recurring daytime charging spikes in AMI usage, then confirm them with time-series and load-shape models before assigning the EV flag.
Continuous Reconciliation

Retroactive updates make reconciliation continuous

Utilities can revise prior data months later, requiring versioned archives and repeatable reconciliation. We preserve every version of every extract and flag changes that exceed a threshold set to your preference.

  • Prior snapshot: archived utility batch, version N
  • Reconcile N ↔ N+1: same account · meter · interval
  • Decision: accept · replace · review · retain
  • Archive + threshold: preserve every version, flag changes ≥ X SD from prior usage
Cogito meter-to-billing reconciliation view
Validation Interface

Results surfaced in a validation user interface

Every result is documented and visible, so coverage gaps, corrections, reconciliation results, and address matching are never a black box.

  • Validation summary with dated PDF exports
  • Coverage view with a source-by-month grid that makes a missing month visible at a glance
  • Reconciliation view showing matched meters, meters without accounts, accounts without meters, and billing variances
  • Corrections view that opens and quantifies every correction file, including how late it arrived and how much usage it restated
  • Address matching reduced to a small, scored review queue of exact matches, fuzzy matches, and uncertain pairs
Cogito coverage and missing months view
Our Process in a Nutshell

We ingest utility feeds, reconcile identity, time, usage, and context — then document and validate every decision

01 · Ingest

File intake + schema checks · account–premise–meter mapping

02 · Time

Billing-cycle normalization · AMI coverage and gaps

03 · Usage

Thresholds and outliers · revisions / estimated reads

04 · Context

Fuel cohort + solar/EV flag · customer / premise metadata

05 · Governance

Reason codes + audit trail · release validation

Result

Clean, defensible, analysis-ready customer data — with evidence behind every change.

Start with clean, defensible data

Clean, reconciled, versioned data is the foundation that makes analytics trustworthy. Talk to us about a data cleaning and validation engagement for your feeds.

Explore More Solutions