What it is
A cohort of comparable systems with their filed figures normalized across years, so the comparison is between numbers that mean the same thing. Systems are matched on the identifier the regulator publishes, not on name, because the same utility is filed under different names by different agencies.
Filings that cannot be relied on are flagged and kept, not silently dropped. A cohort that quietly excludes its awkward members is a cohort that flatters whatever it is compared against.
What you hand us
A list of systems, or the criteria that define the cohort — state, county, connection band, filing year. If you already have a peer set, hand us that and we will tell you which of its members have a filing history worth comparing.
What you get back
One row per system-year, with the loss figure, the year it was filed for, the identifier it was filed under, and a flag on any figure whose series moves too much to rely on. Every row names the archived file and the cell it came from.
What it is built from
The filings each state’s regulator publishes, archived as published, and re-parsed from those bytes. See what we assemble, and where it comes from.
What it does not do
It does not compare systems across states. Two regulators asking two different questions produce two populations, and a figure from one is not a benchmark for the other — a cohort we build stays inside one regulator’s population and says so.
It does not tell you a system is performing badly. This corpus holds a published threshold for one of its states, so for the rest there is nothing here to be over — which is a statement about what we have read, not about what those states have set — and a percentile is a position rather than a verdict.
It is not a live feed. It is a file on a schedule, and the schedule is as often as the regulator publishes, which for most of these is once a year.