The benchmark holds one row for every active community water system in the United States: who it serves, what it spends and collects, what a household pays where a tariff is filed, and what that costs against local income.
No single agency publishes this. It is built from regulatory filings and official statistics, all of them filings rather than audits, and every figure is either as-published or computed from as-published values. The page beside a figure says which.
Accuracy
The hard problem is not any single figure. It is establishing that a utility in a financial record and a water system in a regulatory record are one operating entity, which no agency states. Where that is uncertain, the figure is withheld: ambiguity is refused and never resolved by guessing, and every system is tested against an independent measure of its scale before anything is published for it. A row with no figure is a normal outcome.
No accuracy rate is published yet. A sample of 200 matches is drawn for review, and the draw is addressed by the matching rule itself: change how a match is made and a different 200 are drawn, and any verdict written about the old rule is discarded and not carried forward. No rate is published until the current draw has been read in full.
That mechanism exists because of the failure this kind of number usually hides. An earlier review of 200 matches found a wrong one, and the rule was corrected in response. That was a good change, and behind the one match was a fault that had produced seven across the whole set. But from the moment the rule was corrected, those 200 rows measured nothing: they were the rows the rule had been fitted to. Every one of them then read as correct, which is what fitting to a sample looks like from the inside.
So the sample is not a fixed list that a correction can quietly contaminate. It moves with the rule, and the rows the rule is measured on are rows it has never been adjusted against.
When the rate is published it will state four things, because a bare percentage answers none of them:
- What it is the accuracy of. The join between a government and a water system, and nothing else. Not the figures either record carries.
- The denominator, and that the draw preceded the reading. 200 matches, selected on each government’s own identifier and not on its position.
- Who read them, by name.
- What it does not mean. A sampled rate is not a per-page confidence. It says what share of matches of this kind hold; it does not say that the match on the page in front of you is one of them.
Two mistakes are already known and are named beside it, never absorbed into it. Both produced well-formed rows carrying another utility’s figures, which is the exact failure every cell reference in a data pack exists to make impossible. One of them, a village matched to a district serving 200 people through one connection, turned out to be a fault that had produced seven wrong matches.
Why a figure is blank
Each blank prints its own reason on the page. There are five kinds:
- The underlying estimate is too noisy to divide by, against the collecting agency’s own reliability standard.
- The geography does not match. The system does not serve the same households as the government it belongs to.
- The value is not a measurement. A zero water utility revenue is almost always an unfilled form, and not a utility that collected nothing.
- The connections are not households. Some systems report wholesale delivery points as connections.
- The filed tariff does not reach the volume the benchmark prices.
Most describe what a regulator collects rather than what a utility did.
What the figures were checked against
Against utilities’ own annual reports and audited statements. Income and service population reproduce what the utility and the statistical agency publish. Operating cost agrees to about one per cent with audited water operating expense once depreciation and amortisation are removed from the audited figure; revenue runs about five per cent above the operating revenue subtotal in a utility’s segment statement.
Both gaps are definitional, and they matter to anyone comparing. The reporting basis here excludes depreciation, so it reads roughly a third low against an audited statement for a capital-intensive utility, and its revenue concept is broader. A figure here describes a government’s statutory return and never a utility’s own accounts. That is what makes thousands of systems comparable, and it is why they will not match what a finance office reports.
What is published
Every system serving over 100,000 people, plus a random sample of the rest stratified by size and region, with the roster fixed before the results were read. Every other system has a cohort position and no figure.
The distributions and cohort statistics are computed over all 49,378 systems, not over the published sample. The sample governs which systems are named, not which are counted.
The benchmark data lists what is downloadable, what is not, and how to request the rest.