What it is
An agency publishes a table. It is a spreadsheet with merged headers, or a PDF, or a portal that returns one system at a time. We pull it, parse it, check the totals against the components the same file publishes, and hand you the result as data.
The checking is most of the value. A published total and the components it is made of can disagree, and until somebody adds the components up nothing says so.
What you hand us
The publication, or a link to it. If it is behind a login or a request process, the access — we do not scrape around a gate.
What you get back
One row per record, with the figures as published, the figures as computed from their own components where those differ, and the coordinate each was read from. Re-run on your schedule.
What it is built from
Your source, plus the extraction map that records which header each field came from and which transform was applied. That map is what lets a figure be traced back to a cell rather than to a process.
What it does not do
It does not correct the source. Where a published figure is wrong we report the discrepancy and publish both numbers; we do not substitute our own and we do not quietly repair a total.
It does not read a scanned page. A PDF of an image needs a person, and we will say so before quoting rather than after.
It does not give you an API. You get files on a schedule; there is no endpoint to integrate against and no account to set up.