How the data is built
Methodology
How we turn raw public records into sourced, structured intelligence, and what we do and don't claim.
The pipeline
- Ingest. We pull directly from official public-records sources (see Data Sources), campaign-finance disclosures, IRS filings, business and lobbying registrations, court records, and government datasets.
- Resolve. Donors, committees, companies, and people appear under many name and employer variants. We resolve them into consistent entities so a person or PAC is counted once, not ten times.
- Connect. We link entities into a graph, who gave to whom, who sits on which board, who lobbies for what, so relationships become queryable.
- Analyze. Aggregations, network measures, industry classification, and (where the data supports it) associations between money and official action.
What we claim
We report sourced facts and statistical associations. When we show that a legislator took money from an industry and voted its way, that is a correlation we document, not a claim of a quid pro quo. We say so plainly, every time.
Limits
- Industry and sector tags are assigned from names, employers, and occupations by rule and are approximate.
- Some data lags (state bulk files are periodic); we note the coverage window on every output.
- Public records contain errors; figures should be verified against source records before publication.
Corrections
We fix mistakes promptly and visibly. See our Neutrality Policy for how we handle disputes.