The breach sources we decided not to use
Three of the largest available datasets would have doubled our coverage on day one. We left them out, and this is the reasoning in full.
Coverage is the easiest number to grow in this category and the easiest to grow dishonestly. Several widely traded compilations are aggregations of aggregations, with no provenance and a meaningful share of fabricated rows.
Two failure modes matter to a user. A false positive sends someone to rotate credentials that were never exposed, which burns trust. A stale positive re-alerts on a 2013 breach they already handled, which burns attention.
Our test is boring: can we name the origin, date the event, and describe what fields were in it? If not, it does not ship, however large it is.
The cost is real — our launch coverage is smaller than a competitor who asks fewer questions. We would rather explain a gap than an invented alert.