Methodology

How a page becomes a dated fact, and what is deliberately withheld.

The path a page takes

Capture
The page is fetched and the bytes are kept. This is the only step that cannot be redone later. A page not captured today is gone.
Reading
The capture is parsed into plans, prices, and features. A reading is stamped with the version of the parser that produced it, so a better parser can re-read old captures without erasing what the old one said.
Change
Two consecutive clean readings are compared. Differences become typed changes, each pointing at both readings.
Fact
Changes become dated intervals: what was true, and when we learned it. Both dates are kept, because a capture from the Internet Archive is read years after it was taken.
Measurement
Facts are counted across a market, over a fixed cohort, and published only if coverage is high enough.

What is refused

A reading that might be wrong

Each reading is marked clean, partial, or failed. That is a statement about whether the page was read correctly, not about whether its prices are right. Only clean readings are compared, because a garbled parse and a large price cut produce the same diff.

A measurement without a cohort

Archive coverage improves over time, so counting whoever happens to be present each year produces a rising line that means nothing. Historical measurements use only companies present across the whole window, and nothing is published below 60% coverage or five companies.

A guess about what a feature is

Feature names are matched against a hand-curated list. Nothing is matched approximately. A term that cannot be matched is set aside for a human rather than filed under a near-miss, and it is left out of the counts rather than bundled into "other".

A confidence score

No probability is attached to a claim about the world. A pattern reports how many signals support it, how many contradict it, over how long, and what would refute it.

What this cannot tell you

Where the archive has gaps, it says so.