| How it was counted | Verification tools | Share |
|---|---|---|
| By matching words in the title | ~10 of 99 | 10.1% |
| By reading every tool | 15 of 99 | 15.2% |
How this was built, and what it cannot verify
Honesty is the product
This site makes two kinds of claim and no third kind. Claims about the past are counted from a collection of documents. Claims about the future carry a probability. This page says where the first kind comes from, what had to be corrected along the way, and where the whole thing stops.
Where the numbers come from
Behind the site sits a collection of 661 items on AI and economics: working papers, journal articles, commentary, guides, tools, courses, talks and threads. 530 of them have their full text attached, obtained wherever that could legitimately be done. Some arrived through an automatic watcher of the places this field publishes, the rest were added by hand.
Every figure in the prose is counted from that collection when the page is built, and none is typed in from memory. Every work cited in the margins is looked up by its record, so a reference that does not exist stops the build instead of reaching the reader. A fabricated citation cannot ship quietly.
The collection is also frozen for this release. It keeps growing underneath the site, and a document claiming that every figure is measured cannot have its figures shift overnight while someone is reading it. This release was pinned on 27 July 2026. Re-pinning is a deliberate step, not something a sync does by accident.
Three things that went wrong
The site’s own argument is that cheap results still need checking, so it would be a poor advert for that argument to hide the checks that failed here.
A keyword filter got the headline number wrong. The most quotable statistic on the site is how much of the field’s AI tooling exists to verify something. It was first produced by matching tool titles against words like review, replication and audit. Reading all 99 tools one by one instead, and recording a verdict for each, gives a different answer.
The filter missed 6 verifiers whose names say nothing about what they do, and wrongly caught one literature screener with the word review in its title. What a thing is called is not what it does. The audited count is the one the site reports, and the hand labels ship with the site so anyone can disagree with a row.
Six things were counted twice. The same work can appear as several web pages: a set of skill files that are one bundle, a tool and the repository behind it, a product with an organisation page and a package page. Those were marked by hand as duplicates and then, for a while, the site read past that marking and counted them separately. Three published numbers were wrong as a result. 15 rows now fold together wherever tools are counted.
Several works sat in the wrong year. The site was reading the date each source happened to supply rather than the corrected date the pipeline had already worked out. One eleven month error moved a well known paper half a year along the timeline, and it also switched off the estimator that fills other gaps, so 43 further papers silently lost their dates. The repair replaced guessing with reading: every NBER paper’s issue date is now taken from that paper’s own page. Dated works rose from 188 to 278.
Between two releases the verification share of tools fell. The field did not build fewer referees. The count of verification tools roughly doubled while the number of tools counted nearly tripled, because the automatic watcher started sweeping in general AI infrastructure that is not economics research tooling at all. The share is a fact about this collection. The count is a fact about the field.
How much of the link graph is observable
The link structure in Part 1 is the most load-bearing descriptive claim on this site, and it is computed on the smallest denominator. A share like “half of a paper’s links land on another paper” counts only links that resolve to something already in the collection. Everything else a work points at is invisible to it.
| Register | External links | Land in the corpus | Share observable |
|---|---|---|---|
| commentary / Substack | 2,519 | 156 | 6.2% |
| working paper / article | 2,407 | 69 | 2.9% |
| other | 1,311 | 314 | 24% |
| tool / skill | 760 | 41 | 5.4% |
| guide / how-to | 527 | 31 | 5.9% |
| course | 279 | 34 | 12.2% |
| talk / podcast | 90 | 0 | 0% |
| thread | 19 | 1 | 5.3% |
Corpus wide that is 646 observable links out of 7,912, or 8.2%. The register that carries the most weight in the argument, papers, is the least observable at 2.9%. Recorded talks have 0 observable outbound links, so no statement about how talks link is supported by anything.
The visible slice is also not a random sample of the field’s links. It is precisely the subset that points at material the curator already collected, so a high self linking rate is partly a restatement of what was gathered. This is the same endogeneity the forecast layer guards against with leave one source out checks, applied here to the descriptive half.
What moves when an acquisition source is dropped
The collection is assembled from named acquisition sources. Recomputing the descriptive headline numbers with each source removed shows which of them are facts about the field and which are facts about the collector.
| Measure | All sources | Drop local | Drop upstream |
|---|---|---|---|
| Rows | 605.0 | 177.0 | 428.0 |
| Full text held (%) | 82.3 | 86.4 | 80.6 |
| Papers (% of rows) | 35.7 | 29.4 | 38.3 |
| Tools (% of rows) | 16.4 | 14.1 | 17.3 |
| Paper → paper links (%) | 50.7 | 0.0 | 50.0 |
| …observable edges behind it | 69.0 | 1.0 | 22.0 |
| Tool → tool links (%) | 58.5 | — | 54.5 |
| …observable edges behind it | 41.0 | 0.0 | 22.0 |
| Practice-layer docs | 123.0 | 38.0 | 85.0 |
| Anthropic share, practice layer (%) | 59.3 | 78.9 | 50.6 |
Read the spread across the row, not the level: this table is recomputed by scenario/R/build_guards.R without the duplicate folding build_metrics.R applies, so the “all sources” column is its own baseline rather than the headline figure.
Three readings matter. Full text coverage and register composition are robust, moving a few points at most, so the shape of the collection is not an artifact of one feed. The vendor share is not: the Anthropic share of the practice layer swings from 50.6% to 78.9% depending on which source is dropped, which is a wide enough band that P15 should be read as a claim about this collection rather than about practice. And the self linking rates do not survive the test at all, not because the share moves but because the denominator collapses. Dropping the hand added source leaves 1 observable paper edge and 0 tool edges. There is no measurement left to be robust or fragile.
These guards were measured over the 605 pinned records still present in the database, against the 661 the frozen census names; see the note on the frozen cohort below.
What this page cannot verify
A collection is not a census. This is one curator’s map, and it leans toward methods and tooling because that is what the curator works on. 55.3% of dated works fall in the first half of 2026, which is partly the field flooding and partly the newest material being easiest to find.
The frozen census no longer fully reproduces. Every database scored number on this site is measured over a pinned list of 661 records, so that a forecast cannot be moved by the collection growing underneath it. 56 of those records are no longer in the database, because the curation ledger records explicit removals while the pin is a silent filter. Nothing published here has moved, since the figures are committed files, but a rebuild today would quietly compute them over a smaller cohort. The freeze needs either a re-cut with a stated reason or a loud failure when a pinned record goes missing. Recorded here rather than repaired quietly, because a canary that fires is a decision, not a number to bump.
Some counts are judgments. Every tool was classified by reading it, against one question: does it check another machine’s output? A screener for systematic reviews, a meta analysis extractor and a paper generator that contains a reviewer are genuinely hard cases, and another reader would call two or three of them differently.
The missing dates are not random. 273 items (41.4%) carry no usable date, and they are mostly the living guides, tool pages and homepages that are revised rather than published. Those never had a date and never will. So every chart of time on this site under counts exactly the fast moving layer the argument says matters, and no reweighting fixes it, because the missing dates do not exist. They are shown as a visible residual on the Timeline page rather than quietly dropped.
The story is not in the database. The collection supplies counts, dates, titles and links. Which paper argues what, and what it all means, is reading and interpretation. Those are two different kinds of claim and the site tries not to blur them.
The forecast is not measurement. Nothing past 2028 is measured, and most of it turns on choices institutions make rather than on what the models can do. The error bars are meant to be read as wide. What is promised instead is that the bets are written down in advance, in public, and graded later on the Predictions scoreboard.
What gets in, and what gets published
Anything on AI and economics worth a record goes in, from any source, refereed or not. A Substack guide and an NBER working paper are both records here. Full text goes in wherever it can legitimately be obtained, and where it cannot, the gap is recorded row by row rather than waved away. A page that merely points at content, with nothing of its own to read, is marked as such and licenses no claim about what lies behind it.
The site publishes counts, titles, authors, dates and tags. It never publishes the curator’s private notes or the fetched text of anyone’s document. Titles appear because a title is a bibliographic fact. The documents themselves do not.
Colophon
The prose was drafted and revised with AI coding agents, a disclosure that belongs on a site about exactly that. Every past claim resolves to a record in the collection, every number is recounted when the page is built, the tool classification was done by hand against a fixed list, and the forecast layer is pinned to a manifest that records the freeze date and a checksum for every file a score depends on. Still owed, and owned by the curator: an outside panel to grade the scoreboard, a browsable copy of the collection, and the hand labels published as a dataset in their own right.
If the prose and the underlying collection ever disagree, the collection is right.