A model of verified judgment
The branch as comparative statics on the cost of audit
This is a formal sketch rather than a finished result. It turns the essay’s central claim, that generation got cheap while verification stayed dear, into the smallest model that makes the ending a comparative static rather than a mood. The economics is mine to own and refine. The parameters are illustrative, and the propositions are stated as sketches a referee would still want to break.
The essay’s argument is one sentence long. Producing plausible research became nearly free, while confirming that a result is real stayed costly, so the scarce factor is verified judgment. That sentence has an equilibrium behind it. Below is the smallest version I can write down honestly.
Setup, a market for plausible results
A unit continuum of authors each submit one paper to a gatekeeper, which might be a journal, a funder’s panel or a data editor. A paper has quality \(q \in \{H, L\}\). A high quality paper is a real result. A low quality paper is plausible but false, the paper mill output that reads like research and is not.
Two costs sit on the author. Producing a genuinely high quality result costs \(k\), the real work, including the author’s own checking. Generating a plausible submission costs \(g\), and the premise of the whole essay is that \(g\) has collapsed toward zero. The price of a competent looking draft is now a few dollars of inference. The menu pricing literature is where that premise is measured. Menu Pricing of Large Language Models documents how model providers price generation down a demand curve, which is the supply side of the flood. Verification did not fall the same way, and that asymmetry, with \(g \to 0\) while \(k\) stays bounded away from zero, is the entire engine here.
A published paper returns private value \(V\), meaning a line on the CV, a grant, standing. A low quality paper that is caught returns a penalty \(\varphi \ge 0\), meaning the desk reject, the retraction, the reputational mark. Normalize the outside option, which is not submitting at all, to zero.
Grounding in the corpus
Dirk Bergemann, Alessandro Bonatti, Alex Smolin · Menu Pricing of Large Language Models
Susan Athey, Fiona Scott Morton · Artificial Intelligence, Competition, and Welfare (NBER WP 34444)
Verification technology, the scarce input
The gatekeeper chooses an audit intensity \(a \in [0,1]\), applied per submission. Read \(a\) as the probability that a low quality paper is caught. Take \(\pi(a) = a\) for cleanliness, and assume a true result never fails the audit, so there are no false positives, a simplification I return to. Auditing costs
\[ c(a, \theta) = \theta \, \psi(a), \qquad \psi' > 0,\ \psi'' > 0,\ \psi(0)=0, \]
per submission, with \(\psi\) increasing and convex, so deeper checking costs more than proportionately. The scalar \(\theta > 0\) is the cost of verification technology, human referee hours at one extreme and cheap agentic reproduction, the “referee factory”, at the other. This is the binding input. Verified judgment is what the model is short of, and \(c(a, \theta)\) is its price.
The gatekeeper’s problem, and the branch
Sustaining trust is worth something. Let \(B > 0\) be the per submission value to the gatekeeper of a pool readers can trust, over a pool they cannot, which is the difference between a literature people build on and one they discount to reputation. The gatekeeper sets \(a \ge a^{*}\), and so pays for trust, when the audit it requires is cheaper than the trust it buys.
\[ \theta\,\psi(a^{*}) \;\le\; B \quad\Longleftrightarrow\quad \theta \;\le\; \bar\theta \equiv \frac{B}{\psi(a^{*})} . \]
This is the branch, and it is a comparative static on one parameter.
Proposition 2 (the ending is priced by audit cost, not capability). Hold model capability fixed. There is a threshold \(\bar\theta\) such that the gatekeeper selects the audit ending, where \(a \ge a^{*}\) and verification is mandatory and rewarded, if and only if \(\theta \le \bar\theta\), and the flood ending, where \(a < a^{*}\) and the pool pools, otherwise. The selected ending turns on the cost of verification \(\theta\), not on how good the models are.
Intuition. The same capability jump that drove \(g \to 0\), feeding the flood, also drives \(\theta \to 0\), because agentic reproduction can re-run an analysis, chase a citation, or re-derive a proof at inference cost. So capability pushes both levers. It raises \(a^{*}\), as Proposition 1 shows, so more audit is now required, and it lowers \(\theta\), so that audit is now affordable. Which effect the institution acts on is the ending. The technology does not choose. The journal, the funder, and the data editor choose, by deciding whether machine verification is required and rewarded.
The flood as a capacity constraint
The flood ending has a sharper reading than “gatekeepers declined to audit.” Let realized audit be \[ a_{\text{real}} = \min\!\big(a,\, R/N\big), \] where \(R\) is refereeing capacity and \(N\) the submission volume. With human referees \(R\) is fixed at throughput. As \(g \to 0\), \(N \to \infty\), so \(R/N \to 0\) and \(a_{\text{real}} \to 0\) even for a gatekeeper who wants to audit. At fixed human capacity the audit ending is not merely unchosen, it is impossible. The referee factory, where \(\theta \to 0\), matters precisely because it relaxes the constraint on \(R\). So the two endings map cleanly onto one structural question, which is whether audit capacity scales with the flood.
But machine capacity is not human capacity, and this collection says so. An earlier version of this section let \(R\) grow freely with machine audit. The corpus’s own best evidence points the other way. Toda’s controlled test puts a model against economic theory beyond its knowledge cutoff; the radiology study measures human–AI collaboration directly; and the coding-agent work finds output that is empirically consistent but interpretively vulnerable, which is a precise way of saying the machine produces something a human must still read. Each describes machine checking as a complement terminating in a scarce human sign off, never as a substitute for it. The site’s own forecast ladder already concedes the point, since tier 2 of P5 explicitly requires a human signing off. So write
\[ a_{\text{real}} = \min\!\big(a,\ \delta R / N\big), \qquad 0 < \delta < 1, \]
where \(\delta\) is the fraction of machine flagged work that a human actually clears. The qualitative result survives, since capacity still scales where it did not before, but the flood ending is no longer defeated by capability alone. It is defeated by capability times whatever sign off throughput institutions are willing to staff, and \(\delta\) is a policy variable, not a technical one.
Why the discount is there
Alexis Akira Toda · Can AI Refute Economic Theory? Evidence from Beyond the Knowledge Cutoff (arXiv)
Paul Goldsmith-Pinkham, Chenhao Tan, Alexander K. Zentefis · Human-AI Collaboration in Radiology: The Case of Pulmonary Embolism
Meysam Alizadeh, Fabrizio Gilardi, Mohsen Mosleh, Enkelejda Kasneci · AI Coding Agents in Social Science: Methodologically Diverse, Empirically Consistent, Interpretively Vulnerable
Yiqing Xu, Leo Yang · Scaling Reproducibility: An AI-Assisted Workflow for Large-Scale Replication and Reanalysis (arXiv)
The no false positives assumption is doing more work than a simplification should. Setting false positives to zero is analytically convenient, but it removes the asymmetry that will actually decide adoption: a journal fears rejecting one true paper far more than it fears passing several false ones. Under that loss function a gatekeeper facing an imperfect machine auditor rationally sets \(a\) below the level this model derives, and may decline machine audit entirely at a false positive rate that a symmetric loss would call acceptable. The threshold in Proposition 1 is therefore an upper bound on audit intensity, not a prediction of it.
Proposition 2 presumes an instrument that barely exists. Pricing the ending by audit cost assumes the checking technology is available to be cheapened. On this collection’s own count, 3 of 15 verification tools check a result rather than a manuscript, and not one has been benchmarked against human referees. Proposition 2 should be read as conditional on the results checker of P14 arriving, which is a forecast rather than a fact. When it does not, trust does not vanish so much as move. Readers go on author identity and institution instead of on the audit, because the audit no longer tells them anything. That is the whitelist world the essay names.
A hybrid where trust is priced in a certification market
There is a third region between the two endings. Suppose a private certifier offers audit at price \(p\) per paper while the journal lags at \(a < a^{*}\). High quality authors buy certification to separate. Low quality authors do not, because for them the expected value of inviting scrutiny is below \(p\). Costly certification then re-manufactures the separating equilibrium the journal failed to enforce, so trust becomes a good with a price, sold outside the journal. This may substitute for gatekeeping or corrode it, leaving a two tier literature of certified and merely published work. Which one it does is a market structure question, and the market structure around AI is where Artificial Intelligence, Competition, and Welfare (NBER WP 34444) locates the welfare stakes. The scarce factor of the essay’s title, trust, is in this region literally a priced input.
What would move the result
The value of trust \(B\). If readers do not actually discount a flooded pool, and citation counts keep rising regardless of quality, then \(\bar\theta\) collapses and no audit is adopted at any cost. The audit ending needs a community that punishes pooling.
The penalty \(\varphi\). Enforcement, meaning retraction that bites and reputational memory, lowers \(a^{*}\) directly. Cheap audit and weak penalties are substitutes for sustaining honesty. The cheapest path to the audit ending may run through enforcement rather than intensity.
The asymmetry between \(k\) and \(g\). The whole model lives on \(k\) staying bounded while \(g \to 0\). If genuine verification also gets cheap, so that \(k \to 0\) with \(g\), then \(a^{*} \to 0\) and the problem dissolves. The essay’s wager is that the checked, human audited part of \(k\) is exactly what does not fall. That wager is the model’s load bearing assumption, and its most testable one.
Whether audit is verifiable by machine and mandated. \(\theta\) is realized institutionally rather than technologically. A reproduction agent that exists but is neither required nor rewarded leaves \(\theta\) high in practice.
Limitations, because this is a sketch. Two types, one period, one gatekeeper, no false positives, \(\pi(a)=a\), and no dynamic reputation or gatekeeper competition. Volume enters informally. A full treatment would microfound \(N(g)\) and let \(R\) be chosen. The parameters are illustrative and nothing here is estimated, so the comparative statics are meant to be read as directions rather than magnitudes. The economics is mine to get wrong and to fix. Corrections are the point of publishing a draft rather than a proof.