When AI Generates Methodology Faster Than Humans Can Review It, How Does Certification Close the Gap?
- SFiT Newsroom
- 2 days ago
- 6 min read
[AI Summary]
Mathematics is experiencing what Terence Tao calls "proof indigestion": AI generates proofs far faster than people can review, understand, or absorb them.
The same structural problem has arrived in LCA and product carbon footprinting — and the consequences are more immediate. What piles up in mathematics is papers. What piles up in carbon management is carbon claims that have already gone public.
AI now generates not just results but the methodological choices behind them. What is expanding is not the volume of data but the number of judgments awaiting review.
The way to close the gap is not to add reviewers. It is to move auditability upstream — from a document produced afterwards to a property the data carries from the moment it is created. Let machines check what machines can check, so people only review judgment.

1. The real problem is not that AI is wrong — it is that humans cannot keep up with verification
Terence Tao recently described a situation in which AI produces mathematical proofs in volume, some of them already machine-verified in formal systems such as Lean, yet there are not enough qualified people to read them, understand them, and decide whether they deserve a place in the body of mathematical knowledge. The bottleneck has moved from producing answers to digesting them.
This deserves attention in our field, because it does not describe something peculiar to mathematics. It describes the common structure facing any knowledge-intensive profession confronted with AI: generation scales exponentially, review remains linear.
What differs is the consequence. In mathematics, the backlog sits quietly on a preprint server and the cost is lost efficiency. In carbon management the backlog does not sit quietly. It becomes a number in a client report, a benchmark in a procurement decision, an attachment to a compliance filing. An audit lag is an efficiency problem in mathematics. In certification it is a liability problem.
2. What is expanding is judgments, not data
Discussions of AI's impact on LCA usually stop at "it calculates faster." The reality now is different. AI does not only produce results; it makes methodological choices on the user's behalf — setting system boundaries, selecting allocation methods, filling in missing emission factors, inferring transport distances and grid mixes.
The implication is this. A model that once arrived at a reviewer's desk with perhaps five contested points now carries fifty implicit judgments, most of them unmarked. They sit inside defaults that look reasonable.
AI has not made review lighter. It has expanded the surface area of review by an order of magnitude. This is the pressure our methodology's audit boundary is now absorbing.
3. Three kinds of backlog, translated into LCA terms
Tao describes three forms of proof backlog. Each has a clear counterpart in carbon footprint modelling.
Unreviewed models
The model is built and the number exists, but no qualified person has worked through the assumptions and boundaries. It may already be quoted in a presentation without ever having been examined.
Models that pass formal checks but nobody understands
This is the dangerous category. ILCD validation passes, mass balance closes, units are consistent, the file imports cleanly — every machine-checkable item is green. Yet nobody can explain why a co-product was allocated by mass rather than by economic value. Correct format is not correct modelling, just as formal verification is not significance.
Models that are readable but cannot be reused
The model itself is clear, but it is a one-off deliverable. Assumptions live in email threads and spreadsheets, boundary conditions are buried in report prose, data sources carry no version tags. The next project has to rebuild it. The work was delivered; no asset was created.
4. Why hiring more reviewers does not solve it
The instinctive response to an audit lag is to expand the review team. Structurally, that path does not work, for three reasons.
Review is sequential; generation is parallel. A single model cannot be split across ten reviewers working simultaneously, because judgment requires command of the whole context. Generating ten models requires only ten times the compute. The two curves are fundamentally different.
People qualified to review methodology are already scarce. Those who can judge whether an allocation method is appropriate are the same people who can build the model. Moving them into review reduces modelling capacity; not moving them lets the backlog grow.
Review cannot be handed wholesale to AI. Using AI to review AI-generated models appears to solve the speed problem, but it breaks the chain of accountability. When something goes wrong, no one can explain the assumptions or carry the consequences.
The way out is therefore not to make review faster. It is to reduce what requires human review in the first place.
5. Four levers for strengthening the audit boundary
Lever 1 — Build auditability in at the moment of modelling, not afterwards
Most workflows today run in this order: build the model, produce the report, then assemble an assumptions list for the reviewer. That sequence guarantees a lag, because the explanation is generated retrospectively — and retrospective explanation tends to rationalise a result that has already been reached.
The alternative is to capture each judgment in structured form at the moment it is made: which database and which version this factor came from, what this allocation ratio rests on, what this boundary excludes and on what grounds, and who decided. Review then becomes inspection of an existing record rather than reconstruction of a process.
Lever 2 — Machines check format; people review judgment
Anything a machine can check should not consume expert time: unit consistency, mass and energy balance, expired factor versions, data years outside a plausible range, incomplete mandatory fields, ILCD conformity.
Once these are automated, human attention can concentrate where machines cannot help: whether the boundary is drawn correctly, whether the allocation method is defensible, whether proxy data genuinely represents this process, whether the assumptions are optimistic. That is where professional judgment actually earns its value.
Lever 3 — Risk-tier the review; not every model deserves equal depth
A common weakness in current practice is distributing review intensity evenly. But a model used for internal improvement tracking and a model supporting a public carbon claim that enters a customer's procurement decision carry entirely different risk.
Tiering by use, disclosure audience, financial scale and sensitivity — and concentrating the deepest review on the highest-risk models while others receive sampling or a simplified route — is the only way to raise effective audit coverage without increasing total review hours.
Lever 4 — Accountability and scope of validity travel with the data
Every model should carry two things: a named person able to defend its judgments, and an explicit statement of where the result may and may not be used.
When AI can produce a model at the press of a button, "who made it" stops being a meaningful question. The meaningful question becomes who can explain it and who is answerable for it. And if the scope of validity does not travel with the data, misapplication to an unsuitable context is only a matter of time.
6. Where SFiT sits: we strengthen the data layer, not the judgment
To be clear about scope: across these four levers, SFiT does not intervene in third-party professional judgment. Judgment belongs to the modeller and the verification body.
What we work on is the layer that makes judgment efficiently checkable: data structure, source traceability, version control, format conformity, automatic retention of the audit trail, and integration between systems. That is data-layer integration, not carbon consultancy.
Put differently: if a large share of the audit lag comes from evidence being scattered and slow to cross-check, then structuring the evidence is itself the most direct way to raise review efficiency. Time reviewers no longer spend locating data becomes time they spend exercising judgment.
7. Red flags in AI-generated LCA models
Results only, with no traceable inventory data or factor sources.
A claim that the model "passed system validation," without specifying whether what was validated was format, calculation, or methodological appropriateness.
A highly complete model with no indication anywhere of uncertainty or data gaps.
A modeller who cannot say which assumption, if removed, would reverse the conclusion.
Factor sources named only by database, without version or reference year.
Gap-filled data with no distinction marked between measured and estimated values.
No statement of validity scope, or one drawn so broadly it excludes nothing.
Review comments generated entirely by AI, with no named human sign-off.
8. Three sentences to take away
AI solves modelling capacity, not review capacity.
"Correctly formatted," "computationally sound," "understood by someone," and "safe to rely on" are four different things.
AI output that no human has explained and taken responsibility for is a candidate result, not a basis for a carbon claim.
Scope of this article
This is a methodological commentary on review workflow and data governance. It does not constitute a determination of conformity with any particular standard, and it does not replace the verification procedures required under ISO 14040/14044/14067. References to the situation in mathematics are drawn from publicly available remarks and are used here as a structural analogy, not as a claim about that field.
Raymond Wang · SFiT Corp 2026
Comments