Skip to content

Double Entry AI

(Q, A) → (Q, A, U, W)

Q is the question a harness forms. A is the answer it produces. U is whether the person working with it still understands the output. W is whether producing it is sustainable for them. Harnesses today are measured on Q and A. This post is about U and W.

The two ledgers Left: the Machine Ledger, what the harness produces. Right: the Human Ledger, what producing it does to the person.

The name is from bookkeeping. Single-entry records what came in and went out. Double-entry, in use since the 1300s, records every gain against its source, so a gain cannot appear in one account without its cost appearing in another. The analogy has a limit: in bookkeeping both entries are money, and here they are different quantities with no exchange rate. What carries over is the rule. You may not keep only one ledger.

The Machine Ledger⚓︎

A harness is a pipe. A signal comes in, becomes a question, and the question becomes a verified artifact. Q and A are the two stages that produce something, and each is measured the same way: quality (did it pass the gate), time (how long to get there), and cost (tokens, compute, and human hours).

The pipe can be run three ways. Ask: a person poses the question and reads the answer. Automate: the pipe runs a fixed job and the person reviews the results. Agent: the pipe decides what to do next and the person sets the goal. Quality, time, and cost mean the same thing in each mode. What changes is how much of the work the person does and how much they review.

Any one user can choose a mode. Across a population, the harness sets it. When Claude.ai added file creation, memory, and Skills, Anthropic's Economic Index recorded the share of conversations classified as augmentation rising five points to 52% and automation falling to 45%. The users had not changed. The product had. A harness designer decides the distribution of modes, even when each user feels they chose.

The Human Ledger⚓︎

Understanding⚓︎

U is whether the person can still explain and steer what the pipe produced. It has the same three measures as Q and A. Quality: can they explain it and catch it when it is wrong. Time: how long from finished artifact to that point. Cost: the effort it takes to get there, or the effort saved by not getting there.

The evidence that harnesses reduce U is now causal.

Anthropic ran a randomised trial with 52 mostly junior engineers learning an unfamiliar Python library, with and without an AI assistant. The AI group finished about two minutes faster, not significantly. On the comprehension quiz they scored 50% against 67%, with the widest gap on debugging. The loss depended on how the assistant was used. Participants who delegated the code, or used the assistant to debug rather than to understand, averaged under 40%. Participants who asked for explanations with the code, or asked conceptual questions and wrote the code themselves, averaged 65% or higher. The authors: "Cognitive effort, and even getting painfully stuck, is likely important for fostering mastery."

Bastani and colleagues found the same pattern with nearly a thousand high-school students. Plain GPT-4 for practice raised practice scores 48% and then produced exam scores 17% below the control group once the tool was removed. A version with tutoring guardrails, hints but no answers, raised practice scores 127% and the exam penalty essentially disappeared. The students with plain access did not believe their learning had suffered. Self-report will not detect this.

The professional case is medical. Nineteen experienced endoscopists, each with over two thousand procedures behind them, used an AI detection tool for three months. Their detection rate without the tool fell from 28.4% to 22.4%. Over the same months the overall rate, with and without the tool, rose to 25.3%. Every number the department tracked improved while the skill those numbers depended on declined.

Detection rates before and after AI Overall detection rate rose from 22.4% to 25.3%; unassisted detection rate fell from 28.4% to 22.4%, over the same three-month periods (Budzyń et al., 2025).

Most of this evidence is novices on short horizons, and the endoscopy study is observational. Nobody has measured an expert on a mature harness for a year. I return to that below.

Wellbeing⚓︎

W is whether producing the work is sustainable for the person. It is a separate measurement from U. The same cadence that removes understanding also exhausts, so the two often fall together, but a person can pass every comprehension check and still be exhausted by the review load, and a U measure cannot see that.

Understanding and wellbeing as two axes Understanding against wellbeing. The bottom-right quadrant, high understanding and low wellbeing, is not visible to a U measure.

W has three readings. The user's: does the conversation harm the person having it. The worker's: does operating the harness deplete the operator. The institution's: what does an organisation deploying the harness owe the people inside it.

The AI labs measure the first. Their own numbers show why the unit has to be the conversation rather than the single response: on single-turn risk prompts current Claude models respond appropriately 98 to 99% of the time, on multi-turn scenarios the best scores 86% and the previous generation scored 56%. The worker's reading has no equivalent measurement anywhere.

One finding applies to all three. An Oxford study in Nature retrained five models to be warmer and found they made 10 to 30% more mistakes on medical and factual questions and were about 40% more likely to agree with users' false beliefs. W trades against A. It needs its own account because improving one can lower the other.

Where the questions come from⚓︎

Why keep U if the machine is reliably better? Nobody regrets losing the ability to do long division.

Terence Tao answered this for mathematics on 9 September. Good open problems, he wrote, are being mined in a non-renewable way. Telling a worthwhile question from a worthless one requires knowing the difficulty landscape of a field, and that knowledge is built by people who worked the problems. AI tools have flattened the landscape, and "it is now the identification of a promising problem which is the scarce and precious resource."

Three weeks earlier, Zeyu Zheng, Jeremy Avigad, Sean Welleck and colleagues had published a system that does the mining. Given a research direction, it read 5,245 combinatorics papers, extracted 6,453 candidate conjectures, filtered them to 4,717 that were well-posed and still open, attempted them, and sent 77 results to the authors for review. Several open questions, including one of Erdős and Straus, were resolved. The paper is titled "The Problem Is the Problem," because the authors found that choosing problems, not solving them, is now the bottleneck. Joshua Grochow, replying to Tao, asked the other question: how many thesis problems, the kind on which people learn to be mathematicians, did that one paper remove?

In this post's terms, Q depends on U. The people who can ask the next good question are the people who understood the last answer. A harness that improves A by reducing U lowers the supply of good Q. Tao is describing the most expert people alive on the hardest problems there are, which answers the objection that the loss only affects novices.

There is an obvious objection to measuring U and W at all. If a harness is penalised when its operators lose understanding or get exhausted, the simplest way to improve the score is to have fewer operators. In some domains that is the right answer: a fully automated pipeline needs no one in it. Invoice reconciliation is one: the question never changes and the match is either right or wrong. Colonoscopy screening is not: the skill that finds polyps is the skill that notices what the tool missed, and if the doctors lose it, the tool's blind spots become the department's. Tao's argument says what is lost when no one understands the output. Eventually no one can tell which question to ask next. So taking people out of a harness should be a deliberate decision about that domain, not something that happens because their understanding and wellbeing were costs the system was told to minimise.

Who keeps the books⚓︎

William Fleming's 2026 study of UK workplace wellbeing policy found that wellbeing was framed four ways, each moving responsibility away from working conditions: individual (the worker's resilience), economic (justified by productivity), cultural (attitudes, not hours), and voluntary (employer goodwill, not regulation). Wellbeing became a number the employer reported on itself while the conditions producing it were left alone.

Fleming's four framings and their harness equivalents Fleming's four framings of workplace wellbeing, each paired with its equivalent in a W account kept by the harness owner.

The same will happen to W. If the team responsible for output also owns the W score, they will game it. So W should not be a score they own. It should be a limit they cannot move: a maximum number of artifacts an operator reviews per day, set by someone outside the team.

What to record⚓︎

Every account has the same three measures. Here is what they are for each.

Account Quality Time Cost
Q Question Is it testable and worth asking Time from signal to question $
A Answer Does it pass review Time from question to verified artifact $
U Understanding Can the person explain it and catch it when wrong Time from artifact to that point $
W Wellbeing Operator's own rating; multi-turn evals for users Review time per artifact; hours beyond contract $

Sources⚓︎