Worth reading · February 2026 · Temporal leads

Slower, and sure they were faster.

Experienced open-source developers took 19 percent longer with AI and believed afterward that it had made them 20 percent faster. Seven months later the same group published that its follow-up had become hard to read, and changed the design. Both halves are the finding.

The question

Is what I sell still worth selling? The Temporal question is whether the time I am supposedly saving accrues to me. Underneath it is the principle the whole inquiry runs on: self-report lies at the moment you most need it to hold. On this build the question is whether the agent made this task faster. On the long clock it is whether I can trust my own read of that at all.

Here is a group that measured the gap between the feeling and the clock, on developers who had spent years in the repositories they were working in, and then measured its own measurement and said where it broke.

What they found

The sourceMETR, Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity (10 July 2025) and We are Changing our Developer Productivity Experiment Design (24 February 2026). A task-level randomized trial, and the note that changed it. Read the original →

In early 2025: sixteen experienced contributors to large open-source repositories, working on 246 real issues from their own projects, each issue randomized to AI allowed or not. With AI, issues took 19 percent longer. Before the study the developers expected AI to speed them up by 24 percent. After it, having been slowed down, they still believed it had sped them up by 20 percent.

In February 2026 the same group published what happened next. A late-2025 run with 57 developers and more than 800 tasks estimated 18 percent slower for the original developers, with an interval from 38 percent slower to 9 percent faster, and 4 percent slower for newly recruited ones, from 15 percent slower to 9 percent faster. Then the authors said what the numbers could not: developers who did not want to work without AI had stopped taking part, the pay rate had dropped, and the tasks chosen were the ones people expected AI to help with. They believe developers are likely more sped up in early 2026 than a year before, call their own data only weak evidence for the size of it, and are moving to other designs. Self-reported speedups remain very high, and, in their words, those estimates can be quite unreliable.

The three questions

Every research post runs the same three rows. Which of the nine it names, how good the evidence is, and whether the thing is being measured.

RowStatusEvidence
NamedWhich of the nine questions does this answer? Temporal Output per hour on a defined metric, and hours per week. Epistemic second: the gap between what the developers believed and what the clock said is the self-read the model tells me never to trust.
EvidencedHow good is the evidence? open The 2025 result stands as a floor on how wrong an estimate can be. The follow-up the authors themselves declined to interpret. The question is open in their words, not mine.
MeasuredIs this being read on the Data page? pending Output per hour locks in October 2026. The weekly hours line on Data is the denominator without the numerator; it reads the floor, not the dividend.

What this cannot say:

  1. Whether AI speeds up experienced developers in 2026. The authors' own limit, stated plainly, which is why the row above reads open.
  2. Anything about the developers who would not work without AI. They left the study. The people most sure of the speedup are the least measured, and the design cannot fix that from inside.
  3. Whether time saved on a task became time. A faster task is not a shorter week. Nothing in either paper reaches the dividend.
  4. What faster is worth when the reviewer is the one who was sped up. The slowdown in 2025 was traced largely to reviewing and integrating what the tool produced. The clock and the judgement are the same person.

What I am doing about it

The status says what is missing. These are the moves, on my own work, that would supply it. Each is either read automatically off the footprint or asked of me on a cadence, and each costs something to keep doing.

  1. AutoNever publish a speed claim about my own work without a defined output metric beside it. Hours alone are the floor. A speedup with no numerator is a feeling with a chart.Costs: no headline number, ever.
  2. Asked · WeeklyEstimate the week's speedup before I look at the line, and keep both. The gap between the guess and the reading is the Epistemic signal I care about most, and this trial is the reason.Costs: a weekly record of being wrong in a known direction.
  3. AutoKeep the hours line on the page saying holding, not rising until the quarterly answer says where the hours went. A page that refuses to say I am faster until it can say faster at what.Costs: a page that refuses to say I am faster.
  4. AutoRead the design changes, not just the results. When a group I trust says its data got hard to interpret and shows its work, that is the finding. The standing commitment on this page is the same move made in advance.Costs: publishing my own broken readings with the same visibility as the clean ones.

What I keep

The tools, and the trial. The 2025 result was a floor on how wrong an estimate can be, and I would rather know the floor than the feeling. Nothing in either paper says stop. They say measure, and say when the measurement broke.

What it costs to keep watching is a weekly guess I am on record for, and a definition of output I cannot tune afterward.

Their findings, published July 2025 and February 2026; read by me against the nine dimensions, September 2026. Written with AI assistance; the reading is mine. — Clay

Want the nine readings taken on your own work?

The kit ships with the same instrument I am running on myself. Mentorship installs it with me in the room.

One list, no drip.

New Resources, Research, Data readings, and kit editions. This list is the only announcement I send.