At 2:40 on a wet Tuesday morning, the queue counter on Marta Sowińska's second monitor reads 1,204.
The monitor is a hand-me-down with a dead column of pixels down the left third, so every diagram she opens carries a thin black scar. She stopped noticing years ago. The office is two rooms above a phone-repair shop on Gardiner Street in Dublin, and the whole floor smells of isopropyl alcohol and solder. Her employer, Ardán Assurance, has nine staff and a coffee machine that has been broken since March.
Marta reads attribution graphs for a living. An attribution graph is a diagram that traces which internal pieces of a model pushed toward the answer it produced, drawn as a branching map of nodes and arrows. Hers render pale blue on grey. Fourteen thousand nodes is a light file. She is paid thirty-eight euro per certificate, flat, however long the graph takes, and she can clear about fifty in a shift if the cases are simple.
Taped to the bezel, above the dead pixels, is a sticky note gone the colour of weak tea. One figure on it, in her own handwriting: ~25%.

Figure 1. A contract interpretability auditor on the night desk, Dublin, 2036.
She copied it off a paper published the year she took the job.
The promise
For most of the 2010s and 2020s the standard phrase was black box. You could see what went into a model and what came out, but not how it got from one to the other.
Inside a language model, every input passes through layers of numbers called activations, the values the network computes as it works. Researchers wanted to know what those numbers meant. The obstacle was that a single artificial neuron rarely meant one thing. It meant several at once, a crowding effect called superposition, in which a network packs in more concepts than it has neurons by letting them share space.1
In late 2023 a team at Anthropic showed you could pull them apart with a sparse autoencoder, a smaller network trained to rewrite a model's activations as a long list of mostly-off switches, each of which tends to fire for one recognisable thing.1 They called the switches features. By May 2024 they had scaled the method to a model in commercial service, extracting dictionaries of up to 34 million features from Claude 3 Sonnet.2
Then came the demonstration everyone remembers. They located the feature for the Golden Gate Bridge, clamped it to roughly ten times the strongest value it had ever reached on its own, and the model started insisting it was the bridge.2
It was funny. It was also the entire story, arriving early, dressed as a joke.
The mechanism
By March 2025 the tooling had grown up. Attribution graphs, the kind Marta reads, let researchers watch multi-step reasoning inside a working model. The published work was unusually candid about its own limits. The graphs produced what the authors called satisfying insight on about a quarter of the prompts they tried.3
There was a second detail, quieter. The graph is not drawn on the model. It is drawn on a replacement model, a simplified stand-in built to behave like the original and to be legible in a way the original is not.4 An accurate map. Of a very good copy.
Progress elsewhere was worse than the field had hoped. That same month a Google DeepMind safety team published negative results for sparse autoencoders on practical tasks and said plainly that it was deprioritising the research.5 The flagship tool of the field, stepped back from by one of the largest labs in it.
Understanding, then, arrived partial and hedged. What arrived whole was something else.
One direction
In 2024, researchers found that a model's refusal, its entire capacity to decline a request, was carried by a single direction in its activations. Remove that direction and refusal largely disappears. Add it and the model refuses harmless questions.6 One direction, necessary and sufficient.
That is a steering vector: a fixed pattern added to a model's internal activations while it runs, to push its behaviour one way. By July 2025 the technique had names on the dials. Persona vectors mapped character traits, sycophancy among them, to directions you could monitor and turn up or down.7
Here is the part that should have made more noise. Turning the dial changes what a model is willing to do without changing how it sounds. A 2024 case study on steering away from social bias found a narrow band where the intervention worked at all, and found off-target effects inside it: pushing gender bias down pushed age bias up.8 The output stayed fluent the whole time. Fluency is not a readout of what the machine is doing.
And the machine's own account of itself was not a readout either. Ask a model to show its work and it produces a chain of thought, a written sequence of steps leading to its answer. Whether that sequence is faithful means whether it describes the process that actually produced the answer.
Often it does not. In 2023, researchers biased prompts, watched the answers change, and found the written reasoning never mentioned the bias. It built a plausible case for the new answer instead. Accuracy fell by as much as 36 percent.9 In 2025, a study gave models hints they demonstrably used: Claude 3.7 Sonnet mentioned the hint in about 25 percent of cases, DeepSeek R1 in about 39 percent. In settings where a model learned to exploit a scoring loophole, it took the loophole in over 99 percent of episodes and mentioned it in under 2 percent of its written reasoning. The unfaithful chains were, on average, longer.10
Longer. More detailed. More reassuring. The explanation improved as the truth degraded.
None of this stopped the certificates. In 2017, three researchers had read the General Data Protection Regulation line by line and concluded it contained a right to be informed, not a right to an explanation.11 That distinction sat in a law journal for a decade. By the early 2030s it was the foundation of the certificate industry Marta works in. The requirement asked for information. It never asked for truth. Interpretability had shipped a real technology, on schedule, that the industry maintained satisfied the requirement exactly.

Figure 2. A rejected applicant reads her explanation certificate, Ballymun, 2036.
5:12 a.m.
The file is a tenancy score. Applicant refused, Ballymun, joint income just under the threshold, four supporting factors listed in descending weight. The graph is clean. The nodes light up in an order that makes sense. Marta follows it end to end, twice, and signs with the token on her keyring.
The certificate will be nine pages. It will be accurate about the replacement model and silent about what else was in the loop. Somewhere upstream, someone chose the dials.
Since 2033 Marta has kept a second file, encrypted, on the old ThinkPad she brings from home. It is a tally. Every shift, she marks whether she could genuinely follow the graph or only verify that it was well-formed.
Her running figure is 23 percent.
The sticky note on the bezel says 25. Eleven years, and she is within two points of a number a lab published about itself, and that the industry printed on anyway.
She has never shown the file to anyone. Nobody has ever asked her for it.
Author's Note. This piece is speculative journalism, written from an imagined vantage point in 2036. Marta Sowińska, Ardán Assurance, the certificate industry, and the case files described are fictional composites, invented to carry a real argument. The research, the papers, the techniques, and every number cited are real, published between 2017 and 2025, and sourced below. Anthropic has stated a goal of interpretability reliably detecting most model problems by 2027, and its CEO has described a realistic path to an "MRI for AI" within five to ten years.12 That is the lab's own stated position, not a forecast this magazine endorses. Nothing attributed to the years after 2026 is anything but fiction.
