The Clarity Gap
Field note · Model design
When intelligence gets harder to use
What if smarterfeels less clear?
Opus 4.6 became known as the model that could make difficult things feel understandable. Then came more detail, more precision—and the growing sense that clarity had slipped.
The question
Did capability truly degrade after Opus 4.6—or did increased precision and extra information exceed the amount of context humans can assimilate easily?
20 responses · one controlled incidentThe uneasy upgrade
A model can be right and still make understanding expensive.
Opus 4.6 became known as the most understandable and versatile model in the family. People could ask it to move from code to strategy to explanation and still feel that the important parts remained visible.
Later versions appeared to carry more information. Yet many readers felt a loss of clarity in 4.7 and 4.8. That feeling can be mistaken for declining intelligence. It can also signal something subtler: the answer may contain more truth than the reader can comfortably organize in one pass.
This analysis exists to separate those possibilities. It asks a deliberately narrow question: when four generations explain the same failure, which one best balances coverage, causal structure, policy guidance, and economy?
The result is not a verdict on intelligence, and it is not a human comprehension study. It is a small, machine-verifiable proxy that exposes a useful tension between saying more and making more usable sense.
A human measure
What clarity has to accomplish
A comprehensible answer preserves the causal skeleton while removing the weight that does not help a reader think.
- 01
Causal anchors
Enough structure to know what happened and why.
- 02
Placed context
Detail arrives where it resolves the reader’s question.
- 03
Active restraint
No surplus claims compete for limited working memory.
- 04
Transfer
The reader can retell the cause and predict what comes next.
The comparison
Small by design. Revealing by contrast.
One synthetic duplicate-payment incident was presented through four routes, with five trials for each model. Every answer faced the same essential challenge: reconstruct the failure without burying the reader.
- 20
- responses
- 4
- routes
- 5
- trials each
- 1
- shared incident
The sample supports comparison, not generalization. It can reveal a pattern worth testing; it cannot settle how all people experience these models across all tasks.
The expansion
Every generation asked for more attention.
Mean response length rose steadily. Opus 5 used more than twice the words of 4.6 to address the same synthetic incident.
- Opus 4.6450.8
- Opus 4.7542.8
- Opus 4.8686.6
- Opus 5951.4
as many words in Opus 5 as in 4.6951.4 versus 450.8 on average
Three rounds
The ordering moved. The tension remained.
- Round 014.7 › 4.8 › 4.6 › 5
Initial ordering
- Round 024.7 = 4.6 › 4.8 › 5
Adjusted ordering
- Round 034.7 › 4.6 › 4.8 › 5
Overall balance
- Opus 4.711/15
- Opus 4.610/15
- Opus 4.86.5/15
- Opus 52.5/15
Opus 4.7 finishes first, but only one point separates it from 4.6. That is a proxy win—not proof of a better human reading experience.
Two kinds of winning
Economy and reconstruction pull apart.
The model that most decisively controlled length was not the one that made every causal and policy element easiest for a machine rubric to recover.
- Opus 4.614/15
- Opus 4.711/15
- Opus 4.85/15
- Opus 50/15
4.6 wins this dimension with 14/15. Opus 5 scores zero.
- Opus 59.5/15
- Opus 4.89.5/15
- Opus 4.77.5/15
- Opus 4.63.5/15
Opus 5 and 4.8 share the lead. The proxy was tie-heavy: 18 of 30 decisions ended level.
Decisive overall winners
−241.42words on averageReconstruction winners
+273words on averageThe human-centered reading
Selection is part of intelligence.
More complete is not automatically more comprehensible.
The reconstruction scores reward information that can be found and checked. Human understanding asks a different question: did the answer preserve the details that make the whole thing click? A concise explanation can score lower on explicit reconstruction while leaving a person with a cleaner, more durable mental model.
That is why 4.6’s result remains interesting. Its economy may reflect omission. It may also reflect judgment: selecting the contextual details that do the most causal work, then stopping before surplus information competes for attention.
The verdict, with limits
4.7 takes the proxy.
4.6 keeps the human question open.
Opus 4.7 wins the machine-verifiable overall-balance proxy at 11/15, narrowly ahead of 4.6 at 10/15. It is the most defensible all-round result in this comparison.
It does not prove that 4.7 is more comprehensible to humans than 4.6. No human reader was tested. Opus 4.6 may be more comprehensible precisely because it selects the right contextual details—and leaves the rest out.
For teams choosing a model
Treat assimilation cost as a product metric.
The commercial edge is not the longest answer or the highest coverage score. It is useful intelligence: the amount of correct structure a person can absorb, retain, and apply.
- Can a reader retell the cause after one reading?
- Can they predict what the proposed change will prevent?
- Can they act without searching the answer a second time?