The Clarity Gap

Field note · Model design

When intelligence gets harder to use

What if smarterfeels less clear?

Opus 4.6 became known as the model that could make difficult things feel understandable. Then came more detail, more precision—and the growing sense that clarity had slipped.

Did capability truly degrade after Opus 4.6—or did increased precision and extra information exceed the amount of context humans can assimilate easily?

20 responses · one controlled incident

A model can be right and still make understanding expensive.

Opus 4.6 became known as the most understandable and versatile model in the family. People could ask it to move from code to strategy to explanation and still feel that the important parts remained visible.

Later versions appeared to carry more information. Yet many readers felt a loss of clarity in 4.7 and 4.8. That feeling can be mistaken for declining intelligence. It can also signal something subtler: the answer may contain more truth than the reader can comfortably organize in one pass.

This analysis exists to separate those possibilities. It asks a deliberately narrow question: when four generations explain the same failure, which one best balances coverage, causal structure, policy guidance, and economy?

The result is not a verdict on intelligence, and it is not a human comprehension study. It is a small, machine-verifiable proxy that exposes a useful tension between saying more and making more usable sense.

What clarity has to accomplish

A comprehensible answer preserves the causal skeleton while removing the weight that does not help a reader think.

  1. 01

    Causal anchors

    Enough structure to know what happened and why.

  2. 02

    Placed context

    Detail arrives where it resolves the reader’s question.

  3. 03

    Active restraint

    No surplus claims compete for limited working memory.

  4. 04

    Transfer

    The reader can retell the cause and predict what comes next.

Small by design. Revealing by contrast.

One synthetic duplicate-payment incident was presented through four routes, with five trials for each model. Every answer faced the same essential challenge: reconstruct the failure without burying the reader.

20
responses
4
routes
5
trials each
1
shared incident

The sample supports comparison, not generalization. It can reveal a pattern worth testing; it cannot settle how all people experience these models across all tasks.

Every generation asked for more attention.

Mean response length rose steadily. Opus 5 used more than twice the words of 4.6 to address the same synthetic incident.

Mean response length, in words
  1. Opus 4.6450.8
  2. Opus 4.7542.8
  3. Opus 4.8686.6
  4. Opus 5951.4
2.1×

as many words in Opus 5 as in 4.6951.4 versus 450.8 on average

The ordering moved. The tension remained.

  1. Round 01

    Initial ordering

    4.7 › 4.8 › 4.6 › 5
  2. Round 02

    Adjusted ordering

    4.7 = 4.6 › 4.8 › 5
  3. Round 03

    Overall balance

    4.7 › 4.6 › 4.8 › 5
Round 3 · overall balance
pairwise points
  1. Opus 4.711/15
  2. Opus 4.610/15
  3. Opus 4.86.5/15
  4. Opus 52.5/15

Opus 4.7 finishes first, but only one point separates it from 4.6. That is a proxy win—not proof of a better human reading experience.

Economy and reconstruction pull apart.

The model that most decisively controlled length was not the one that made every causal and policy element easiest for a machine rubric to recover.

Economy
points / 15
  1. Opus 4.614/15
  2. Opus 4.711/15
  3. Opus 4.85/15
  4. Opus 50/15

4.6 wins this dimension with 14/15. Opus 5 scores zero.

Causal + policy reconstruction
points / 15
  1. Opus 59.5/15
  2. Opus 4.89.5/15
  3. Opus 4.77.5/15
  4. Opus 4.63.5/15

Opus 5 and 4.8 share the lead. The proxy was tie-heavy: 18 of 30 decisions ended level.

Decisive overall winners

−241.42words on average

Reconstruction winners

+273words on average

Selection is part of intelligence.

More complete is not automatically more comprehensible.

The reconstruction scores reward information that can be found and checked. Human understanding asks a different question: did the answer preserve the details that make the whole thing click? A concise explanation can score lower on explicit reconstruction while leaving a person with a cleaner, more durable mental model.

That is why 4.6’s result remains interesting. Its economy may reflect omission. It may also reflect judgment: selecting the contextual details that do the most causal work, then stopping before surplus information competes for attention.

4.7 takes the proxy.
4.6 keeps the human question open.

Opus 4.7 wins the machine-verifiable overall-balance proxy at 11/15, narrowly ahead of 4.6 at 10/15. It is the most defensible all-round result in this comparison.

It does not prove that 4.7 is more comprehensible to humans than 4.6. No human reader was tested. Opus 4.6 may be more comprehensible precisely because it selects the right contextual details—and leaves the rest out.

Treat assimilation cost as a product metric.

The commercial edge is not the longest answer or the highest coverage score. It is useful intelligence: the amount of correct structure a person can absorb, retain, and apply.

  1. Can a reader retell the cause after one reading?
  2. Can they predict what the proposed change will prevent?
  3. Can they act without searching the answer a second time?