The Code/X ArchiveView on X
Lon Lundgren

@Lon

After Anthropic made Fable 5 permanently available in subscription plans, I noticed a large drop in performance. The model felt dumber, and I couldn't explain why.

Measured five different ways, August delivered dramatically fewer thinking tokens than July.
Image from the post
821291.9K486
Lon Lundgren

@Lon

This wasn't a one-time drop. Reasoning fell over the entire period, and fluctuated across multi-day episodes.

Some of these fluctuations aligned with specific product announcements and releases.

I started to see how the model could feel great one day, and terrible the next.
Image from the post
2519222
Lon Lundgren

@Lon

I was consistently using an xhigh or max effot level, but when I looked deeper, I found that most invocations to the model were receiving little to no thinking tokens at all.

And when longer thinking runs did happen, they almost never reached published benchmark levels.
Image from the post
2513811
Lon Lundgren

@Lon

This unfolded across a six week data capture and analysis odyssey, and led to a number of surprising findings.

Next time the model feels dumber, don't ask if the model was "nerfed".

Ask about the inference regime you were served, instead.
x.com/Lon/status/210…
4516954
End of thread