The Code/X ArchiveView on X
Alexandr Wang

@alexandr_wang

Replying to @SemiAnalysis_

this is a silly argument. GPT-5.6 Sol was 88.8% on TerminalBench 2.1 and 37.3% on TerminalBench 4.0. That doesn't make it a more benchmaxxed model than GPT-6 Astra even though the diff is larger. TerminalBench 4.0 is simply a much harder eval that isn't saturated, whereas TerminalBench 2.1 is.

We don't claim Muse Spark 1.3 is as strong as Astra or Fable 5.1, but it is significantly more cost-effective. Our future models will compete more directly with those models.
54361.3K99