this is a silly argument. GPT-5.6 Sol was 88.8% on TerminalBench 2.1 and 37.3% on TerminalBench 4.0. That doesn't make it a more benchmaxxed model than GPT-6 Astra even though the diff is larger. TerminalBench 4.0 is simply a much harder eval that isn't saturated, whereas TerminalBench 2.1 is.
We don't claim Muse Spark 1.3 is as strong as Astra or Fable 5.1, but it is significantly more cost-effective. Our future models will compete more directly with those models.
Alexandr Wang
@alexandr_wang
Replying to @SemiAnalysis_
12:27 AM UTC · Sep 8, 2026 · 73.1K Views
54361.3K99