The Code/X ArchiveView on X
Kun Chen

@kunchenguid

used grok 4.6 for a full day of real work, here's my unbiased review

i used it as my @myfirstmate - the orchestrator agent that i directly talk to that manages all my other agents - and also as a worker agent directly on some implementation tasks

first, the good:

yes it's fast and it's great. everyone knows this so i'm not going to repeat it. i however noticed something else more subtle but extremely important - judgment

an orchestrator often has to make judgment calls, like "is this a low-stake change that should be done quickly, or should we really safeguard it", "does this decision need to escalate to the user", "should this task keep going or let's step back". this really tests whether a model was trained with the right rewards

that's where many other models fall apart. sol will remodel my kitchen before making me breakfast. opus 5 will go down rabbit holes and come back with an essay but no rabbit

throughout the whole day, grok 4.6 almost made every decision right for me. i deliberately gave it a lot of autonomy - it merged things when i would have. it asked me for approval for things i actually cared about. it stopped workers from over-engineering. it often recommended ship first fix later when it's safe to do so

chef's kiss

now, the bad:

i have the highest tier supergrok heavy subscription, and it's just not giving enough quota. a full day of mostly single-threaded usage exhausted a whole week of quota. the same work will translate to about 20-30% of the $200 subscription from anthropic or openai using opus or sol

without pricing being on par, i can totally stand behind the model but i can't really recommend the subscription purchase for consumers

overall, grok is on a great trajectory of delivering the most useful model in the world that's fast, pleasant, and capable. i just hope the team realizes how much this pricing problem is holding back their model + grok build harness' adoption, and figure out how to get it fixed
163952.8K337