Seven Releases With Grok: Fast, Honest, and What the Diff Did Not Show
In my last post I cancelled Ollama. The money running out early was enough on its own, but the part I wrote about was how many review rounds the work needed before I could merge it. That gives me a simple test for any model: how many times does the work go around review, and does each fix pass make things better or add new problems? This post runs the same test on Grok. Starting on September 21, it implemented four releases of my Zone 2 Trainer app and two releases of my personal assistant app. On a seventh release I swapped the roles, and Grok reviewed code that Claude wrote. Grok did well on that test. That is the short version. The longer version is about where the bugs that did get through were hiding: not in what the changed lines did, but in what they assumed. The Setup The models are grok-4.6 and grok-4.7 from xAI, on a SuperGrok subscription. grok-4.7 came out right after the first release, so it did the rest. On the other side was Claude. It wrote the issue lists and...