Posts

Showing posts with the label code-review

Seven Releases With Grok: Fast, Honest, and What the Diff Did Not Show

Image
In my last post I cancelled Ollama. The money running out early was enough on its own, but the part I wrote about was how many review rounds the work needed before I could merge it. That gives me a simple test for any model: how many times does the work go around review, and does each fix pass make things better or add new problems? This post runs the same test on Grok. Starting on September 21, it implemented four releases of my Zone 2 Trainer app and two releases of my personal assistant app. On a seventh release I swapped the roles, and Grok reviewed code that Claude wrote. Grok did well on that test. That is the short version. The longer version is about where the bugs that did get through were hiding: not in what the changed lines did, but in what they assumed. The Setup The models are grok-4.6 and grok-4.7 from xAI, on a SuperGrok subscription. grok-4.7 came out right after the first release, so it did the rest. On the other side was Claude. It wrote the issue lists and...

A Convincing Review of the Wrong Path

Image
This article follows Executor Model, Reviewer Model: Why Cross-Vendor Code Review Actually Works , where I wrote about using a second model to validate review findings. This time, the missing check was not another opinion. It was running the code. I am the only user of my personal assistant app, so every issue in its tracker is one I filed after watching something fail on my phone. One was the fasting timer. I told it in chat that I had started a fast at seven in the evening, and nothing started. I handed that issue and eight others to Claude. It read the source, quoted the right files, and explained why the timer could not fail the way I described. I closed all nine issues that afternoon. A few days later I asked Claude to check again, but this time it had to run the reported inputs. Four verdicts flipped. The second pass took six and a half minutes instead of two and a half. Those extra four minutes recovered four real bugs and found two more I had never noticed to file. T...