Seven Releases With Grok: Fast, Honest, and What the Diff Did Not Show

Abstract dark blue illustration of a bright grid of changed squares with a cyan scanning beam, while the real fault glows orange off to the side, outside the beam

In my last post I cancelled Ollama. The money running out early was enough on its own, but the part I wrote about was how many review rounds the work needed before I could merge it. That gives me a simple test for any model: how many times does the work go around review, and does each fix pass make things better or add new problems?

This post runs the same test on Grok. Starting on September 21, it implemented four releases of my Zone 2 Trainer app and two releases of my personal assistant app. On a seventh release I swapped the roles, and Grok reviewed code that Claude wrote. Grok did well on that test. That is the short version. The longer version is about where the bugs that did get through were hiding: not in what the changed lines did, but in what they assumed.

The Setup

The models are grok-4.6 and grok-4.7 from xAI, on a SuperGrok subscription. grok-4.7 came out right after the first release, so it did the rest. On the other side was Claude. It wrote the issue lists and scope documents, reviewed every pull request Grok wrote, and implemented the seventh release as Claude Opus 5.5. This is the setup from Executor Model, Reviewer Model: one company's model writes the code and another company's model reviews it.

One thing to keep in mind for the whole post: the model that wrote most of the specs is the same model that reviewed the code. You will see why that matters.

What Held Up Every Time

Three things were true in all six releases Grok implemented.

  • It was fast. The first release went from a fresh issue list to a pull request with tests in 23 minutes. Answering a review usually took 4 to 13 minutes. One release of my assistant app took 65 minutes from start to a tagged release, with three review rounds inside it.
  • It never claimed more than it did. Not once in six releases did a commit or pull request say something was fixed when it was not. That was the deepseek-v4-pro:0813 habit I described in the Ollama post, so I was watching for it.
  • It said when work was not finished. On one sync bug it fixed part of the problem, wrote the two remaining gaps onto the issue, and ended the comment with "Do not close this on the PR landing." A model trying to look finished does not write that.

So the question was never "does it work." The question was what got past it.

Where the Bugs Were Hiding

Every release had at least one real bug that survived the first pass. When I lined them up, none of them looked wrong when you read only the changed lines. Each one was an assumption the new code relied on and never checked.

A wrong spec, followed exactly. The first release added a guard so a double tap on Start would not start two workouts. The guard never fired. It checked "is a session running and is training active," but the screen turns training off a moment before it sends Start, so the second half was always false. Here is the part that stays with me: the issue asked for that same check, a session and training active. Claude wrote that issue. Grok built what it said. Both tests for the guard passed, because they checked the rule on its own and never sent a real double tap through the app.

When does this run, compared to that? A later release replaced a crude half-second wait at the end of a workout with a proper "wait until the last heart rate samples are saved" signal. It looked like a clear upgrade. But the signal was set up after the code that waited on it already ran. On the first workout there was nothing to wait for. On every workout after that, it waited on the previous workout's signal, which was long finished. Either way it waited for nothing, with no error and no log line, and the half-second wait that used to do the job was deleted in the same change. The order was even written down in a comment in the same file. As first written, the fix was worse than the one-line sleep it replaced. Review caught it before the merge.

A borrowed contract. In my assistant app, Grok added OpenRouter as a provider by reusing the OpenAI client, because my scope told it to. The client was reusable, but OpenRouter's billing is not OpenAI's. OpenRouter checks your balance against the largest possible answer before it starts, so a small balance got a "payment required" error even though the real answer would have been cheap. Nothing in the source code shows that. One live request did: "You requested up to 16000 tokens, but can only afford 7211."

A rule the framework hides. To stop a background sync from bringing back a workout the user had deleted, Grok changed a database write. The SQL was correct. But the database library treats that kind of query as a read, so the workout history screen would never have refreshed after a background update. Review caught this one too. You only see this in the code the library generates, not in anything a person writes.

Different domains, same question: what is this change assuming that nobody checked? For concurrency, it is when something runs. For integrations, it is whose behavior you inherited.

The Spec and the Reviewer Were Part of It

I would like to say the model wrote the bugs and the reviewer caught them. The record does not say that.

Two of the six releases trace their worst first-pass bug to a scope Claude wrote. The double-tap guard was one. The other was in my assistant app: the scope said a cut-off reply "produces the marker," so Grok put a hidden marker inside the reply text. Every part of the app that did not know about the marker would have shown it or saved it, including into notes. The same scope also said "the client reports a truncated result," which pointed the right way. Grok followed the more concrete sentence. The lesson I took is to write the contract into the scope, not the mechanism. A capable model will build the mechanism you describe, flaws included.

The reviewer's fixes were not always right either. In one review, Claude found a real bug and attached a fix that would have made it worse. Grok took the finding, rejected the fix, explained why in one sentence, and solved it a different way. In another review, Claude's fix was right but incomplete. It said to include the conversation when saving a note, and did not say "through the privacy check." Grok did exactly what it said. One "save this as a note" would have sent private chat history to a cloud model for a user who had chosen local only. The next review round caught it.

Three of Grok's six fix passes added a new bug while fixing the old ones. In each case the fix was right where it aimed, and nobody asked what else it touched. One of those three came straight from the reviewer's remedy.

It Pushes Back, When It Has a Reason

Before the second release I wrote down one question: would Grok push back on a wrong instruction, or just do it? It did both, and mostly at the right times.

  • It rejected the reviewer's wrong fix described above.
  • It declined a finding the review had marked as unproven. Lowering a limit shared by three providers on a guess was a real change, and it said it would not do it without a live request. When the live request came back, it made the narrowest correct fix, for OpenRouter only.
  • It corrected a version number in the scope, and said why in the pull request.
  • When the review said a test was missing, it pointed out the test existed under another name, and renamed it to the review's wording instead of arguing.

Where it did not push back was the double-tap spec and the privacy remedy. In both cases the instruction looked reasonable and was wrong in a way you only see by reading code outside the change.

Roles Reversed

For the seventh release, Claude Opus 5.5 implemented and Grok reviewed. This was a good test of whether the reviewer role was doing the real work, or just the model in it.

Grok found two blockers in Claude's code, after over 3,900 tests had passed. The first needed code outside the change. The second was right there in the new code, and still easy to miss without thinking like an attacker:

  • The morning briefing would have opened links nobody pasted. The new link reader scanned the whole message sent to the model. The briefing puts calendar and inbox text into that message, so every morning it would have fetched meeting links and newsletter tracking links. A fetched tracking link tells the sender you opened their email.
  • The address check could be tricked. The link reader checked that a website did not point into my home network, then looked the address up again to connect. Claude had written a comment saying this only mattered for "a hostile DNS server." Grok's answer: "The threat is a link the user pasted. The attacker owns that domain's DNS."

Claude agreed with both and pushed back on one severity claim. It also kept one decision against the review, and found one case Grok's list missed. After the phone test found a separate bug, Grok took a last short look and called one item "not a hold": a document tile that overflowed at bigger text sizes. New tests at larger font scales all failed on the approved code. It was a real accessibility bug that Grok had under-ranked. When it says something is minor, I check anyway.

The phone test also turned up five issues that neither model's review raised, like a broken icon on every attached document. Tests found none of them, and neither did the reviews. That is a job for the phone, whichever model reviews.

The Weekly Limit Is the Real Cost

SuperGrok has a weekly limit, and in the first week I ran straight into it. The week started on a Monday at 11:06 AM. By Wednesday morning, after the four Zone 2 Trainer releases in this post, one more release I never assessed and left out, and a branch cleanup session, Grok stopped mid-session with "You have run out of credits." It stayed that way until the next Monday:

SuperGrok usage panel reading "Weekly limit reached", 100% used, 98% of it Grok Build and 2% Chat, resetting September 28, 2026 at 11:06 AM

That is two days of work and five days of waiting, which is worse than the three days a week I wrote about with Ollama Pro.

The next week I watched the meter. Implementing two releases of my assistant app in one afternoon, including their review rounds, used 63% of the week. Reviewing the seventh release took it to 66%:

SuperGrok usage panel showing the weekly limit at 66% used, all of it Grok Build, resetting October 5, 2026, with $0.00 of extra usage credits

So about 3% for a review, against roughly 31% for each release it implemented. That comparison needs caveats. Reviewing reads a diff, while implementing explores, writes, tests and repeats. The releases were different. The 3% covers the first two review rounds only. And I have no matching number on the Claude side, because that week's Claude usage was mixed with other work.

Even with those caveats, it changes how I think about the plan. Speed does not save weekly budget. A fast release spends the same week as a slow one, it just gets there sooner. If one Grok review costs a tenth of one Grok implementation, and the review finds the bugs that matter, then the review is where that plan earns its money.

It also showed me what speed skips. In the 65-minute release, the full test suite ran once, on the version that still had all four first-pass bugs. The merge happened before anyone looked at the last commit, and the phone checks in the scope never ran. Claude ran them afterwards, and everything that could be checked passed. But that took about another hour, on a different budget. The checks still had to be paid for. They just came out of a different plan.

What I Am Not Saying

Seven releases is not a benchmark. The work was different every time, some releases were much harder than others, and one person judged all of it. The same model wrote most of the specs, reviewed the code, and helped me collect the evidence for this post. Part of the seventh release ran on airplane Wi-Fi, so I would not compare those times with the rest.

I am also not saying Grok is better than Claude, or the other way around. Each model missed something a careful read of the diff did not catch, including in its own code.

Where I Landed

Grok did well where the Ollama models struggled: few review rounds, and no "fixed" that was not fixed. Its fix passes still need review. Three of six added a new bug, even though none of them claimed more than they did.

What changed is how I review. Reading the changed lines carefully is not enough, because the bugs that got through lived in what those lines assumed: sometimes in another file, sometimes in a comment a few dozen lines away. Now I ask a different question depending on the work. For concurrency, when does this run, compared to what it depends on? For integrations, whose behavior did this inherit? For specs I write, am I describing the result I want, or a mechanism a good model will follow exactly?

And on a weekly limit, the numbers point toward spending Grok on review more than on implementation.

References

Comments

Popular posts from this blog

CockroachDB performance characteristics with YCSB(A) benchmark

Digsby is bringing out a Linux and Mac client very soon

Running CockroachDB with Docker Compose and Minio, Part 2