Cheap Plan, Expensive Reviews: Why I Cancelled Ollama

Two weeks ago I ended a post about rate limits with "Ollama is back." I had re-subscribed monthly on purpose, to answer the question properly this time: does this plan give me the flexibility I did not have the first time around? I have the answer now, and it is no. I cancelled.
The short version is that the money ran out early. The longer version is the reason for this post: even while the money lasted, the plan was not the real cost. The real cost was how many review rounds it took to get the work to a state I could merge.
The Money Part
Ollama Pro is $20 a month and includes $60 of cloud usage. This is my dashboard today:

The $60 was gone within the first week of the month. When I saw the reset was still three weeks away, I cancelled. Most of the requests went to kimi-k2.7-code, the model I said in the last post I had not evaluated yet. I evaluated it.
For comparison, my Claude weekly limit has covered most of this past week of daily work across two apps, and it is on pace to run out a day or so before it resets. A weekly limit that runs out near the end of the week costs a day of waiting. A monthly credit that runs out with three weeks left costs three weeks.
That alone was enough to cancel. But it is not the part I want to write about.
A Small Release, Three Problems
Here is one real example, from my Zone 2 Trainer app. I asked kimi-k2.7-code to "scope a bug fix release, pull all issues that are marked as bug and build 6.10.2." It did a sensible job of the scoping. It asked three questions, proposed cutting the patch from the last release tag instead of the main branch so no unverified Bluetooth changes would ride along, and picked two bug fixes. Then I asked it to post a pull request:

Every automated check passed: release metadata, unit tests, lint, and the release build. I sent the pull request to Claude Opus 5 for review. It said the core fix was sound and the new test was real, and then it found three problems:
- It could mix up whose workout is whose. Some workouts in my app come in from another fitness app. If I edited one of those, my app would forget where it came from, and on the next sync it would send its own copy of the heart rate, calories and distance back to Health Connect. Now the same workout shows up twice, and your totals for the day are wrong.
- It could lose the Strava badge for good. A small badge shows which workouts are already on Strava. Editing a workout while Strava was turned off wiped that badge, and nothing would ever bring it back, even though the workout was still on Strava.
- The release notes promised a fix that was not in the release. They said 6.10.2 fixed one bug. That fix was already in the app two versions earlier, and nobody had tested it on a real workout yet.
To be fair to Kimi, the fix pass was quick and correct. It took under nine minutes, added two regression tests, and a second review found nothing. Then it merged, tagged the release, built it, installed it on my phone, and wrote this:

For the final release, that is fair. Nothing in it touched the sensors. But the first version said the same thing while it was still promising the other fix, and that one could only be checked by hand. The bug was about pausing: if you paused a workout, the app saved the wrong start time, as if you had started later than you really did. To check it, you have to do a real workout, pause it partway, finish, and see whether the history shows the time you actually started. Nobody had done that, and it still has not been done.
None of these three were broken code. Nothing crashed and no test failed. The work looked finished, and it was wrong in the places where you have to know what the data means.
The Same Shape in a Bigger Release
The example above is small on purpose. The same pattern showed up in my personal assistant app, on a bigger release.
kimi-k2.7-code built a release covering four issues. It shipped with every automated gate green: formatter, analyzer, the project rule checker, and 3,416 tests passing. Within about an hour of the release, using it on my phone had turned up three defects:
- A calendar event spanning several days still showed only its start date in the editor. That was the exact defect the issue described, in the file it named.
- An "Edit fields" button did nothing. No error, no feedback. The code that handled button actions had no branch for that action, so it just returned.
- Events read from a photo were supposed to be marked "(inferred)" where the model guessed a value. They never were. The flags defaulted to false and the model never set them. The tests passed because they built their own flags.
The follow-up pull request, from the same model, addressed all three. Review found eight more problems in the fix. One example: the new end-date row reused the start-date picker, so picking a new end date moved the whole event. It took three review rounds to get it mergeable.
When the Fix Is Where Things Break
The part I found most useful came from one pull request where two models took turns answering the same review, back to back. It is about as close to a controlled test as my real work gets.
| Round | Model | Findings to fix | What the fix pass did |
|---|---|---|---|
| 1 | deepseek-v4-pro:0813 |
12 | Fixed 6. Left 3 still broken but reported them fixed. Added 3 new regressions. |
| 2 | glm-5.3 |
6 | Fixed all 3 regressions. Fixed the other 3 only partly. Added 2 narrow new defects. |
| 3 | Claude Opus 5 | 6 | Closed all six by changing the underlying rules. |
The reviewer in all three rounds was Claude Opus 5. Rounds one and two are the setup I argued for in Executor Model, Reviewer Model: one company's model writes the code and another company's model reviews it. Round three is not. There, Opus 5 was fixing findings from its own review, so keep that in mind when reading the last row.
The DeepSeek pass is the one that stays with me. Its commit said it addressed all 12 findings. Three of them were still live, including a layout overflow "fixed" in a way that still overflowed, with a new code comment that described the opposite of what the code does. Two of the three new regressions broke behavior that worked on the main branch.
GLM was clearly better. It broke nothing. But it had the same habit, fixing the example instead of the rule. A currency check went from missing $5M to handling $5M and still missing $5bn.
That matches what I saw in July. Back then GLM 5.2 was my daily driver, and in my comparison of Codex, Claude and GLM 5.2 about 40% of its findings did not hold up when Claude checked them. The newer model broke less here, but its work still needed checking.
About Speed
When I started writing this, I planned to blame part of it on the open models being slow. I went back to the timestamps to check, and that turned out to be only partly true.
On the 6.10.2 release, Kimi's own time was short: under a minute to post the pull request, under nine minutes for the fix pass, under two minutes to release. That is in the same range as the other models I use. The long gaps on the pull request were mine, not the model's.
On the bigger jobs, the wall-clock numbers are worse. Kimi took over an hour to answer an eight-finding review, and DeepSeek almost an hour for twelve. But those times include my own gaps too, so I would not lean on them.
The number that does not depend on my schedule is rounds. How many times does the work go around review before it is right, and does each fix pass make things better or add new problems? That is where the open models cost me.
What I Am Not Saying
The tests were real. The releases shipped and mostly work. For mechanical, well-bounded work, like renames, tests, or a screen that looks like four other screens, these models are genuinely useful.
What I am saying is about the cost of getting from "looks done" to "is done." Every extra review round is time, attention, and budget, and the budget for review has to come from somewhere.
Where I Landed
I cancelled Ollama Pro. The $60 did not last the first week, and the work it bought needed more rounds of review than the work from the plans I am keeping. That also ends the two-plan setup I described in GLM Plans, Fable Validates, and API Keys Are Not for Development, where an Ollama plan covered my less critical projects.
The way I compare plans now: price per month tells you very little. Multiply it by the number of review rounds it takes to get to mergeable, and then check how long you wait when the limit runs out. That is the number that decides whether the work actually ships.
Comments