Posts

Showing posts with the label code review

Cheap Plan, Expensive Reviews: Why I Cancelled Ollama

Image
Two weeks ago I ended a post about rate limits with "Ollama is back." I had re-subscribed monthly on purpose, to answer the question properly this time: does this plan give me the flexibility I did not have the first time around? I have the answer now, and it is no. I cancelled. The short version is that the money ran out early. The longer version is the reason for this post: even while the money lasted, the plan was not the real cost. The real cost was how many review rounds it took to get the work to a state I could merge. The Money Part Ollama Pro is $20 a month and includes $60 of cloud usage. This is my dashboard today: The $60 was gone within the first week of the month. When I saw the reset was still three weeks away, I cancelled. Most of the requests went to kimi-k2.7-code , the model I said in the last post I had not evaluated yet. I evaluated it. For comparison, my Claude weekly limit has covered most of this past week of daily work across two apps, and i...

Executor Model, Reviewer Model: Why Cross-Vendor Code Review Actually Works

Image
I thought one of my projects was in reasonable shape. I had been working in it for months, it was on version 4.24.3, and Claude had been through most of it more than once. Then I pointed Codex GPT-5.6 Sol at it with Ultra High reasoning, asked it to evaluate the codebase and file issues for anything it found, and it came back with 42 new issues in about ten minutes. Some of them were release blocking. I had Claude Opus 4.8 go through all 42 looking for false positives, and it confirmed zero. Every issue was real. That result is what convinced me to stop having one vendor's model check its own family's work. The Setup: One Model Writes, a Different Vendor Reviews The workflow is simple to state. Implement with one model, working from the most critical issues to the least critical and matching model strength to issue criticality: the highest-confidence model takes the most critical work, and the cheaper, faster models take the routine fixes further down the list. Then valid...