I Replaced $300/Month in OpenAI API Costs With a Local Stack. Here's Exactly What That Took.
Replacing $300/mo OpenAI API costs with local Ollama stack
Hardware cost amortization. Here's what that actually involved.
The Setup
Latency tradeoffs
What I Built
Hardware cost amortization
Hardware cost amortization. This is where most implementations go wrong — not in the concept, but in the specifics that documentation tends to gloss over. The version that's actually running in production looks different from what the README describes.
Latency tradeoffs
Latency tradeoffs. This is where most implementations go wrong — not in the concept, but in the specifics that documentation tends to gloss over. The version that's actually running in production looks different from what the README describes.
Which models replaced
Which models replaced which. This is where most implementations go wrong — not in the concept, but in the specifics that documentation tends to gloss over. The version that's actually running in production looks different from what the README describes.
What Broke First
Litellm routing. This is the part that took the most time to figure out — not the implementation itself, but the failure mode you only encounter at the wrong moment. The fix is in the tooling, not the concept.
What's Running in Production
The stack: llm, ollama, openai, cost. Monitoring via Zabbix with a custom alert template. The rule I apply to everything in production: if it can't self-report a failure, it doesn't go in.
Key Takeaways
- Hardware — hardware cost amortization
- Latency — latency tradeoffs
- Which — which models replaced which
Questions or a different approach? Reply to The Operator's Edge newsletter — I read every response.
Live Life Automated covers the broader philosophy behind how I think about automation and building systems that don't require constant babysitting. Available at mfitz.net.