Vigyata.AI
Is this your channel?

Claude Code + Ollama = Free Forever

225.4K views· 4,026 likes· 9:15· Jan 25, 2026

🛍️ Products Mentioned (5)

🚀 Access ALL video resources & get personalized help in my community: https://www.skool.com/agentic-labs 🦙 Download Ollama: https://ollama.com 💬 My AI voice-to-text software (Wispr Flow): https://wisprflow.ai/r?WALTER67 ☕ Buy me a coffee: https://www.buymeacoffee.com/leonvanzyl 💵 Donate using PayPal: https://www.paypal.com/ncp/payment/EKRQ8QSGV6CWW In this video you will learn how to use local models in Claude Code completely free using Ollama. Run AI coding agents offline on your own machine with no API costs or subscriptions. You will learn how to install Ollama, download recommended models like Qwen3-Coder, GPT-OSS, and GLM 4.7 Flash, and configure Claude Code to use local LLMs. Perfect for vibe coding and agentic coding on consumer grade hardware with full access to Claude Code features like subagents, MCP servers, and slash commands. ⏰ TIMESTAMPS: 00:00 Claude Code local models demo 01:05 Install Ollama on your machine 01:47 Download models from Ollama 02:15 Recommended models for Claude Code 02:30 GPT OSS 20B VRAM requirements 02:42 Qwen3 Coder best local coding model 02:49 GLM 4.7 Flash model overview 05:30 Configure Claude Code with Ollama 06:06 Test offline mode no API key 06:36 Local models vs Sonnet and Opus 07:10 Troubleshooting Ollama errors 08:01 Context window size comparison #claudecode #vibecoding #agenticcoding

About This Video

In this video I show you how I’m using local models inside Claude Code completely free, using Ollama. I demo it by building a real-time weather app that hits a live API, pulls accurate data for any city, and wraps it in nice UI animations—while the actual coding agent runs locally on my machine. No subscriptions, no API costs, and you can run securely offline. The best part is you still get the full Claude Code experience: slash commands, MCP servers, sub-agents, and switching between plan/edit modes—everything just works on consumer-grade hardware. I walk you through installing Ollama, pulling models, and then launching Claude Code the “special” way so it doesn’t ask for an Anthropic login or API key. I also break down which models to use and why: GPT-OSS 20B is the smallest, Qwen3-Coder is my current favorite (better results for me), and GLM 4.7 Flash is a solid option—especially if you’ve got the VRAM for the bigger quant. I’ll also cover troubleshooting (like 404s or version mismatches) and how context window sizes differ, because that directly affects response quality once your sessions get large.

Frequently Asked Questions

🎬 More from Leon van Zyl