Local AI · Topic hub
Local LLMs
Llama, Qwen, DeepSeek, GLM, Kimi and Gemma — running frontier models on your own box.
35 articles across guides, comparisons, tools
Guides22
All guides →
GuidesRTX 5090 AI TOPS Explained: What 3352 Actually Measures
NVIDIA's 3352 AI TOPS is sparse FP4, not the precision your local model runs. The full ladder, the 16x gap, and when TOPS is the wrong spec to buy on.
Local LLMs on Apple Silicon: Unified Memory Wins
A 64GB MacBook loads models a 24GB RTX 4090 cannot touch. What unified memory buys you, where MLX still beats everything, and which stack to actually run.
GuidesLocal LLMs for Coding: What Actually Works (2026)
Which local coding models are real vs toys, wiring Ollama's OpenAI API into your editor, real latency numbers, and the privacy case for client code.
GuidesOllama Agent Mode (2026): What Typing 'ollama' Does Now
Ollama v0.32 turned the bare 'ollama' command into a local agent that codes, edits files and runs skills. What changed, how we use it, and how to opt out.
GuidesNemotron 3 Nano 4B: NVIDIA's Edge Model on Your RTX Card
A 4B hybrid Mamba-Transformer model NVIDIA built for Jetson and RTX hardware, not data centers. Architecture, benchmarks, license, and the run command.
GuidesQwen3.8-Max: What's Confirmed and What Isn't Yet (2026)
Alibaba's 2.4T-parameter Qwen3.8-Max went live August 3 as a hosted API only. Open weights for it and a 27B sibling are due next week: the real spec sheet.
GuidesNVIDIA RTX Spark for Local AI (2026): What It Actually Runs
128GB unified memory, a 20-core Grace CPU, a Blackwell GPU on one chip. What RTX Spark can actually run locally, what it costs, and who should wait to buy.
GuidesVoxtral: Mistral's Offline Transcription Model, Tested
Mistral's Voxtral runs speech-to-text fully offline under Apache 2.0, in 3B and 24B sizes. VRAM needs, how it differs from Whisper, and the browser build.
GuidesMiniMax M3 Locally: The 428B Model's Real Hardware Cost
MiniMax's open-weight M3 posts frontier coding scores with a 1M-token context. What self-hosting a 428B MoE model takes, and who should just use the API.
GuidesBuild a Fully Offline Voice Assistant With Home Assistant
Faster-whisper, Piper, and Ollama wired into Home Assistant Assist, with real latency numbers for built-in intents, GPU tier and CPU-only Raspberry Pi hardware.
GuidesDeepSeek V4 Local Guide: Real VRAM Needs in 2026
DeepSeek V4 shipped MIT-licensed open weights. But no stable Ollama build loads it yet — here's what actually runs locally, on what hardware, and what doesn't.
GuidesLocal AI on a Laptop: What Your Machine Really Runs (2026)
No desktop GPU needed: what 8GB, 16GB and 32GB laptops genuinely run in 2026 — Gemma 4 12B in 16GB RAM, the E4B option, and honest speed expectations.
GuidesKimi K3 Locally: The 2.8T Hardware Reality Check (2026)
Moonshot's Kimi K3 tops open-model charts, with weights landing July 27. What it actually takes to run 2.8T parameters at home — and what to run instead.
GuidesGLM-5.2 Local Setup: VRAM, Quants and Real Hardware Paths
Z.ai's GLM-5.2 leads every open-weights coding benchmark. Here's the honest VRAM math per quant, the three hardware paths that work, and who should bother.
GuidesQwen3.6-27B Local Setup: A 27B Model That Beats a 397B One
Alibaba's Qwen3.6-27B fits on one RTX 4090 and edges its own 397B-parameter predecessor on coding benchmarks. Our Ollama setup, VRAM notes and first tokens/sec.
GuidesLLM Quantization Explained: Q4 vs Q8 in Practice
What quantization actually does to local models, GGUF quant names decoded, the real quality cost of Q4, and when stepping up to Q6 or Q8 is worth the VRAM.
GuidesGemma 4 12B: Run Google's New Multimodal Model Locally
Google DeepMind's Gemma 4 12B runs text, image, audio and video natively on a 16GB machine. What's confirmed at launch, and the Ollama setup.
GuidesUsed RTX 3090 for AI in 2026: The 24GB Bargain Guide
Why a used RTX 3090 is still the smartest VRAM-per-dollar buy for local AI, what 24GB actually unlocks, ex-mining risks, and the day-one tests that protect you.
GuidesHow to Run Llama Locally: 3 Ways Ranked (2026)
Ollama's one command, LM Studio's GUI, or raw llama.cpp — three ways to run Llama locally, ranked from daily use, with the VRAM table per model size.
GuidesOllama Complete Guide (2026): Install to Daily Use
Everything we know from running Ollama for real work: install, picking models and quants, the API, context-length tuning, and the mistakes that waste VRAM.
GuidesThe 9 Best Local AI Tools We Actually Run (2026)
We run local AI daily on our own RTX 4080 rig. These are the 9 tools that survived — LLM runtimes, image and video pipelines — plus the 5 we cut and why.
GuidesBest GPU for Local AI in 2026: A VRAM-First Guide
The GPU guide written from a rig that renders AI daily: why VRAM beats speed, what each budget tier really runs, and the used cards that embarrass new ones.
Comparisons10
All comparisons →
ComparisonsLM Studio Bionic vs Ollama: Agent App or DIY Local Stack?
LM Studio's Bionic agent and Ollama solve different local AI problems. We compare agent features, GGUF and MLX engines, cloud spillover, and privacy controls.
ComparisonsRTX 5070 Ti vs RTX 3090 for Local AI: 16GB New or 24GB Used?
The 5070 Ti's 16GB warranty versus a used 3090's 24GB ceiling. Bandwidth, tokens/sec, power and price compared — plus the one question that settles it.
ComparisonsRTX 3090 vs Radeon AI PRO R9700 (2026): 24GB or 32GB?
AMD's new $1,299 workstation card just launched with 32GB. We compare it against the used RTX 3090 we actually recommend most for local AI VRAM per dollar.
ComparisonsRTX 5090 vs 4090 for Local AI: Is 32GB Worth It?
The 5090's 32GB and 1,792 GB/s against the 4090's 24GB — what the bandwidth gap actually does to tokens per second, and when the used 4090 still wins.
ComparisonsUsed RTX 3090 vs New RTX 4060 Ti 16GB for Local AI (2026)
24GB used vs 16GB new, verified: bandwidth, tokens/sec on 8B-32B models, power draw, current pricing, and which model tier each card actually wins.
ComparisonsKimi K3 vs GLM-5.2 (2026): Open-Weight Heavyweights Compared
Moonshot's 2.8T-parameter Kimi K3 against Z.ai's 744B GLM-5.2 — benchmarks, token pricing, agentic coding and what it takes to run each one yourself.
ComparisonsUsed RTX 3090 vs RTX 5060 Ti 16GB for Local AI (2026)
24GB used vs 16GB new, tested for local LLMs: memory bandwidth, real tokens/sec, power draw, and which model sizes each card can actually hold.
ComparisonsOllama vs llama.cpp (2026): Which Local LLM Tool Wins?
We run both on our rig. Ollama wraps llama.cpp for one-command ease; llama.cpp trades that away for raw speed and control. When each earns its place.
ComparisonsLocal AI vs Cloud APIs: The Real Cost Math (2026)
We run both daily. When local hardware beats per-token APIs, when cloud wins, and the break-even math nobody shows: electricity, amortization, rental.
ComparisonsOllama vs LM Studio (2026): Same-Hardware Verdict
We run both on the same 16GB RTX 4080. Where Ollama's one-command workflow wins, where LM Studio's GUI and offload sliders win, and where we keep both.
Tools3
All tools →
Toolsllama.cpp — Tool Hub: Facts, Setup & Our Verdict
Everything about llama.cpp in one place: what the engine does, GGUF basics, the server flags that matter, and when to drop down from Ollama.
ToolsOllama — Tool Hub: Facts, Best Tutorials & Verdict
Everything about Ollama in one place: what it does best, VRAM realities, its limits, the best video tutorials, and links to our tested guides.
ToolsLM Studio — Tool Hub: Facts, Tutorials & Verdict
Everything about LM Studio in one place: the desktop way to run LLMs locally, its VRAM-fit model browser, limits, best tutorials, and our tested guides.