Local AI · Topic hub
GPUs & Hardware
VRAM maths, RTX buying guides and what actually predicts local performance.
31 articles across guides, comparisons
Guides24
All guides →
GuidesRTX 5090 AI TOPS Explained: What 3352 Actually Measures
NVIDIA's 3352 AI TOPS is sparse FP4, not the precision your local model runs. The full ladder, the 16x gap, and when TOPS is the wrong spec to buy on.
GuidesVRAM for Training vs Inference: The 8x Byte Ledger
Training costs 16 bytes per parameter; inference costs 2. Here is the full byte ledger, why the multiplier is a range, and which GPU it rules out.
GuidesCloud GPU Pricing: When Renting Beats Buying a Card
H100s span $1.49 to $6.98 an hour across providers. Here's the break-even maths against buying, and the billing details that decide your real bill.
GuidesWhy Are GPUs Used for AI? The Plain-English Explainer
Why GPUs, not CPUs, run AI: parallel cores, matrix math and tensor cores, benchmarked with real VRAM and bandwidth specs from the RTX 4080 we render on daily.
GuidesRun Stable Diffusion Locally: Setup That Works in 2026
SD1.5 vs SDXL VRAM math, the ComfyUI desktop install, your first workflow, where checkpoints come from, and the errors that eat everyone's first afternoon.
GuidesLocal LLMs for Coding: What Actually Works (2026)
Which local coding models are real vs toys, wiring Ollama's OpenAI API into your editor, real latency numbers, and the privacy case for client code.
GuidesStable Audio 3 Local Setup: Run the Open Weights on Your GPU
Stable Audio 3's open-weight models run on consumer hardware. The Hugging Face repos, the stable-audio-3 library, license limits, and what a 4080 can expect.
GuidesPoolside Laguna XS 2.1: A 33B Coding Model for One GPU
Poolside's Laguna XS 2.1 activates just 3B of its 33B parameters per token, enough for real agentic coding on a single consumer GPU. Specs and setup.
GuidesNemotron 3 Nano 4B: NVIDIA's Edge Model on Your RTX Card
A 4B hybrid Mamba-Transformer model NVIDIA built for Jetson and RTX hardware, not data centers. Architecture, benchmarks, license, and the run command.
GuidesWan 2.2 Animate: Character Animation on Your GPU (2026)
Wan 2.2 Animate transfers a real performance onto any character image, motion and expressions included. ComfyUI setup, VRAM tiers, and mode choice.
GuidesRun MiniMax H3 Locally: Open Weights in ComfyUI (2026)
MiniMax H3's open weights are live. What you actually download, the 42.5GB ComfyUI footprint, the 12GB claim, and which parts stay hosted-only.
GuidesRun LTX-2 Locally: AI Video With Audio on 16GB (2026)
LTX-2 generates video and synced audio in one pass on your own GPU. The real VRAM tiers, the 2.3 update, ComfyUI setup, and where 16GB cards hit limits.
GuidesNVIDIA RTX Spark for Local AI (2026): What It Actually Runs
128GB unified memory, a 20-core Grace CPU, a Blackwell GPU on one chip. What RTX Spark can actually run locally, what it costs, and who should wait to buy.
GuidesOpen-Weight AI Video Models You Can Actually Run
MiniMax H3 going open-weight restarts this argument. What runs on 8GB, what needs 24GB, and what the quantization actually costs you in quality.
GuidesHunyuan 3 Local Guide: Can You Actually Run 295B at Home?
Tencent open-sourced Hunyuan 3 — 295B total, 21B active, Apache-2.0. We break down the real VRAM math, the GGUF quant path, and who should skip it entirely.
GuidesMiniMax M3 Locally: The 428B Model's Real Hardware Cost
MiniMax's open-weight M3 posts frontier coding scores with a 1M-token context. What self-hosting a 428B MoE model takes, and who should just use the API.
GuidesDeepSeek V4 Local Guide: Real VRAM Needs in 2026
DeepSeek V4 shipped MIT-licensed open weights. But no stable Ollama build loads it yet — here's what actually runs locally, on what hardware, and what doesn't.
GuidesKimi K3 Locally: The 2.8T Hardware Reality Check (2026)
Moonshot's Kimi K3 tops open-model charts, with weights landing July 27. What it actually takes to run 2.8T parameters at home — and what to run instead.
GuidesGLM-5.2 Local Setup: VRAM, Quants and Real Hardware Paths
Z.ai's GLM-5.2 leads every open-weights coding benchmark. Here's the honest VRAM math per quant, the three hardware paths that work, and who should bother.
GuidesQwen3.6-27B Local Setup: A 27B Model That Beats a 397B One
Alibaba's Qwen3.6-27B fits on one RTX 4090 and edges its own 397B-parameter predecessor on coding benchmarks. Our Ollama setup, VRAM notes and first tokens/sec.
GuidesRun Wan 2.2 Locally: Free AI Video on Your Own GPU (2026)
Our exact local Wan 2.2 setup on an RTX 4080: GGUF quantization, the Triton + SageAttention + TeaCache speed stack, and the mistakes that waste a weekend.
GuidesGemma 4 12B: Run Google's New Multimodal Model Locally
Google DeepMind's Gemma 4 12B runs text, image, audio and video natively on a 16GB machine. What's confirmed at launch, and the Ollama setup.
GuidesUsed RTX 3090 for AI in 2026: The 24GB Bargain Guide
Why a used RTX 3090 is still the smartest VRAM-per-dollar buy for local AI, what 24GB actually unlocks, ex-mining risks, and the day-one tests that protect you.
GuidesBest GPU for Local AI in 2026: A VRAM-First Guide
The GPU guide written from a rig that renders AI daily: why VRAM beats speed, what each budget tier really runs, and the used cards that embarrass new ones.
Comparisons7
All comparisons →
ComparisonsRTX 5070 Ti vs RTX 3090 for Local AI: 16GB New or 24GB Used?
The 5070 Ti's 16GB warranty versus a used 3090's 24GB ceiling. Bandwidth, tokens/sec, power and price compared — plus the one question that settles it.
ComparisonsRTX 3090 vs Radeon AI PRO R9700 (2026): 24GB or 32GB?
AMD's new $1,299 workstation card just launched with 32GB. We compare it against the used RTX 3090 we actually recommend most for local AI VRAM per dollar.
ComparisonsRTX 5090 vs 4090 for Local AI: Is 32GB Worth It?
The 5090's 32GB and 1,792 GB/s against the 4090's 24GB — what the bandwidth gap actually does to tokens per second, and when the used 4090 still wins.
ComparisonsUsed RTX 3090 vs New RTX 4060 Ti 16GB for Local AI (2026)
24GB used vs 16GB new, verified: bandwidth, tokens/sec on 8B-32B models, power draw, current pricing, and which model tier each card actually wins.
ComparisonsUsed RTX 3090 vs RTX 5060 Ti 16GB for Local AI (2026)
24GB used vs 16GB new, tested for local LLMs: memory bandwidth, real tokens/sec, power draw, and which model sizes each card can actually hold.
ComparisonsLocal AI vs Cloud APIs: The Real Cost Math (2026)
We run both daily. When local hardware beats per-token APIs, when cloud wins, and the break-even math nobody shows: electricity, amortization, rental.
ComparisonsOllama vs LM Studio (2026): Same-Hardware Verdict
We run both on the same 16GB RTX 4080. Where Ollama's one-command workflow wins, where LM Studio's GUI and offload sliders win, and where we keep both.