AI Video Sensei
Just in

Local AI · Topic hub

GPUs & Hardware

VRAM maths, RTX buying guides and what actually predicts local performance.

18 articles across guides, comparisons

Open-Weight AI Video Models You Can Actually RunGuides

Open-Weight AI Video Models You Can Actually Run

MiniMax H3 going open-weight restarts this argument. What runs on 8GB, what needs 24GB, and what the quantization actually costs you in quality.

1 Aug5 min read
Hunyuan 3 Local Guide: Can You Actually Run 295B at Home?Guides

Hunyuan 3 Local Guide: Can You Actually Run 295B at Home?

Tencent open-sourced Hunyuan 3 — 295B total, 21B active, Apache-2.0. We break down the real VRAM math, the GGUF quant path, and who should skip it entirely.

31 Jul5 min read
MiniMax M3 Locally: The 428B Model's Real Hardware CostGuides

MiniMax M3 Locally: The 428B Model's Real Hardware Cost

MiniMax's open-weight M3 posts frontier coding scores with a 1M-token context. What self-hosting a 428B MoE model takes, and who should just use the API.

27 Jul5 min read
DeepSeek V4 Local Guide: Real VRAM Needs in 2026Guides

DeepSeek V4 Local Guide: Real VRAM Needs in 2026

DeepSeek V4 shipped MIT-licensed open weights. But no stable Ollama build loads it yet — here's what actually runs locally, on what hardware, and what doesn't.

25 Jul4 min read
Kimi K3 Locally: The 2.8T Hardware Reality Check (2026)Guides

Kimi K3 Locally: The 2.8T Hardware Reality Check (2026)

Moonshot's Kimi K3 tops open-model charts, with weights landing July 27. What it actually takes to run 2.8T parameters at home — and what to run instead.

20 Jul4 min read
GLM-5.2 Local Setup: VRAM, Quants and Real Hardware PathsGuides

GLM-5.2 Local Setup: VRAM, Quants and Real Hardware Paths

Z.ai's GLM-5.2 leads every open-weights coding benchmark. Here's the honest VRAM math per quant, the three hardware paths that work, and who should bother.

20 Jul3 min read
Qwen3.6-27B Local Setup: A 27B Model That Beats a 397B OneGuides

Qwen3.6-27B Local Setup: A 27B Model That Beats a 397B One

Alibaba's Qwen3.6-27B fits on one RTX 4090 and edges its own 397B-parameter predecessor on coding benchmarks. Our Ollama setup, VRAM notes and first tokens/sec.

18 Jul5 min read
Run Wan 2.2 Locally: Free AI Video on Your Own GPU (2026)Guides

Run Wan 2.2 Locally: Free AI Video on Your Own GPU (2026)

Our exact local Wan 2.2 setup on an RTX 4080: GGUF quantization, the Triton + SageAttention + TeaCache speed stack, and the mistakes that waste a weekend.

17 Jul5 min read
Guides

Gemma 4 12B: Run Google's New Multimodal Model Locally

Google DeepMind's Gemma 4 12B runs text, image, audio and video natively on a 16GB machine. What's confirmed at launch, and the Ollama setup.

17 Jul5 min read
Used RTX 3090 for AI in 2026: The 24GB Bargain GuideGuides

Used RTX 3090 for AI in 2026: The 24GB Bargain Guide

Why a used RTX 3090 is still the smartest VRAM-per-dollar buy for local AI, what 24GB actually unlocks, ex-mining risks, and the day-one tests that protect you.

16 Jul3 min read
Best GPU for Local AI in 2026: A VRAM-First GuideGuides

Best GPU for Local AI in 2026: A VRAM-First Guide

The GPU guide written from a rig that renders AI daily: why VRAM beats speed, what each budget tier really runs, and the used cards that embarrass new ones.

10 Jul3 min read
RTX 5070 Ti vs RTX 3090 for Local AI: 16GB New or 24GB Used?Comparisons

RTX 5070 Ti vs RTX 3090 for Local AI: 16GB New or 24GB Used?

The 5070 Ti's 16GB warranty versus a used 3090's 24GB ceiling. Bandwidth, tokens/sec, power and price compared — plus the one question that settles it.

31 Jul5 min read
RTX 3090 vs Radeon AI PRO R9700 (2026): 24GB or 32GB?Comparisons

RTX 3090 vs Radeon AI PRO R9700 (2026): 24GB or 32GB?

AMD's new $1,299 workstation card just launched with 32GB. We compare it against the used RTX 3090 we actually recommend most for local AI VRAM per dollar.

27 Jul5 min read
RTX 5090 vs 4090 for Local AI: Is 32GB Worth It?Comparisons

RTX 5090 vs 4090 for Local AI: Is 32GB Worth It?

The 5090's 32GB and 1,792 GB/s against the 4090's 24GB — what the bandwidth gap actually does to tokens per second, and when the used 4090 still wins.

25 Jul4 min read
Used RTX 3090 vs New RTX 4060 Ti 16GB for Local AI (2026)Comparisons

Used RTX 3090 vs New RTX 4060 Ti 16GB for Local AI (2026)

24GB used vs 16GB new, verified: bandwidth, tokens/sec on 8B-32B models, power draw, current pricing, and which model tier each card actually wins.

22 Jul5 min read
Comparisons

Used RTX 3090 vs RTX 5060 Ti 16GB for Local AI (2026)

24GB used vs 16GB new, tested for local LLMs: bandwidth, tokens/sec, power draw, and which model classes each one actually unlocks.

19 Jul4 min read
Local AI vs Cloud APIs: The Real Cost Math (2026)Comparisons

Local AI vs Cloud APIs: The Real Cost Math (2026)

We run both daily. When local hardware beats per-token APIs, when cloud wins, and the break-even math nobody shows — electricity, amortization and rental included.

17 Jul3 min read
Ollama vs LM Studio (2026): Same-Hardware VerdictComparisons

Ollama vs LM Studio (2026): Same-Hardware Verdict

We run both on the same 16GB RTX 4080. Where Ollama's one-command workflow wins, where LM Studio's GUI and offload sliders win, and the setup where we keep both.

Updated 27 Jul4 min read

More in Local AI