AI Video Sensei
Just in

Derek Holt

Local AI & Hardware Writer

Runs the site's local-inference rig and benchmarks every GPU, quant, and speed-stack claim on it personally before it goes in a guide. Will not shut up about VRAM bandwidth.

Beat: Local AI

Latest from Derek

RTX 3090 vs Radeon AI PRO R9700 (2026): 24GB or 32GB?Comparisons

RTX 3090 vs Radeon AI PRO R9700 (2026): 24GB or 32GB?

AMD's new $1,299 workstation card just launched with 32GB. We compare it against the used RTX 3090 we actually recommend most for local AI VRAM per dollar.

2026-07-27
Kokoro vs Chatterbox vs XTTS: Best Local TTS in 2026Comparisons

Kokoro vs Chatterbox vs XTTS: Best Local TTS in 2026

A same-question comparison of the three local voice models that actually matter: what each one clones, what license lets you ship, and which one your hardware can even run.

2026-07-27
faster-whisper vs OpenAI Whisper: 4x Speed, Same Accuracy?Comparisons

faster-whisper vs OpenAI Whisper: 4x Speed, Same Accuracy?

SYSTRAN's own benchmark table doesn't hit the 4x it advertises. I did the division on their published numbers so you don't have to take the headline on faith.

2026-07-27
XTTS Voice Cloning Locally: Clone Any Voice, Fully OfflineGuides

XTTS Voice Cloning Locally: Clone Any Voice, Fully Offline

XTTS clones a voice from a 6-second sample on your own GPU, fully offline. The correct 2026 install, real VRAM numbers, and where long-form narration quality actually breaks down.

2026-07-27
Voxtral: Mistral's Offline Transcription Model, TestedGuides

Voxtral: Mistral's Offline Transcription Model, Tested

Mistral's Voxtral runs speech-to-text fully offline under Apache 2.0, in 3B and 24B sizes. VRAM needs, how it differs from Whisper, and the browser build.

2026-07-27
Run Whisper Locally: Free Offline Transcription (2026)Guides

Run Whisper Locally: Free Offline Transcription (2026)

faster-whisper install to first transcript, the large-v3-vs-turbo call I actually make, and what a real file costs on a GPU versus a CPU, with no API bill involved.

2026-07-27
Piper TTS: Fast Offline Voice Synthesis on a Raspberry PiGuides

Piper TTS: Fast Offline Voice Synthesis on a Raspberry Pi

Piper runs real-time neural text-to-speech on a Raspberry Pi with no GPU at all. The install, the license change nobody mentions, and which quality tier actually fits a Pi 4 versus a Pi 5.

2026-07-27
MiniMax M3 Locally: The 428B Model's Real Hardware CostGuides

MiniMax M3 Locally: The 428B Model's Real Hardware Cost

MiniMax's open-weight M3 posts frontier coding scores with a 1M-token context. What self-hosting a 428B MoE model takes, and who should just use the API.

2026-07-27
Local Voice AI: Whisper, TTS & Offline Assistants (2026)Guides

Local Voice AI: Whisper, TTS & Offline Assistants (2026)

The complete map of the self-hosted voice stack: Whisper and faster-whisper for speech-to-text, Kokoro/Piper/Chatterbox/XTTS for voices, Ollama for the brain, and how they wire together offline.

2026-07-27
LM Studio Bionic: Local AI Agents That Touch Your FilesGuides

LM Studio Bionic: Local AI Agents That Touch Your Files

LM Studio shipped Bionic, a separate Mac app that turns open models into file-touching agents. What it does, what it can't, and whether it replaces your chat app.

2026-07-27
Kokoro TTS Locally: The 82M-Parameter Voice Model SetupGuides

Kokoro TTS Locally: The 82M-Parameter Voice Model Setup

Why a tiny Apache-2.0 model with no voice cloning is still the narration tool I install first, plus the actual pip-install-to-first-spoken-line steps.

2026-07-27
Build a Fully Offline Voice Assistant With Home AssistantGuides

Build a Fully Offline Voice Assistant With Home Assistant

The complete faster-whisper, Piper, and Ollama pipeline wired into Home Assistant Assist, with real latency numbers across built-in intents, a GPU tier, and CPU-only Raspberry Pi hardware.

2026-07-27
RTX 5090 vs 4090 for Local AI: Is 32GB Worth It?Comparisons

RTX 5090 vs 4090 for Local AI: Is 32GB Worth It?

The 5090's 32GB and 1,792 GB/s against the 4090's 24GB — what the bandwidth gap actually does to tokens per second, and when the used 4090 still wins.

2026-07-25
DeepSeek V4 Local Guide: Real VRAM Needs in 2026Guides

DeepSeek V4 Local Guide: Real VRAM Needs in 2026

DeepSeek V4 shipped MIT-licensed open weights. But no stable Ollama build loads it yet — here's what actually runs locally, on what hardware, and what doesn't.

2026-07-25
Local AI on a Laptop: What Your Machine Really Runs (2026)Guides

Local AI on a Laptop: What Your Machine Really Runs (2026)

No desktop GPU needed: what 8GB, 16GB and 32GB laptops genuinely run in 2026 — Gemma 4 12B in 16GB RAM, the E4B option, and honest speed expectations.

2026-07-23
Used RTX 3090 vs New RTX 4060 Ti 16GB for Local AI (2026)Comparisons

Used RTX 3090 vs New RTX 4060 Ti 16GB for Local AI (2026)

24GB used vs 16GB new, verified: bandwidth, tokens/sec on 8B-32B models, power draw, current pricing, and which model tier each card actually wins.

2026-07-22
Kimi K3 vs GLM-5.2 (2026): Open-Weight Heavyweights ComparedComparisons

Kimi K3 vs GLM-5.2 (2026): Open-Weight Heavyweights Compared

Moonshot's 2.8T-parameter Kimi K3 against Z.ai's 744B GLM-5.2 — benchmarks, token pricing, agentic coding and what it takes to run each one yourself.

2026-07-20
Kimi K3 Locally: The 2.8T Hardware Reality Check (2026)Guides

Kimi K3 Locally: The 2.8T Hardware Reality Check (2026)

Moonshot's Kimi K3 tops open-model charts, with weights landing July 27. What it actually takes to run 2.8T parameters at home — and what to run instead.

2026-07-20
GLM-5.2 Local Setup: VRAM, Quants and Real Hardware PathsGuides

GLM-5.2 Local Setup: VRAM, Quants and Real Hardware Paths

Z.ai's GLM-5.2 leads every open-weights coding benchmark. Here's the honest VRAM math per quant, the three hardware paths that work, and who should bother.

2026-07-20
Tools

llama.cpp — Tool Hub: Facts, Setup & Our Verdict

Everything about llama.cpp in one place: what the engine does, GGUF basics, the server flags that matter, and when to drop down from Ollama.

2026-07-19
Comparisons

Used RTX 3090 vs RTX 5060 Ti 16GB for Local AI (2026)

24GB used vs 16GB new, tested for local LLMs: bandwidth, tokens/sec, power draw, and which model classes each one actually unlocks.

2026-07-19
Ollama vs llama.cpp (2026): Which Local LLM Tool Wins?Comparisons

Ollama vs llama.cpp (2026): Which Local LLM Tool Wins?

We run both on our rig. Ollama wraps llama.cpp for one-command ease; llama.cpp trades that away for raw speed and control. When each earns its place.

2026-07-18
Qwen3.6-27B Local Setup: A 27B Model That Beats a 397B OneGuides

Qwen3.6-27B Local Setup: A 27B Model That Beats a 397B One

Alibaba's Qwen3.6-27B fits on one RTX 4090 and edges its own 397B-parameter predecessor on coding benchmarks. Our Ollama setup, VRAM notes and first tokens/sec.

2026-07-18
Local AI vs Cloud APIs: The Real Cost Math (2026)Comparisons

Local AI vs Cloud APIs: The Real Cost Math (2026)

We run both daily. When local hardware beats per-token APIs, when cloud wins, and the break-even math nobody shows — electricity, amortization and rental included.

2026-07-17
LLM Quantization Explained: Q4 vs Q8 in PracticeGuides

LLM Quantization Explained: Q4 vs Q8 in Practice

What quantization actually does to local models, GGUF quant names decoded, the real quality cost of Q4, and when stepping up to Q6 or Q8 is worth the VRAM.

2026-07-17
Guides

Gemma 4 12B: Run Google's New Multimodal Model Locally

Google DeepMind's Gemma 4 12B runs text, image, audio and video natively on a 16GB machine. What's confirmed at launch, and the Ollama setup.

2026-07-17
Used RTX 3090 for AI in 2026: The 24GB Bargain GuideGuides

Used RTX 3090 for AI in 2026: The 24GB Bargain Guide

Why a used RTX 3090 is still the smartest VRAM-per-dollar buy for local AI, what 24GB actually unlocks, ex-mining risks, and the day-one tests that protect you.

2026-07-16
How to Run Llama Locally: 3 Ways Ranked (2026)Guides

How to Run Llama Locally: 3 Ways Ranked (2026)

Ollama's one command, LM Studio's GUI, or raw llama.cpp — three ways to run Llama on your own machine, ranked from our daily use, with the VRAM table per model size.

2026-07-16
Ollama — Tool Hub: Facts, Best Tutorials & VerdictTools

Ollama — Tool Hub: Facts, Best Tutorials & Verdict

Everything about Ollama in one place: what it does best, VRAM realities, its limits, the best video tutorials, and links to our tested guides.

2026-07-10
LM Studio — Tool Hub: Facts, Tutorials & VerdictTools

LM Studio — Tool Hub: Facts, Tutorials & Verdict

Everything about LM Studio in one place: the desktop way to run LLMs locally, its VRAM-fit model browser, limits, best tutorials, and our tested guides.

2026-07-10
ComfyUI — Tool Hub: Facts, Best Tutorials & VerdictTools

ComfyUI — Tool Hub: Facts, Best Tutorials & Verdict

Everything about ComfyUI in one place: the node-based local studio for image and video generation, VRAM needs, limits, best tutorials, and our workflow library.

2026-07-10
10 ComfyUI Workflows We Run Weekly (2026 Library)Prompts

10 ComfyUI Workflows We Run Weekly (2026 Library)

Copy-paste ComfyUI workflow recipes from a rig that renders daily: SDXL and Flux image pipelines, Wan video with the full speed stack, upscaling and batch tricks.

2026-07-10
Ollama vs LM Studio (2026): Same-Hardware VerdictComparisons

Ollama vs LM Studio (2026): Same-Hardware Verdict

We run both on the same 16GB RTX 4080. Where Ollama's one-command workflow wins, where LM Studio's GUI and offload sliders win, and the setup where we keep both.

2026-07-10
Ollama Complete Guide (2026): Install to Daily UseGuides

Ollama Complete Guide (2026): Install to Daily Use

Everything we know from running Ollama for real work: install, picking models and quants, the API, context-length tuning, and the mistakes that waste VRAM.

2026-07-10
The 9 Best Local AI Tools We Actually Run (2026)Guides

The 9 Best Local AI Tools We Actually Run (2026)

We run local AI daily on our own RTX 4080 rig. These are the 9 tools that survived — LLM runtimes, image and video pipelines — plus the 5 we uninstalled and why.

2026-07-10
Best GPU for Local AI in 2026: A VRAM-First GuideGuides

Best GPU for Local AI in 2026: A VRAM-First Guide

The GPU guide written from a rig that renders AI daily: why VRAM beats speed, what each budget tier really runs, and the used cards that embarrass new ones.

2026-07-10