Just in

Local AI · Topic hub

Ollama & LM Studio

The two runners most people start with — setup, model management and where each wins.

16 articles across guides, comparisons, tools

MacBook open on a minimalist desk at blue hour with teal light trails streaming from the screen, beneath a large glowing LM Studio emblemGuides

Local LLMs on Apple Silicon: Unified Memory Wins

A 64GB MacBook loads models a 24GB RTX 4090 cannot touch. What unified memory buys you, where MLX still beats everything, and which stack to actually run.

14 Aug6 min read
Cinematic local AI hardware illustration for: Local LLMs for Coding: What Actually Works (2026)Guides

Local LLMs for Coding: What Actually Works (2026)

Which local coding models are real vs toys, wiring Ollama's OpenAI API into your editor, real latency numbers, and the privacy case for client code.

12 Aug9 min read
Cinematic local AI hardware illustration for: Ollama Agent Mode (2026): What Typing 'ollama' Does NowGuides

Ollama Agent Mode (2026): What Typing 'ollama' Does Now

Ollama v0.32 turned the bare 'ollama' command into a local agent that codes, edits files and runs skills. What changed, how we use it, and how to opt out.

10 Aug6 min read
Cinematic local AI hardware illustration for: LM Studio Bionic: Local AI Agents That Touch Your FilesGuides

LM Studio Bionic: Local AI Agents That Touch Your Files

LM Studio shipped Bionic, a Mac app that turns open models into file-touching agents. What it does, what it can't, and whether it replaces your chat app.

27 Jul5 min read
Cinematic local AI hardware illustration for: Build a Fully Offline Voice Assistant With Home AssistantGuides

Build a Fully Offline Voice Assistant With Home Assistant

Faster-whisper, Piper, and Ollama wired into Home Assistant Assist, with real latency numbers for built-in intents, GPU tier and CPU-only Raspberry Pi hardware.

27 Jul5 min read
Cinematic local AI hardware illustration for: Qwen3.6-27B Local Setup: A 27B Model That Beats a 397B OneGuides

Qwen3.6-27B Local Setup: A 27B Model That Beats a 397B One

Alibaba's Qwen3.6-27B fits on one RTX 4090 and edges its own 397B-parameter predecessor on coding benchmarks. Our Ollama setup, VRAM notes and first tokens/sec.

18 Jul5 min read
Cinematic local AI hardware illustration for: Gemma 4 12B: Run Google's New Multimodal Model LocallyGuides

Gemma 4 12B: Run Google's New Multimodal Model Locally

Google DeepMind's Gemma 4 12B runs text, image, audio and video natively on a 16GB machine. What's confirmed at launch, and the Ollama setup.

17 Jul6 min read
Cinematic local AI hardware illustration for: How to Run Llama Locally: 3 Ways Ranked (2026)Guides

How to Run Llama Locally: 3 Ways Ranked (2026)

Ollama's one command, LM Studio's GUI, or raw llama.cpp — three ways to run Llama locally, ranked from daily use, with the VRAM table per model size.

16 Jul3 min read
Cinematic local AI hardware illustration for: Ollama Complete Guide (2026): Install to Daily UseGuides

Ollama Complete Guide (2026): Install to Daily Use

Everything we know from running Ollama for real work: install, picking models and quants, the API, context-length tuning, and the mistakes that waste VRAM.

Updated 10 Aug5 min read
Cinematic local AI hardware illustration for: The 9 Best Local AI Tools We Actually Run (2026)Guides

The 9 Best Local AI Tools We Actually Run (2026)

We run local AI daily on our own RTX 4080 rig. These are the 9 tools that survived — LLM runtimes, image and video pipelines — plus the 5 we cut and why.

Updated 7 Aug4 min read

More in Local AI