- How a local model works and what limits its speed
- Can you run local AI models on a regular Ryzen VPS?
- What speed to expect
- What makes sense on a CPU VPS
- Quick start on a VPS
- GPU servers for serious workloads
- Hourly GPU rental and dedicated GPU servers
- Private options that accept crypto
- Mini PCs for local models at home
- What counts as comfortable
- Mini PCs that already handle vibe coding and images
- Mini PC with an external GPU over OCuLink
- AMD Ryzen AI Max+ 395 (Strix Halo) — all in one box
- Mac mini with M5 Pro
- Compact PCs with a built-in RTX
- NVIDIA DGX Spark
- Mid-range: Ryzen 7 and Ryzen AI 9 mini PCs with 32 GB
- Budget options
- What to avoid
- What to choose
- Local models for image generation
- Interfaces
- Local models for video generation
- Models for vibe coding
- The stack: engine plus editor
- Summary table: task → model → minimum hardware
- Common mistakes
- Where to start today
- Useful links
- FAQ
- Related reading
Local AI models are neural networks that run on your own computer or server instead of someone else’s cloud. Your chats, code and photos stay with you, there is no subscription, and the model keeps working without internet. The price is hardware: the smarter the model, the more memory it needs.
This is an honest guide without marketing promises: what you can realistically run on a regular Ryzen VPS, when you really need a GPU, which mini PC to buy for home, and which models are current for chat, images, video and vibe coding. Model versions and prices were checked in fall 2026. At the end there is a summary table: task → model → minimum hardware.
How a local model works and what limits its speed
A model is a large file with weights. You download it — usually from Hugging Face, and for MIT and Apache-2.0 models also by torrent via Pirate Face — and run it with an inference engine: Ollama, LM Studio or llama.cpp. All three expose an OpenAI-compatible API, so chat UIs, bots and code editors can connect to them.
- Memory size decides whether a model runs at all. With Q4 quantization (weights compressed to about 4 bits) a rough estimate is 0.6 GB per billion parameters plus 1–4 GB for context. An 8B model needs 5–6 GB, a 30B model 18–20 GB, a 120B model around 65 GB.
- Memory bandwidth decides speed. Every new token requires reading the active weights. That is why a GPU with 900+ GB/s answers many times faster than a CPU with regular DDR5.
- MoE models are faster than their size suggests. In a Mixture of Experts model only part of the weights works for each token: about 3B in Qwen3-Coder 30B, about 5B in gpt-oss-120b. You need memory for the whole model, but speed is close to a small one.
- Images and video are a different story. Diffusion models are limited by compute as well as memory. You need a GPU, ideally NVIDIA: most guides and optimizations target CUDA.
Can you run local AI models on a regular Ryzen VPS?
Yes, but only small language models. A regular VPS has no GPU, so the CPU does all the work. Modern Ryzen chips handle it reasonably well: fast cores with AVX2/AVX-512, and llama.cpp and Ollama are well optimized for CPUs. The limit is memory: there is not much of it, and bandwidth is shared with other tenants on the host.
As an example, here is the Ryzen VPS lineup at AlfaHost (Netherlands, dedicated vCPU, NVMe), as of this writing:
| Plan | vCPU | RAM | NVMe | Price per month | What it can run |
|---|---|---|---|---|---|
| RZ-4 | 2 | 4 GB | 50 GB | $13.50 | 1–2B models, embeddings, Whisper small |
| RZ-5 | 4 | 8 GB | 70 GB | $20 | up to 4B comfortably, 8–9B slowly |
| RZ-6 | 8 | 12 GB | 100 GB | $40 | 8–9B models with decent context |
| RZ-7 | 12 | 16 GB | 150 GB | $55 | 8–14B models, gpt-oss-20b just fits |
| Custom | up to 16 | up to 32 GB | up to 1 TB | calculator | gpt-oss-20b and MoE models up to ~30B in Q4 |
Availability changes from time to time, so check current plans on the VPS page. A useful detail: the panel lets you set the CPU mode to host-passthrough so the VM sees all CPU instructions — llama.cpp benefits noticeably.
What speed to expect
A rough guide for 4–8 Ryzen vCPUs with Q4 quantization. These are estimates based on memory bandwidth and user reports; real numbers depend on the CPU generation, memory speed and noisy neighbors.
| Model | Q4 file | RAM with headroom | CPU generation speed |
|---|---|---|---|
| Qwen3.5 0.8B | 1 GB | 2 GB | 20–40 tokens/s |
| Qwen3.5 2B | 2.7 GB | 4 GB | 10–20 tokens/s |
| Qwen3.5 4B | 3.4 GB | 6 GB | 6–12 tokens/s |
| Qwen3.5 9B, Llama 3.1 8B | 5–6.6 GB | 8–10 GB | 3–6 tokens/s |
| gpt-oss-20b (MoE) | 14 GB | 16–20 GB | 5–10 tokens/s on 8+ vCPU |
Five tokens per second is roughly slow reading speed. Fine for background jobs, sluggish for live chat. Long prompts are a separate pain on CPUs: feed the model a 5–10k-token document and the first word of the answer may take a minute or two.
What makes sense on a CPU VPS
- A Telegram or Matrix bot with a small model for template answers and message classification.
- Background summarization, translation and tagging when you do not need instant answers.
- Embeddings and search over your own notes (RAG): embedding models are small and fast even on CPUs.
- Audio transcription with whisper.cpp — small and medium models run acceptably on a CPU.
- A test bench: try the API, prompts and integrations before paying for a GPU.
What not to do on a CPU VPS. Dense 30B+ models will crawl at 1–2 tokens per second. Image generation technically runs, but one SDXL image takes minutes and FLUX even longer. Video on a CPU is practically unusable: a single short clip can take hours.
Quick start on a VPS
Installing Ollama and the first model takes two commands:
curl -fsSL https://ollama.com/install.sh | sh
ollama run qwen3.5:4b
By default Ollama listens on 127.0.0.1:11434 only and has no password. Keep it that way: an API exposed to the internet becomes a free server for anyone. Connect from your PC through an SSH tunnel and harden the server with our VPS security checklist.
ssh -N -L 11434:127.0.0.1:11434 user@your-vps
# http://localhost:11434 on your PC now reaches Ollama on the server
GPU servers for serious workloads
If you want a 27–35B model for coding, image generation or video, you need a GPU. The key spec is video memory (VRAM): the model has to fit entirely, otherwise speed drops several times.
| VRAM | Typical cards | What fits |
|---|---|---|
| 16 GB | RTX 4060 Ti 16 GB, RTX 5070 Ti, RTX 5080 | LLMs up to ~24B in Q4, SDXL and FLUX in GGUF, Wan 2.2 5B |
| 24 GB | RTX 3090, RTX 4090 | Qwen3.6 27B and Qwen3-Coder 30B, FLUX dev in FP8, Wan 2.2 14B in FP8 |
| 32 GB | RTX 5090 | the same with room for context, LTX-2 in FP8 comfortably |
| 48 GB | RTX 6000 Ada, A6000 | 70B models in Q4, long context, image LoRA training |
| 80–96 GB | A100 80 GB, H100, RTX PRO 6000 | gpt-oss-120b, 70B in Q8, video LoRA training |
Hourly GPU rental and dedicated GPU servers
For experiments, hourly rental is the cheapest option: start a server, work, delete it, and pay only for the time used. Prices below are a guide as of this writing:
- RunPod — on-demand GPU pods. An RTX 5090 costs about $0.69/hour on Community Cloud or $0.99/hour on Secure Cloud; an RTX PRO 6000 with 96 GB is about $1.69–2.09/hour.
- Vast.ai — a marketplace of GPUs from many hosts. Often the cheapest, but reliability and network speed vary from host to host.
- Hetzner GPU servers — dedicated machines in Germany and Finland. GEX44 with an RTX 4000 SFF Ada (20 GB) is about €234/month; GEX131 with an RTX PRO 6000 Blackwell Max-Q (96 GB) is about €1,197/month plus setup.
A simple rule: if you need a GPU for fewer than 15–20 hours a week, renting is almost always cheaper than buying. If the model must run 24/7 (a bot, a team service), compare a monthly dedicated GPU server with your own hardware.
Prepare a plan before renting: which image (Ubuntu with CUDA), which models to download, where to save results. Otherwise half of the paid hour goes to driver setup and downloading tens of gigabytes of weights.
Private options that accept crypto
If you prefer not to tie GPU rental to a card and an ID, several providers take cryptocurrency and only need an email or a wallet. Most of them are decentralized marketplaces where GPUs come from independent owners and small data centers. Marketplace prices move hourly, so treat these as a guide:
- Clore.ai — a GPU marketplace with per-minute billing. Email sign-up, payment in BTC or the CLORE token. An RTX 4090 is roughly $0.20–0.50/hour; spot is cheaper but another renter can outbid you.
- Akash Network — a decentralized cloud. With Console Air and your own Keplr wallet you pay in AKT or USDC with no account on the site. An RTX 4090 is about $0.35–0.50/hour, an H100 about $2.60.
- io.net — a GPU network on Solana. Pay in USDC (2% fee) or the IO token; sign in with email or GitHub. An RTX 4090 is about $0.25–0.37/hour.
- Vast.ai and RunPod from the list above also take crypto, but through payment processors: Vast.ai uses BitPay and Crypto.com and does not refund crypto-bought credits, and RunPod asks for KYC before your first crypto payment.
How private is it really? Blockchain payments are pseudonymous, not anonymous: every transaction is public and can often be traced back to the exchange where you bought the coins. And on a marketplace your workload runs on someone else’s machine, whose owner can technically access the disk and memory. Keep personal data, API keys and confidential documents off rented hosts — for sensitive work, your own mini PC (below) is the safer choice. Crypto payments may also have tax implications depending on where you live.
Mini PCs for local models at home
A home server for local AI models pays off if you use it every day. But “it runs the model” and “it is comfortable to work with” are different things. Three specs decide the experience, and mini PCs differ a lot on each:
- Memory size decides whether the model fits entirely. If it doesn’t, speed drops several times.
- Memory bandwidth decides generation speed — how many tokens per second you see on screen.
- GPU compute decides how fast the model reads your request before it starts answering, and how fast images render.
The third point is often overlooked, yet for vibe coding it matters most. An agent such as Cline or Aider sends the model 10–30k tokens on every step: system instructions, open files, edit history. If the GPU processes prompts at 100 tokens per second, every agent step starts with a 2–5 minute pause. Image generation is compute-bound too, not memory-bound.
What counts as comfortable
- Vibe coding: a Qwen3-Coder 30B or Qwen3.6 35B-A3B class model or better, generation of 30+ tokens per second and prompt processing of 500–1,000+ tokens per second, so a 20k-token context is read in 20–40 seconds rather than minutes.
- Images: SDXL at 1024×1024 in under 10–20 seconds per image, FLUX in under 30–60 seconds. Slower works, but iterating on prompts gets tedious.
By these criteria mini PCs fall into two groups: machines that are already good for daily coding and image work, and budget options for chat, autocomplete and getting started.
Mini PCs that already handle vibe coding and images
| Option | Price | Vibe coding | Images | Video |
|---|---|---|---|---|
| Mini PC + RTX 3090 / 5060 Ti over OCuLink | about $1,100–2,000 | excellent for models up to 30B | best value | yes, Wan 2.2 and LTX in FP8 |
| Ryzen AI Max+ 395, 64–128 GB | $2,000–2,200 for 128 GB | excellent, including 80–120B models | acceptable: SDXL ~20 s, FLUX ~1–1.5 min | slow |
| Mac mini M5 Pro, 48–64 GB | from $1,699 | good for models up to 35B | moderate | slow |
| ASUS ROG NUC 16 with RTX 5080 Laptop | from $3,449 | good for models up to 20–30B | fast | yes, within 16 GB |
| NVIDIA DGX Spark | $4–5k | excellent, models up to 120B | acceptable: SDXL ~12 s, FLUX ~35 s | yes, but not fast |
Mini PC with an external GPU over OCuLink
The best-value way to get both fast vibe coding and fast image generation at home. Many mini PCs with a Ryzen 7 8845HS / H 255 or Ryzen AI 9 HX 370 have an OCuLink port (PCIe 4.0 x4). It connects to a dock with its own power supply and a regular desktop graphics card. By day it is a quiet mini PC; for heavy work a full RTX card kicks in. Four PCIe lanes are enough for inference: the model loads into VRAM once and then runs at almost full card speed.
A rough budget: a mini PC with OCuLink and 32 GB, a dock such as the AOOSTAR AG02 with a built-in PSU, and a card — a new RTX 5060 Ti 16 GB or a used RTX 3090 24 GB. Used 3090 prices roughly doubled in 2026 precisely because of demand for local AI, so compare offers carefully.
The RTX 3090 combo is the most versatile: 24 GB fits Qwen3-Coder 30B or Qwen3.6 35B-A3B entirely in a 4-bit quant. Generation runs above 100 tokens per second and long contexts are processed in seconds. SDXL renders an image in about 5–6 seconds, FLUX in FP8 in 20–30 seconds, with room for Wan 2.2 and LTX video. The RTX 5060 Ti 16 GB is cheaper and quieter. It fits gpt-oss-20b entirely, while 30B models run with part of the experts offloaded to system RAM (the --n-cpu-moe option in llama.cpp) — slower, but still comfortable. For images 16 GB covers SDXL and FLUX in FP8 or GGUF.
Downsides: two boxes on the desk instead of one, the dock gets loud under load, and OCuLink is not hot-pluggable — connect the card with the PC turned off.
AMD Ryzen AI Max+ 395 (Strix Halo) — all in one box
GMKtec EVO-X2, Framework Desktop, Beelink GTR9 Pro and similar machines. The key feature is up to 128 GB of fast LPDDR5X (~256 GB/s), with 96 GB assignable to the Radeon 8060S iGPU, and more on Linux. It is the best compact option for vibe coding: in llama.cpp Vulkan benchmarks Qwen3-Coder 30B generates 70–100 tokens per second and processes prompts at about 1,400 tokens per second, with almost no slowdown up to 8k tokens of context. Qwen3-Coder-Next 80B does about 60 tokens per second, gpt-oss-120b about 50. Dense 70B models manage only about 5 tokens per second.
Strix Halo handles images, but without records. According to AMD, on ROCm 7.2 SDXL at 1024×1024 takes about 18–25 seconds and full-precision FLUX.1 dev about 80 seconds. On the other hand, 96 GB fits models and pipelines that won’t squeeze into a 16–24 GB card — for example full-precision FLUX with ControlNet, or Wan 2.2 14B for video, though clips take a long time. Some new ComfyUI optimizations land on NVIDIA first.
The 128 GB version costs roughly $2,000–2,200 in official stores. 64 GB is enough for vibe coding with 30–35B models plus images; 80–120B models need the 128 GB version. Memory is soldered, so you can’t add more later.
Mac mini with M5 Pro
In August 2026 Apple refreshed the Mac mini: an M6 model (up to 32 GB unified memory, 170 GB/s, from $899) and an M5 Pro model (24, 48 or 64 GB, 307 GB/s, from $1,699). The big change for AI is a neural accelerator in every GPU core: long prompts are processed 3–4 times faster than on the M4 Pro, which is exactly what you feel when working with an agent. Ollama and LM Studio support MLX, and 48–64 GB comfortably runs Qwen3.6 27B/35B and Devstral Small 2. The limit is the 64 GB ceiling: models like gpt-oss-120b no longer fit.
Images work through ComfyUI or Draw Things. Early tests show the M6 rendering a FLUX image in about 35 seconds versus a minute and a half on the M4; the M5 Pro is faster thanks to more GPU cores, but still far from an RTX card. A base M6 with 16–24 GB is tight for vibe coding: it fits models up to 14B or gpt-oss-20b.
Compact PCs with a built-in RTX
If you don’t want an external dock, some mini PCs have a discrete GPU inside. The most powerful is the roughly 3-liter ASUS ROG NUC 16: Core Ultra 9 290HX Plus and up to an RTX 5080 Laptop GPU with 16 GB GDDR7. For images it is fast — FLUX in FP8 fits entirely — and for code 16 GB covers gpt-oss-20b and 30B models with partial offload. The downside is price: from $3,449 for the RTX 5080 version. A mobile RTX 5080 is also noticeably weaker than the desktop card, so on price-to-speed a mini PC with OCuLink and a desktop GPU wins.
NVIDIA DGX Spark
DGX Spark is a desktop “supercomputer” built on the GB10 chip with 128 GB unified memory and the full CUDA stack. Generation speed is close to Strix Halo, but it processes long prompts noticeably faster: about 1,700 tokens per second of prompt processing and 55 of generation on gpt-oss-120b. It renders images roughly 2.5 times faster than Strix Halo: SDXL in about 12 seconds, FLUX.1 dev in about 35 once warmed up. All new ComfyUI optimizations and fine-tuning tools work out of the box. At roughly $4–5k it is a choice for developers who specifically need CUDA and 128 GB of memory in one box.
Mid-range: Ryzen 7 and Ryzen AI 9 mini PCs with 32 GB
The most common mini PCs of 2026 use the eight-core Ryzen 7 H 255, 8845HS or 7840HS with Radeon 780M graphics — essentially the same chip under different names. One step up are the Ryzen AI 9 365, HX 370 and HX 470 with Radeon 880M/890M, as in the Minisforum AI X1 Pro or GMKtec EVO-X1. The iGPU uses regular system memory, so capacity and dual channel matter most: pick a model with two SO-DIMM slots and install at least 32 GB. RAM prices are high in 2026, so a ready-made 32 GB configuration is often cheaper than a barebone plus separately bought memory. The 50+ TOPS NPU is barely used for LLMs yet — it is a marketing number, not generation speed.
What you get. With LM Studio or llama.cpp on the Vulkan backend these machines run 7–14B models and MoE models like gpt-oss-20b and Qwen3-Coder 30B, generating 20–30 tokens per second on short prompts. That is enough for chat, code questions and editor autocomplete. But prompt processing on a Radeon 780M is about 400 tokens per second on an empty context and 60–130 at 32k tokens. So agentic vibe coding with long contexts means multi-minute pauses on every step. SDXL works but is several times slower than on an RTX card; FLUX and video are for experiments.
The main advantage of this class is the upgrade path: many models have OCuLink, so the mini PC can later become the external-GPU setup described above. Examples: Beelink SER8, GMKtec K8 Plus and AOOSTAR GEM12 Max on the 8845HS, or GMKtec EVO-X1 on the HX 370.
Budget options
Good for trying local models and running a small home assistant on 4–9B models such as Qwen3.5 4B/9B or the compact Gemma 4 variants. Not for vibe coding or images, but fine for chat, translation and working with text.
- A Ryzen 5 7640HS mini PC with 24 GB (e.g. Firebat F1) — the best pick at this level: 24 GB fits even gpt-oss-20b, and the Radeon 760M works via Vulkan.
- Intel N150 boxes with 16 GB — only for 1–4B models and light server duty. If the budget allows, pay a bit more for a Ryzen.
What to avoid
- 8–12 GB configurations and soldered 16 GB memory — it can’t be upgraded, and for models it is the key resource.
- Suspiciously cheap listings promising huge RAM and SSD for the price of an N150 — reviews mention old hardware arriving instead.
- Buying “for AI” because of the NPU: it barely affects local LLM or image generation speed yet.
- Jetson boards as a home server: the Orin Nano Super with 8 GB only runs small models, and Jetson AGX Thor is built for robotics and costs more than DGX Spark.
What to choose
- Vibe coding and images at a sensible price — a mini PC with OCuLink and 32 GB plus an RTX 3090 or RTX 5060 Ti 16 GB in a dock.
- Vibe coding with large models in one quiet box — Ryzen AI Max+ 395 with 64–128 GB. Images work too, just slower than on an RTX.
- Silence and the Apple ecosystem — Mac mini M5 Pro with 48–64 GB.
- CUDA, fine-tuning and models up to 120B — DGX Spark.
- Chat and code autocomplete — a Ryzen 7 H 255 or HX 370 mini PC with 32 GB, with the option to add a GPU over OCuLink later.
- Try it with zero spend — your current computer with 16 GB RAM and 4–9B models.
A mini PC is easy to keep running 24/7 as a home server — for example next to Proxmox on a mini PC. To reach it from outside without open ports, use OpenZiti.
Local models for image generation
| Model | Strengths | VRAM | License |
|---|---|---|---|
| SDXL | huge library of LoRAs and styles, runs almost anywhere | 6–8 GB | open (OpenRAIL++) |
| FLUX.1 [schnell] / [dev] | precise prompt following, photorealism | ~7 GB in GGUF Q4, ~12 GB in FP8, 24 GB uncompressed | schnell — Apache 2.0, dev — non-commercial |
| FLUX.2 [klein] 4B | new in 2026: fast, 4 steps, good quality | from 8 GB in GGUF | Apache 2.0 |
| FLUX.2 [dev] | top quality and editing | 24 GB+ with quantization | non-commercial |
| Z-Image Turbo | fast photorealism in 8 steps | up to 16 GB | Apache 2.0 |
| Qwen-Image / Qwen-Image-2.1 | readable text in images, editing | 12–16 GB quantized | Apache 2.0 |
Qwen-Image-2.1 came out on September 20, 2026 with day-one ComfyUI support. If you are just starting, pick SDXL or FLUX.2 [klein]: they suit 8–12 GB cards best.
Interfaces
- ComfyUI — a node-based workflow editor. Harder at first, but new models (video included) land here first, and you can drag ready-made workflows into the window.
- Forge — the familiar A1111-style interface, but faster and lighter on memory. Supports FLUX.
- AUTOMATIC1111 — the classic with tons of extensions, but slower development and higher VRAM use.
- Fooocus — the simplest “type a prompt, get an image” option. SDXL-based models only, no newer architectures.
Local models for video generation
Video is the most demanding task. Even on an RTX 4090 a 5-second 720p clip takes from a minute to ten minutes, and you need 32–64 GB of system RAM for offloading parts of the model.
| Model | Highlights | Minimum VRAM | Comfortable |
|---|---|---|---|
| Wan 2.2 TI2V-5B | text and image to video, 720p, Apache 2.0 | ~8 GB with offloading and GGUF | 12–16 GB |
| Wan 2.2 14B (T2V, I2V) | the most realistic motion among open models | 12–16 GB in FP8/GGUF | 24 GB |
| LTX-2 (versions 2.3 and 2.5) | the fastest, generates video with audio | ~10–16 GB in GGUF/FP8 | 24–32 GB |
| HunyuanVideo 1.5 | 8.3B, cinematic motion | ~14 GB with offloading | 24 GB |
Newer Wan versions (2.5, 2.6) are mostly available through the developer’s cloud; Wan 2.2 has open downloadable weights. For LTX-2.5 the developers officially list 32 GB VRAM as the minimum. HunyuanVideo uses Tencent’s own license with restrictions — read it if you plan commercial use. All three run in ComfyUI.
If you need video a couple of times a month, renting an RTX 4090 or A100 for a few hours is more sensible than buying an expensive card.
Models for vibe coding
Vibe coding means describing a task in plain words while an AI agent reads the project, edits files and runs commands. A local model for this has to handle tool calls and long context. Current options as of fall 2026 (sizes from the Ollama library):
| Model | Type | Q4 size | Where to run |
|---|---|---|---|
| Qwen3.6 27B / 35B | dense / MoE (3B active) | 18–24 GB | 24 GB GPU, 32–48 GB Mac |
| Qwen3-Coder 30B | MoE, 3.3B active | 19 GB | 24 GB GPU, 32 GB Mac |
| Devstral Small 2 (24B) | dense, Apache 2.0 | 15 GB | 16–24 GB GPU |
| gpt-oss-20b | MoE, ~3.6B active | 14 GB | 16 GB memory, even without a GPU |
| GLM-4.7-Flash | 30B-class MoE | 19 GB | 24 GB GPU |
| Qwen3-Coder-Next | MoE 80B, 3B active | 52 GB | 64–128 GB unified memory |
| gpt-oss-120b | MoE, ~5B active | 65 GB | Strix Halo, DGX Spark, 80 GB GPU |
| Devstral 2 (123B) | dense | ~75 GB | 128 GB unified memory, slow |
| DeepSeek-V4-Flash | MoE 284B, 13B active | ~103 GB in 3-bit | 128 GB+, easier via API |
For editor autocomplete use a separate small and fast model such as qwen2.5-coder:1.5b: it suggests a line in a fraction of a second even on a laptop.
The stack: engine plus editor
- Engine: Ollama (simplest), LM Studio (friendly GUI) or
llama-serverfrom llama.cpp (most knobs). All provide an API at something likehttp://localhost:11434or:1234. - Continue — an extension for VS Code and JetBrains: chat, edits and autocomplete with a local model.
- Cline — an agent in VS Code: reads files, proposes changes and runs commands after your approval.
- Aider — a terminal agent that works through git and commits every change.
- Cursor accepts a custom OpenAI-compatible base URL, but requests go through Cursor’s servers, so it cannot see
localhost. You need an externally reachable endpoint with authentication, and the setup is no longer fully local.
Example Continue config (~/.continue/config.yaml):
name: Local
version: 0.0.1
schema: v1
models:
- name: Qwen3.6 27B
provider: ollama
model: qwen3.6:27b
roles: [chat, edit, apply]
- name: Autocomplete
provider: ollama
model: qwen2.5-coder:1.5b
roles: [autocomplete]
Aider with the same model is a single command:
aider --model ollama_chat/qwen3.6:27b
An important setting: agents need long context, and Ollama’s default is modest. Raise it, for example with OLLAMA_CONTEXT_LENGTH=32768, or the model will “forget” the start of a file and get confused. And be realistic: local models on 24 GB handle scripts, small projects and routine work well, but fall behind top cloud models in large codebases.
Summary table: task → model → minimum hardware

| Task | Models | Minimum | Comfortable |
|---|---|---|---|
| Chat, bots, writing | Qwen3.5 4B–9B, Gemma 4 E4B, Llama 3.1 8B | CPU VPS 4 vCPU / 8 GB RAM | any 8 GB GPU or 16 GB Mac |
| Document search (RAG) | embedding model + 4–9B LLM | 4 vCPU / 8 GB RAM | 8–12 GB GPU |
| Audio transcription | Whisper (whisper.cpp) | 2–4 vCPU | any GPU |
| Vibe coding | Qwen3.6 27B, Qwen3-Coder 30B, Devstral Small 2, gpt-oss-20b | 16 GB GPU or 32 GB unified memory | 24–32 GB GPU, 48–64 GB Mac |
| Large models | gpt-oss-120b, Qwen3-Coder-Next, Devstral 2 | 64 GB unified memory | Ryzen AI Max+ 395 / DGX Spark with 128 GB |
| Images | SDXL, FLUX.2 [klein], Z-Image Turbo, Qwen-Image | NVIDIA 8 GB | NVIDIA 16–24 GB |
| Video | Wan 2.2, LTX-2, HunyuanVideo 1.5 | NVIDIA 12–16 GB + 32 GB RAM | NVIDIA 24–32 GB + 64 GB RAM |
| LoRA training | SDXL, FLUX, Wan | 24 GB GPU | rented A100/H100 80 GB |
Common mistakes
- Looking at parameter count instead of memory. A model that does not fit in VRAM spills into system RAM and slows down 5–10 times.
- Buying a CPU VPS for images and video. CPUs are not suitable for diffusion models — go straight for a GPU.
- Exposing Ollama or ComfyUI to the internet without a password. Use an SSH tunnel, a VPN or a reverse proxy with authentication.
- Forgetting about context. Long context eats gigabytes of memory on top of the model size.
- Buying hardware before experimenting. Try the model on an hourly rental or your current PC first — you will see how much memory you actually need.
- Ignoring licenses. FLUX.1 [dev] and FLUX.2 [dev] forbid commercial use without a separate license, and HunyuanVideo has its own restrictions.
Where to start today
- Install Ollama or LM Studio on your computer and run Qwen3.5 4B, or gpt-oss-20b if you have 16 GB of memory.
- For a 24/7 bot, take a Ryzen VPS with 4–8 vCPU and 8–16 GB RAM, for example at AlfaHost, and keep the API private.
- For images and video, rent an RTX 4090 for a couple of hours and try workflows in ComfyUI.
- For code, connect Continue or Cline to a local model and compare it with a cloud assistant on your own tasks.
- Only then decide whether you need a home mini PC and how much memory it should have.
Useful links
- Ollama model library · LM Studio · llama.cpp
- What is Hugging Face · Pirate Face: models via torrent
- ComfyUI · Forge
- Wan 2.2 · Qwen-Image-2.1
- Continue · Cline · Aider
- Ryzen VPS at AlfaHost · RunPod GPU pricing
FAQ
Can I run an AI model on a VPS without a GPU?
Yes, small 1–9B language models via Ollama or llama.cpp. On 4–8 Ryzen vCPUs they produce roughly 3–20 tokens per second depending on size. Models above 20B, images and video need a GPU.
How much memory does a local model need?
With Q4 quantization, about 0.6 GB per billion parameters plus 1–4 GB for context. An 8B model takes 5–6 GB, 30B takes 18–20 GB, 120B around 65 GB.
Can I generate images on a CPU?
Technically yes, but one SDXL image takes minutes, and video on a CPU is practically unusable. Images need an NVIDIA GPU with 8 GB or more, video 12–16 GB or more.
Mac mini or a Ryzen AI Max mini PC for home?
A Mac mini M5 Pro with up to 64 GB is a quiet option for models up to 35B. Ryzen AI Max+ 395 with 128 GB is cheaper per gigabyte and runs 120B-class models. For images and video both trail an RTX card: if you need both code and images, a mini PC with OCuLink and an RTX 3090 or RTX 5060 Ti 16 GB in an external dock is better value.
Which local model is best for coding?
As of fall 2026: with 24 GB VRAM, Qwen3.6 27B or Qwen3-Coder 30B; with 16 GB, Devstral Small 2 or gpt-oss-20b; with 128 GB unified memory, Qwen3-Coder-Next or gpt-oss-120b.
Can I use a local model in Cursor?
Partly: Cursor accepts a custom OpenAI-compatible URL, but requests go through its servers, so localhost will not work. For a fully local setup, Continue, Cline or Aider are easier.
Is it safe to expose Ollama to the internet?
No. The API has no password by default, so anyone could use your server. Connect through an SSH tunnel or VPN, and if you need external access, put a reverse proxy with authentication in front.
Is it cheaper to rent a GPU or buy one?
If you need a GPU for fewer than 15–20 hours a week, renting is almost always cheaper. For 24/7 workloads a dedicated server or your own hardware wins.
Related reading
- What is Hugging Face: where to find models
- Pirate Face: torrent mirror for open AI models
- VPS security on Ubuntu and Debian
- OpenZiti: reach your home server without open ports
- Proxmox on a mini PC
- Services: self-hosted setup on a VPS
In short: local AI models for bots and background jobs run fine on a regular Ryzen VPS, as long as they are small language models. Coding with 27–35B models needs a 24 GB GPU or a 48–64 GB Mac, large MoE models need 128 GB of unified memory, and images and video need an NVIDIA card — your own or rented. Ask about your own setup in the comments.
No email, no trackers — just the update feed.








