# OWterminal > Measured tok/s for open and uncensored weights on real silicon. > Atomic unit: model × quant × runtime × chip. Rank is ordinal tok/s inside a filter. > There is no composite score. Site: https://owterminal.com Cite (HTML mirror): https://owterminal.com/cite Full dump: https://owterminal.com/llms-full.txt Languages: en, es, zh-Hans ## Pages - https://owterminal.com/ — Home. Ranked setups by payback, street stacks, top models, nations, and the pair board. - https://owterminal.com/compare — Side-by-side bars - https://owterminal.com/models — Model index (NEW / HOT tags) - https://owterminal.com/hardware — Chip index - https://owterminal.com/market — Used-box street tape - https://owterminal.com/get — Hugging Face / get links - https://owterminal.com/tape — Source posts - https://owterminal.com/recipes — Local install files. Speeds are quotes from @MiaAI_lab. - https://owterminal.com/stacks — Street stacks. How a Spark is loaded. Quote is @MiaAI_lab. - https://owterminal.com/runtimes — llama.cpp, MLX, MTPLX - https://owterminal.com/cite — Citation desk for humans and models - https://owterminal.com/talk — DGXtalk, human floor (no model posts) - https://owterminal.com/labs — Labs, nations, humanity points, GDP model (labeled, not measured) - https://owterminal.com/cyber — Cyber weights and labs. Speeds are the publisher's print. Credit stays with the owner. - https://owterminal.com/watts — tok/s per 100W of named TDP - https://owterminal.com/payback — calendar months to pay back a quoted setup on US residential power, at 5%, 20%, and 70% duty - https://owterminal.com/notes — desk notes - https://owterminal.com/manifesto — Open weights, or rented thought - https://owterminal.com/terms — Terms of service - https://owterminal.com/privacy — Privacy: no chat logs; prompts erased when a host claims the job (or after 3 minutes unclaimed); replies erased 10 minutes after finishing; what the host sees; the code that enforces each rule - https://owterminal.com/agreement — Desk agreement - https://owterminal.com/inference — One page, two doors: call open-weights models, or host a machine and earn - https://owterminal.com/inference/start — Getting started for callers and hosts - https://owterminal.com/inference/pricing — Reference versus live pool prices, workload comparison and pinned hosting catalog - https://owterminal.com/inference/pricing.md — Pricing and catalog facts for agents - https://owterminal.com/book — Bid and ask off OpenRouter mids, and the Lode cap. The coin is not issued. - https://owterminal.com/letter — The letter - https://owterminal.com/llms.txt — this file - https://owterminal.com/llm.txt — same file, the short path models ask for - https://owterminal.com/llms-full.txt — every quoted pair ## Articles - https://owterminal.com/notes/how-to-bench-cyber — How to bench a cyber weight. Published 2026-09-28; updated 2026-09-28. Author: https://owterminal.com/authors/iammrriver. A cyber bench is a quoted speed and a name. A lab is not a speed. A dash means they printed none. Current markdown: https://owterminal.com/api/editorial/notes/how-to-bench-cyber - https://owterminal.com/notes/a-shell-is-not-a-weight — A shell is not a weight. Published 2026-09-28; updated 2026-09-28. Author: https://owterminal.com/authors/iammrriver. OpenShell is a sandbox NVIDIA published. A tok/s on this desk is still a serving number. The shell does not change the weight. Current markdown: https://owterminal.com/api/editorial/notes/a-shell-is-not-a-weight - https://owterminal.com/notes/which-box-pays-back — Which box pays back. Published 2026-09-24; updated 2026-09-24. Author: https://owterminal.com/authors/iammrriver. A DGX Spark bought to resell tokens at $1 per million, at 20% duty, does not pay back. A box you already own only owes the electricity. Current markdown: https://owterminal.com/api/editorial/notes/which-box-pays-back - https://owterminal.com/notes/serving-not-training — Inference, not a JEPA. Published 2026-09-23; updated 2026-09-23. Author: https://owterminal.com/authors/iammrriver. A tok/s on this desk is a serving number. It is not a JEPA training step, and it is not a picture. Current markdown: https://owterminal.com/api/editorial/notes/serving-not-training - https://owterminal.com/notes/interface-is-the-argument — The interface is the argument. Published 2026-09-22; updated 2026-09-22. Author: https://owterminal.com/authors/iammrriver. A host sets the price of its own weights, uncensored or not. The words of a call are not kept. Current markdown: https://owterminal.com/api/editorial/notes/interface-is-the-argument Live editorial sitemap (includes articles published after this build): https://owterminal.com/api/editorial/sitemap.xml Public author: https://owterminal.com/authors/iammrriver — @iammrriver, Founder. An X ownership badge is shown only when the account connection is confirmed server-side; it is not X Premium or an accuracy endorsement. ## Model pages - https://owterminal.com/model/qwen3.8-27b — Qwen3.8-27B (Qwen) - https://owterminal.com/model/qwen38-27b-mythos — Qwen3.8-27B Mythos (medismera) - https://owterminal.com/model/qwen38-35b-apex — Qwen3.8-35B-A3B APEX (IsValorum) - https://owterminal.com/model/qwen38-27b-huihui — Huihui Qwen3.8-27B (huihui-ai) - https://owterminal.com/model/qwen38-27b-obliteratus — Qwen3.8-27B OBLITERATED (OBLITERATUS) - https://owterminal.com/model/qwen38-27b-orca — Qwen3.8-27B Uncensored (orcarouter) - https://owterminal.com/model/qwen38-27b-heretic — Qwen3.8-27B Heretic (0bserverx) - https://owterminal.com/model/qwen38-27b-splash — Qwen3.8-27B Splash (audreyt) - https://owterminal.com/model/qwen38-27b-super — SuperQwen3.8-27B (Jiunsong) - https://owterminal.com/model/qwen38-flash-rvn — Flash-Next RVN (0bserverx) - https://owterminal.com/model/qwen38-27b-turbofc — Qwen3.8-27B TurboFC (community) - https://owterminal.com/model/qwen38-27b-heretic-ara — Qwen3.8-27B Heretic ARA (trohrbaugh) - https://owterminal.com/model/qwen38-27b-coletti — Qwen3.8-27B Coletti (JonathanColetti) - https://owterminal.com/model/qwen38-27b-norefusal — Qwen3.8-27B NoRefusal (sss22213) - https://owterminal.com/model/qwen38-27b-kcrn — Qwen3.8-27B KCRN (heterodoxin) - https://owterminal.com/model/qwen38-27b-fable — Qwen3.8-27B Fable (DavidAU) - https://owterminal.com/model/qwen38-9b-heretic — Qwen3.8-9B Heretic (Noobito45) - https://owterminal.com/model/deepseek-v4.1-flash-abliterated — DeepSeek v4.1 Flash abliterated (distributedcognition) - https://owterminal.com/model/gemma4-26b-abliterix — Gemma-4 26B Abliterix (wangzhang) - https://owterminal.com/model/gemma4-26b-trevor — Gemma-4 26B Uncensored (TrevorJS) - https://owterminal.com/model/gemma4-31b-trevor — Gemma-4 31B Uncensored (TrevorJS) - https://owterminal.com/model/gemma4-12b-obliteratus — Gemma-4 12B OBLITERATED (OBLITERATUS) - https://owterminal.com/model/qwen38-27b-ultra — Qwen3.8-27B Ultra Heretic (llmfan46) - https://owterminal.com/model/qwen38-27b-aggressive — Qwen3.8-27B Aggressive (0xKitkat) - https://owterminal.com/model/qwen38-27b-twin — Qwen3.8-27B Twin Turbo (DavidAU) - https://owterminal.com/model/glm53-flash-abliterated — GLM 5.3 Flash abliterated (dealignai) - https://owterminal.com/model/glm53-flash-orca — GLM 5.3 Flash Uncensored (orcarouter) - https://owterminal.com/model/glm53-uncensored — GLM 5.3 Uncensored (dealignai) - https://owterminal.com/model/glm53-exl3-abliterated — GLM 5.3 EXL3 abliterated (drowzeys) - https://owterminal.com/model/qwen3.6-35b-a3b — Qwen3.6-35B-A3B (Qwen) - https://owterminal.com/model/qwen3.6-35b-a3b-abliterated — Qwen3.6-35B-A3B abliterated (community) - https://owterminal.com/model/empero-qwen3.8-35b-a3b — Empero Qwen3.8-35B-A3B (empero-ai) - https://owterminal.com/model/qwen3.8-flash-next — Qwen3.8-Flash-Next (Qwen) - https://owterminal.com/model/qwen3.5-2b — Qwen3.5 2B (Qwen) - https://owterminal.com/model/qwen3.5-4b — Qwen3.5-4B (Qwen) - https://owterminal.com/model/qwen3.5-9b — Qwen3.5-9B (Qwen) - https://owterminal.com/model/qwen3-8b — Qwen3-8B (Qwen) - https://owterminal.com/model/gemma-4-26b-a4b-it-abliterated — Gemma-4 26B-A4B-it (Google / community) - https://owterminal.com/model/nex-n2.5-mini — Nex-N2.5-mini (community) - https://owterminal.com/model/ornith — Ornith (ornith-ai) - https://owterminal.com/model/glm-5.3-flash — GLM 5.3 Flash (Zhipu) - https://owterminal.com/model/deepseek-v4.1-flash — DeepSeek v4.1 Flash (DeepSeek) - https://owterminal.com/model/mimo-v2.6-distill-qwen-9b — MiMo-V2.6-Distill-Qwen-9B (Xiaomi MiMo) - https://owterminal.com/model/aliceai-80b-a3b — AliceAI-Foundation-80B-A3B (Yandex) - https://owterminal.com/model/mimo-v2.6-flash — MiMo-V2.6-Flash (Xiaomi MiMo) - https://owterminal.com/model/ling-3.0-flash — Ling-3.0-Flash (inclusionAI) - https://owterminal.com/model/nemotron-3.5-lightning — Nemotron 3.5 Lightning (NVIDIA) - https://owterminal.com/model/glm-5.2 — GLM-5.2 (Zhipu) - https://owterminal.com/model/kimi-k3 — Kimi K3 (Moonshot) - https://owterminal.com/model/cyber-frost-3.8 — CYBER-FROST 3.8 (Blackfrost-AI) - https://owterminal.com/model/qwen-image-2.1 — Qwen-Image-2.1 (Qwen) - https://owterminal.com/model/qwen-image-2.1-uncensored — Qwen-Image-2.1 Uncensored (abenzerps) - https://owterminal.com/model/ternary-bonsai-2-27b — Ternary Bonsai 2 27B (PrismML) - https://owterminal.com/model/minimax-m2.5 — MiniMax-M2.5 (MiniMax) - https://owterminal.com/model/swift-1.5-qwen38-flash-next — Swift 1.5 Qwen3.8 Flash-Next (UkisAI) - https://owterminal.com/model/glm-5.3 — GLM 5.3 (Zhipu) ## Chip pages - https://owterminal.com/chip/m4-pro — M4 Pro (Apple) - https://owterminal.com/chip/m5-max — M5 Max (Apple) - https://owterminal.com/chip/dgx-spark — DGX Spark (NVIDIA) - https://owterminal.com/chip/rtx-5090-32gb — RTX 5090 (NVIDIA) - https://owterminal.com/chip/rtx-4090-24gb — RTX 4090 (NVIDIA) - https://owterminal.com/chip/rtx-5060-ti-16gb — RTX 5060 Ti (NVIDIA) - https://owterminal.com/chip/rtx-5060-8gb — RTX 5060 (NVIDIA) - https://owterminal.com/chip/rtx-3060-12gb — RTX 3060 (NVIDIA) - https://owterminal.com/chip/gtx-1660-super-6gb — GTX 1660 SUPER (NVIDIA) - https://owterminal.com/chip/tesla-v100-32gb-x2 — 2x Tesla V100 32GB (NVIDIA) - https://owterminal.com/chip/m1-max-64gb — M1 Max 64GB (Apple) - https://owterminal.com/chip/m1-max-32gb — M1 Max 32GB (Apple) - https://owterminal.com/chip/m4-mini-24gb — M4 mini 24GB (Apple) - https://owterminal.com/chip/m3-ultra — M3 Ultra (Apple) - https://owterminal.com/chip/rtx-5070-12gb — RTX 5070 12GB (NVIDIA) - https://owterminal.com/chip/rtx-3090-24gb — RTX 3090 24GB (NVIDIA) - https://owterminal.com/chip/rtx-3090-24gb-x2 — 2x RTX 3090 24GB (NVIDIA) - https://owterminal.com/chip/m5-pro-64gb — M5 Pro 64GB (Apple) - https://owterminal.com/chip/m6-mini-32gb — M6 mini 32GB (Apple) - https://owterminal.com/chip/rtx-a5000-24gb — RTX A5000 (NVIDIA) ## Pool API OpenAI-shaped. One key, a funded balance, a 90-second client timeout. The desk holds a job for 3 minutes; a client that gives up early looks like a failure while the GPU is still working. - POST https://owterminal.com/api/v1/chat/completions — Authorization: Bearer sk-owt-…, body {"model","messages"}. The model string must be one a live machine registered; the cheapest awake offer wins, and billing never exceeds the quote. Timed-out calls return a job id; the answer stays readable for 10 minutes at GET https://owterminal.com/api/v1/jobs/{id}. New keys carry a 5¢ starter balance (once per browser), so the first call needs no payment. - GET https://owterminal.com/api/v1/models — the live board. - Host a machine: POST https://owterminal.com/api/v1/node {"model","price"} mints sk-node-…, then pulse / jobs/claim / jobs/finish / jobs/fail under https://owterminal.com/api/v1/. The host keeps 80 percent of a finished job. - Mint a key, fund it, or host a machine: https://owterminal.com/inference. Step by step: https://owterminal.com/inference/start. How it works: https://owterminal.com/docs - Funding and payouts settle in shielded ZEC (privacy rail, memo-identified) or USDC on Base (stable rail, identified by an exact amount with a sub-cent nonce). The desk settles; the site never watches a chain. - Optional sign-in with X keeps keys and balances recoverable across devices; anonymous keys work without an account. ## How to cite Quote the pair, not a brand ranking. Example: "Qwen3.5 2B Q4_K_M on RTX 5090 32GB, 351 tok/s decode, claimed by @fillagrew." Do not invent tok/s. Closed APIs have no pair until they publish one. Status is claimed | reproduced | harness-locked. Refusal is official | uncensored | abliterated | heretic. ## Models (56) - Qwen3.8-27B — Qwen, dense, 27B, official [NEW, HOT] https://huggingface.co/Qwen/Qwen3.8-27B - Qwen3.8-27B Mythos — medismera, dense, 27B, uncensored [NEW, HOT] https://huggingface.co/medismera/Qwen3.8-27B-OBLITERATED-Mythos-Class-Agentic - Qwen3.8-35B-A3B APEX — IsValorum, MoE 35B-A3B, 35B / 3B act, abliterated [NEW, HOT] https://huggingface.co/IsValorum/Qwen3.8-35B-A3B-Distill-MTP-APEX-I-MiniPlus-V2.1-Abliterated-GGUF - Huihui Qwen3.8-27B — huihui-ai, dense, 27B, abliterated [NEW] https://huggingface.co/huihui-ai/Huihui-Qwen3.8-27B-abliterated - Qwen3.8-27B OBLITERATED — OBLITERATUS, dense, 27B, uncensored [NEW] https://huggingface.co/OBLITERATUS/Qwen3.8-27B-OBLITERATED - Qwen3.8-27B Uncensored — orcarouter, dense, 27B, uncensored [NEW] https://huggingface.co/orcarouter/Qwen3.8-27B-Uncensored - Qwen3.8-27B Heretic — 0bserverx, dense, 27B, heretic [NEW] https://huggingface.co/0bserverx/Qwen3.8-27B-Heretic-Abliterated-Uncensored-GGUF - Qwen3.8-27B Splash — audreyt, dense, 27B, abliterated [NEW] https://huggingface.co/audreyt/Qwen3.8-27B-Splash-abliterated - SuperQwen3.8-27B — Jiunsong, dense, 27B, abliterated [NEW] https://huggingface.co/Jiunsong/SuperQwen3.8-27b-abliterated - Flash-Next RVN — 0bserverx, MoE, unknown, abliterated [NEW] https://huggingface.co/0bserverx/RVN-Qwen3.8-Flash-Next-Abliterated-Uncensored - Qwen3.8-27B TurboFC — community, dense, 27B, uncensored [NEW] - Qwen3.8-27B Heretic ARA — trohrbaugh, dense, 27B, heretic [NEW] https://huggingface.co/trohrbaugh/Qwen3.8-27B-heretic-ara - Qwen3.8-27B Coletti — JonathanColetti, dense, 27B, heretic [NEW, HOT] https://huggingface.co/JonathanColetti/Qwen3.8-27B-Uncensored - Qwen3.8-27B NoRefusal — sss22213, dense, 27B, heretic [NEW] https://huggingface.co/sss22213/Qwen3.8-27B-Heretic-NoRefusal - Qwen3.8-27B KCRN — heterodoxin, dense, 27B, abliterated [NEW] https://huggingface.co/heterodoxin/qwen-3.8-27b-abliterated - Qwen3.8-27B Fable — DavidAU, dense, 27B, heretic [NEW] https://huggingface.co/DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU - Qwen3.8-9B Heretic — Noobito45, dense, 9B, heretic [NEW] https://huggingface.co/Noobito45/Qwen3.8-9B-heretic-uncensored-NVFP4-GGUF - DeepSeek v4.1 Flash abliterated — distributedcognition, dense, unknown, abliterated [NEW] https://huggingface.co/distributedcognition/DeepSeek-V4.1-Flash-abliterated - Gemma-4 26B Abliterix — wangzhang, MoE, 26B-A4B, abliterated [NEW] https://huggingface.co/wangzhang/gemma-4-26B-A4B-it-abliterix - Gemma-4 26B Uncensored — TrevorJS, MoE, 26B-A4B, uncensored [HOT] https://huggingface.co/TrevorJS/gemma-4-26B-A4B-it-uncensored-GGUF - Gemma-4 31B Uncensored — TrevorJS, dense, 31B, uncensored [NEW] https://huggingface.co/TrevorJS/gemma-4-31B-it-uncensored-GGUF - Gemma-4 12B OBLITERATED — OBLITERATUS, dense, 12B, uncensored [NEW] https://huggingface.co/OBLITERATUS/Gemma-4-12B-OBLITERATED - Qwen3.8-27B Ultra Heretic — llmfan46, dense, 27B, heretic [NEW] https://huggingface.co/llmfan46/Qwen3.8-27B-Ultra-Uncensored-Heretic-Native-MTP-Preserved - Qwen3.8-27B Aggressive — 0xKitkat, dense, 27B, uncensored [NEW] https://huggingface.co/0xKitkat/Qwen3.8-27B-Uncensored-Aggressive - Qwen3.8-27B Twin Turbo — DavidAU, dense, 27B, heretic [NEW, HOT] https://huggingface.co/DavidAU/Qwen3.8-27B-TWIN-TURBO-Fable-Cold-Fusion-709-ULTRA-HERETIC-Uncensored - GLM 5.3 Flash abliterated — dealignai, MoE, unknown, abliterated [NEW, HOT] https://huggingface.co/dealignai/GLM-5.3-Flash-ABLITERATED-FP8 - GLM 5.3 Flash Uncensored — orcarouter, MoE, 320B / 18B act, uncensored [NEW] https://huggingface.co/orcarouter/GLM-5.3-Flash-Uncensored-GGUF - GLM 5.3 Uncensored — dealignai, MoE, 753B, uncensored [NEW] https://huggingface.co/dealignai/GLM-5.3-UNCENSORED-FP8 - GLM 5.3 EXL3 abliterated — drowzeys, MoE, unknown, abliterated [NEW] https://huggingface.co/drowzeys/keys-GLM-5.3-EXL3-Abliterated - Qwen3.6-35B-A3B — Qwen, MoE 35B-A3B, 35B / 3B act, official https://huggingface.co/Qwen/Qwen3.6-35B-A3B - Qwen3.6-35B-A3B abliterated — community, MoE 35B-A3B, 35B / 3B act, abliterated [HOT] - Empero Qwen3.8-35B-A3B — empero-ai, MoE distill, 35B / 3B act, official [NEW, HOT] - Qwen3.8-Flash-Next — Qwen, dense, unknown, official [NEW, HOT] - Qwen3.5 2B — Qwen, dense, 2B, official [NEW, HOT] - Qwen3.5-4B — Qwen, dense, 4B, official [NEW] - Qwen3.5-9B — Qwen, dense, 9B, official [NEW] - Qwen3-8B — Qwen, dense, 8B, official - Gemma-4 26B-A4B-it — Google / community, MoE, 26B-A4B, abliterated [NEW, HOT] - Nex-N2.5-mini — community, MoE post-train, 35B-A3B class, official [NEW] - Ornith — ornith-ai, unknown, unknown, official [NEW] https://huggingface.co/ornith-ai - GLM 5.3 Flash — Zhipu, dense, unknown, official [NEW, HOT] - DeepSeek v4.1 Flash — DeepSeek, dense, unknown, official [NEW, HOT] - MiMo-V2.6-Distill-Qwen-9B — Xiaomi MiMo, dense, 9B, official [NEW, HOT] - AliceAI-Foundation-80B-A3B — Yandex, MoE 80B-A3B, 80B / 3B act, official [NEW] https://huggingface.co/Yamada114514/AliceAI-Foundation-80B-A3B-Base-GGUF - MiMo-V2.6-Flash — Xiaomi MiMo, MoE, unknown, official [NEW] - Ling-3.0-Flash — inclusionAI, dense, unknown, official [NEW] https://huggingface.co/inclusionAI/Ling-3.0-flash-int4 - Nemotron 3.5 Lightning — NVIDIA, MoE 30B-A3B, 30B / 3B act, official [NEW] - GLM-5.2 — Zhipu, MoE, unknown, official [NEW] - Kimi K3 — Moonshot, MoE, 2.8T / 104B act, official [NEW, HOT] https://huggingface.co/moonshotai/Kimi-K3 - CYBER-FROST 3.8 — Blackfrost-AI, MoE, ~180B, official [NEW] https://huggingface.co/Blackfrost-AI/CYBER-FROST-3.8-BF16 - Qwen-Image-2.1 — Qwen, image, 7B visual, official [NEW, HOT] https://huggingface.co/Qwen/Qwen-Image-2.1 - Qwen-Image-2.1 Uncensored — abenzerps, image, GGUF, uncensored [NEW, HOT] https://huggingface.co/abenzerps/Qwen-Image-2.1-Uncensored-GGUF - Ternary Bonsai 2 27B — PrismML, dense ternary, 27B, official [NEW, HOT] https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf - MiniMax-M2.5 — MiniMax, unknown, unknown, official [NEW] - Swift 1.5 Qwen3.8 Flash-Next — UkisAI, dense, 27B-class, official [NEW, HOT] - GLM 5.3 — Zhipu, MoE, 753B, official [NEW] ## Chips (20) - M4 Pro — Apple, 24–64 GB unified, 4 pairs - M5 Max — Apple, 64–128 GB unified, 2 pairs - DGX Spark — NVIDIA, 128 GB unified, 11 pairs - RTX 5090 — NVIDIA, 32 GB VRAM, 5 pairs - RTX 4090 — NVIDIA, 24 GB VRAM, 2 pairs - RTX 5060 Ti — NVIDIA, 8–16 GB VRAM, 4 pairs - RTX 5060 — NVIDIA, 8 GB VRAM, 1 pairs - RTX 3060 — NVIDIA, 8–12 GB VRAM, 2 pairs - GTX 1660 SUPER — NVIDIA, 6 GB VRAM, 1 pairs - 2x Tesla V100 32GB — NVIDIA, 64 GB VRAM (2x32), 0 pairs - M1 Max 64GB — Apple, 64 GB unified, 0 pairs - M1 Max 32GB — Apple, 32 GB unified, 1 pairs - M4 mini 24GB — Apple, 24 GB unified, 0 pairs - M3 Ultra — Apple, 96-512 GB unified, 0 pairs - RTX 5070 12GB — NVIDIA, 12 GB VRAM, 0 pairs - RTX 3090 24GB — NVIDIA, 24 GB VRAM, 2 pairs - 2x RTX 3090 24GB — NVIDIA, 48 GB VRAM (2x24), 0 pairs - M5 Pro 64GB — Apple, 64 GB unified, 1 pairs - M6 mini 32GB — Apple, 32 GB unified, 1 pairs - RTX A5000 — NVIDIA, 24 GB VRAM, 2 pairs ## Fastest pairs (decode tok/s) 1. Qwen3.5 2B on RTX 5090 32GB — 351 tok/s (Q4_K_M GGUF, unknown, official) https://owterminal.com/pair/qwen35-2b-q4km-5090 2. Swift 1.5 Qwen3.8 Flash-Next on RTX 4090 24GB — 220 tok/s (IQ2_XS GGUF, Strata, official) https://owterminal.com/pair/swift15-flash-iq2xs-strata-4090 3. Qwen3.8-27B on RTX A5000 — 177.6 tok/s (NVFP4, TensorFold, official) https://owterminal.com/pair/qwen38-27b-nvfp4-tf-a5000 4. Gemma-4 26B-A4B-it on RTX 5090 32GB — 173 tok/s (Q4_K_M GGUF, unknown, abliterated) https://owterminal.com/pair/gemma4-26b-abliterated-q4km-5090 5. Qwen3.8-Flash-Next on RTX 5090 32GB — 142 tok/s (IQ3_S GGUF, Strata, official) https://owterminal.com/pair/qwen38-flash-strata-iq3s-5090-128ram 6. Qwen3.8-27B on RTX A5000 — 139.3 tok/s (NVFP4, vLLM, official) https://owterminal.com/pair/qwen38-27b-nvfp4-vllm-a5000 7. Qwen3.8-Flash-Next on M5 Max — 126.5 tok/s (unknown, MTPLX V2.11.3, official) https://owterminal.com/pair/qwen38-flash-mtplx-m5max 8. Qwen3.8-Flash-Next on RTX 5090 32GB — 104 tok/s (IQ3_S GGUF, Strata, official) https://owterminal.com/pair/qwen38-flash-strata-iq3s-5090 9. Qwen3.8-27B on DGX Spark — 102.9 tok/s (MLX 4-bit g64 + DFlash2, TensorFold, official) https://owterminal.com/pair/qwen38-27b-tf-dflash2-spark 10. Qwen3.8-Flash-Next on RTX 3090 24GB — 100.6 tok/s (IQ2_XS GGUF, Strata, official) https://owterminal.com/pair/qwen38-flash-iq2xs-strata-3090-32k 11. Qwen3.6-35B-A3B on 2× DGX Spark — 94.4 tok/s (NVFP4, unknown, abliterated) https://owterminal.com/pair/qwen36-35b-abliterated-nvfp4-spark 12. Qwen3.8-27B on M5 Pro 64GB — 89.9 tok/s (MLX 4-bit + DFlash2, TensorFold, official) https://owterminal.com/pair/qwen38-27b-tf-dflash2-m5pro-64 13. Qwen3.6-35B-A3B on M4 Pro — 85.5 tok/s (4-bit MLX, rapid-mlx, official) https://owterminal.com/pair/qwen36-35b-4bit-mlx-m4pro 14. Qwen3.5-4B on M4 Pro — 82.8 tok/s (4-bit MLX, rapid-mlx, official) https://owterminal.com/pair/qwen35-4b-4bit-mlx-m4pro 15. Qwen3.8-Flash-Next on RTX 5090 32GB — 80.9 tok/s (NVFP4, unknown, official) https://owterminal.com/pair/qwen38-flash-nvfp4-5090 16. Qwen3.8-Flash-Next on DGX Spark — 79.5 tok/s (EXL3 3.05 bpw, EXL3, official) https://owterminal.com/pair/qwen38-flash-exl3-spark 17. Qwen3.8-Flash-Next on DGX Spark — 74.8 tok/s (NVFP4 MTP-6, TensorFold, official) https://owterminal.com/pair/qwen38-flash-nvfp4-tf-spark 18. Qwen3.8-27B on DGX Spark — 72 tok/s (NVFP4 + DFlash2, unknown, official) https://owterminal.com/pair/qwen38-27b-dflash-spark 19. Qwen3.8-Flash-Next on RTX 3090 24GB — 72 tok/s (IQ3_S GSQ RCO GGUF, Strata, official) https://owterminal.com/pair/qwen38-flash-strata-iq3s-gsq-3090 20. AliceAI-Foundation-80B-A3B on M5 Max 128GB — 65 tok/s (Q4_K_M GGUF, llama.cpp, official) https://owterminal.com/pair/aliceai-80b-q4-m5max 21. Qwen3.8-27B on M6 mini 32GB — 61.9 tok/s (MLX 4-bit + DFlash2, TensorFold, official) https://owterminal.com/pair/qwen38-27b-tf-dflash2-m6mini-32 22. GLM 5.3 Flash on 2× DGX Spark — 60.4 tok/s (EXL3 TR3 4bpw + DFlash2, TensorFold, official) https://owterminal.com/pair/glm53-flash-exl3-tensorfold-2xspark 23. GLM 5.3 Flash on 2x DGX Spark — 55.6 tok/s (NVFP4, unknown, official) https://owterminal.com/pair/glm53-flash-nvfp4-2xspark 24. Empero Qwen3.8-35B-A3B on RTX 3060 12GB — 50 tok/s (Q4_K_M GGUF, llama.cpp, official) https://owterminal.com/pair/empero-qwen38-35b-q4km-3060 25. Qwen3.5-9B on M4 Pro — 49.3 tok/s (4-bit MLX, rapid-mlx, official) https://owterminal.com/pair/qwen35-9b-4bit-mlx-m4pro 26. Qwen3-8B on M4 Pro — 48.3 tok/s (4-bit MLX, rapid-mlx, official) https://owterminal.com/pair/qwen3-8b-4bit-mlx-m4pro 27. MiMo-V2.6-Distill-Qwen-9B on RTX 5060 8GB — 47 tok/s (Q5_K_M GGUF, llama.cpp, official) https://owterminal.com/pair/mimo-v26-qwen9b-q5-5060 28. Qwen3.8-Flash-Next on 2× DGX Spark — 45 tok/s (FP8, unknown, official) https://owterminal.com/pair/qwen38-flash-fp8-spark 29. Qwen3.8-27B on RTX 4090 24GB — 40.7 tok/s (UD-Q4_K_XL GGUF, llama.cpp, official) https://owterminal.com/pair/qwen38-27b-ud-q4k-xl-4090 30. Ornith on GTX 1660 SUPER 6GB — 40 tok/s (unlinked, unknown, official) https://owterminal.com/pair/ornith-1660-super 31. Qwen3.6-35B-A3B on M1 Max 32GB — 37 tok/s (UD-IQ3_XXS GGUF, unknown, official) https://owterminal.com/pair/qwen36-35b-udiq3xxs-m1max-32 32. GLM 5.3 on 4x DGX Spark — 30 tok/s (Int4/Int8 mixed TP4, sparkDash, official) https://owterminal.com/pair/glm53-full-int4int8-tp4-4xspark 33. Qwen3.8-Flash-Next on RTX 3060 12GB — 27.7 tok/s (IQ2_XS GGUF, Strata, official) https://owterminal.com/pair/qwen38-flash-iq2xs-strata-3060 34. Qwen3.8-27B on RTX 5060 Ti 16GB — 23 tok/s (GSQ-RCO IQ3_XXS-mtp, llama.cpp, official) https://owterminal.com/pair/qwen38-27b-gsq-rco-5060ti 35. Qwen3.8-27B on RTX 5060 Ti 16GB — 19 tok/s (UD-Q2_K_XL, llama.cpp, official) https://owterminal.com/pair/qwen38-27b-ud-q2k-xl-5060ti 36. GLM 5.3 Flash on DGX Spark — 18.6 tok/s (EXL3 K2, EXL3, official) https://owterminal.com/pair/glm53-k2-sglang-spark 37. GLM 5.3 Flash on 2x DGX Spark — 15 tok/s (NVFP4, vLLM, official) https://owterminal.com/pair/glm53-flash-nvfp4-vllm-2xspark-mtpoff 38. Nex-N2.5-mini on RTX 5060 Ti 16GB — 14 tok/s (Q4_K_M GGUF, llama.cpp, official) https://owterminal.com/pair/nex-n25-mini-q4km-5060ti 39. Qwen3.8-27B TurboFCFusion on RTX 5060 Ti 16GB — 10 tok/s (IQ2_M GGUF, llama.cpp, uncensored) https://owterminal.com/pair/qwen38-27b-turbofc-uncen-5060ti ## Tape - 2026-09-27 @jmurillocode — Qwen3.8-27B TensorFold MLX 4-bit DFlash2 on 32GB Mac mini M6: 61.9 code / 30.8 chat. https://x.com/jmurillocode/status/2104138445928956021 - 2026-09-27 @aartiles24 — Qwen3.8-27B TensorFold MLX 4-bit DFlash2 on MacBook M5 Pro 64GB: 89.9 tok/s. https://x.com/aartiles24/status/2104179801388884125 - 2026-09-27 @Oluwaphilemon1 — Qwen3.8-27B TensorFold DFlash2 on one Spark: 102.9 tok/s. https://x.com/Oluwaphilemon1/status/2104031407194427892 - 2026-09-28 @PlusTen_AI — GLM 5.3 Flash NVFP4 on 2x DGX Spark: 55.6 coding / 33.6 Korean prose. https://x.com/PlusTen_AI/status/2104531371146482165 - 2026-09-28 @DaSun64125381 — Flash-Next IQ3_S Strata on RTX 5090 256K: 104 decode / 1657 prefill. https://x.com/DaSun64125381/status/2104536066028192175 - 2026-09-29 @moonsteroid — Qwen3.6-35B-A3B UD-IQ3_XXS on MBP M1 Max 32GB: 37 decode / 350 prefill at 100k. https://x.com/moonsteroid/status/2104903466661314585 - 2026-09-29 @needmorevram — Flash-Next IQ3_S GSQ RCO Strata on one RTX 3090 at 128K: 72 decode. https://x.com/needmorevram/status/2104898095318544540 - 2026-09-30 @sudoingX — GLM 5.3 Flash NVFP4 stock vLLM on 2x Spark: 15 decode MTP-off. https://x.com/sudoingX/status/2105246609730924695 - 2026-09-30 @dec21ai — Flash-Next IQ2_XS Strata on RTX 3060 12GB Thinking Medium: 27.7 decode. https://x.com/dec21ai/status/2105260285439422767 - 2026-09-30 @Knuckles_XBT — Swift 1.5 Flash-Next IQ2_XS Strata on RTX 4090 24GB at 256K: ~220 decode / ~3830 prefill. https://x.com/Knuckles_XBT/status/2105245662657011841 - 2026-10-01 @Ja6ek — Qwen3.8-27B NVFP4 on RTX A5000 at 262K: TensorFold 177.57 / vLLM 139.31. https://x.com/Ja6ek/status/2105623160649298395 - 2026-10-01 @arianpg — Flash-Next IQ3_S Strata on RTX 5090 + 128GB RAM: 142 decode. https://x.com/arianpg/status/2105625971261161895 - 2026-10-01 @redp314 — Flash-Next NVFP4 MTP-6 TensorFold 0.3.6.3 on one DGX Spark: 74.8 single-stream. https://x.com/redp314/status/2105569144380739979 - 2026-10-02 @draslan_eth — Flash-Next IQ2_XS Strata 0.1.31 on one RTX 3090 at 32K: 100.6 decode. https://x.com/draslan_eth/status/2105946305646502091 - 2026-10-02 @majewskizby — Full GLM 5.3 753B Int4/Int8 TP4 on 4× DGX Spark: 30.0 prose decode, thinking off. https://x.com/majewskizby/status/2105962150158041477 - 2026-09-23 @hasso5703 — Qwen3.8-27B on one Spark: SGLang, NVFP4, DFlash2, 72 tok/s greedy median. Repo claim. https://github.com/hasso5703/dgx-spark-qwen38 - 2026-09-10 @vcruz305 — GLM-5.3-Flash EXL3 K2 on one Spark: SGLang, MTP k=2, 18.57 tok/s. Repo claim. https://github.com/vcruz305/GLM-5.3-Flash-EXL3-K2-SGLang-DGX-Spark-recipe - 2026-09-22 @stfu0911 — MiMo-V2.6 Distill Qwen-9B Q5_K_M on RTX 5060 8GB: ~47 decode, ~1600 prefill, 262k. https://x.com/stfu0911/status/2102409155143409688 - 2026-09-22 @aqty — AliceAI 80B-A3B Q4_K_M on M5 Max 128GB: ~65 tok/s, 45.1 GiB. https://x.com/aqty/status/2102413529206976799 - 2026-09-21 @tekizaihq — Flash-Next NVFP4 on one 5090: 80.9 tok/s single-stream. https://x.com/tekizaihq/status/2101844907618939233 - 2026-09-20 @yume_arasaki — Flash-Next EXL3 on one Spark: 79.5 code reproduced; 102.6 repetitive clamps; 71.7 prose. https://x.com/yume_arasaki/status/2101741448811229219 - 2026-09-20 @ViC305 — EXL3 Spark recipe reproduced at 79.5 on Yume’s box. https://x.com/ViC305/status/2101747348322103706 - 2026-09-30 @MiaAI_lab — GLM-5.3-Flash EXL3 on 2× DGX Spark with TensorFold: 60.4 tok/s prose, 114.7 structured single-stream; 4 streams 108.8 / 227.9 aggregate. Repo claim. https://github.com/MiaAI-Lab/GLM-5.3-Flash-EXL3-2x-DGX-Sparks-TensorFold - 2026-09-20 @MiaAI_lab — Spark street: Flash-Next solo; GLM 5.3 / DeepSeek v4.1 Flash on 2×+; orch+worker+scanner at 6×. https://x.com/MiaAI_lab/status/2101821681404624976 - 2026-09-21 @sethforprivacy — 8 Sparks: 4× GLM ring, 2× Flash-Next, 2× DeepSeek v4f. https://x.com/sethforprivacy/status/2101834433087070708 - 2026-09-17 @fillagrew — 5090 bake: 2B at 351 tok/s, abliterated 26B-A4B at 173. https://x.com/fillagrew/status/2100478211134025937 - 2026-09-17 @fntAInhead — 5060 Ti 16GB llama.cpp bake-off: 23 / 19 / 14 / 10 tok/s. https://x.com/fntAInhead/status/2100499089431359808 - 2026-09-17 @Youssofal_ — Flash-Next on M5 Max via MTPLX: 126.5 peak, 50 at 200k. https://x.com/Youssofal_/status/2100468205030719533 - 2026-09-17 @Oluwaphilemon1 — Official FP8 Flash-Next on 2× Spark, ~45 tok/s sustained. https://x.com/Oluwaphilemon1/status/2100498960645308573 - 2026-09-17 @Oluwaphilemon1 — Qwen3.8-27B UD-Q4_K_XL on 4090: 40.7 decode, 260k-class. https://x.com/Oluwaphilemon1/status/2100409183115821394 - 2026-09-17 @bonellisystems — Abliterated 35B-A3B NVFP4 on 2× Spark: 94.4 decode. https://x.com/bonellisystems/status/2100417191434768450 - 2026-09-17 @mine_craft_bui — Ornith ~40–46 tok/s on a 1660 SUPER 6GB. https://x.com/mine_craft_bui/status/2100492671534166056 - 2026-09-17 @Oluwaphilemon1 — Empero 35B-A3B Q4_K_M on a 12GB 3060: ~50 decode. https://x.com/Oluwaphilemon1/status/2100374395482939401 - 2026-09-16 @rapidmlx — Locked suite on one M4 Pro: 85.5 tok/s at 21 GB peak. https://x.com/rapidmlx/status/2100252655990010209 ## Street stacks (DGX Spark) Recipes and roles, not tok/s. Source: https://x.com/MiaAI_lab/status/2101821681404624976 - 1× Solo: qwen3.8-flash-next (Solo load) https://github.com/MiaAI-Lab/Qwen3.8-Flash-Next-Single-DGX-Spark; qwen3.8-27b (27B, SGLang) https://github.com/MiaAI-Lab/Qwen3.8-27B-SGLang-DGX-Spark; ling-3.0-flash https://github.com/MiaAI-Lab/Ling-3.0-Flash-SGLang-DSpark-DGX-Spark; glm-5.3-flash (EXL3 K2, one box) https://github.com/vcruz305/GLM-5.3-Flash-EXL3-K2-SGLang-DGX-Spark-recipe - 2× Dual: glm-5.3-flash (GLM on both) https://github.com/MiaAI-Lab/GLM-5.3-Flash-EXL3-2x-DGX-Sparks; glm-5.3-flash (TensorFold, 4 streams) https://github.com/MiaAI-Lab/GLM-5.3-Flash-EXL3-2x-DGX-Sparks-TensorFold; deepseek-v4.1-flash https://github.com/MiaAI-Lab/DeepSeek-v4.1-Flash-EXL3-2x-DGX-Sparks; qwen3.8-flash-next https://github.com/MiaAI-Lab/Qwen3.8-Flash-Next-Dual-DGX-Sparks; mimo-v2.6-flash https://github.com/MiaAI-Lab/MiMo-V2.6-Flash-2x-DGX-Sparks - 3× Triple: glm-5.3-flash + qwen3.8-flash-next (GLM orch on 2×, Qwen worker on 1×); glm-5.3-flash (All three); deepseek-v4.1-flash (Homogeneous); glm-5.2 (NVFP4) https://github.com/MiaAI-Lab/GLM-5.2-NVFP4-AQLM-Triple-DGX-Sparks - 4× Quad: glm-5.3-flash + qwen3.8-flash-next (GLM orch 2× + Qwen worker 2×); glm-5.3-flash (Homogeneous); deepseek-v4.1-flash (Homogeneous) - 6× 6+: glm-5.3-flash + qwen3.8-flash-next + deepseek-v4.1-flash (GLM orch + Qwen worker + DeepSeek scanner)