OpenWeightsTerminal

Cite

/llms.txt · /llms-full.txt · 39 pairs · 56 models · 20 chips

Measured pairs

  1. Qwen3.5 2B · 351 tok/s

    RTX 5090 32GB · unknown · Q4_K_M GGUF

    @fillagrew

  2. Swift 1.5 Qwen3.8 Flash-Next · 220 tok/s

    RTX 4090 24GB · Strata · IQ2_XS GGUF

    @Knuckles_XBTGet

  3. Qwen3.8-27B · 177.6 tok/s

    RTX A5000 · TensorFold · NVFP4

    @Ja6ekGet

  4. Gemma-4 26B-A4B-it · 173 tok/s

    RTX 5090 32GB · unknown · Q4_K_M GGUF

    @fillagrew

  5. Qwen3.8-Flash-Next · 142 tok/s

    RTX 5090 32GB · Strata · IQ3_S GGUF

    @arianpgGet

  6. Qwen3.8-27B · 139.3 tok/s

    RTX A5000 · vLLM · NVFP4

    @Ja6ek

  7. Qwen3.8-Flash-Next · 126.5 tok/s

    M5 Max · MTPLX V2.11.3 · unknown

    @Youssofal_

  8. Qwen3.8-Flash-Next · 104 tok/s

    RTX 5090 32GB · Strata · IQ3_S GGUF

    @DaSun64125381Get

  9. Qwen3.8-27B · 102.9 tok/s

    DGX Spark · TensorFold · MLX 4-bit g64 + DFlash2

    @Oluwaphilemon1HF

  10. Qwen3.8-Flash-Next · 100.6 tok/s

    RTX 3090 24GB · Strata · IQ2_XS GGUF

    @draslan_ethGet

  11. Qwen3.6-35B-A3B · 94.4 tok/s

    2× DGX Spark · unknown · NVFP4

    @bonellisystems

  12. Qwen3.8-27B · 89.9 tok/s

    M5 Pro 64GB · TensorFold · MLX 4-bit + DFlash2

    @aartiles24HF

  13. Qwen3.6-35B-A3B · 85.5 tok/s

    M4 Pro · rapid-mlx · 4-bit MLX

    @rapidmlxHF

  14. Qwen3.5-4B · 82.8 tok/s

    M4 Pro · rapid-mlx · 4-bit MLX

    @rapidmlx

  15. Qwen3.8-Flash-Next · 80.9 tok/s

    RTX 5090 32GB · unknown · NVFP4

    @tekizaihqGet

  16. Qwen3.8-Flash-Next · 79.5 tok/s

    DGX Spark · EXL3 · EXL3 3.05 bpw

    @yume_arasaki

  17. Qwen3.8-Flash-Next · 74.8 tok/s

    DGX Spark · TensorFold · NVFP4 MTP-6

    @redp314Get

  18. Qwen3.8-27B · 72 tok/s

    DGX Spark · unknown · NVFP4 + DFlash2

    @hasso5703Get

  19. Qwen3.8-Flash-Next · 72 tok/s

    RTX 3090 24GB · Strata · IQ3_S GSQ RCO GGUF

    @needmorevramGet

  20. AliceAI-Foundation-80B-A3B · 65 tok/s

    M5 Max 128GB · llama.cpp · Q4_K_M GGUF

    @aqtyHF

  21. Qwen3.8-27B · 61.9 tok/s

    M6 mini 32GB · TensorFold · MLX 4-bit + DFlash2

    @jmurillocodeHF

  22. GLM 5.3 Flash · 60.4 tok/s

    2× DGX Spark · TensorFold · EXL3 TR3 4bpw + DFlash2

    @MiaAI_labGet

  23. GLM 5.3 Flash · 55.6 tok/s

    2x DGX Spark · unknown · NVFP4

    @PlusTen_AIGet

  24. Empero Qwen3.8-35B-A3B · 50 tok/s

    RTX 3060 12GB · llama.cpp · Q4_K_M GGUF

    @Oluwaphilemon1

  25. Qwen3.5-9B · 49.3 tok/s

    M4 Pro · rapid-mlx · 4-bit MLX

    @rapidmlx

  26. Qwen3-8B · 48.3 tok/s

    M4 Pro · rapid-mlx · 4-bit MLX

    @rapidmlx

  27. MiMo-V2.6-Distill-Qwen-9B · 47 tok/s

    RTX 5060 8GB · llama.cpp · Q5_K_M GGUF

    @stfu0911HF

  28. Qwen3.8-Flash-Next · 45 tok/s

    2× DGX Spark · unknown · FP8

    @Oluwaphilemon1

  29. Qwen3.8-27B · 40.7 tok/s

    RTX 4090 24GB · llama.cpp · UD-Q4_K_XL GGUF

    @Oluwaphilemon1HF

  30. Ornith · 40 tok/s

    GTX 1660 SUPER 6GB · unknown · unlinked

    @mine_craft_buiHF

  31. Qwen3.6-35B-A3B · 37 tok/s

    M1 Max 32GB · unknown · UD-IQ3_XXS GGUF

    @moonsteroidHF

  32. GLM 5.3 · 30 tok/s

    4x DGX Spark · sparkDash · Int4/Int8 mixed TP4

    @majewskizby

  33. Qwen3.8-Flash-Next · 27.7 tok/s

    RTX 3060 12GB · Strata · IQ2_XS GGUF

    @dec21aiGet

  34. Qwen3.8-27B · 23 tok/s

    RTX 5060 Ti 16GB · llama.cpp · GSQ-RCO IQ3_XXS-mtp

    @fntAInheadHF

  35. Qwen3.8-27B · 19 tok/s

    RTX 5060 Ti 16GB · llama.cpp · UD-Q2_K_XL

    @fntAInheadHF

  36. GLM 5.3 Flash · 18.6 tok/s

    DGX Spark · EXL3 · EXL3 K2

    @vcruz305Get

  37. GLM 5.3 Flash · 15 tok/s

    2x DGX Spark · vLLM · NVFP4

    @sudoingX

  38. Nex-N2.5-mini · 14 tok/s

    RTX 5060 Ti 16GB · llama.cpp · Q4_K_M GGUF

    @fntAInhead

  39. Qwen3.8-27B TurboFCFusion · 10 tok/s

    RTX 5060 Ti 16GB · llama.cpp · IQ2_M GGUF

    @fntAInhead

Models

  • Qwen3.8-27B — Qwen · dense · 27B · new, hot · https://huggingface.co/Qwen/Qwen3.8-27B
  • Qwen3.8-27B Mythos — medismera · dense · 27B · new, hot · https://huggingface.co/medismera/Qwen3.8-27B-OBLITERATED-Mythos-Class-Agentic
  • Qwen3.8-35B-A3B APEX — IsValorum · MoE 35B-A3B · 35B / 3B act · new, hot · https://huggingface.co/IsValorum/Qwen3.8-35B-A3B-Distill-MTP-APEX-I-MiniPlus-V2.1-Abliterated-GGUF
  • Huihui Qwen3.8-27B — huihui-ai · dense · 27B · new · https://huggingface.co/huihui-ai/Huihui-Qwen3.8-27B-abliterated
  • Qwen3.8-27B OBLITERATED — OBLITERATUS · dense · 27B · new · https://huggingface.co/OBLITERATUS/Qwen3.8-27B-OBLITERATED
  • Qwen3.8-27B Uncensored — orcarouter · dense · 27B · new · https://huggingface.co/orcarouter/Qwen3.8-27B-Uncensored
  • Qwen3.8-27B Heretic — 0bserverx · dense · 27B · new · https://huggingface.co/0bserverx/Qwen3.8-27B-Heretic-Abliterated-Uncensored-GGUF
  • Qwen3.8-27B Splash — audreyt · dense · 27B · new · https://huggingface.co/audreyt/Qwen3.8-27B-Splash-abliterated
  • SuperQwen3.8-27B — Jiunsong · dense · 27B · new · https://huggingface.co/Jiunsong/SuperQwen3.8-27b-abliterated
  • Flash-Next RVN — 0bserverx · MoE · unknown · new · https://huggingface.co/0bserverx/RVN-Qwen3.8-Flash-Next-Abliterated-Uncensored
  • Qwen3.8-27B TurboFC — community · dense · 27B · new
  • Qwen3.8-27B Heretic ARA — trohrbaugh · dense · 27B · new · https://huggingface.co/trohrbaugh/Qwen3.8-27B-heretic-ara
  • Qwen3.8-27B Coletti — JonathanColetti · dense · 27B · new, hot · https://huggingface.co/JonathanColetti/Qwen3.8-27B-Uncensored
  • Qwen3.8-27B NoRefusal — sss22213 · dense · 27B · new · https://huggingface.co/sss22213/Qwen3.8-27B-Heretic-NoRefusal
  • Qwen3.8-27B KCRN — heterodoxin · dense · 27B · new · https://huggingface.co/heterodoxin/qwen-3.8-27b-abliterated
  • Qwen3.8-27B Fable — DavidAU · dense · 27B · new · https://huggingface.co/DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU
  • Qwen3.8-9B Heretic — Noobito45 · dense · 9B · new · https://huggingface.co/Noobito45/Qwen3.8-9B-heretic-uncensored-NVFP4-GGUF
  • DeepSeek v4.1 Flash abliterated — distributedcognition · dense · unknown · new · https://huggingface.co/distributedcognition/DeepSeek-V4.1-Flash-abliterated
  • Gemma-4 26B Abliterix — wangzhang · MoE · 26B-A4B · new · https://huggingface.co/wangzhang/gemma-4-26B-A4B-it-abliterix
  • Gemma-4 26B Uncensored — TrevorJS · MoE · 26B-A4B · hot · https://huggingface.co/TrevorJS/gemma-4-26B-A4B-it-uncensored-GGUF
  • Gemma-4 31B Uncensored — TrevorJS · dense · 31B · new · https://huggingface.co/TrevorJS/gemma-4-31B-it-uncensored-GGUF
  • Gemma-4 12B OBLITERATED — OBLITERATUS · dense · 12B · new · https://huggingface.co/OBLITERATUS/Gemma-4-12B-OBLITERATED
  • Qwen3.8-27B Ultra Heretic — llmfan46 · dense · 27B · new · https://huggingface.co/llmfan46/Qwen3.8-27B-Ultra-Uncensored-Heretic-Native-MTP-Preserved
  • Qwen3.8-27B Aggressive — 0xKitkat · dense · 27B · new · https://huggingface.co/0xKitkat/Qwen3.8-27B-Uncensored-Aggressive
  • Qwen3.8-27B Twin Turbo — DavidAU · dense · 27B · new, hot · https://huggingface.co/DavidAU/Qwen3.8-27B-TWIN-TURBO-Fable-Cold-Fusion-709-ULTRA-HERETIC-Uncensored
  • GLM 5.3 Flash abliterated — dealignai · MoE · unknown · new, hot · https://huggingface.co/dealignai/GLM-5.3-Flash-ABLITERATED-FP8
  • GLM 5.3 Flash Uncensored — orcarouter · MoE · 320B / 18B act · new · https://huggingface.co/orcarouter/GLM-5.3-Flash-Uncensored-GGUF
  • GLM 5.3 Uncensored — dealignai · MoE · 753B · new · https://huggingface.co/dealignai/GLM-5.3-UNCENSORED-FP8
  • GLM 5.3 EXL3 abliterated — drowzeys · MoE · unknown · new · https://huggingface.co/drowzeys/keys-GLM-5.3-EXL3-Abliterated
  • Qwen3.6-35B-A3B — Qwen · MoE 35B-A3B · 35B / 3B act · https://huggingface.co/Qwen/Qwen3.6-35B-A3B
  • Qwen3.6-35B-A3B abliterated — community · MoE 35B-A3B · 35B / 3B act · hot
  • Empero Qwen3.8-35B-A3B — empero-ai · MoE distill · 35B / 3B act · new, hot
  • Qwen3.8-Flash-Next — Qwen · dense · unknown · new, hot
  • Qwen3.5 2B — Qwen · dense · 2B · new, hot
  • Qwen3.5-4B — Qwen · dense · 4B · new
  • Qwen3.5-9B — Qwen · dense · 9B · new
  • Qwen3-8B — Qwen · dense · 8B
  • Gemma-4 26B-A4B-it — Google / community · MoE · 26B-A4B · new, hot
  • Nex-N2.5-mini — community · MoE post-train · 35B-A3B class · new
  • Ornith — ornith-ai · unknown · unknown · new · https://huggingface.co/ornith-ai
  • GLM 5.3 Flash — Zhipu · dense · unknown · new, hot
  • DeepSeek v4.1 Flash — DeepSeek · dense · unknown · new, hot
  • MiMo-V2.6-Distill-Qwen-9B — Xiaomi MiMo · dense · 9B · new, hot
  • AliceAI-Foundation-80B-A3B — Yandex · MoE 80B-A3B · 80B / 3B act · new · https://huggingface.co/Yamada114514/AliceAI-Foundation-80B-A3B-Base-GGUF
  • MiMo-V2.6-Flash — Xiaomi MiMo · MoE · unknown · new
  • Ling-3.0-Flash — inclusionAI · dense · unknown · new · https://huggingface.co/inclusionAI/Ling-3.0-flash-int4
  • Nemotron 3.5 Lightning — NVIDIA · MoE 30B-A3B · 30B / 3B act · new
  • GLM-5.2 — Zhipu · MoE · unknown · new
  • Kimi K3 — Moonshot · MoE · 2.8T / 104B act · new, hot · https://huggingface.co/moonshotai/Kimi-K3
  • CYBER-FROST 3.8 — Blackfrost-AI · MoE · ~180B · new · https://huggingface.co/Blackfrost-AI/CYBER-FROST-3.8-BF16
  • Qwen-Image-2.1 — Qwen · image · 7B visual · new, hot · https://huggingface.co/Qwen/Qwen-Image-2.1
  • Qwen-Image-2.1 Uncensored — abenzerps · image · GGUF · new, hot · https://huggingface.co/abenzerps/Qwen-Image-2.1-Uncensored-GGUF
  • Ternary Bonsai 2 27B — PrismML · dense ternary · 27B · new, hot · https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf
  • MiniMax-M2.5 — MiniMax · unknown · unknown · new
  • Swift 1.5 Qwen3.8 Flash-Next — UkisAI · dense · 27B-class · new, hot
  • GLM 5.3 — Zhipu · MoE · 753B · new

Hardware

  • M4 Pro — Apple · 24–64 GB unified · 4 pairs
  • M5 Max — Apple · 64–128 GB unified · 2 pairs
  • DGX Spark — NVIDIA · 128 GB unified · 11 pairs
  • RTX 5090 — NVIDIA · 32 GB VRAM · 5 pairs
  • RTX 4090 — NVIDIA · 24 GB VRAM · 2 pairs
  • RTX 5060 Ti — NVIDIA · 8–16 GB VRAM · 4 pairs
  • RTX 5060 — NVIDIA · 8 GB VRAM · 1 pairs
  • RTX 3060 — NVIDIA · 8–12 GB VRAM · 2 pairs
  • GTX 1660 SUPER — NVIDIA · 6 GB VRAM · 1 pairs
  • 2x Tesla V100 32GB — NVIDIA · 64 GB VRAM (2x32) · 0 pairs
  • M1 Max 64GB — Apple · 64 GB unified · 0 pairs
  • M1 Max 32GB — Apple · 32 GB unified · 1 pairs
  • M4 mini 24GB — Apple · 24 GB unified · 0 pairs
  • M3 Ultra — Apple · 96-512 GB unified · 0 pairs
  • RTX 5070 12GB — NVIDIA · 12 GB VRAM · 0 pairs
  • RTX 3090 24GB — NVIDIA · 24 GB VRAM · 2 pairs
  • 2x RTX 3090 24GB — NVIDIA · 48 GB VRAM (2x24) · 0 pairs
  • M5 Pro 64GB — Apple · 64 GB unified · 1 pairs
  • M6 mini 32GB — Apple · 32 GB unified · 1 pairs
  • RTX A5000 — NVIDIA · 24 GB VRAM · 2 pairs

Tape

  • 2026-09-27 @jmurillocode — Qwen3.8-27B TensorFold MLX 4-bit DFlash2 on 32GB Mac mini M6: 61.9 code / 30.8 chat. https://x.com/jmurillocode/status/2104138445928956021
  • 2026-09-27 @aartiles24 — Qwen3.8-27B TensorFold MLX 4-bit DFlash2 on MacBook M5 Pro 64GB: 89.9 tok/s. https://x.com/aartiles24/status/2104179801388884125
  • 2026-09-27 @Oluwaphilemon1 — Qwen3.8-27B TensorFold DFlash2 on one Spark: 102.9 tok/s. https://x.com/Oluwaphilemon1/status/2104031407194427892
  • 2026-09-28 @PlusTen_AI — GLM 5.3 Flash NVFP4 on 2x DGX Spark: 55.6 coding / 33.6 Korean prose. https://x.com/PlusTen_AI/status/2104531371146482165
  • 2026-09-28 @DaSun64125381 — Flash-Next IQ3_S Strata on RTX 5090 256K: 104 decode / 1657 prefill. https://x.com/DaSun64125381/status/2104536066028192175
  • 2026-09-29 @moonsteroid — Qwen3.6-35B-A3B UD-IQ3_XXS on MBP M1 Max 32GB: 37 decode / 350 prefill at 100k. https://x.com/moonsteroid/status/2104903466661314585
  • 2026-09-29 @needmorevram — Flash-Next IQ3_S GSQ RCO Strata on one RTX 3090 at 128K: 72 decode. https://x.com/needmorevram/status/2104898095318544540
  • 2026-09-30 @sudoingX — GLM 5.3 Flash NVFP4 stock vLLM on 2x Spark: 15 decode MTP-off. https://x.com/sudoingX/status/2105246609730924695
  • 2026-09-30 @dec21ai — Flash-Next IQ2_XS Strata on RTX 3060 12GB Thinking Medium: 27.7 decode. https://x.com/dec21ai/status/2105260285439422767
  • 2026-09-30 @Knuckles_XBT — Swift 1.5 Flash-Next IQ2_XS Strata on RTX 4090 24GB at 256K: ~220 decode / ~3830 prefill. https://x.com/Knuckles_XBT/status/2105245662657011841
  • 2026-10-01 @Ja6ek — Qwen3.8-27B NVFP4 on RTX A5000 at 262K: TensorFold 177.57 / vLLM 139.31. https://x.com/Ja6ek/status/2105623160649298395
  • 2026-10-01 @arianpg — Flash-Next IQ3_S Strata on RTX 5090 + 128GB RAM: 142 decode. https://x.com/arianpg/status/2105625971261161895
  • 2026-10-01 @redp314 — Flash-Next NVFP4 MTP-6 TensorFold 0.3.6.3 on one DGX Spark: 74.8 single-stream. https://x.com/redp314/status/2105569144380739979
  • 2026-10-02 @draslan_eth — Flash-Next IQ2_XS Strata 0.1.31 on one RTX 3090 at 32K: 100.6 decode. https://x.com/draslan_eth/status/2105946305646502091
  • 2026-10-02 @majewskizby — Full GLM 5.3 753B Int4/Int8 TP4 on 4× DGX Spark: 30.0 prose decode, thinking off. https://x.com/majewskizby/status/2105962150158041477
  • 2026-09-23 @hasso5703 — Qwen3.8-27B on one Spark: SGLang, NVFP4, DFlash2, 72 tok/s greedy median. Repo claim. https://github.com/hasso5703/dgx-spark-qwen38
  • 2026-09-10 @vcruz305 — GLM-5.3-Flash EXL3 K2 on one Spark: SGLang, MTP k=2, 18.57 tok/s. Repo claim. https://github.com/vcruz305/GLM-5.3-Flash-EXL3-K2-SGLang-DGX-Spark-recipe
  • 2026-09-22 @stfu0911 — MiMo-V2.6 Distill Qwen-9B Q5_K_M on RTX 5060 8GB: ~47 decode, ~1600 prefill, 262k. https://x.com/stfu0911/status/2102409155143409688
  • 2026-09-22 @aqty — AliceAI 80B-A3B Q4_K_M on M5 Max 128GB: ~65 tok/s, 45.1 GiB. https://x.com/aqty/status/2102413529206976799
  • 2026-09-21 @tekizaihq — Flash-Next NVFP4 on one 5090: 80.9 tok/s single-stream. https://x.com/tekizaihq/status/2101844907618939233
  • 2026-09-20 @yume_arasaki — Flash-Next EXL3 on one Spark: 79.5 code reproduced; 102.6 repetitive clamps; 71.7 prose. https://x.com/yume_arasaki/status/2101741448811229219
  • 2026-09-20 @ViC305 — EXL3 Spark recipe reproduced at 79.5 on Yume’s box. https://x.com/ViC305/status/2101747348322103706
  • 2026-09-30 @MiaAI_lab — GLM-5.3-Flash EXL3 on 2× DGX Spark with TensorFold: 60.4 tok/s prose, 114.7 structured single-stream; 4 streams 108.8 / 227.9 aggregate. Repo claim. https://github.com/MiaAI-Lab/GLM-5.3-Flash-EXL3-2x-DGX-Sparks-TensorFold
  • 2026-09-20 @MiaAI_lab — Spark street: Flash-Next solo; GLM 5.3 / DeepSeek v4.1 Flash on 2×+; orch+worker+scanner at 6×. https://x.com/MiaAI_lab/status/2101821681404624976
  • 2026-09-21 @sethforprivacy — 8 Sparks: 4× GLM ring, 2× Flash-Next, 2× DeepSeek v4f. https://x.com/sethforprivacy/status/2101834433087070708
  • 2026-09-17 @fillagrew — 5090 bake: 2B at 351 tok/s, abliterated 26B-A4B at 173. https://x.com/fillagrew/status/2100478211134025937
  • 2026-09-17 @fntAInhead — 5060 Ti 16GB llama.cpp bake-off: 23 / 19 / 14 / 10 tok/s. https://x.com/fntAInhead/status/2100499089431359808
  • 2026-09-17 @Youssofal_ — Flash-Next on M5 Max via MTPLX: 126.5 peak, 50 at 200k. https://x.com/Youssofal_/status/2100468205030719533
  • 2026-09-17 @Oluwaphilemon1 — Official FP8 Flash-Next on 2× Spark, ~45 tok/s sustained. https://x.com/Oluwaphilemon1/status/2100498960645308573
  • 2026-09-17 @Oluwaphilemon1 — Qwen3.8-27B UD-Q4_K_XL on 4090: 40.7 decode, 260k-class. https://x.com/Oluwaphilemon1/status/2100409183115821394
  • 2026-09-17 @bonellisystems — Abliterated 35B-A3B NVFP4 on 2× Spark: 94.4 decode. https://x.com/bonellisystems/status/2100417191434768450
  • 2026-09-17 @mine_craft_bui — Ornith ~40–46 tok/s on a 1660 SUPER 6GB. https://x.com/mine_craft_bui/status/2100492671534166056
  • 2026-09-17 @Oluwaphilemon1 — Empero 35B-A3B Q4_K_M on a 12GB 3060: ~50 decode. https://x.com/Oluwaphilemon1/status/2100374395482939401
  • 2026-09-16 @rapidmlx — Locked suite on one M4 Pro: 85.5 tok/s at 21 GB peak. https://x.com/rapidmlx/status/2100252655990010209

Street stacks

Recipes on DGX Spark. Not tok/s pairs. Source @MiaAI_lab.

  • 1× Solo — qwen3.8-flash-next (Solo load); qwen3.8-27b (27B, SGLang); ling-3.0-flash; glm-5.3-flash (EXL3 K2, one box)
  • 2× Dual — glm-5.3-flash (GLM on both); glm-5.3-flash (TensorFold, 4 streams); deepseek-v4.1-flash; qwen3.8-flash-next; mimo-v2.6-flash
  • 3× Triple — glm-5.3-flash + qwen3.8-flash-next (GLM orch on 2×, Qwen worker on 1×); glm-5.3-flash (All three); deepseek-v4.1-flash (Homogeneous); glm-5.2 (NVFP4)
  • 4× Quad — glm-5.3-flash + qwen3.8-flash-next (GLM orch 2× + Qwen worker 2×); glm-5.3-flash (Homogeneous); deepseek-v4.1-flash (Homogeneous)
  • 6× 6+ — glm-5.3-flash + qwen3.8-flash-next + deepseek-v4.1-flash (GLM orch + Qwen worker + DeepSeek scanner)

Back to Home · Manifesto · Terms