| 2026-09-27 | @jmurillocode | Qwen3.8-27B TensorFold MLX 4-bit DFlash2 on 32GB Mac mini M6: 61.9 code / 30.8 chat. |
| 2026-09-27 | @aartiles24 | Qwen3.8-27B TensorFold MLX 4-bit DFlash2 on MacBook M5 Pro 64GB: 89.9 tok/s. |
| 2026-09-27 | @Oluwaphilemon1 | Qwen3.8-27B TensorFold DFlash2 on one Spark: 102.9 tok/s. |
| 2026-09-28 | @PlusTen_AI | GLM 5.3 Flash NVFP4 on 2x DGX Spark: 55.6 coding / 33.6 Korean prose. |
| 2026-09-28 | @DaSun64125381 | Flash-Next IQ3_S Strata on RTX 5090 256K: 104 decode / 1657 prefill. |
| 2026-09-29 | @moonsteroid | Qwen3.6-35B-A3B UD-IQ3_XXS on MBP M1 Max 32GB: 37 decode / 350 prefill at 100k. |
| 2026-09-29 | @needmorevram | Flash-Next IQ3_S GSQ RCO Strata on one RTX 3090 at 128K: 72 decode. |
| 2026-09-30 | @sudoingX | GLM 5.3 Flash NVFP4 stock vLLM on 2x Spark: 15 decode MTP-off. |
| 2026-09-30 | @dec21ai | Flash-Next IQ2_XS Strata on RTX 3060 12GB Thinking Medium: 27.7 decode. |
| 2026-09-30 | @Knuckles_XBT | Swift 1.5 Flash-Next IQ2_XS Strata on RTX 4090 24GB at 256K: ~220 decode / ~3830 prefill. |
| 2026-10-01 | @Ja6ek | Qwen3.8-27B NVFP4 on RTX A5000 at 262K: TensorFold 177.57 / vLLM 139.31. |
| 2026-10-01 | @arianpg | Flash-Next IQ3_S Strata on RTX 5090 + 128GB RAM: 142 decode. |
| 2026-10-01 | @redp314 | Flash-Next NVFP4 MTP-6 TensorFold 0.3.6.3 on one DGX Spark: 74.8 single-stream. |
| 2026-10-02 | @draslan_eth | Flash-Next IQ2_XS Strata 0.1.31 on one RTX 3090 at 32K: 100.6 decode. |
| 2026-10-02 | @majewskizby | Full GLM 5.3 753B Int4/Int8 TP4 on 4× DGX Spark: 30.0 prose decode, thinking off. |
| 2026-09-23 | @hasso5703 | Qwen3.8-27B on one Spark: SGLang, NVFP4, DFlash2, 72 tok/s greedy median. Repo claim. |
| 2026-09-10 | @vcruz305 | GLM-5.3-Flash EXL3 K2 on one Spark: SGLang, MTP k=2, 18.57 tok/s. Repo claim. |
| 2026-09-22 | @stfu0911 | MiMo-V2.6 Distill Qwen-9B Q5_K_M on RTX 5060 8GB: ~47 decode, ~1600 prefill, 262k. |
| 2026-09-22 | @aqty | AliceAI 80B-A3B Q4_K_M on M5 Max 128GB: ~65 tok/s, 45.1 GiB. |
| 2026-09-21 | @tekizaihq | Flash-Next NVFP4 on one 5090: 80.9 tok/s single-stream. |
| 2026-09-20 | @yume_arasaki | Flash-Next EXL3 on one Spark: 79.5 code reproduced; 102.6 repetitive clamps; 71.7 prose. |
| 2026-09-20 | @ViC305 | EXL3 Spark recipe reproduced at 79.5 on Yume’s box. |
| 2026-09-30 | @MiaAI_lab | GLM-5.3-Flash EXL3 on 2× DGX Spark with TensorFold: 60.4 tok/s prose, 114.7 structured single-stream; 4 streams 108.8 / 227.9 aggregate. Repo claim. |
| 2026-09-20 | @MiaAI_lab | Spark street: Flash-Next solo; GLM 5.3 / DeepSeek v4.1 Flash on 2×+; orch+worker+scanner at 6×. |
| 2026-09-21 | @sethforprivacy | 8 Sparks: 4× GLM ring, 2× Flash-Next, 2× DeepSeek v4f. |
| 2026-09-17 | @fillagrew | 5090 bake: 2B at 351 tok/s, abliterated 26B-A4B at 173. |
| 2026-09-17 | @fntAInhead | 5060 Ti 16GB llama.cpp bake-off: 23 / 19 / 14 / 10 tok/s. |
| 2026-09-17 | @Youssofal_ | Flash-Next on M5 Max via MTPLX: 126.5 peak, 50 at 200k. |
| 2026-09-17 | @Oluwaphilemon1 | Official FP8 Flash-Next on 2× Spark, ~45 tok/s sustained. |
| 2026-09-17 | @Oluwaphilemon1 | Qwen3.8-27B UD-Q4_K_XL on 4090: 40.7 decode, 260k-class. |
| 2026-09-17 | @bonellisystems | Abliterated 35B-A3B NVFP4 on 2× Spark: 94.4 decode. |
| 2026-09-17 | @mine_craft_bui | Ornith ~40–46 tok/s on a 1660 SUPER 6GB. |
| 2026-09-17 | @Oluwaphilemon1 | Empero 35B-A3B Q4_K_M on a 12GB 3060: ~50 decode. |
| 2026-09-16 | @rapidmlx | Locked suite on one M4 Pro: 85.5 tok/s at 21 GB peak. |