You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Qwen3.8 27B enters 22nd, and the docs are worth +21.8
Alibaba shipped qwen/qwen3.8-27b on 2026-08-14, a dense 27B vision-language
model with public weights. 202 merged runs, both conditions, $19.74, zero
provider retries and one transport error that was retried into a clean cell.
Baseline winner is @low at 18.1 +/-4.7 (n=41), rank 22 of 27.
Pinned --provider-order akashml, and not as a preference: two of the model's
three endpoints sat at status -2 for the whole session, so the only healthy one
was also the only one that keeps this row a single bf16 price basis instead of
a blend with two fp8 providers. audit.py now asserts that pin, as it does for
the 2.4T sibling. Serving is the weak point and the entry says so: the runs
decoded at a median 30.6 tok/s against the suite's 84-95.
The probe cut two tiers. @xHigh is advertised but unusable, with all three cells
past the 900s budget including 1735s on h1_component in a single turn and 1096s
on e1_counter, the easiest task in the suite. @disabled is accepted here, unlike
the 2.4T sibling, and was swept as the bracket floor.
The headline finding is the documentation lift: +21.8, from 18.1 to 39.9, with
correctness going 43.6 -> 93.6 at @low. Qwen3.6-27B gained +22.0. Two 27B models
a generation apart landing within 0.2 of each other is the strongest evidence in
the study that the substitution law belongs to the size class and not to one
checkpoint, and the local-inference chart now shows the two bars side by side.
The mechanism is visible twice: the tool raises the solve rate AND shortens runs
(median 864s -> 436s), which drops the over-budget rate from 37% to 11%.
Worth flagging for the record, because it decided the published variant. 37% of
baseline runs exceeded the 900s model-time budget. At RAW solve rates @low and
@Medium tie at 46%; the budget takes three solves from @Medium and one from @low
and produces 43.6 against 37.8. Eight runs across both conditions passed every
hidden test between 913s and 1005s and are scored as failures, and
m1_erc20_capped@medium did it four separate times. Qwen3.6-27B published at 30%
over budget, so this is in line with precedent rather than new, but it is the
clearest case the dataset has for revisiting the budget for small models.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Copy file name to clipboardExpand all lines: results/model_meta.json
+50Lines changed: 50 additions & 0 deletions
Original file line number
Diff line number
Diff line change
@@ -913,6 +913,56 @@
913
913
"BF16": 4893.2
914
914
}
915
915
}
916
+
},
917
+
"qwen/qwen3.8-27b": {
918
+
"name": "Qwen: Qwen3.8 27B",
919
+
"context_length": 262144,
920
+
"hugging_face_id": "Qwen/Qwen3.8-27B",
921
+
"price_per_m": {
922
+
"input": 0.45,
923
+
"output": 3.2,
924
+
"cache_read": null,
925
+
"cache_write": null
926
+
},
927
+
"pricing_note": "as listed 2026-08-16 on akashml, the bf16 endpoint the runs were pinned to; the model shipped 2026-08-14 and did not exist at the snapshot date. The two fp8 endpoints are cheaper ($0.40/$3.00 on chutes) but both sat at status -2 through the sweep",
928
+
"type": "dense",
929
+
"params_total": "27B",
930
+
"params_active": "27B",
931
+
"arch": {
932
+
"layers": 64
933
+
},
934
+
"gguf": {
935
+
"repo": "unsloth/Qwen3.8-27B-GGUF",
936
+
"IQ4_XS": 15.7,
937
+
"Q4_K_M": 17.1,
938
+
"Q6_K": 22.9,
939
+
"Q8_0": 29.0,
940
+
"BF16": 54.7,
941
+
"files": {
942
+
"UD-IQ2_XXS": 9.0,
943
+
"UD-IQ2_M": 10.3,
944
+
"UD-Q2_K_XL": 10.7,
945
+
"UD-IQ3_XXS": 11.9,
946
+
"Q3_K_S": 12.6,
947
+
"UD-Q3_K_XL": 13.4,
948
+
"Q3_K_M": 13.8,
949
+
"IQ4_XS": 15.7,
950
+
"Q4_0": 16.1,
951
+
"Q4_K_S": 16.1,
952
+
"IQ4_NL": 16.3,
953
+
"Q4_K_M": 17.1,
954
+
"Q4_1": 17.5,
955
+
"UD-Q4_K_XL": 17.9,
956
+
"Q5_K_S": 19.3,
957
+
"Q5_K_M": 19.8,
958
+
"UD-Q5_K_XL": 20.2,
959
+
"Q6_K": 22.9,
960
+
"UD-Q6_K_XL": 25.9,
961
+
"Q8_0": 29.0,
962
+
"UD-Q8_K_XL": 31.5,
963
+
"BF16": 54.7
964
+
}
965
+
}
916
966
}
917
967
},
918
968
"params_sources": "Lab model cards / HF repo metadata, fetched 2026-07-24: MiMo 1.02T/42B and MiniMax 428B/23B from HF READMEs; DeepSeek 1.6T/49B from HF README (safetensors shows 862B due to FP4/FP8 mixed packing); GLM total 753B from HF safetensors, active ~40B third-party consensus (no lab statement); K3 2.8T/104B lab-stated on the model card, published 2026-07-27 as moonshotai/Kimi-K3; Hy3 295B/21B and Qwen 27B dense lab-stated; V4 Flash 0731 304B from its own card and safetensors tensor count, 13B active from the preview card for the identical architecture (43 layers, 256 experts, 6 routed + 1 shared per token). Muse Glimmer 30B 29.78B dense from the HF safetensors index on meta-models/Muse-Glimmer-30B (29,776,626,688 params, no expert keys in config.json), fetched 2026-08-11. DeepSeek V4 Pro 0813 has no published weights repo and DeepSeek's own docs are unreachable from here, so its parameter counts are left null rather than copied from the preview checkpoint."
0 commit comments