Measured 8 September 2026
llama.cpp against halogen-flash-server.
The same model, the same twenty tasks, the same order: 17 against 18 tasks solved, but 486 against 157 minutes.
Qwen3.8-Flash-Next on the HP ZBook Ultra G1a
- 886tokens/s
- Prompt processingon a 69,000-token context
- 39.9tokens/s
- Output on long contextthe same 69,000 tokens ahead of it
- 39.3tokens/s
- Output in chatshort questions, under 800 prompt tokens
- 157minutes
- for all twenty tasks339,134 tokens generated
Measured with halogen-flash-server. Means over the tasks in each group, measured on the run's own answers, not on filler prompts. Temperature 0, prompt cache off, 70 W GPU budget.
Open the measurement report

