Notes tagged
#benchmark
5 posts
- 3 min read
GPT-5.6 Luna costs 10x less than Claude Sonnet 5 — and scored exactly the same
Ran my own benchmark against six closed models and four local ones — and found two cases where the top models silently refused completely benign tasks.
- 4 min read
The coach assistant: 32 sources, 70 questions, 77.4%
The corpus grew from 9 to 32 sources. The re-run scored 77.4% on 70 questions, with 10/10 grounding in the new module.
- 4 min read
60 coach questions: why niche AI needs its own benchmark
First basketball slice of LII Sport Bench: 60 questions, 9 public RFB Academy sources, and the first numbers for the RAG assistant.
- 5 min read
Russian open models on my own Blackwell — and where NVFP4 breaks
An open RU-model ladder on an RTX PRO 6000: the dense 27B takes the top, a 12B lands half a step behind at half the size, and NVFP4 on Blackwell came up for only one of three.
- 2 min read
Fixed NVFP4 on Blackwell — it was missing CUDA 12.9
Two Blackwell bugs turned out to be one CUDA-toolkit gap, not a vLLM bug. Both models run NVFP4 now; memory ran 4-5GB heavier than the morning estimate.