Notes tagged
#llm
7 posts
- 1 min read
One model, two circuits: own iron and cloud
We're a Yandex Cloud partner now. One Qwen3.6, two delivery circuits — own iron at Selectel and cloud via the Yandex Cloud API. Both 152-FZ-attested.
- 5 min read
Russian open models on my own Blackwell — and where NVFP4 breaks
An open RU-model ladder on an RTX PRO 6000: the dense 27B takes the top, a 12B lands half a step behind at half the size, and NVFP4 on Blackwell came up for only one of three.
- 2 min read
Fixed NVFP4 on Blackwell — it was missing CUDA 12.9
Two Blackwell bugs turned out to be one CUDA-toolkit gap, not a vLLM bug. Both models run NVFP4 now; memory ran 4-5GB heavier than the morning estimate.
- 3 min read
DeepSeek V4 Flash on my architect bench: 0.717 and ₽0.022 per solved task
DeepSeek V4 Flash scored 0.717 on my architect bench: below the top-3, above V4-Pro at 4.8× lower price — and for the first time realistic on 2×RTX PRO 6000.
- 2 min read
The boring bet: the frontier arrived where I've been building
I bet on verifiable AI — grounded in the client's corpus, cited, with an honest refusal — back when it sounded niche. I went through dozens of the year's biggest AI interviews: the frontier describes exactly that bet. And I have the one thing the voices on stage don't — a number.
- 2 min read
The full turn: the model is an API call
The full turn: the model is an API call for kopecks. Two circuits — public on CPU, sovereign on-site. The edge lives in the layer above the model.
- 2 min read
Picking the model for the hardware I have
Picking a model for the trusted-AI build on 24 GB. The bench beat intuition: a MoE beat a dense model twice its size; 'thinking' only got in the way.