
Dev Radar: Open weights and on-prem inference - the sovereign stack gets easier
Sovereign AI rests on three things an organisation can own: model weights, an inference server and hardware. All three have moved since late July. Here is what matters for private deployment decisions, without the optimism of a launch post.
Models: permissive licences, but read each one
DeepSeek released DeepSeek-V4.1-Flash on 10 September: a 552B-parameter mixture-of-experts model with native visual understanding and a context of up to one million tokens, with weights on Hugging Face under the MIT licence. DeepSeek says its KV cache needs a quarter of the HBM and an eighth of the SSD storage of the previous generation. Alibaba's Qwen team released open weights for Qwen3.8-27B in mid-August: a dense multimodal model under Apache 2.0 with 27B parameters, text, image and video input and a native context of 262,000 tokens. A 27B-parameter model asks far less of hardware than a 552B-parameter one, so the two suit deployments of different sizes.
Not every open-weight release is open source in licence terms. Moonshot's Kimi K3 and Z.ai's GLM-5.3 are published under custom licences. Legal review of each model licence belongs inside model selection, not after it.
One operational detail supports the sovereign case. DeepSeek retired the V4-Flash and V4-Flash-Vision-Exp API models, whose names temporarily route to V4.1-Flash, and since 14 September deepseek-v4-pro requests also route to V4.1-Flash until a V4.1-Pro launches. On a hosted API, the model behind a name can change. Weights you host change when your team changes them.
Serving: vLLM 0.30 and Ollama 0.35
The vLLM project released v0.30.0 on 22 September. It adds support for DeepSeek-V4.1-Flash and GLM-5.3-Flash, a weight-cache daemon for fast engine restarts and Gumbel-max watermarking. A security fix bounds validation-error response bodies, closing an about 5,300x response amplification. There are breaking changes too: scale-out endpoints now have to be enabled explicitly with the enable-scale-out flag. The earlier v0.29.0 (9 September) made Model Runner V2 the default and deprecated V1, with removal targeted for v0.32. Ollama released v0.35.0 on 28 September; v0.34.4 (23 September) improved Qwen 3.8 prompt processing on Apple Silicon.
On-prem inference stacks ship roughly every two weeks, with breaking changes. A sovereign deployment needs pinned versions and upgrade testing as routine.
Hardware: a second rack-scale supplier
NVIDIA reported second-quarter fiscal 2027 revenue of $96.2 billion on 26 August, $89.0 billion of it from Data Center, and said its Vera Rubin platform is in full production. AMD said on 23 July that its Helios rack-scale system, with 72 Instinct MI455X GPUs per rack, is in production. AMD estimates up to 30% more tokens per dollar than an NVIDIA Vera Rubin NVL72 rack; NVIDIA claims 10x agent throughput against its Grace Blackwell platform. Both figures are vendor claims. For organisations planning on-prem capacity there is now a credible second rack-scale supplier, but supply goes first to hyperscalers and frontier labs: OpenAI expects to bring Helios online from the fourth quarter of 2026, and Anthropic plans to deploy up to 2 gigawatts.
What it means for sovereign deployment decisions
Start from the data classification and the workload, not from the largest model.
Put licence review at the start of model selection.
Pin inference server versions, test upgrades against your own evaluation set and schedule security releases such as vLLM 0.30 for any exposed endpoint.
Plan hardware with lead time in mind: new platforms reach the largest buyers first.
Keep a hosted model as a cost and quality reference, so the sovereign choice is made with open eyes.
Easier does not mean easy: the stack still needs an owner.
