
Air-gapped assistant for a public institution
Year:
2026
Service:
Sovereign AI
Industry:
Public Sector
Team:
5 specialists, 10 weeks
Reference scenario: a Polish public institution gets an assistant for case files and official correspondence that runs on open-weight models, including Polish ones, in its own server room. No document leaves the building and a person approves every draft.
Introduction
A Polish public institution handles thousands of cases a year: applications, appeals, legal opinions and correspondence with other offices. Before a case officer can draft a reply, they search statutes, internal procedures and earlier decisions. Some staff have started pasting fragments of documents into public chatbots, and the security team cannot see it.
In this scenario the institution asks for an assistant it can run itself. The requirements are short: Polish good enough for official writing, answers that cite their sources, and no data outside the building. The PLLuM project site lists a locally installed assistant for office staff among its intended uses. The question is less whether it can be done than how to do it so that the security officer, the data protection officer and the IT department will all sign it off.

Challenge
Three constraints shape the work.
Nothing leaves the network. Case files contain personal data and legally protected information. The security policy rules out external model APIs. In the scenario the institution is also covered by the national cybersecurity system act as amended to implement NIS2 (amendment in force since 3 April 2026), so supply chain risk must be managed for every installed component.
Polish first. General models write fluent Polish, but official style, references to legal provisions and inflected proper names expose weaknesses. Quality has to be shown on the institution's own questions, not on public benchmarks.
Disconnected systems age. The typical failure is designing updates last. Model weights, container images and security patches need a controlled way in: signed artefacts, a written procedure and an audit trail.
Then the regulatory frame. The AI Act's AI literacy duty has applied since 2 February 2025 and its transparency duties since 2 August 2026. Poland's act on AI systems entered into force on 11 August 2026; from 28 October 2026 its supervisor, KRiBSI, can carry out inspections and issue individual opinions.
Solution
Two decisions come first.
Model choice by blind test, not by reputation. With the institution's lawyers and case officers we build an evaluation set of 200 real questions. Three open-weight candidates (Bielik-PL-11B-v3.0-Instruct, a PLLuM model and Mistral Small 4) are then compared in blind A/B tests, the method Japan's Digital Agency is applying to domestic models in its government AI platform. Licences are checked before testing, because some PLLuM variants are non-commercial.
Updates designed before day one. The offline update route is built and rehearsed before the first user sees the system.
What gets built:
A disconnected Kubernetes cluster in the institution's server room, serving models with vLLM behind an OpenAI-compatible gateway with keys and quotas per department.
An offline registry and model mirror with signed artefacts and a written update procedure.
Permission-aware indexing of procedures, legislation and the case archive; hybrid lexical and vector search with a reranker; every answer points to its source passages.
Document tasks run by an agent with a narrow set of tools exposed through an internal MCP gateway: summarise a file, draft a reply, compare versions. The agent reads and drafts; a case officer approves every outgoing document.
An evaluation harness run on every model or prompt change, MLflow tracing, adversarial scanning with Garak and a red-team exercise on prompt injection hidden in internal documents.
Complete action logs kept on site.
For routine, high-volume steps the design also evaluates the compressed Bielik-PL-Minitron-7B-v3.0-Instruct, reported by its authors as 33% smaller and up to 50% faster while retaining about 90% of the full model's quality. Fine-tuning is left out of the first release: retrieval and evaluation come first.
Stack: Kubernetes, vLLM, open-weight models, vector store and reranker, MCP, MLflow, Garak.

Result
This is a reference scenario, so the outcome is expressed as acceptance criteria, not historic results.
Quality target: at least 90% of answers on a held-out part of the question set rated by the institution's reviewers as correct and fully cited; an answer without a source counts as a failure.
Isolation criterion: network tests confirm no outbound connection from the cluster, and the offline update is rehearsed end to end twice before go-live.
Oversight criterion: no document leaves the system without a named approver in the log.
At the end the institution owns the cluster configuration as code, model weights it can inspect and replace, the evaluation set, the runbooks and a governance pack: the AI system inventory entry, the AI Act classification record, a data protection assessment prepared with the 35-question UODO checklist for the public sector, and guidance for staff. The next step in the scenario is a second department, added without new infrastructure.

