On-prem AI for a precision manufacturer

Year:

2026

Service:

Sovereign AI

Industry:

Manufacturing

Team:

5 specialists, 12 weeks

Reference scenario: a Japanese precision manufacturer puts decades of engineering knowledge behind a bilingual assistant running on its own hardware. Drawings and process data stay on site, and domestic and open-weight models are compared in blind tests.

Introduction

Decades of engineering knowledge at a Japanese precision manufacturer sit in design standards, defect reports, process change notices and customer specifications. Almost all of it is in Japanese, much of it in scanned PDFs, and the engineers who know where to look are approaching retirement. Overseas plants and sales offices send their questions to headquarters in English and wait days for an answer.

The scenario: the company wants an assistant that answers engineering questions in Japanese or English, shows the source passage, and never sends a drawing or a process parameter outside its own walls. The timing is no accident. Japan's Digital Agency is trialling domestic foundation models on a domestic cloud in its government AI platform, and its method, blind comparison on real work, can be reused by any company weighing domestic models against global ones.

Challenge

IP protection before convenience. Drawings, tolerances and process recipes are the company's competitive position. Legal and the head of engineering rule out external model APIs for this data and want protection against privileged insiders too.

Two languages, one truth. Translating the corpus into English would create a second, unreviewed version of every standard. Engineers need answers in their own language, with the Japanese original as the authority.

Model choice has become a procurement question. The government trial names tsuzumi 2 from NTT DATA, Takane 32B from Fujitsu and PLaMo 2.0 Prime from Preferred Networks, running on SAKURA Cloud. The board asks which would serve the company better: these or a foreign open-weight model. Nobody has evidence from the company's own documents.

Governance without hard law. Japan's AI Promotion Act carries no penalties, so the working reference is the voluntary AI Guidelines for Business, version 1.2 of 31 March 2026, which added AI agents. The European subsidiary also uses the assistant, so the EU AI Act's transparency duties, applicable since 2 August 2026, become a design requirement.

Solution

The design starts from two questions: where the data may live, and what evidence supports the model choice.

Core on-prem, domestic cloud as an option. Design data is served from GPU servers in the company's own data room. The architecture stays portable (containers, open model formats, infrastructure as code), so less sensitive workloads can later move to a domestic cloud such as SAKURA Cloud without a redesign.

Blind evaluation on the company's documents. Senior engineers write a bilingual set of 300 questions with reference answers, drawn from documents cleared for evaluation. Five models go into blind A/B tests: the three domestic models from the government trial and two open-weight models, Mistral Small 4 and a Llama-family model. Licence and deployment terms are checked first; a model that cannot run inside the company's boundary does not go forward for design data.

What gets built:

  • Model serving with vLLM on Kubernetes behind an internal gateway with keys and quotas per team.

  • Confidential virtual machines for the tenant that holds drawings. NVIDIA's own benchmark of September 2026 reports 96.1 to 98.2% of baseline throughput for confidential LLM inference; the pilot measures the overhead on the company's hardware.

  • Ingestion with OCR for scanned reports, permission-aware indexing, hybrid search with a reranker. The index keeps the Japanese original; the answer appears in the user's language with the source passage beside it.

  • An agent with read-only tools exposed through MCP servers for the document management system and the drawing register. Anything that goes to a customer is approved by an engineer.

  • MLflow tracing, an evaluation harness run on every change, and complete action logs.

Stack: Kubernetes, vLLM, vector store and reranker, MCP, MLflow, Intel TDX confidential virtual machines with NVIDIA GPUs.

Result

In the scenario the pilot is accepted against criteria agreed before the build:

  • Quality target: at least 85% of answers to held-out questions rated by senior engineers as correct and correctly sourced, with Japanese and English scores reported separately.

  • Data boundary criterion: no design document and no query leaves the company network; network tests and a review of the gateway logs confirm it.

  • Response time target: overseas plants get a sourced answer within minutes, where the modelled baseline is several days of correspondence with headquarters.

The manufacturer owns everything that carries knowledge: the index, the evaluation set, the blind-test results for each model, the prompts and agent definitions, and the deployment code. The serving layer does not depend on one model, so the winner can be replaced when a better domestic or open-weight model appears, and the same test set shows whether it really is better. Governance records refer to the AI Guidelines for Business.

Book a readiness call.

Bring one process, product or function where AI should help. We will suggest the most practical next step.

Book a readiness call.

Bring one process, product or function where AI should help. We will suggest the most practical next step.