Energy and compute orchestration, safety first

Year:

2026

Service:

Custom AI Systems

Industry:

Energy

Team:

Senior-led build with coding agents

An energy management system for a Central European site that combines solar generation, battery storage, water-cooled compute and heat recovery. The optimiser only proposes; a deterministic safety layer decides and keeps the site safe without it.

Introduction

An energy investor operating an on-site generation, storage and compute installation in Central Europe wanted one system to decide, every quarter of an hour, where each kilowatt-hour should go: into the battery, into the grid, or into compute whose waste heat is put to use.

The installation combines 21 kWp of solar generation, about 20 kWh of battery storage, two 5.5 kW water-cooled compute units and a heat recovery loop that feeds a ground heat exchanger.

We designed and built the energy management system around one principle: safety before autonomy. An optimiser recalculates the plan every 15 minutes, but it can only propose. A deterministic execution and safety layer on the site controller decides, and it is built to keep the plant safe even if the optimiser falls silent.

Challenge

Flexible compute is an attractive load for surplus solar power, but it is also a physical risk. A water-cooled unit that draws 5.5 kW without coolant flow heats its own circuit towards the alarm threshold. The ground loop that absorbs the recovered heat has temperature limits of its own. And the economics shift every 15 minutes with the market price for exported energy, the solar forecast and the battery's state of charge.

Energy optimisation projects tend to stall here for one of two reasons. Either the clever part is wired directly to the hardware, so a software fault, a stale forecast or a lost connection becomes a plant fault. Or the control logic is tested only on the happy path, and the first real pump failure exposes it.

The investor wanted neither. The brief was strict: physical safety ranks above economics, missing data counts as an unsafe state rather than as 'OK', and nothing may depend on the optimiser being alive. Every claim about the system had to be reproducible by someone other than its author.

Solution

The architecture separates advice from action.

The optimiser is a Python 3.12 service. Every 15 minutes it computes a plan by dynamic programming from market prices in 15-minute slots, a weather-based solar forecast, the tariff model and the battery's state of charge, backed by a probabilistic planner and a digital twin of the site. It publishes recommendations and a heartbeat over MQTT. It has no access to hardware.

The execution and safety layer runs on an industrial edge controller on an open automation platform. Every path that can start a compute unit passes through one interlock. An independent watchdog cuts power to any unit that draws power without its cooling conditions met. Alarms latch and survive a restart. If the optimiser goes silent, the layer falls back to safe rules.

What was built:

  • Optimiser with dynamic programming, stochastic planning and a digital twin

  • Deterministic state machine for compute, pumps and ground-loop thermal protection

  • Read-only telemetry integration for the compute units that refuses any write command

  • Device and signal simulators, including 29 fault scenarios such as a welded contactor, a frozen probe or a pump that stops pumping

  • Three quality gates in CI: unit tests, strict configuration checks, and behaviour tests that run the real automation core in simulated time against a thermal model of the plant

The system was built with coding agents under a strict, written verification workflow. Agents that did not author a change verify it; one such review found a critical flaw that had passed the author's own tests. A change to safety logic needs a numbered client decision, a positive test and a mutation test. Lessons learned go into a written log.

Result

Status: Control system delivered and running on the site controller in advisory mode; staged hardware commissioning.

The latest independent verification round, run by a verifier agent that did not write the code, recorded:

  • Gate 1: 1,441 unit tests passed

  • Gate 2: configuration checks with 0 errors

  • Gate 3: 712 behaviour tests passed and 0 failed on the real automation core in simulated time

  • Mutation testing: 30 of 38 deliberately injected faults were caught by the tests

Known defects are not hidden behind a green build. They stay in every run as strict expected failures until they are fixed or formally accepted. The control strategy was validated against a one-year backtest on measured site data.

The investor holds the complete repository: optimiser, safety layer, simulators, test suites, the numbered decision log and the lessons learned log. Hardware commissioning is staged, and the project rule is plain: no safety code reaches the controller before it has passed independent verification.

Book a readiness call.

Bring one process, product or function where AI should help. We will suggest the most practical next step.

Book a readiness call.

Bring one process, product or function where AI should help. We will suggest the most practical next step.