
Agentic Operations Desk for freight exceptions
Year:
2026
Service:
Custom AI Systems
Industry:
Logistics
Team:
5 specialists, 10 weeks (typical)
Reference scenario: an Agentic Operations Desk for a European freight operator. Agents triage shipment exceptions, gather facts through scoped tools and prepare resolutions; planners approve every action that costs money or makes a commitment to a customer.
Introduction
The freight operator in this scenario moves containers between North Sea and Baltic ports and inland terminals by rail and road. Its exception desk is where the margin leaks. A vessel arrives late, a customs hold appears, a rail slot is missed, a document is wrong, and a planner starts opening five systems to find out what happened and what it will cost.
The operator asks for a custom build on our Agentic Operations Desk accelerator: agents that do the investigation and prepare the resolution, and people who decide. The engagement runs as a 10-week Production Sprint and ends in a supervised pilot on six exception types, with an evaluation set, approval gates and an audit trail that the operations director can read.

Challenge
The desk in the scenario handles a few hundred exception events on a normal day. Each follows the same pattern: read the alert, find the shipment in the transport management system, check the carrier and terminal feeds, check customs status, work out the options, write to the customer. Most of the time goes on gathering facts, not on judgement. Events that arrive at night wait for the morning shift while storage and demurrage charges keep running.
An earlier chatbot pilot has stalled. It could answer questions about a shipment but had no authority to act, and nobody could say how often it was right.
The operator also fears the opposite failure: one large agent with broad write access to bookings. Published measurements support that caution. On software tasks, METR measures a time horizon of about 12 hours at 50 percent success for public frontier models, but only about 1.5 hours at 80 percent. Reliable automation therefore needs short, checkable steps. Gartner predicted in 2025 that over 40 percent of agentic AI projects would be cancelled by the end of 2027.
Solution
The design answers both of the operator's fears: an assistant that cannot act, and an agent that acts too freely.
Bounded steps, not one agent. The exception process is cut into steps short enough to check: classify, investigate, propose, draft, execute. An orchestrator with durable execution holds the state of each case, so a restart or a slow carrier API never loses work.
Tools through MCP, read by default. Every system of record is exposed through an MCP server with scoped permissions. Agents read freely within their scope. Anything that writes goes through an approval gate.
Evaluation before autonomy. Before the pilot, historic exceptions are turned into an evaluation set with expected outcomes, and the desk runs in shadow mode beside the planners.
What is built in the scenario:
Triage, investigation, resolution and communication agents under one orchestrator
MCP servers for the transport management system, carrier and terminal feeds, customs status and email
An approval console: planners approve rebookings, cost commitments and customer messages; low-risk internal notes pass automatically
A frontier model for planning, with smaller open-weight models hosted in an EU region for classification and extraction
OpenTelemetry-based tracing of every tool call, with cost and latency budgets per step
A policy layer that blocks actions outside an agent's scope, and escalation to a designated person when confidence is low or a budget is exceeded
Customer-facing messages are sent by planners after review. Where customers talk to the system directly, it states that it is an AI system, in line with Article 50 of the AI Act, which has applied since 2 August 2026.

Result
The pilot is accepted against criteria agreed before the build, not against a demo.
Target: at least 85 percent end-to-end task success on the evaluation set for the six exception types
Target: human touch time per exception cut by half within the pilot scope
Acceptance criterion: no write action in a system of record without a recorded approval, checked on the full audit trail
Acceptance criterion: cost per completed case stays under the ceiling set for each exception type
These are design targets for the scenario, not audited results.
At the end of the sprint the operator owns the source code, the MCP servers, the evaluation set and harness, the approval policies, the traces and a runbook for adding the next exception type. The planners keep the decisions. The desk takes the searching, the comparing and the first draft, and it leaves evidence for every step it took.

