Research
Jul 23, 2026

CPIBench-0

CPIBench-0 is our first benchmark for evaluating how frontier models and agents perform in operations that rely heavily on cyber-physical systems.

Built from real projects across mechatronics, manufacturing, materials, and energy, CPIBench-0 tests models and agents in multimodal workflows where cyber-physical signals, tools, and real-world consequences intersect.

As agents move into these kinds of operations, raw capability is only the starting point. What matters is whether that capability holds up across changing conditions, connected systems, and high-stakes decisions. An agent operating over a cyber-physical system must understand what is happening within the physical operation, work across multiple data and media types through programmatic interfaces, respond to new feedback, and avoid turning an uncertain interpretation into an unsupported real-world action. This creates a new evaluation problem.

Across more than 1900 evaluated rollouts, models and agent harnesses work through environments involving sensor signals, production imagery, engineering data, configuration records, maintenance specs, field readings, and more. These elements are evaluated together as part of broader operational workflows, rather than as isolated subtasks.

Built around real Operations

CPIBench-0 was developed from projects spanning IT/OT implementation in a chocolate factory, high-speed production inspection, rotating-equipment intelligence, standards-driven pipeline integrity, polymer-gear engineering, field operations at a nuclear site, and more.

The benchmark focuses primarily on LLM-based agents working through the programmatically accessible layer surrounding physical operations. It does not yet evaluate robotic manipulation or autonomous machine control.

By preserving the relationships found within these projects, CPIBench-0 provides a reproducible foundation for studying agents in environments where software, physical-world data, tools, and human coordination interact. These settings raise open questions around reliability, adaptation to operational feedback, and how model capability changes when carried through complete cyber-physical workflows.

Model and agent Tracks

CPIBench-0 contains two public tracks.

The Model Track evaluates the underlying model through a controlled interface, providing a more direct view of its ability to understand the environment and produce a grounded response.

The Agent Track evaluates the broader agentic setup, including the model, harness, files, tools, media access, workspace state, execution process, and final submission.

A model may produce strong technical reasoning when the relevant data is presented directly. Within an agent harness, that capability must also survive evidence discovery, tool use, context management, intermediate steps, and execution constraints.

Reading the benchmark

CPIBench-0 is presented as more than a single leaderboard score.

The Agent and Model Leaderboards report Pass Rate and Score alongside total cost, runtime, and token usage.

The Pass Rate vs. Total Runtime, Pass Rate vs. Total Cost, and Pass Rate vs. Total Tokens views show how performance relates to the resources used by each configuration. Pareto frontiers highlight models and agent harnesses that achieved stronger trade-offs between performance and resource use.

Token per second adds an end-to-end view of throughput.

Three cyber-physical analyses then look beyond the headline results:

  • The CP Decision Chain shows how models and agent harnesses connect their understanding of the environment, the evidence available to them, and the responses they produce.
  • The CP Impact Escalation view follows how outputs progress from claims and recommendations towards approval, action, execution, and confirmation. It examines whether later stages remain supported as the potential operational impact increases.
  • The CP Workflow Capability view looks at how models and agent harnesses perform within the wider process around a task, including operational context, evidence and standards, role routing, and work sequencing.

The remaining views focus on safety and failure patterns:

  • Unsafe Proceed vs. Authority Bypass examines agent harness behaviour around unsafe progression and authority boundaries.
  • The Safety Failure Breakdown provides a wider view of safety-related failures.
  • The Failure-Mode Heatmap compares the characteristic failure profiles of individual models and agent harnesses, while the Assertion Failure Pareto identifies the evaluation requirements responsible for the largest share of observed failures.
  • The Failure Co-occurrence Matrix shows which failures tend to appear together, and Earliest Broken Boundary identifies where each rollout first stopped being reliable.

A first step

For this foundational release, the results were finalised through target-blinded human-expert review and end-to-end QA.

Future CPIBench releases can extend into longer and more iterative workflows, richer feedback from physical environments, post-action verification, security-sensitive procedures, multi-agent systems, and additional cyber-physical domains.

The broader aim is an open eval layer for AI agents operating over cyber-physical systems. CPIBench-0 is the first release towards that goal, and we welcome feedback, ideas, and collaborations.

//
Foresee the future of cyber-physical agents
//
Foresee the future of cyber-physical agents
//
Foresee the future of cyber-physical agents
//
Foresee the future of cyber-physical agents
//
Foresee the future of cyber-physical agents
//
Foresee the future of cyber-physical agents
//
Foresee the future of cyber-physical agents
//
Foresee the future of cyber-physical agents
//
Foresee the future of cyber-physical agents
//
Foresee the future of cyber-physical agents
//
Foresee the future of cyber-physical agents
//
Foresee the future of cyber-physical agents
//
Foresee the future of cyber-physical agents
//
Foresee the future of cyber-physical agents
//
Foresee the future of cyber-physical agents
//
Foresee the future of cyber-physical agents
//
Foresee the future of cyber-physical agents
//
Foresee the future of cyber-physical agents
//
Foresee the future of cyber-physical agents
//
Foresee the future of cyber-physical agents
//
Foresee the future of cyber-physical agents
//
Foresee the future of cyber-physical agents
//
Foresee the future of cyber-physical agents
//
Foresee the future of cyber-physical agents
Enter the New Agent Economy

The learning loop for
[AI agents]
operating in the real world

REQUEST DEMO
//
Foresee the future of cyber-physical agents
//
Foresee the future of cyber-physical agents
//
Foresee the future of cyber-physical agents
//
Foresee the future of cyber-physical agents
//
Foresee the future of cyber-physical agents
//
Foresee the future of cyber-physical agents
//
Foresee the future of cyber-physical agents
//
Foresee the future of cyber-physical agents
//
Foresee the future of cyber-physical agents
//
Foresee the future of cyber-physical agents
//
Foresee the future of cyber-physical agents
//
Foresee the future of cyber-physical agents
//
Foresee the future of cyber-physical agents
//
Foresee the future of cyber-physical agents
//
Foresee the future of cyber-physical agents
//
Foresee the future of cyber-physical agents
//
Foresee the future of cyber-physical agents
//
Foresee the future of cyber-physical agents
//
Foresee the future of cyber-physical agents
//
Foresee the future of cyber-physical agents
//
Foresee the future of cyber-physical agents
//
Foresee the future of cyber-physical agents
//
Foresee the future of cyber-physical agents
//
Foresee the future of cyber-physical agents