Evidence, evaluation and experiments

How do we know an agent is any good?

Claims about agentic AI are easy to make and much harder to prove. This pillar focuses on the evidence behind those claims, moving beyond capability demonstrations, benchmark scores, and vendor narratives to ask a more practical question: what can the agent actually do, under what conditions, and compared with what baseline?

We examine peer-reviewed research, industry experiments, and first-party results (including the research behind REASONTOCHAIN) to understand what has actually been measured. Evaluation goes beyond accuracy to consider cost, latency, data and tool grounding, constraint handling, reliability, hallucination, coherence, and human auditability. We also examine the limitations behind the numbers: sample size, evaluation design, baselines, model versions, and potential sources of bias.

The objective is simple: build the measurement literacy needed to distinguish meaningful evidence from an impressive demo and to assess whether an agent is ready for real supply chain work.