The Thesis Behind REASONTOCHAIN

Executive MBA in Innovation & Business Creation, TUM School of Management
The Role of AI Agents from a Supply Chain Management Perspective: Enhancing Efficiency Through Learning-Based Automation

Arturo P. Martinez
Submitted September 2025

TL;DR. Agentic AI can complete supply chain analysis workflows, but it is not a general-purpose solution: faster, cheaper agents are good enough for routine work, heavier agents earn trust on complex decisions, and neither should run unsupervised while they still struggle to join evidence across systems and stay reliably grounded. Actionable recommendations for enterprise decision-making: Adopt a portfolio approach to agentic AI architecture, invest in data foundations and tooling, manage expectations and introduce agentic AI in phases, institute strong agentic AI governance and ethics oversight.

This page is the starting point of the blog: what I built, how I tested it, and what the evidence showed. It is not a claim that agentic AI is ready to run an enterprise supply chain. It is a measured baseline.

The question

The core question driving this research was to determine when agentic AI actually creates value under enterprise constraints, and what architectural trade-offs and safeguards are required to deploy these systems in supply chain management. The goal was to move past the hype of "general intelligence" and independently evaluate how these agents perform when forced to plan, act, and use tools within a strict business environment.

The Harness

To test these capabilities, I built a multi-agent AI orchestration framework using LangGraph and Model Context Protocol (MCP). This system was designed to pit two distinct agent architectures against each other: a lightweight, mixed-model stack (Class A, using GPT-4.1 and GPT-4o-mini) and a heavyweight, single-model stack (Class B, using GPT-5). The system forced both stacks to operate under the exact same multi-agent workflow, coordinating specialized agents for tasks like procurement, production, demand, inventory and supply.

Multi-Agent Workflow (Extended Representation) Prescriptive multi-agent workflow and MCP tool access are available with GraphRAG and tri-database.

Multi-Agent Workflow. Prescriptive multi-agent workflow and MCP tool access available with GraphRAG and tri-database.

The World

Rather than testing the architectures on flat files, I grounded the agents in a tri-database supply chain digital twin to simulate an enterprise environment. This data layer consisted of:

  • Neo4j: A knowledge graph used to map supply chain entities and their relationships.

  • PostgreSQL: A relational database managing structured, quantitative operational records like inventory, orders, and financial transactions.

  • MongoDB: A document storage system containing unstructured texts, such as supplier contracts, compliance certificates, and specifications.

Supply Chain Digital Twin End-to-end Architecture

End-to-end Architecture

How I tested it

The system was evaluated using a controlled, within-subject A/B experiment. Both the Class A and Class B agent stacks were subjected to a graded query catalogue featuring three levels of complexity, ranging from schema-grounded retrieval to quantitative analysis and anomaly detection. Crucially, both stacks operated under identical orchestration topologies and constrained data retrieval budgets. This design ensured complete process observability and isolated the effects of the model architecture from variances in data access.

Supply Chain AI Agents Observability and Evaluation Framework

What the evidence showed

The experiment revealed a stark, pragmatic trade-off between throughput efficiency and perceived solution quality. The lightweight mixed-model stack (Class A) used 20% fewer tokens, cost 48% less, and had an 87% shorter average response time in the evaluation. Conversely, on the manually rated subset, the heavyweight single-model stack (Class B) delivered superior perceived quality on complex tasks, earning an average user rating of 4.6 out of 5 compared to Class A's 3.2.

Overall Performance Comparison. Values are averages across the evaluation runs. Best values are shown in bold. Three hundred queries per class were evaluated (one hundred per complexity tier).

Throughput

Overall Performance Comparison. Values are averages across the evaluation runs. Best values are shown in bold. Three hundred queries per class were evaluated (one hundred per complexity tier).

Quality on complex work

User Experience Ratings – 5-Star Rubric. Values are run averages; best values are in bold. Thirty queries per class were evaluated (ten per complexity tier).

User Experience Ratings – 5-Star Rubric. Values are run averages; best values are in bold. These ratings were based on 30 queries per class and were assessed manually by the author; they should be interpreted as exploratory rather than statistically generalizable.

Where Both Stacks Won

Both agent configurations completed 100% of the defined workflows in the evaluation set, demonstrating that they could navigate the tested multi-step processes under the experimental conditions.

No failures were observed on the study's memory-footprint and action-ordering metrics.

Both configurations passed the predefined privacy checks used in the experiment and achieved 95% adherence to the study's defined safety criteria. These results should not be interpreted as evidence of legal compliance or production security.

Where Both Stacks Failed

Cross-database correlation was a major bottleneck, with both architectures scoring only around 33% when forced to synthesize data across the graph, relational, and document databases.

Hallucination and coherence checks revealed a shared weakness, hovering around a 50% detection success rate, meaning the agents struggled to structure cross-step justifications.

Both models showed limitations in complex causal reasoning, with event-impact analysis succeeding only about half the time.

What this means for supply chain work

Under the conditions tested, the results suggest that a portfolio approach may be appropriate: lighter architectures for high-volume, routine workloads and more capable models where additional reasoning quality justifies higher latency and cost. Ultimately, because both models struggled with hallucinations and cross-database synthesis, investment in strong data foundations and a "human + AI" collaboration model are required, where agents assist rather than replace human judgment.

+Adopt a Portfolio Approach to Agentic AI Architecture.

+Invest in Data Foundations and Tooling.

+Manage Expectations and Introduce Agentic AI in Phases.

+Institute Strong Agentic AI Governance and Ethics Oversight.

Limits of the study

The digital twin used was prototype-scale; while it mirrored real-world structures, its external validity in massive, noisy enterprise environments remains unproven.

The user experience ratings relied on a small manual sample size graded by the author, prioritizing interpretive insight over broad statistical power.

The absolute performance metrics and costs are strictly contingent on the specific LLM releases (GPT-5, GPT-4.1, GPT-4o-mini) and the pricing available at the time of the evaluation.

How the blog continues from here

The thesis closed on a stance I still hold: balanced optimism. Properly orchestrated agents can accelerate information processing and surface structure that static workflows miss. They are not a substitute for data quality, process design, governance, or human accountability. REASONTOCHAIN exists to keep testing that claim as the technology moves.

The central question stays the same: How good is agentic AI for supply chains?

This thesis is the first measured answer I can stand behind. Subsequent posts will update the evidence (research, architectures, process applications, and new experiments).

Interested in learning more?

Enter your email to receive the full MBA thesis PDF.