The Thesis Behind REASONTOCHAIN
Executive MBA in Innovation & Business Creation, TUM School of Management
The Role of AI Agents from a Supply Chain Management Perspective: Enhancing Efficiency Through Learning-Based Automation
Arturo P. Martinez
Submitted September 2025
TL;DR. Agentic AI can complete supply chain analysis workflows, but it is not a general-purpose solution: faster, cheaper agents are good enough for routine work, heavier agents earn trust on complex decisions, and neither should run unsupervised while they still struggle to join evidence across systems and stay reliably grounded. Actionable recommendations for enterprise decision-making: Adopt a portfolio approach to agentic AI architecture, invest in data foundations and tooling, manage expectations and introduce agentic AI in phases, institute strong agentic AI governance and ethics oversight.
This page is the starting point of the blog: what I built, how I tested it, and what the evidence showed. It is not a claim that agentic AI is ready to run an enterprise supply chain. It is a measured baseline.
The question
The core question driving this research was to determine when agentic AI actually creates value under enterprise constraints, and what architectural trade-offs and safeguards are required to deploy these systems in supply chain management. The goal was to move past the hype of "general intelligence" and independently evaluate how these agents perform when forced to plan, act, and use tools within a strict business environment.
The Harness
To test these capabilities, I built a multi-agent AI orchestration framework using LangGraph and Model Context Protocol (MCP). This system was designed to pit two distinct agent architectures against each other: a lightweight, mixed-model stack (Class A, using GPT-4.1 and GPT-4o-mini) and a heavyweight, single-model stack (Class B, using GPT-5). The system forced both stacks to operate under the exact same multi-agent workflow, coordinating specialized agents for tasks like procurement, production, demand, inventory and supply.
Multi-Agent Workflow. Prescriptive multi-agent workflow and MCP tool access available with GraphRAG and tri-database.
The World
Rather than testing the architectures on flat files, I grounded the agents in a tri-database supply chain digital twin to simulate an enterprise environment. This data layer consisted of:
Neo4j: A knowledge graph used to map supply chain entities and their relationships.
PostgreSQL: A relational database managing structured, quantitative operational records like inventory, orders, and financial transactions.
MongoDB: A document storage system containing unstructured texts, such as supplier contracts, compliance certificates, and specifications.
End-to-end Architecture
How I tested it
The system was evaluated using a controlled, within-subject A/B experiment. Both the Class A and Class B agent stacks were subjected to a graded query catalogue featuring three levels of complexity, ranging from schema-grounded retrieval to quantitative analysis and anomaly detection. Crucially, both stacks operated under identical orchestration topologies and constrained data retrieval budgets. This design ensured complete process observability and isolated the effects of the model architecture from variances in data access.
What the evidence showed
The experiment revealed a stark, pragmatic trade-off between throughput efficiency and perceived solution quality. The lightweight mixed-model stack (Class A) used 20% fewer tokens, cost 48% less, and had an 87% shorter average response time in the evaluation. Conversely, on the manually rated subset, the heavyweight single-model stack (Class B) delivered superior perceived quality on complex tasks, earning an average user rating of 4.6 out of 5 compared to Class A's 3.2.
Throughput
Overall Performance Comparison. Values are averages across the evaluation runs. Best values are shown in bold. Three hundred queries per class were evaluated (one hundred per complexity tier).
Quality on complex work
User Experience Ratings – 5-Star Rubric. Values are run averages; best values are in bold. These ratings were based on 30 queries per class and were assessed manually by the author; they should be interpreted as exploratory rather than statistically generalizable.
Where Both Stacks Won
Both agent configurations completed 100% of the defined workflows in the evaluation set, demonstrating that they could navigate the tested multi-step processes under the experimental conditions.
No failures were observed on the study's memory-footprint and action-ordering metrics.
Both configurations passed the predefined privacy checks used in the experiment and achieved 95% adherence to the study's defined safety criteria. These results should not be interpreted as evidence of legal compliance or production security.
Where Both Stacks Failed
Cross-database correlation was a major bottleneck, with both architectures scoring only around 33% when forced to synthesize data across the graph, relational, and document databases.
Hallucination and coherence checks revealed a shared weakness, hovering around a 50% detection success rate, meaning the agents struggled to structure cross-step justifications.
Both models showed limitations in complex causal reasoning, with event-impact analysis succeeding only about half the time.
What this means for supply chain work
Under the conditions tested, the results suggest that a portfolio approach may be appropriate: lighter architectures for high-volume, routine workloads and more capable models where additional reasoning quality justifies higher latency and cost. Ultimately, because both models struggled with hallucinations and cross-database synthesis, investment in strong data foundations and a "human + AI" collaboration model are required, where agents assist rather than replace human judgment.
+Adopt a Portfolio Approach to Agentic AI Architecture.
+Invest in Data Foundations and Tooling.
+Manage Expectations and Introduce Agentic AI in Phases.
+Institute Strong Agentic AI Governance and Ethics Oversight.
Limits of the study
The digital twin used was prototype-scale; while it mirrored real-world structures, its external validity in massive, noisy enterprise environments remains unproven.
The user experience ratings relied on a small manual sample size graded by the author, prioritizing interpretive insight over broad statistical power.
The absolute performance metrics and costs are strictly contingent on the specific LLM releases (GPT-5, GPT-4.1, GPT-4o-mini) and the pricing available at the time of the evaluation.
How the blog continues from here
The thesis closed on a stance I still hold: balanced optimism. Properly orchestrated agents can accelerate information processing and surface structure that static workflows miss. They are not a substitute for data quality, process design, governance, or human accountability. REASONTOCHAIN exists to keep testing that claim as the technology moves.
The central question stays the same: How good is agentic AI for supply chains?
This thesis is the first measured answer I can stand behind. Subsequent posts will update the evidence (research, architectures, process applications, and new experiments).
Interested in learning more?
Enter your email to receive the full MBA thesis PDF.