15–90x Less Token Usage: The TPC-H Benchmark for Governed Enterprise AI
TL;DRConnecting an AI agent directly to enterprise systems through MCP gives it access — not the business meaning it needs to combine that data correctly. Our TPC-H benchmark quantifies the gap: agents using raw MCP servers burn 15–90x more tokens per question than agents using the Semantic Execution Harness, and worse, they can return a confidently wrong answer even when every call succeeds and the SQL runs cleanly. The Harness resolves business definitions before the model is ever called, cutting cost dramatically while getting every question right. The takeaway for anyone evaluating agentic AI in production: access to data isn't the same as understanding it, and that gap shows up as both a bigger bill and a real accuracy risk.
TPC-H — split across Postgres and object storage in MinIO, all 22 queries rewritten as a business user would actually ask them — is the test bench. An agent connected directly to raw MCP servers has to rediscover the schema from scratch on every question: about 21 turns and 92,300 tokens on average. The Semantic Execution Harness resolves business meaning before the model is ever called, cutting that to 1,000–2,000 tokens. The bigger finding is accuracy, not cost: three of the 22 queries return valid, successfully-executed SQL that's still wrong — a missing relationship overstates revenue by 66.7%, an ambiguous join overstates profit by 116.7%, and an undefined metric can honestly read as either 10% or 91%. Through the Harness, all 22 questions come back correct; a small local model alone gets 15 of 22, and cloud escalation for the hard ones closes it to 22 of 22. A cross-system example — renewal risk pulled from a CRM, billing, and Slack — returns full citations and an audit trail, and even catches its own summary underselling a competitive risk. The larger point: because business definitions and policy live outside the model, it can be swapped at any time without reopening a security review.
Read on LinkedIn