thinking sense · Labs

Engineering at
thinking sense.

Depth in both engineering and enterprise AI adoption is our focus. Deep technical writing on ontology auto-discovery, federated query execution, small model fine-tuning, and the architecture of governed enterprise AI. By the team that shipped Oracle Autonomous Database at 500,000-database scale — and is now building the intelligence layer above your data.

Ontology & Semantic LayerFederated ExecutionModel Fine-Tuning & SLMsEnterprise ArchitectureData Governance

Written by

The thinking sense engineering team

Team background

Oracle Database · HP/Autonomy · Intel · Sun · Bell Labs · 4C Insights

BenchmarkSeptember 24, 2026

15–90x Less Token Usage: The TPC-H Benchmark for Governed Enterprise AI

TL;DRConnecting an AI agent directly to enterprise systems through MCP gives it access — not the business meaning it needs to combine that data correctly. Our TPC-H benchmark quantifies the gap: agents using raw MCP servers burn 15–90x more tokens per question than agents using the Semantic Execution Harness, and worse, they can return a confidently wrong answer even when every call succeeds and the SQL runs cleanly. The Harness resolves business definitions before the model is ever called, cutting cost dramatically while getting every question right. The takeaway for anyone evaluating agentic AI in production: access to data isn't the same as understanding it, and that gap shows up as both a bigger bill and a real accuracy risk.

TPC-H — split across Postgres and object storage in MinIO, all 22 queries rewritten as a business user would actually ask them — is the test bench. An agent connected directly to raw MCP servers has to rediscover the schema from scratch on every question: about 21 turns and 92,300 tokens on average. The Semantic Execution Harness resolves business meaning before the model is ever called, cutting that to 1,000–2,000 tokens. The bigger finding is accuracy, not cost: three of the 22 queries return valid, successfully-executed SQL that's still wrong — a missing relationship overstates revenue by 66.7%, an ambiguous join overstates profit by 116.7%, and an undefined metric can honestly read as either 10% or 91%. Through the Harness, all 22 questions come back correct; a small local model alone gets 15 of 22, and cloud escalation for the hard ones closes it to 22 of 22. A cross-system example — renewal risk pulled from a CRM, billing, and Slack — returns full citations and an audit trail, and even catches its own summary underselling a competitive risk. The larger point: because business definitions and policy live outside the model, it can be swapped at any time without reopening a security review.

Read on LinkedIn
Industry ThesisSeptember 25, 2026

Accelerating the Diffusion

TL;DRPacing the frontier is a reasonable short-term response to a diffusion problem — but it's a response to the wrong variable. Diffusion lags when an enterprise has no data-driven understanding of its own business, when the token cost of a trustworthy answer makes broad deployment uneconomical, and when there's no governance layer to clear a security review. Absent all three, what's left is trial-and-error — bolting a frontier model onto an existing process and hoping it holds. That's exactly where AI ROI misses the mark. History says the fix was never to slow the invention; every prior wave — electricity, computers, the internet — closed its lag by building the infrastructure that let it be absorbed faster, and each wave the lag got shorter. The exponential of AI isn't going to stop. The only lever that has ever actually worked is accelerating diffusion, and that's the race thinking sense is built to run.

Dario Amodei's "We Must Pace the Frontier" set off a real debate: should the labs deliberately slow model capability for safety's sake? But pacing the frontier is a response to a symptom — the underlying condition is diffusion lag, and it's worth being specific about why it happens. Diffusion lags when there's no governed ontology connecting what a model can do to what "revenue" or "at-risk customer" actually means inside a business. It lags when the token cost of re-discovering the same schema and business logic on every question makes broad deployment uneconomical. And it lags when there's no governance layer enforcing who can see what, so every serious use case stalls in a security review instead of shipping. History is emphatic that the fix isn't slowing the invention: electricity took roughly 75 years to show up in factory output, computers took about 50, the internet took 25 — each wave compressing because the absorption infrastructure got built faster, not because progress paused. Blackstone's Jon Gray made the current version of this case with portfolio data rather than theory: 14 companies in Blackstone's own portfolio saw revenue grow 21x in a year, and Anthropic and OpenAI's combined run-rate revenue has reached $105 billion — proof the lag isn't a law of nature. Our own TPC-H benchmark shows what closing it is worth: 92,000-plus tokens and three confidently wrong answers when a model has to rediscover a business from scratch, versus 1,000–2,000 tokens and 22 of 22 correct, with citations and an audit trail, when a governed ontology resolves business meaning first. The exponential of AI will not stop, and pacing the frontier might be a defensible short-term reflex — but every technology adoption wave in history has been resolved by accelerating diffusion, not slowing invention, and each time the enterprise's ability to absorb the new capability got measurably better at scale.

Read on LinkedIn
Semantic EvolutionSeptember 1, 2026

How to Train a Small Model for Databases — and Why It Makes the Ontology Smarter Every Time

TL;DRThe Semantic Evolution Harness doesn't just discover relationships — it gets better at discovering them. Fine-tuning a small model on verified investigation trajectories creates a virtuous loop: more data investigated → sharper ontology → higher accuracy → better investigations → repeat. Don't teach the model the database. Teach it how to interrogate one.

Enterprise ontology has always been expensive because someone had to build it by hand. The alternative is to fine-tune a small language model (1B–7B parameters) on how to investigate a database — not on the data itself. Give it a thin toolset: inspect_schema, column_stats, compare_groups, find_join_path, run_sql. Ask it a business question. It investigates — 10 to 20 tool calls, each one a decision about where to look next — and returns an evidence-backed answer. The database becomes external working memory; the context window no longer bounds how much data the model can reason over. Every successful investigation produces something more valuable than the answer: a trajectory — every observation, tool choice, result, and next step. That trajectory is the training signal. Generate synthetic databases where ground truth is known (churn driven by usage decline + support problems + price increases + noise), run a teacher model through investigations, auto-score against ground truth, and you have millions of verified trajectories. Distill those into the SLM via LoRA/PEFT. The model doesn't learn what causes churn — it learns which evidence to seek, which operations to run, and in what order. Then the loop compounds. Every real investigation produces measurable outcomes: did it discover the actual relationship, did the SQL execute, did it pick the right tables, was the conclusion statistically supported, did it respect security policies? Those outcomes become reinforcement signals. Each pass through the enterprise makes the model sharper. Sharper models discover better ontology relationships. Better relationships produce more accurate answers. More accurate answers generate more verified trajectories. The Semantic Evolution Harness is the name for this loop — the mechanism by which thinking sense's ontology layer doesn't just bootstrap from what's there on day one, but evolves to higher accuracy as more of the enterprise is investigated.

Read on LinkedIn
Architecture Deep DiveSeptember 1, 2026

How thinking sense and OmniGate Work: A Technical Architecture Reference for Enterprise Architects

TL;DRThe moment you register a backend, thinking sense starts learning it — no data dictionary, no manual modelling, no consulting project. Four signal sources build the enterprise ontology from what's actually there: schemas and foreign keys instantly, value-overlap joins asynchronously, real query history continuously, and policy documents on upload. Every inferred relationship lands in a human review queue. The ontology grows from use.

Most enterprise AI initiatives stall on the same prerequisite: a model cannot reason across systems it doesn't semantically understand. thinking sense removes that bottleneck with auto-discovery. Register a backend and four independent pipelines fire immediately — FK relationships are extracted mechanically (no LLM, no guessing), real row data is sampled to catch joins no foreign key declares, the audit log of executed NL2SQL is mined continuously for patterns analysts actually use, and uploaded policy documents are parsed to produce computed business terms a schema alone cannot express. None of these writes to the live ontology directly. Structural certainties can auto-accept; everything inferred — value-overlap joins, mined query patterns, document-derived terms — lands in a reviewable queue. A bad join degrades one answer. A bad access-policy grant is a security decision. This project never lets an inference make the second kind of call on its own. The ontology sits inside a clean two-layer stack: thinking sense (understand, govern, plan, learn) above OmniGate (connect, federate, route, translate, secure, cache, execute). Four wire-protocol frontends — Postgres, MySQL, Oracle wire, and gRPC/MCP for AI agents — converge on one canonical Statement record before touching any backend-specific code, so a single firewall, router, access-control layer, and NL2SQL translator serve every protocol. Federation reaches any JDBC, REST, or S3 backend across 60+ named connectors, with per-backend timing captured in Federation Plans and rollup suggestions generated automatically from repeated questions. Cluster mode via Apache Ignite is opt-in; the free edition runs single-node with the full feature surface.

Read the full technical reference →
Market SignalAugust 28, 2026

The Market Is Beginning to Move Toward the Harness

TL;DRNVIDIA's acquisitions, Palantir's ontology proof, and the economics of distillation all point in the same direction: AI outcome = Model × Harness × Enterprise Knowledge × Compute. The harness is becoming the persistent layer.

NVIDIA's moves — reported acquisition of Hugging Face for open-model distribution, Kumo AI for relational foundation models over structured data, Poolside for coding intelligence, and heavy investment in distillation and synthetic data — collectively signal a shift: enterprise AI is moving from consuming a handful of giant foundation models toward assembling, specializing, and operating portfolios of intelligence. That requires a control layer. Palantir demonstrated that an ontology between enterprise data and operational applications creates enormous value — but at the cost of moving all data into their platform first. The next step is whether ontology construction can become self-evolving, federated execution can continuously discover relationships, and verified executions can distill into reusable task-specific models. Jevons' paradox ensures better harnesses expand the inference market — make a trusted enterprise answer 10× cheaper and companies ask 100× more questions. The harness compounds with use: it captures definitions, permissions, execution history, verified workflows, and distilled task-specific intelligence that competitors cannot replicate. Turning abundant rented intelligence into proprietary, compounding enterprise intelligence — that is the durable advantage.

Read on LinkedIn
Enterprise RiskAugust 21, 2026

From Breach Notice to 24-Hour Action Plan — In One Question

TL;DRSix people and most of a day, or one analyst and one sentence. thinking sense turns a third-party breach notice into a governed, evidence-backed 24-hour action plan — without a single ticket to the CMDB team.

When a critical identity-verification vendor reports unauthorized access, answering "what are we exposed to?" means a war room today: CMDB tickets, SecOps log pulls, manual timestamp reconciliation, policy hunting. The problem isn't missing data — it's missing meaning. thinking sense builds the exposure chain from live enterprise data: maps the vendor to 2 direct applications, 4 downstream business services, and 3 restricted data types (SSN Last 4, Driver License Number, Beneficial Owner Identifier). The critical distinction: a CMDB says what could be exposed; API logs say what did interact — and only one survives contact with a regulator. Policy isn't retrieved, it's executed, producing a timestamped action plan at 30 min, 4 hours, 8 hours, and 24 hours. The same setup answers an unscripted datacenter fire question without rebuilding anything. The architectural thesis: build it like a database, not an agent loop — declarative intent → semantic plan → governed execution. You don't move the enterprise to the AI. You teach the AI the enterprise.

Read on LinkedIn
Document AIAugust 20, 2026

What If You Could JOIN a PDF With Your Enterprise Data — Using SQL?

TL;DRthinking sense AI Extract turns a scanned document in S3 into a live, queryable SQL table — no ETL, no pipeline, no copy/paste. And the LLM doesn't get the final word on what it extracted.

A scanned purchase order lands in S3. Instead of building an ingestion pipeline, a business user asks: "What's the amount due on this PO, and does it match what we expected?" thinking sense turns that into SQL, calls PDF_EXTRACT() — a real SQL table function — extracts the fields, and joins them with live enterprise data. The critical detail: LLMs can silently 'correct' data, rounding numbers or reformatting dates. thinking sense independently verifies every extracted value against the actual document text, so you don't just get a number — you get Amount Due: $3,902.15 → GROUNDED ✓. OCR kicks in automatically for scanned or photographed PDFs. The pipeline is: PDF → AI Extract → Ground → JOIN → Answer. Documents stop being files that AI reads. They become live, queryable enterprise data.

Read on LinkedIn
Data GovernanceAugust 16, 2026

How thinking sense Generates (and Governs) Its Ontology

TL;DRHand-building an enterprise ontology is a multi-quarter project that finishes just as the schema drifts again — thinking sense makes it a byproduct of registering a data source.

Most "let AI query our data" initiatives stall on the same unglamorous prerequisite: a model can't reason about a database it doesn't semantically understand. thinking sense fires three pipelines at registration — structural FK chaining (deterministic, no model call), embedding-based join detection across undeclared relationships, and TabPFN-2 column-level PII classification — to auto-generate a semantic ontology in weeks instead of months. The key design asymmetry: join suggestions above a confidence threshold can auto-apply to the live ontology; PII suggestions never can, categorically — different failure modes demand different trust thresholds.

Read on LinkedIn
Data InfrastructureAugust 9, 2026

Smarter Models Are Not Good Enough. It Needs a Harness for Your Data.

TL;DRA model that writes correct SQL is necessary. It's nowhere near sufficient.

Every AI-writes-SQL demo hides the real engineering: which database do you run that query against, what happens when a hundred agents ask at once, and how does the model know "revenue" means one specific column? thinking sense is the harness layer — connection pooling for burst agent workloads, caching for the repeated state-checks agentic workflows make constantly, routing that decides which backend a question belongs to, and semantic grounding that maps business vocabulary across sources before any model call. Federation is a meaning problem before it's a plumbing problem.

Read on LinkedIn
Engineering LeadershipAugust 1, 2026

We Built a Platform That Runs 500,000 Databases, Executes Over 100 Billion SQLs, and Serves Over 100 Billion API Calls. Here's What It Took.

TL;DRThe most consequential architectural decision in Oracle Autonomous Database wasn't about storage or compute — it was location transparency.

Starting with ten engineers and an impossible mandate, the team built a platform running 500,000 databases executing 100 billion SQL statements and 100 billion API calls every hour. Location transparency — customers connect to logical endpoints, never physical machines — made zero-downtime patching, live rebalancing, and hardware replacement invisible to workloads at that scale. The real lesson: operational automation doesn't limit innovation, it creates the headroom for it — and the culture built through 3am incidents scales better than any architecture review.

Read on LinkedIn
AI InfrastructureJuly 26, 2026

The Database Didn't Need to Change. Until AI Agents Arrived.

TL;DRAI agents impose requirements no traditional database deployment was designed for — and PostgreSQL's own write model points to the solution.

AI agents fan out across parallel sub-tasks, clone working state in milliseconds, and need instant crash recovery without replay procedures — requirements traditional deployments were never built for. Separating compute from storage at the database layer maps naturally onto PostgreSQL's append-only MVCC writes: blocks are written once and immutable, object storage is exactly the right substrate, and cloning becomes a pointer swap instead of a data copy. The result is decades of database correctness given a new home — purpose-built for the stateless, elastic demands of AI agent infrastructure.

Read on LinkedIn
Database EngineeringJune 30, 2026

Making the PostgreSQL Optimizer Smarter with Small Language Models (SLMs)

TL;DRSmall, specialized language models can fix PostgreSQL's query planner at the source — without replacing it.

PostgreSQL's optimizer fails when statistical guesses diverge from reality: bad cardinality estimates, data skew, and anti-patterns like NOT IN can turn millisecond queries into 37-second disasters. Six SLM integration points — pre-optimizer rewriter, cardinality oracle, join-order re-ranker, root-cause classifier, real-time plan critic, and continuous learning loop — each close a specific slice of that gap, running in-process at sub-millisecond latency. The result is a learned, retrainable optimizer layer that gets smarter as your workload evolves, deployable in advisor mode first with zero production risk.

Read on LinkedIn

What the market is saying

Investor SignalSeptember 24, 2026

Where's the Beef? What a Twenty-Year-Old Playbook Says About AI's ROI Gap

Jon Gray — President & COO, Blackstone

TL;DREvery technology revolution gets asked the same question at the same moment: where’s the beef? Electrification got asked in the 1870s. The internet got asked in the 1990s. AI is getting asked now, with over a trillion dollars in committed infrastructure spend and skeptics wondering if enterprise ROI will ever show up. Blackstone’s Jon Gray recently gave the most convincing answer we’ve seen — not a forecast, but portfolio evidence: real companies already showing measurable productivity and revenue gains from AI, at a growth rate with no precedent in enterprise software history. His sharper point is the one that matters for us: the bottleneck was never the invention. It’s the twenty-year lag between when a technology exists and when the institutions around it rebuild to use it. That’s not a reason for caution. It’s the most optimistic reading of the data available — the lag is a known, closable quantity, and closing it faster, safely, is the whole reason thinking sense exists.

Where we were: in the 1870s, electricity existed years before it changed a single factory’s output, because output didn’t move until factories were physically rebuilt around it — new floor plans, new machinery, new ways of organizing work. The invention was never the constraint. The reorganization was. Where we are: Jon Gray, Blackstone’s President and COO, made the same argument to more than 100 top investors this month, backed by data instead of theory. Fourteen companies in Blackstone’s own portfolio saw revenue surge 21x in a single year. Anthropic and OpenAI’s combined annualized revenue has reached $105 billion. Gray’s reframe is the important part: the constraint today isn’t enterprise demand for AI, and it isn’t model capability — both are real and compounding. The constraint is supply of what AI needs to actually deliver: power and compute at the macro level, and inside the enterprise, a governed layer that lets a frontier model touch real business data safely enough to be trusted with an answer. Blackstone is putting its capital at the first constraint — the physical infrastructure, the picks and shovels. Where we’re going: thinking sense is built for the second one. The Governed Ontology Layer and OmniGate are the enterprise’s version of the factory rebuild — the reorganization that lets AI’s capability actually show up as measured productivity, the same way standardized wiring and rebuilt floor plans let electricity show up in output a century ago. History says this lag closes; it always has. It just takes something on the inside of the institution doing the rebuilding. That’s the bet: AI is one of the biggest productivity unlocks in a generation, and the enterprises that close their own twenty-year lag fastest — safely, with governance built in from the start — are the ones who get to compound on it first.

Watch: Jon Gray, Blackstone — “Where’s the Beef?”
Category ValidationAugust 16, 2026

Building the Trusted Platform for the Agentic Enterprise

Rohan Kumar — President & Chief Platform and Engineering Officer, Salesforce

TL;DRSalesforce’s President independently names the AI Control Plane as the foundational layer every enterprise needs — with trusted governance, context, and action sitting on top.

When the President of one of the world’s largest enterprise software companies publishes a piece naming your category — unprompted, independently — the thesis has become a consensus. Rohan Kumar’s article describes an “Enterprise AI Harness” built on an AI Control Plane, with trusted governance as a first-class architectural concern. That is the same architecture thinking sense has been building since day one.

Read on LinkedIn
Market SignalAugust 17, 2026

Why Stripe Paid $7B+ for Model Routing

OpenRouter — How Model Routing Works: Providers, Fallbacks & Auto Router

TL;DRStripe paid $7B+ for economic routing — cost, latency, and failover across 400+ public models. What it cannot buy is grounded routing: routing enterprise intent through a semantic graph to a governed set of enterprise-approved models — frontier or open, cloud or edge — and returning an auditable answer.

OpenRouter’s routing is genuinely intelligent at the infrastructure layer — inverse-square price weighting, throughput optimization, automatic failover, and an Auto Router that classifies prompts by task type using community spend signals. That is real engineering, and Stripe paid $7B+ to own it. But economic routing optimizes which provider answers the request. Grounded routing determines what the question means — resolving intent against the enterprise semantic graph, compiling a governed execution plan, and producing an answer with an audit trail that survives a model swap. The $7B proves routing is infrastructure. The gap above it is the category.

Read the routing deep-dive
Investor SignalSeptember 3, 2026

The Incumbents Are Coming

Seema Amble — Partner, Andreessen Horowitz (a16z)

TL;DRIncumbents own the record. Model providers control the front door. But “the job is bigger than the record” — the work crosses systems, parties, and companies. The agent that wins will learn both the profession and the institution. The institution learning layer is still unclaimed.

Amble defines four agent types by how much judgment each requires: retrieval, process, policy, and principal agents. Incumbents are moving up the hierarchy — but only within their record. A general-purpose agent like Claude can reach across systems, but hits a hard problem: “each application may describe the same customer, contract, or transaction differently.” Without an entity resolution layer, cross-system intelligence fails at the semantic seam. The a16z thesis maps directly to thinking sense’s position: the Governed Ontology Layer resolves entity conflicts before any model sees the data, holds the institution’s policies across all systems, and compounds institution-specific intelligence through RLBUF — learning that belongs to the enterprise, not to a model provider or a vertical AI company. The incumbents own the record. The labs control the model. thinking sense is what the enterprise builds to own the intelligence layer in between.

Read on a16z
Investor SignalAugust 28, 2026

There Will Be No God Model

Ashu Garg — General Partner, Foundation Capital

TL;DRNo single model wins. The moat is the learning loop and workflow ownership — not model weights. The application layer’s opportunity is the compounding flywheel built on top of a model, not inside one.

Garg’s thesis: models improve faster than businesses reorganize, so the application layer captures the friction — and whoever owns the learning loop compounds it into a defensible position. “Proprietary data matters when it gives you a proprietary way to get better.” The test for defensibility: what does your system learn from every use that a competitor cannot replicate? For thinking sense, it is the investigation trajectories — every ontology inference, confirmed relationship, corrected metric, and executed plan — that distill into task-specific SLMs and feed back through RLBUF. Garg’s Cursor analogy maps directly: Cursor learns which model fits which task by observing accept/reject signals. thinking sense learns which data relationships hold, which policies govern them, and which execution paths produce trusted outcomes — from your analysts’ own corrections. The harness is where the learning happens. The model is an input to it.

Read on LinkedIn
Investor SignalSeptember 1, 2026

Models Are Not a Moat

Navin Chaddha — Managing Partner, Mayfield

TL;DRModels become programmable ingredients. Value moves above the model: proprietary data, context and memory, evaluation systems, agent infrastructure, and workflow ownership. “The model is an input. The loop is the company.”

Chaddha names four structural shifts: model intelligence becomes abundant (NVIDIA → Hugging Face, $12.9B); orchestration becomes the strategic question; open weights rewrite the economics of capability; and the moat migrates above and below the model — to proprietary data, context and memory, evaluations, agent infrastructure, and workflow ownership. The quote that defines the category: “A prompt can be copied. A model can be replaced. A deeply integrated context and memory layer, built from proprietary interactions and outcomes, is much harder to reproduce.” For enterprises, the prescription is to own context, evaluations, and model optionality — not the model itself. That is precisely what the Governed Ontology Layer provides: durable enterprise context that compounds with every verified interaction, evaluated by the business users closest to the outcomes, routing seamlessly across frontier or open models as the landscape shifts.

Read on LinkedIn

Want to see how these ideas fit together architecturally?