Scattered grey dots labelled Raw Text-to-SQL at 21 percent accuracy on the left, an orange arrow, and a clean orange connected graph labelled Semantic Execution at 100 percent accuracy on the right.

OpenAI Dots Have a Data Problem. Here Is the Fix.

A Specialist Dot posts Q3 revenue into your CFO's Slack at 6 a.m., an hour before the board call. The number is wrong, and nothing flagged it. OpenAI just shipped always-on autonomous agents that reason, browse, and collaborate across Slack and Teams. The moment they query your enterprise database, accuracy drops below 21%. The data layer is the bottleneck.

Capability Raw LLM Text-to-SQL Semantic Execution Layer (Colrows)
Accuracy on enterprise schemas 0-21% on private corporate data (BEAVER benchmark)1 100% on modeled queries (dbt 2026 benchmark)2
Failure mode Silent wrong answers. Query returns data, but the numbers are wrong. Graceful refusal. If the model is missing, the system returns an error, not a guess.
Multi-source queries Agent guesses join paths across databases. Hallucinated joins break governance. Compiler proves valid join paths across 16+ engines before SQL is generated.
Governance Prompt-level only. Agent can bypass UI restrictions via direct API calls. Compile-time RBAC/ABAC. If the user lacks access, the SQL plan is never created.
Audit trail None. No record of which schema the model used or why it chose a specific join. Full trace: generated SQL, confidence score, semantic sources, and timestamp logged.

What OpenAI Dots Actually Are

On September 29, 2026, OpenAI announced Dots at DevDay. Each Dot is a persistent, always-on autonomous agent powered by GPT-6 Astra3.

Each Dot runs on its own dedicated cloud computer with its own browser. It operates 24/7, independent of whether your laptop is open or your phone is charged. It connects to over 4,000 applications through plugins. It carries context across ChatGPT, Slack, and Microsoft Teams without losing the thread of a conversation3.

For enterprise deployments, OpenAI introduced Specialist Dots. These are agents assigned dedicated organizational roles (procurement, invoice processing, contract review) with their own IT-managed credentials and system access. They integrate with Microsoft Agent 365 for governance3.

The capability is real. A Specialist Dot can monitor a CRM, pull historical usage data, draft a renewal proposal, and route it to legal. All without a human prompt.

One question went unanswered at DevDay: where does the Dot get its data, and how do you know the numbers are correct?

The Text-to-SQL Accuracy Cliff

On standardized academic benchmarks, modern LLMs look brilliant at translating natural language into SQL. GPT-4o and o1-preview score between 86.6% and 91.2% on the Spider 1.0 dataset1. Spider uses small, cleanly normalized databases with readable column names.

Enterprise databases look nothing like Spider.

The BEAVER benchmark uses private corporate schemas with cryptic column names, undocumented business rules, and overlapping entity definitions. On BEAVER, GPT-4o scored exactly 0.0% in retrieval settings and 2.0% without retrieval. Advanced agentic frameworks using o1-preview managed 21.3%1. Roughly four out of five queries returned wrong data that looked correct.

For a deeper analysis of this accuracy gap, the gap between benchmark performance and production reality has three specific drivers.

Scale and context dilution

Enterprise data estates span thousands of tables and tens of thousands of columns. Injecting that schema into a context window dilutes relevance. The model guesses which subset of data matters before writing a line of SQL. A wrong guess at this stage is unrecoverable1.

Hidden semantics

When a user asks for "Q3 Revenue," the database cannot tell the model which of four revenue columns is the one approved by the finance department, whether to use calendar or fiscal quarters, or how to resolve entity overlaps between billing systems and CRMs1. The critical business logic lives outside the physical schema.

No verification

A generated SQL query that parses correctly and returns rows looks identical to a correct query. If the model joins two tables incorrectly and silently double-counts active subscriptions, the database executes it without complaint4. The failure mode is a confidently delivered wrong number, indistinguishable from the right one.

Now imagine that failure mode running 24/7 inside a Specialist Dot that autonomously generates reports, monitors KPIs, and feeds numbers into Slack channels your CFO reads every morning.

Why Knowledge Graphs Are Not Enough

The first generation of solutions treats the database schema as a graph to traverse rather than a document to embed. Systems like FalkorDB map primary and foreign keys as explicit edges, turning a silent schema into an interconnected web5.

It works, partially. In controlled studies, representing an enterprise database as a knowledge graph lifted GPT-4's zero-shot accuracy from 16.7% to 54.2%1. That is a 3x improvement. It is also a 46% error rate on enterprise queries.

The limitation is structural. A knowledge graph maps what tables are related. It does not encode which revenue column is the approved metric, how fiscal quarters are calculated, or which join path the finance team has validated. As the distinction between semantic layers and execution layers makes clear: a knowledge graph maps relationships. A semantic execution layer computes what is true.

The Missing Layer: Compile-Time Semantic Execution

The dbt Labs 2026 benchmark tested what happens when you place a deterministic semantic layer between the LLM and the database. Accuracy jumped from 64.5% (raw text-to-SQL) to 100.0% for modeled queries2. When the semantic layer could not answer a question because the data model was missing, it returned an error instead of guessing2.

For anything going to an auditor, a board deck, or an automated workflow, the gap between "wrong answer delivered confidently" and "I cannot answer this yet" separates a tool from a liability.

A Day in the Life: The Q3 Revenue Question

Make it concrete. A Specialist Dot acting as a sales engineer gets a Slack message: "What was Q3 revenue for the North America enterprise segment?" Here is what happens on each path.

The raw text-to-SQL path

The schema has four columns that could be revenue: gross_bookings, rev_recognized, net_rev_fy, and billing_amount, spread across a CRM and a billing system. The model picks one by name similarity. It joins orders to invoices on a plausible key and silently double-counts partially refunded deals. It filters on a calendar quarter when finance closes on a fiscal one. The query parses. Rows come back. A confident number lands in Slack. Nothing errors, and no one can see that three separate decisions were guesses.

The Colrows path

The Dot sends the same intent through the semantic execution layer. The meaning layer resolves "revenue" to the finance-approved definition: recognized revenue, net of refunds, on the fiscal calendar. The structure layer proves the one validated join path and rejects the path that would double-count. Compile-time governance appends the segment and region predicate for the user's scope. The compiler emits dialect-perfect SQL, runs it, and logs the definition, the join path, and the generated SQL to the audit trail. Ask the same question tomorrow and you get the same answer. If no approved model covers that segment yet, the system returns an error instead of a guess.

Same question. One path produces a number nobody can defend. The other produces a number you can put in front of an auditor.

A deterministic semantic execution layer that encodes business logic, proves join paths, and enforces governance produces more accurate analytics than any model upgrade. Context, not reasoning power, is the bottleneck. Fix the Context, Not the Model.

Colrows is built on this principle. It resolves meaning at compile time, before SQL reaches any warehouse.

Traditional semantic layers (Cube.js, dbt Semantic Layer/MetricFlow) resolve meaning at presentation time. Data engineers manually author YAML or JavaScript files to define metrics, name entities, and declare joins6. These work well for static dashboards. They break when the number of business concepts grows faster than the engineering team can model them, or when data spans multiple warehouses.

Colrows resolves meaning at compile time. It autonomously ingests data catalogs, dbt models, BI tool definitions, and documentation to build a typed semantic graph6. That graph has three substrates:

Meaning Layer (Ontologies): Enterprise vocabulary, business logic, hierarchies, and synonyms. This is where "net revenue" gets a single, auditable definition6.

Structure Layer (Semantic Knowledge Graph): Physical table-to-table relationships, explicit joins, and cardinality. Colrows mathematically proves valid join paths before execution. An AI agent cannot hallucinate a false relationship because the compiler rejects unproven paths6.

Behavior Layer (Statistical Profile): Empirical data distributions, value frequencies, and query access patterns. This layer detects schema drift automatically and self-heals the graph as underlying databases change6.

The compiler emits dialect-perfect SQL across 16+ warehouse engines (Snowflake, Databricks, BigQuery, Postgres, and others) optimized for the target system6.

How Autonomous Agents Connect to Colrows

The bridge between an AI agent and a semantic execution layer is the Model Context Protocol (MCP). MCP has become the standard interface connecting AI applications to external tools, with 97 million monthly SDK downloads and 10,000 active servers as of 20267.

OpenAI already supports MCP in the Agents SDK and Responses API for developers building custom agents8. Dots currently connect to external applications through OpenAI's plugin ecosystem. As MCP adoption continues across the industry (and given OpenAI's existing investment in the protocol), the path from Dots to MCP-connected tools like Colrows is a matter of when, not if.

Today, any developer building agents on OpenAI's Agents SDK, Anthropic's Claude, Cursor, or LlamaIndex can connect to Colrows through a single MCP connector endpoint9.

The connection uses OAuth 2.0 Authorization-Code Grant with mandatory PKCE (S256) for public clients9. Short-lived access tokens expire in 15 minutes. Refresh tokens rotate on every use with a 30-day lifecycle9.

The transport is standard, so there is little to build. An agent speaks JSON-RPC 2.0 over HTTPS to the Colrows MCP endpoint and completes a one-time OAuth 2.0 consent flow to authorize access9. Any MCP-capable client (OpenAI's Agents SDK, Claude, Cursor, LlamaIndex) connects by pointing at the endpoint URL. There is no SDK to adopt and no SQL layer to write.

Once authenticated, the agent discovers exactly six read-only tools via the MCP tools/list method9:

MCP Tool What It Does
listDatastores Returns available governed data sources
getSemanticContext Retrieves business definitions, metrics, and entity relationships
executeQuery Compiles natural language intent into governed SQL and returns results
searchVerifiedQueries Finds human-approved, pinned SQL for tier-one metrics
executeVerifiedQuery Runs a pinned query by ID and version. 100% accuracy guarantee.
getQueryStatus Checks execution progress for long-running queries

None of these tools accept raw SQL text. Any request containing a sql field is rejected by the server9. The LLM is physically isolated from the database. It expresses intent. The compiler generates SQL. Hallucinated queries cannot execute.

Governance for Agents That Never Sleep

A Specialist Dot running 24/7 with deep access to internal systems changes the threat model. If governance is enforced only at the presentation layer (hiding a dashboard tab, restricting a BI tool menu), an autonomous agent bypasses it by querying underlying data through APIs.

Colrows pushes security into the compiler. If an agent lacks authorization for a data point, the SQL plan is never generated9.

Compile-time access control

Colrows enforces both Role-Based (RBAC) and Attribute-Based (ABAC) access control at compile time6. When the agent authenticates via MCP, Colrows evaluates the user's identity against a multi-scope hierarchy: global, datastore, persona, and user6.

Row and column-level predicates are injected directly into the Abstract Syntax Tree before SQL generation6. If an agent operating on behalf of a European regional manager queries global sales, Colrows appends WHERE region = 'Europe' at compile time. The agent never sees the unauthorized rows because the query that would have returned them was never created.

PII redaction

Redaction operates independently of access control. A Dot running churn analysis needs access to the customers table for accurate counts, but it does not need plain-text email addresses or government ID numbers. Colrows applies versioned redaction policies (partial masking, cryptographic hashing) per column10. Because the hashes are stable, the agent can still use redacted columns as JOIN keys without exposing PII to the model or to OpenAI's servers10.

Verified Queries for board-level metrics

For regulatory filings, board decks, and public financial disclosures, probabilistic execution is not acceptable. Colrows supports Verified Queries: human-approved, pinned SQL stored by unique ID and version11. When an agent invokes searchVerifiedQueries and executeVerifiedQuery via MCP, the compiler is bypassed entirely. The stored, immutable SQL runs as written9.

This gives you the flexibility of a natural language agent interface with the rigidity of traditional data engineering on the queries that cannot afford to be wrong.

When Is a Warehouse-Native Option Enough?

Be honest about where you need this. If your entire data estate lives in one warehouse and your agents only ever query that warehouse, a warehouse-native option can be enough. Snowflake's Cortex Analyst answering natural language over Snowflake Semantic Views, or Databricks' equivalent, keeps meaning close to the data and adds no extra layer to run.

The limit is the boundary of the engine. These objects are scoped to their own warehouse. The moment recognized revenue lives in Snowflake, product events in Databricks, and entitlements in Postgres, a warehouse-native definition cannot prove a join across all three or enforce one governed metric across engines. That fragmented estate, not the single-vendor one, is where a multi-engine semantic execution layer earns its place. If you are weighing the native route first, the trade-offs of warehouse-native NL-to-SQL are worth reading before you commit.

What This Means for Enterprise Data Teams

The arrival of always-on agents changes the economics of data access. When every employee has a Dot that can ask questions of the data estate continuously, the volume of queries hitting your databases increases by orders of magnitude. Static dashboards were a compromise. They pre-aggregated data because ad-hoc querying was too complex for most business users.

Autonomous agents remove that constraint. But they replace it with a harder one: accuracy at scale.

Without a governed semantic layer, more agents means more wrong answers delivered faster. With a deterministic execution layer, more agents means more people making decisions from the same auditable, governed source of truth.

The role of data engineering shifts accordingly. Less time building ETL pipelines to move data into pre-aggregated tables for BI tools. More time curating the enterprise ontologies, defining business rules, and maintaining the semantic graph that agents rely on. The bottleneck moves from physical data movement to logical knowledge representation.

The cost structure shifts too. In raw text-to-SQL workflows, agents consume massive context windows to ingest schemas, burn reasoning tokens to guess join paths, and run expensive retry loops when queries fail1. When a semantic execution layer handles the structural planning, the agent passes a short natural language string and the compiler does the heavy computation natively, cutting both token usage per query and cost per correct answer. We cover this economics in detail in why brittle semantic layers bleed capital.

Frequently Asked Questions

What are OpenAI Dots?

Dots are persistent, always-on autonomous agents OpenAI announced at DevDay on September 29, 2026. Each Dot runs 24/7 on its own cloud computer with its own browser, connects to thousands of applications, and carries context across ChatGPT, Slack, and Microsoft Teams. Specialist Dots take dedicated organizational roles with IT-managed credentials.

Why do AI agents give wrong answers on enterprise databases?

Modern LLMs score above 86% on clean academic text-to-SQL benchmarks but collapse on real corporate schemas. On the BEAVER benchmark of private enterprise data, GPT-4o scored near 0% and the best agentic framework reached 21.3%. Enterprise schemas have thousands of tables, cryptic names, hidden business rules, and no verification, so a wrong join silently returns a confident but incorrect number.

How does a semantic execution layer fix text-to-SQL accuracy?

A deterministic semantic execution layer resolves meaning at compile time before SQL reaches the warehouse. It encodes business definitions, proves valid join paths, and enforces governance. In the dbt Labs 2026 benchmark, placing a semantic layer between the LLM and the database raised accuracy from 64.5% to 100% on modeled queries, and returned an error instead of guessing when the model was missing.

How do OpenAI agents connect to a semantic layer like Colrows?

Through the Model Context Protocol (MCP). Any agent built on OpenAI's Agents SDK, Claude, Cursor, or LlamaIndex connects to Colrows through a single MCP connector endpoint using OAuth 2.0 with PKCE. The agent discovers six read-only tools and expresses intent in natural language. It never sends raw SQL; the compiler generates governed SQL, so hallucinated queries cannot execute.

How is governance enforced for agents that run continuously?

Colrows pushes security into the compiler. Role-based and attribute-based access controls and row and column predicates are injected into the query plan before SQL is generated, so an agent never sees unauthorized rows. Redaction policies mask PII per column while keeping stable hashes for joins, and Verified Queries let humans pin immutable SQL for board-level metrics.

Do I need a semantic execution layer if I only use Snowflake?

If your whole estate is in one warehouse and your agents only query that warehouse, a warehouse-native option like Snowflake Cortex Analyst over Semantic Views can be enough. You need a multi-engine semantic execution layer once definitions, governance, and proven join paths have to hold across more than one engine, which is the common enterprise case.

Sources

  1. The Text-to-SQL Accuracy Cliff: 91% to 21% on Real Data, Medium (2026). Original academic sources: BEAVER (arXiv:2409.02038), Spider 2.0 (arXiv:2411.07763).
  2. Semantic Layer vs. Text-to-SQL: 2026 Benchmark Update, dbt Docs
  3. Introducing Dots, OpenAI (September 29, 2026)
  4. Why Text-to-SQL Accuracy Drops on Real Schemas, Colrows Blog (2026)
  5. Text-to-SQL with Knowledge Graphs: Multi-Hop Queries, FalkorDB
  6. Colrows: Semantic Execution Layer for Enterprise AI Agents
  7. Model Context Protocol, Wikipedia
  8. Model Context Protocol (MCP), OpenAI Agents SDK
  9. MCP Integration: Connect AI Agents to Colrows, Colrows Docs
  10. Redaction Policies: Mask Sensitive Values, Colrows Docs
  11. Verified Queries: Curated Trusted Examples, Colrows Docs

Your agents are only as accurate as the data layer beneath them.

Book a technical architecture review and see how Colrows makes autonomous agents query governed truth, not hallucinated guesses.