NVDA 230.86 ▲1.09%GOOGL 338.24 ▼1.70%MSFT 512.80 ▼0.02%AMD 615.73 ▲0.65%INTC 120.00 ▼0.19%TSMC 459.20 ▲0.66%AMZN 248.23 ▼0.37%META 725.93 ▲0.10%AAPL 330.32 ▼0.81%PLTR 190.04 ▲1.60%
Markets at last close

Apps

Context is becoming the reliability layer for AI BI

·1 min read

Enterprise text-to-SQL systems are moving beyond simple prompts toward architectures that narrow what models must infer. Uber’s QueryGPT and LinkedIn’s SQL Bot use multi-agent flows for intent detection, prompt enhancement, table retrieval, column pruning, generation, and correction, reflecting a core lesson: at enterprise scale, finding the right tables can matter as much as writing valid SQL.

Reported gains are promising but not equivalent to controlled accuracy benchmarks. Uber said query authoring fell from about 10 minutes to about 3 minutes, reached about 300 daily active users, and had 78% of users report less time writing queries from scratch. LinkedIn reported strong user satisfaction, while a later internal benchmark put correct or close-to-correct responses at 53%, underscoring the gap between usefulness and correctness.

Semantic layers and curated context are presented as the main reliability mechanisms. Snowflake reported GPT-4o with a single prompt at 51% on an internal evaluation and 90%+ accuracy for Cortex Analyst with a semantic model, while dbt Labs reported in-scope Semantic Layer accuracy of 100% for two 2026 models in a small vendor benchmark. A separate abstract found a 4 KB semantic document improved accuracy by +17 to +23 percentage points and made model choice within tier less important.

Pipeline generation remains less proven. ELT-Bench evaluated 100 pipelines, 835 source tables, and 203 data models, with the best setup correctly generating only 3.9% of data models. The practical takeaway is to prioritize metadata, governed metric definitions, owned evaluation sets, visible SQL, and clear refusals over unconstrained generation.

Originally reported by dev.toRead the source →
Related coverage