Text-to-ES|QL benchmark: results and analysis

BIRD Mini-Dev · 500 questions · single-shot, zero-shot, with evidence · Execution Accuracy · Elasticsearch 9.5
Scoring Lenient counts a query as correct when its returned data reproduces the exact gold rows, ignoring presentation: extra columns, column order, and columns the model space-joined that the gold kept separate.
Results Explorer
Charts
Analysis & failure modes

Execution Accuracy

modelbase+focused+full
SQL single-call reference (GPT-4, BIRD): ~48%

Failure categories

modelvariantcorrectempty_resultjoin_errorother_exec_errorsyntax_errorunknown_field_or_indexwrong_result
claude-opus-4-8no_skills163190018143157
claude-opus-4-8skills_focused1742000884214
claude-opus-4-8skills_full17026108411208
claude-sonnet-4-6no_skills81260018121191
claude-sonnet-4-6skills_focused1062901869269
claude-sonnet-4-6skills_full10729017811274
gpt-5.4-minino_skills65193023670107
gpt-5.4-miniskills_focused93242015035196
gpt-5.4-miniskills_full95222013742202
gpt-5.5no_skills228250270238
gpt-5.5skills_focused1662701790227
gpt-5.5skills_full1642302690242
Categories are the strict failure type; leniency does not recategorize.
Select a row to see the side-by-side.

The 2026 picture

Single-shot, zero-shot text-to-ES|QL on BIRD Mini-Dev (500 questions), run against Elasticsearch 9.5. The headline: GPT-5.5 writes correct ES|QL 59% of the time under lenient scoring with nothing but a schema, comfortably above what GPT-4 scored writing SQL on the same benchmark (~48%). For the best model the syntax barrier is effectively gone; what is left is semantic correctness.

Modelno skillfocused skill (~33k)full skill (~65k)
gpt-5.559.0%48.8%50.4%
claude-opus-4-840.4%50.0%49.0%
claude-sonnet-4-630.2%39.2%39.6%
gpt-5.4-mini19.8%30.4%31.6%

Lenient EX. Strict BIRD EX for the same cells: gpt-5.5 45.6 / 33.2 / 32.8, opus 32.6 / 34.8 / 34.0, sonnet 16.2 / 21.2 / 21.4, mini 13.0 / 18.6 / 19.0. Flip the switch above.

Calibration reference: single-call zero-shot text-to-SQL (GPT-4) is ~48% on the same questions. BIRD's public leaderboard tops out higher, but those entries are engineered pipelines with schema linking and self-correction, not one call with one prompt.

What "lenient" forgives. Three presentation differences, never a wrong row: extra returned columns, column order, and concatenation (the model space-joins fields the gold query kept separate, e.g. first_name, last_name returned as one full_name). That last rule recovers 88 results that were correct and scored zero on column split alone. Every runs-but-wrong candidate was re-executed in full against the live cluster, so there is no offline sample gap in these numbers.

Two findings about the skill doc

1. The skill helps weak models and hurts the strongest. Opus, Sonnet, and mini each gain between 9.0 and 10.6 lenient points from the focused skill. GPT-5.5 loses 10.2. It already knows ES|QL (7 syntax errors out of 500 at base), so it had no grammar problem for a reference to fix.

2. More reference is not better, but it is not worse either. The focused bundle (~33k tokens) and the full one (~65k) land within 0.4 to 1.6 points of each other for every model, which is noise at 500 questions. Whatever a reference is worth here is delivered by its first few pages. Curate the skill; there is nothing to gain from dumping the whole manual.
The GPT-5.5 drop is our edit, not the skill. Both bundles are pinned to agent-skills commit e0d6b02 with one local change: because this run targets 9.5, we rewrote the join-key guidance to promote the 9.2 equality predicate (ON source == lookup) over RENAME. Re-running GPT-5.5 with the focused skill exactly as published, same cluster and same 500 questions, scores 59.0%, level with its own base. So the unmodified skill costs GPT-5.5 nothing and the entire 10.2-point drop comes from that one edit: it pushed RENAME usage down from 47% of GPT-5.5's joins to 11%, and Found ambiguous reference errors up from 120 to 706 across the run while the join bucket fell from 939 to 354. Almost exactly a wash. A language feature only pays once the guidance says when not to reach for it.

Why the weaker numbers are low

The models are not bad at the logic of these questions. They trip on a handful of ES|QL-specific mechanics and on the strictness of execution-accuracy scoring. The single biggest cause is one join rule.

#1 cause: LOOKUP JOIN key names. BIRD joins tables on differently-named columns (races.id = results.race_id). ES|QL's bare LOOKUP JOIN form requires the join key to have the same name in both indices (... | LOOKUP JOIN races ON race_id). The model writes a SQL-shaped join and ES|QL rejects it: Unknown column [race_id] in right side of join (or mismatched input '=' when it also writes ON a = b). This is ~45% of Opus's 337 base failures. 9.2's equality predicate lifts the restriction, but only when the key name is unambiguous; when it exists on both sides, RENAME before the join is still the fix.

Failure taxonomy (no-skill, claude-opus-4-8)

Root cause~countWhat happens
LOOKUP JOIN same-name-key134SQL join on different key names; query won't run
Join runs but mis-collapses rows~73wrong base table / one-to-many fan-out corrupts aggregates
Missing DISTINCT24gold uses SELECT DISTINCT; prediction returns duplicates
Ratio / integer division / CAST24int/int truncates; model invents DIVIDE() (not an ES|QL function)
Value / case / literal mismatch~15exact stored value not replicated (e.g. 'east Bohemia', 'VYBER')
Invented funcs / subqueries / prose leak~27DIVIDE,YEAR, SQL subqueries, chain-of-thought emitted as query text

Field-grounding (won't execute) vs logic-but-valid is almost exactly 50/50. The grounding half throws precise errors, so an execution-feedback loop would fix most of it. The other half runs silently wrong, so a loop alone gives no signal.

The skill doc buys validity, not correctness

Adding Elastic's ES|QL reference to the prompt slashes syntax errors, but those queries then fail one stage later on schema grounding or semantics. The failure mix shifts from "doesn't run" to "runs but wrong."

categoryopus no-skillopus skillsonnet no-skillsonnet skill
syntax_error188818186
unknown_field_or_index1434219
wrong_result157214191269
empty_result19202629

"Runs-but-wrong" share of all failures climbs from 52%→72% (Opus) and 52%→76% (Sonnet). The skill all but eliminates field-grounding errors (Opus 143→4, Sonnet 21→9), and for Sonnet it halves the syntax errors its SQL habits produce (single = instead of ==, CASE WHEN, single quotes, COUNT(DISTINCT)). Opus already had the syntax mostly right, so for it the trade shows up as a rise in syntax errors instead: those are the ambiguous-reference failures the join edit introduced.

wrong_result: right idea, wrong tuple

The hardest bucket: the query executes cleanly but returns the wrong rows. Main reasons:

What would move the number

Scope: the headline table, the per-category counts, the skill findings and the GPT-5.5 control are all re-derived from this run (Elasticsearch 9.5, 6,000 scored queries, plus a 500-query control for GPT-5.5 on the unmodified skill), as is every wrong_result share marked "measured", computed over the runs-but-wrong bucket (executed cleanly, wrong data under lenient scoring, n=1,984). Where a share bounds the exposure rather than establishing the cause, both figures are given. The failure taxonomy table above is the one section still carried over from a manual read of the earlier 9.1 run; its underlying buckets barely moved between runs (wrong_result 2,414 to 2,525, empty_result 283 to 289), so the structure holds, but treat its counts as indicative.