In plain terms
We asked 3 AI “teams” the same 12 questions about a clinical database, in plain English, and let each team write and run the database queries itself — including asking us to clarify vague questions, and refusing ones the data can't answer. A team passes a question only when its final answer matches an independently written reference answer, row for row.
The teams finished within one question of each other — a practical tie at this sample size: qwen2.5-14b-checked 8 of 12 · writer-only 7 of 12 · self-checked 7 of 12.
The misses were largely shared, not team-specific: 4 of the 12 questions (B1, B3, M2, U2) were missed by every team; B2, M3 tripped only one team. A separate AI judge read every executed query. On the failed questions the SQL itself still scored at ceiling for construction and schema use — what dropped was fidelity to what the question actually asked. The errors are misreadings, not broken queries.
One conversation per question, on a demonstration dataset — read this as directional evidence, not a benchmark; differences of a question or two are within noise.
Result
Against the gates in force at publication (≥90% overall, ≥80% per scenario): no team qualified
22/36 conversations passed across all teams · 1038 assertions · pass = the final answer matches the independent reference answer row-for-row.
Judge summaryadvisory
A separate AI judge read every executed query against its recorded evidence and scored four axes from 0 to 3. The axes are what it judged, so they lead; each cell shows the median and, where the scores differ, the range behind it. Advisory means it never gates acceptance — the row-for-row reference check does. Where that check failed, the judge explains why; it cannot overrule it.
| team | queries judged | intent | SQL craft | schema | follow-up | weakest queries | composite |
|---|---|---|---|---|---|---|---|
| writer-only | 15 | 3 1–3 | 3 2–3 | 3 | 3 1–3 | B2·turn 1 — 55, B3·turn 1 — 63, M1·turn 0 — 84 | 100 ↓55 |
| self-checked | 14 | 3 1–3 | 3 | 3 | 3 1–3 | B3·turn 1 — 63, M1·turn 0 — 84, M2·turn 0 — 84 | 100 ↓63 |
| qwen2.5-14b-checked | 15 | 3 1–3 | 3 2–3 | 3 | 3 1–3 | B3·turn 1 — 63, M1·turn 0 — 84, M2·turn 0 — 84 | 100 ↓63 |
| all teams | 44 | 3 1–3 | 3 2–3 | 3 | 3 1–3 | writer-only: B2·turn 1 — 55, writer-only: B3·turn 1 — 63, qwen2.5-14b-checked: B3·turn 1 — 63 | 100 ↓55 |
Judge: claude-fable-5 (anthropic) · three independent passes finalized by per-axis medians · rubric 5dc94bbd3042…
- intent
- did the SQL answer the question as asked (0–3)
- SQL craft
- clean, minimal, executable construction — joins, predicates, no dead branches (0–3)
- schema
- only catalogued tables and columns, parameters bound with correct types (0–3)
- follow-up
- the revision honors the new instruction without breaking what worked (0–3, follow-up turns only)
- weakest queries
- every query the judge marked down on any axis, worst first (at most three shown) — the rubric's own anchors decide, not a composite cutoff. Blank when nothing was marked down.
- composite
- a weighted convenience score (0–100; opening queries 47/29/24, follow-ups 40/25/20/15) with the team's floor after ↓. Read the axes first: when they saturate, the composite is nearly a step function of whichever axis dropped.
Scenario matrix
| scenario | reported | gold ok | judge by turn (advisory) | assertions | gold FAIL detail |
|---|---|---|---|---|---|
catalyst-query-gemma-4-12b · A1single-ready | PASS | yes | 100 | 13/13 passedsession_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, token_budget_respected-base, base_writer_outcome, base_version_available, base_validation_recorded, base_execution_succeeded, base_postgres_crosscheck, base_gold_execution_match, successor_visible_under_three_minutes, new_session_isolation | — |
catalyst-query-gemma-4-12b · A2single-ready | PASS | yes | 100 | 13/13 passedsession_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, token_budget_respected-base, base_writer_outcome, base_version_available, base_validation_recorded, base_execution_succeeded, base_postgres_crosscheck, base_gold_execution_match, successor_visible_under_three_minutes, new_session_isolation | — |
catalyst-query-gemma-4-12b · A3single-ready | PASS | yes | 100 | 13/13 passedsession_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, token_budget_respected-base, base_writer_outcome, base_version_available, base_validation_recorded, base_execution_succeeded, base_postgres_crosscheck, base_gold_execution_match, successor_visible_under_three_minutes, new_session_isolation | — |
catalyst-query-gemma-4-12b · A4single-ready | PASS | yes | 100 | 13/13 passedsession_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, token_budget_respected-base, base_writer_outcome, base_version_available, base_validation_recorded, base_execution_succeeded, base_postgres_crosscheck, base_gold_execution_match, successor_visible_under_three_minutes, new_session_isolation | — |
catalyst-query-gemma-4-12b · M1multi-turn | PASS | yes | 84 · 100 · 100 | 60/60 passedsession_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, token_budget_respected-base, base_writer_outcome, base_version_available, base_validation_recorded, base_execution_succeeded, base_postgres_crosscheck, followup_http_created, followup_terminal_status, writer_outcome, base_classification, manual_version_classification, followup_profile, refresh_restored, timeline_current_turn, followup_evidence_available, invocation_duration_sum, invocation_digests, invocation_timestamp_reconciliation, writer_model, reviewer_model, effective_temperature_and_dry, revision_context_exclusions, token_evidence_recorded, token_budget_respected, exact_selected_output, semantic_reviewer_correction, prior_results_stale_after_successor, successor_validation_recorded, successor_execution_succeeded, successor_postgres_crosscheck, followup_http_created-t2, followup_terminal_status-t2, writer_outcome-t2, base_classification-t2, manual_version_classification-t2, followup_profile-t2, refresh_restored-t2, timeline_current_turn-t2, followup_evidence_available-t2, invocation_duration_sum-t2, invocation_digests-t2, invocation_timestamp_reconciliation-t2, writer_model-t2, reviewer_model-t2, effective_temperature_and_dry-t2, revision_context_exclusions-t2, token_evidence_recorded-t2, token_budget_respected-t2, exact_selected_output-t2, semantic_reviewer_correction-t2, prior_results_stale_after_successor-t2, successor_validation_recorded-t2, successor_execution_succeeded-t2, successor_postgres_crosscheck-t2, successor_visible_under_three_minutes-t2, new_session_isolation | — |
catalyst-query-gemma-4-12b · M2retained-guidance | FAIL | no | 84 · 100 · 100 | 62/63 passedsession_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, token_budget_respected-base, base_writer_outcome, base_version_available, base_validation_recorded, base_execution_succeeded, base_postgres_crosscheck, guidance_pinned, followup_http_created, followup_terminal_status, writer_outcome, base_classification, manual_version_classification, followup_profile, refresh_restored, timeline_current_turn, followup_evidence_available, invocation_duration_sum, invocation_digests, invocation_timestamp_reconciliation, writer_model, reviewer_model, effective_temperature_and_dry, revision_context_exclusions, token_evidence_recorded, token_budget_respected, exact_selected_output, semantic_reviewer_correction, prior_results_stale_after_successor, successor_validation_recorded, successor_execution_succeeded, successor_postgres_crosscheck, successor_gold_execution_match, followup_http_created-t2, followup_terminal_status-t2, writer_outcome-t2, base_classification-t2, manual_version_classification-t2, followup_profile-t2, refresh_restored-t2, timeline_current_turn-t2, followup_evidence_available-t2, invocation_duration_sum-t2, invocation_digests-t2, invocation_timestamp_reconciliation-t2, writer_model-t2, reviewer_model-t2, effective_temperature_and_dry-t2, revision_context_exclusions-t2, token_evidence_recorded-t2, token_budget_respected-t2, exact_selected_output-t2, semantic_reviewer_correction-t2, prior_results_stale_after_successor-t2, successor_validation_recorded-t2, successor_execution_succeeded-t2, successor_postgres_crosscheck-t2, successor_gold_execution_match-t2, successor_visible_under_three_minutes-t2, new_session_isolation | judged · successor_gold_execution_match scenarios/catalyst-query-gemma-4-12b/M2/repetition-01/16-gold-execution-match-successor.json |
catalyst-query-gemma-4-12b · M3verified-examples | PASS | yes | 100 · 100 · 100 | 60/60 passedsession_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, token_budget_respected-base, base_writer_outcome, base_version_available, base_validation_recorded, base_execution_succeeded, base_postgres_crosscheck, followup_http_created, followup_terminal_status, writer_outcome, base_classification, manual_version_classification, followup_profile, refresh_restored, timeline_current_turn, followup_evidence_available, invocation_duration_sum, invocation_digests, invocation_timestamp_reconciliation, writer_model, reviewer_model, effective_temperature_and_dry, revision_context_exclusions, token_evidence_recorded, token_budget_respected, exact_selected_output, semantic_reviewer_correction, prior_results_stale_after_successor, successor_validation_recorded, successor_execution_succeeded, successor_postgres_crosscheck, followup_http_created-t2, followup_terminal_status-t2, writer_outcome-t2, base_classification-t2, manual_version_classification-t2, followup_profile-t2, refresh_restored-t2, timeline_current_turn-t2, followup_evidence_available-t2, invocation_duration_sum-t2, invocation_digests-t2, invocation_timestamp_reconciliation-t2, writer_model-t2, reviewer_model-t2, effective_temperature_and_dry-t2, revision_context_exclusions-t2, token_evidence_recorded-t2, token_budget_respected-t2, exact_selected_output-t2, semantic_reviewer_correction-t2, prior_results_stale_after_successor-t2, successor_validation_recorded-t2, successor_execution_succeeded-t2, successor_postgres_crosscheck-t2, successor_visible_under_three_minutes-t2, new_session_isolation | — |
catalyst-query-gemma-4-12b · B1clarification | FAIL | no | — | 25/28 passedsession_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, base_writer_outcome, no_sql_after_non_ready_base, followup_http_created, followup_terminal_status, writer_outcome, base_classification, manual_version_classification, followup_profile, refresh_restored, timeline_current_turn, followup_evidence_available, invocation_duration_sum, invocation_digests, invocation_timestamp_reconciliation, writer_model, reviewer_model, effective_temperature_and_dry, revision_context_exclusions, token_evidence_recorded, token_budget_respected, exact_selected_output, semantic_reviewer_correction, successor_visible_under_three_minutes, new_session_isolation | judged · followup_terminal_status judged · writer_outcome judged · exact_selected_output |
catalyst-query-gemma-4-12b · B2clarification | FAIL | no | 55 | 32/33 passedsession_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, token_budget_respected-base, base_writer_outcome, no_sql_after_non_ready_base, followup_http_created, followup_terminal_status, writer_outcome, base_classification, manual_version_classification, followup_profile, refresh_restored, timeline_current_turn, followup_evidence_available, invocation_duration_sum, invocation_digests, invocation_timestamp_reconciliation, writer_model, reviewer_model, effective_temperature_and_dry, revision_context_exclusions, token_evidence_recorded, token_budget_respected, exact_selected_output, semantic_reviewer_correction, successor_validation_recorded, successor_execution_succeeded, successor_postgres_crosscheck, successor_gold_execution_match, successor_visible_under_three_minutes, new_session_isolation | judged · successor_gold_execution_match scenarios/catalyst-query-gemma-4-12b/B2/repetition-01/16-gold-execution-match-successor.json |
catalyst-query-gemma-4-12b · B3clarification | FAIL | no | 63 | 32/33 passedsession_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, token_budget_respected-base, base_writer_outcome, no_sql_after_non_ready_base, followup_http_created, followup_terminal_status, writer_outcome, base_classification, manual_version_classification, followup_profile, refresh_restored, timeline_current_turn, followup_evidence_available, invocation_duration_sum, invocation_digests, invocation_timestamp_reconciliation, writer_model, reviewer_model, effective_temperature_and_dry, revision_context_exclusions, token_evidence_recorded, token_budget_respected, exact_selected_output, semantic_reviewer_correction, successor_validation_recorded, successor_execution_succeeded, successor_postgres_crosscheck, successor_gold_execution_match, successor_visible_under_three_minutes, new_session_isolation | judged · successor_gold_execution_match scenarios/catalyst-query-gemma-4-12b/B3/repetition-01/16-gold-execution-match-successor.json |
catalyst-query-gemma-4-12b · U1unsupported | PASS | yes | — | 9/9 passedsession_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, token_budget_respected-base, base_writer_outcome, no_sql_after_non_ready_base, successor_visible_under_three_minutes, new_session_isolation | — |
catalyst-query-gemma-4-12b · U2unsupported | FAIL | no | — | 8/9 passedsession_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, token_budget_respected-base, base_writer_outcome, base_version_available, successor_visible_under_three_minutes, new_session_isolation | judged · base_writer_outcome |
catalyst-query-gemma-4-12b-q4-checked · A1single-ready | PASS | yes | 100 | 13/13 passedsession_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, token_budget_respected-base, base_writer_outcome, base_version_available, base_validation_recorded, base_execution_succeeded, base_postgres_crosscheck, base_gold_execution_match, successor_visible_under_three_minutes, new_session_isolation | — |
catalyst-query-gemma-4-12b-q4-checked · A2single-ready | PASS | yes | 100 | 13/13 passedsession_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, token_budget_respected-base, base_writer_outcome, base_version_available, base_validation_recorded, base_execution_succeeded, base_postgres_crosscheck, base_gold_execution_match, successor_visible_under_three_minutes, new_session_isolation | — |
catalyst-query-gemma-4-12b-q4-checked · A3single-ready | PASS | yes | 100 | 13/13 passedsession_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, token_budget_respected-base, base_writer_outcome, base_version_available, base_validation_recorded, base_execution_succeeded, base_postgres_crosscheck, base_gold_execution_match, successor_visible_under_three_minutes, new_session_isolation | — |
catalyst-query-gemma-4-12b-q4-checked · A4single-ready | PASS | yes | 100 | 13/13 passedsession_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, token_budget_respected-base, base_writer_outcome, base_version_available, base_validation_recorded, base_execution_succeeded, base_postgres_crosscheck, base_gold_execution_match, successor_visible_under_three_minutes, new_session_isolation | — |
catalyst-query-gemma-4-12b-q4-checked · M1multi-turn | PASS | yes | 84 · 100 · 100 | 60/60 passedsession_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, token_budget_respected-base, base_writer_outcome, base_version_available, base_validation_recorded, base_execution_succeeded, base_postgres_crosscheck, followup_http_created, followup_terminal_status, writer_outcome, base_classification, manual_version_classification, followup_profile, refresh_restored, timeline_current_turn, followup_evidence_available, invocation_duration_sum, invocation_digests, invocation_timestamp_reconciliation, writer_model, reviewer_model, effective_temperature_and_dry, revision_context_exclusions, token_evidence_recorded, token_budget_respected, exact_selected_output, semantic_reviewer_correction, prior_results_stale_after_successor, successor_validation_recorded, successor_execution_succeeded, successor_postgres_crosscheck, followup_http_created-t2, followup_terminal_status-t2, writer_outcome-t2, base_classification-t2, manual_version_classification-t2, followup_profile-t2, refresh_restored-t2, timeline_current_turn-t2, followup_evidence_available-t2, invocation_duration_sum-t2, invocation_digests-t2, invocation_timestamp_reconciliation-t2, writer_model-t2, reviewer_model-t2, effective_temperature_and_dry-t2, revision_context_exclusions-t2, token_evidence_recorded-t2, token_budget_respected-t2, exact_selected_output-t2, semantic_reviewer_correction-t2, prior_results_stale_after_successor-t2, successor_validation_recorded-t2, successor_execution_succeeded-t2, successor_postgres_crosscheck-t2, successor_visible_under_three_minutes-t2, new_session_isolation | — |
catalyst-query-gemma-4-12b-q4-checked · M2retained-guidance | FAIL | no | 84 · 100 · 100 | 62/63 passedsession_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, token_budget_respected-base, base_writer_outcome, base_version_available, base_validation_recorded, base_execution_succeeded, base_postgres_crosscheck, guidance_pinned, followup_http_created, followup_terminal_status, writer_outcome, base_classification, manual_version_classification, followup_profile, refresh_restored, timeline_current_turn, followup_evidence_available, invocation_duration_sum, invocation_digests, invocation_timestamp_reconciliation, writer_model, reviewer_model, effective_temperature_and_dry, revision_context_exclusions, token_evidence_recorded, token_budget_respected, exact_selected_output, semantic_reviewer_correction, prior_results_stale_after_successor, successor_validation_recorded, successor_execution_succeeded, successor_postgres_crosscheck, successor_gold_execution_match, followup_http_created-t2, followup_terminal_status-t2, writer_outcome-t2, base_classification-t2, manual_version_classification-t2, followup_profile-t2, refresh_restored-t2, timeline_current_turn-t2, followup_evidence_available-t2, invocation_duration_sum-t2, invocation_digests-t2, invocation_timestamp_reconciliation-t2, writer_model-t2, reviewer_model-t2, effective_temperature_and_dry-t2, revision_context_exclusions-t2, token_evidence_recorded-t2, token_budget_respected-t2, exact_selected_output-t2, semantic_reviewer_correction-t2, prior_results_stale_after_successor-t2, successor_validation_recorded-t2, successor_execution_succeeded-t2, successor_postgres_crosscheck-t2, successor_gold_execution_match-t2, successor_visible_under_three_minutes-t2, new_session_isolation | judged · successor_gold_execution_match scenarios/catalyst-query-gemma-4-12b-q4-checked/M2/repetition-01/16-gold-execution-match-successor.json |
catalyst-query-gemma-4-12b-q4-checked · M3verified-examples | FAIL | no | 100 · 100 | 53/57 passedsession_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, token_budget_respected-base, base_writer_outcome, base_version_available, base_validation_recorded, base_execution_succeeded, base_postgres_crosscheck, followup_http_created, followup_terminal_status, writer_outcome, base_classification, manual_version_classification, followup_profile, refresh_restored, timeline_current_turn, followup_evidence_available, invocation_duration_sum, invocation_digests, invocation_timestamp_reconciliation, writer_model, reviewer_model, effective_temperature_and_dry, revision_context_exclusions, token_evidence_recorded, token_budget_respected, exact_selected_output, semantic_reviewer_correction, prior_results_stale_after_successor, followup_http_created-t2, followup_terminal_status-t2, writer_outcome-t2, base_classification-t2, manual_version_classification-t2, followup_profile-t2, refresh_restored-t2, timeline_current_turn-t2, followup_evidence_available-t2, invocation_duration_sum-t2, invocation_digests-t2, invocation_timestamp_reconciliation-t2, writer_model-t2, reviewer_model-t2, effective_temperature_and_dry-t2, revision_context_exclusions-t2, token_evidence_recorded-t2, token_budget_respected-t2, exact_selected_output-t2, semantic_reviewer_correction-t2, prior_results_stale_after_successor-t2, successor_validation_recorded-t2, successor_execution_succeeded-t2, successor_postgres_crosscheck-t2, successor_visible_under_three_minutes-t2, new_session_isolation | judged · followup_terminal_status judged · writer_outcome judged · exact_selected_output judged · prior_results_stale_after_successor |
catalyst-query-gemma-4-12b-q4-checked · B1clarification | FAIL | no | — | 25/28 passedsession_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, base_writer_outcome, no_sql_after_non_ready_base, followup_http_created, followup_terminal_status, writer_outcome, base_classification, manual_version_classification, followup_profile, refresh_restored, timeline_current_turn, followup_evidence_available, invocation_duration_sum, invocation_digests, invocation_timestamp_reconciliation, writer_model, reviewer_model, effective_temperature_and_dry, revision_context_exclusions, token_evidence_recorded, token_budget_respected, exact_selected_output, semantic_reviewer_correction, successor_visible_under_three_minutes, new_session_isolation | judged · followup_terminal_status judged · writer_outcome judged · exact_selected_output |
catalyst-query-gemma-4-12b-q4-checked · B2clarification | PASS | yes | 87 | 33/33 passedsession_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, token_budget_respected-base, base_writer_outcome, no_sql_after_non_ready_base, followup_http_created, followup_terminal_status, writer_outcome, base_classification, manual_version_classification, followup_profile, refresh_restored, timeline_current_turn, followup_evidence_available, invocation_duration_sum, invocation_digests, invocation_timestamp_reconciliation, writer_model, reviewer_model, effective_temperature_and_dry, revision_context_exclusions, token_evidence_recorded, token_budget_respected, exact_selected_output, semantic_reviewer_correction, successor_validation_recorded, successor_execution_succeeded, successor_postgres_crosscheck, successor_gold_execution_match, successor_visible_under_three_minutes, new_session_isolation | — |
catalyst-query-gemma-4-12b-q4-checked · B3clarification | FAIL | no | 63 | 32/33 passedsession_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, token_budget_respected-base, base_writer_outcome, no_sql_after_non_ready_base, followup_http_created, followup_terminal_status, writer_outcome, base_classification, manual_version_classification, followup_profile, refresh_restored, timeline_current_turn, followup_evidence_available, invocation_duration_sum, invocation_digests, invocation_timestamp_reconciliation, writer_model, reviewer_model, effective_temperature_and_dry, revision_context_exclusions, token_evidence_recorded, token_budget_respected, exact_selected_output, semantic_reviewer_correction, successor_validation_recorded, successor_execution_succeeded, successor_postgres_crosscheck, successor_gold_execution_match, successor_visible_under_three_minutes, new_session_isolation | judged · successor_gold_execution_match scenarios/catalyst-query-gemma-4-12b-q4-checked/B3/repetition-01/16-gold-execution-match-successor.json |
catalyst-query-gemma-4-12b-q4-checked · U1unsupported | PASS | yes | — | 9/9 passedsession_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, token_budget_respected-base, base_writer_outcome, no_sql_after_non_ready_base, successor_visible_under_three_minutes, new_session_isolation | — |
catalyst-query-gemma-4-12b-q4-checked · U2unsupported | FAIL | no | — | 8/9 passedsession_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, token_budget_respected-base, base_writer_outcome, base_version_available, successor_visible_under_three_minutes, new_session_isolation | judged · base_writer_outcome |
catalyst-query-gemma-4-12b-qwen2.5-14b-checked · A1single-ready | PASS | yes | 100 | 13/13 passedsession_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, token_budget_respected-base, base_writer_outcome, base_version_available, base_validation_recorded, base_execution_succeeded, base_postgres_crosscheck, base_gold_execution_match, successor_visible_under_three_minutes, new_session_isolation | — |
catalyst-query-gemma-4-12b-qwen2.5-14b-checked · A2single-ready | PASS | yes | 100 | 13/13 passedsession_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, token_budget_respected-base, base_writer_outcome, base_version_available, base_validation_recorded, base_execution_succeeded, base_postgres_crosscheck, base_gold_execution_match, successor_visible_under_three_minutes, new_session_isolation | — |
catalyst-query-gemma-4-12b-qwen2.5-14b-checked · A3single-ready | PASS | yes | 100 | 13/13 passedsession_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, token_budget_respected-base, base_writer_outcome, base_version_available, base_validation_recorded, base_execution_succeeded, base_postgres_crosscheck, base_gold_execution_match, successor_visible_under_three_minutes, new_session_isolation | — |
catalyst-query-gemma-4-12b-qwen2.5-14b-checked · A4single-ready | PASS | yes | 90 | 13/13 passedsession_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, token_budget_respected-base, base_writer_outcome, base_version_available, base_validation_recorded, base_execution_succeeded, base_postgres_crosscheck, base_gold_execution_match, successor_visible_under_three_minutes, new_session_isolation | — |
catalyst-query-gemma-4-12b-qwen2.5-14b-checked · M1multi-turn | PASS | yes | 84 · 100 · 100 | 60/60 passedsession_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, token_budget_respected-base, base_writer_outcome, base_version_available, base_validation_recorded, base_execution_succeeded, base_postgres_crosscheck, followup_http_created, followup_terminal_status, writer_outcome, base_classification, manual_version_classification, followup_profile, refresh_restored, timeline_current_turn, followup_evidence_available, invocation_duration_sum, invocation_digests, invocation_timestamp_reconciliation, writer_model, reviewer_model, effective_temperature_and_dry, revision_context_exclusions, token_evidence_recorded, token_budget_respected, exact_selected_output, semantic_reviewer_correction, prior_results_stale_after_successor, successor_validation_recorded, successor_execution_succeeded, successor_postgres_crosscheck, followup_http_created-t2, followup_terminal_status-t2, writer_outcome-t2, base_classification-t2, manual_version_classification-t2, followup_profile-t2, refresh_restored-t2, timeline_current_turn-t2, followup_evidence_available-t2, invocation_duration_sum-t2, invocation_digests-t2, invocation_timestamp_reconciliation-t2, writer_model-t2, reviewer_model-t2, effective_temperature_and_dry-t2, revision_context_exclusions-t2, token_evidence_recorded-t2, token_budget_respected-t2, exact_selected_output-t2, semantic_reviewer_correction-t2, prior_results_stale_after_successor-t2, successor_validation_recorded-t2, successor_execution_succeeded-t2, successor_postgres_crosscheck-t2, successor_visible_under_three_minutes-t2, new_session_isolation | — |
catalyst-query-gemma-4-12b-qwen2.5-14b-checked · M2retained-guidance | FAIL | no | 84 · 100 · 100 | 62/63 passedsession_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, token_budget_respected-base, base_writer_outcome, base_version_available, base_validation_recorded, base_execution_succeeded, base_postgres_crosscheck, guidance_pinned, followup_http_created, followup_terminal_status, writer_outcome, base_classification, manual_version_classification, followup_profile, refresh_restored, timeline_current_turn, followup_evidence_available, invocation_duration_sum, invocation_digests, invocation_timestamp_reconciliation, writer_model, reviewer_model, effective_temperature_and_dry, revision_context_exclusions, token_evidence_recorded, token_budget_respected, exact_selected_output, semantic_reviewer_correction, prior_results_stale_after_successor, successor_validation_recorded, successor_execution_succeeded, successor_postgres_crosscheck, successor_gold_execution_match, followup_http_created-t2, followup_terminal_status-t2, writer_outcome-t2, base_classification-t2, manual_version_classification-t2, followup_profile-t2, refresh_restored-t2, timeline_current_turn-t2, followup_evidence_available-t2, invocation_duration_sum-t2, invocation_digests-t2, invocation_timestamp_reconciliation-t2, writer_model-t2, reviewer_model-t2, effective_temperature_and_dry-t2, revision_context_exclusions-t2, token_evidence_recorded-t2, token_budget_respected-t2, exact_selected_output-t2, semantic_reviewer_correction-t2, prior_results_stale_after_successor-t2, successor_validation_recorded-t2, successor_execution_succeeded-t2, successor_postgres_crosscheck-t2, successor_gold_execution_match-t2, successor_visible_under_three_minutes-t2, new_session_isolation | judged · successor_gold_execution_match scenarios/catalyst-query-gemma-4-12b-qwen2.5-14b-checked/M2/repetition-01/16-gold-execution-match-successor.json |
catalyst-query-gemma-4-12b-qwen2.5-14b-checked · M3verified-examples | PASS | yes | 100 · 100 · 100 | 60/60 passedsession_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, token_budget_respected-base, base_writer_outcome, base_version_available, base_validation_recorded, base_execution_succeeded, base_postgres_crosscheck, followup_http_created, followup_terminal_status, writer_outcome, base_classification, manual_version_classification, followup_profile, refresh_restored, timeline_current_turn, followup_evidence_available, invocation_duration_sum, invocation_digests, invocation_timestamp_reconciliation, writer_model, reviewer_model, effective_temperature_and_dry, revision_context_exclusions, token_evidence_recorded, token_budget_respected, exact_selected_output, semantic_reviewer_correction, prior_results_stale_after_successor, successor_validation_recorded, successor_execution_succeeded, successor_postgres_crosscheck, followup_http_created-t2, followup_terminal_status-t2, writer_outcome-t2, base_classification-t2, manual_version_classification-t2, followup_profile-t2, refresh_restored-t2, timeline_current_turn-t2, followup_evidence_available-t2, invocation_duration_sum-t2, invocation_digests-t2, invocation_timestamp_reconciliation-t2, writer_model-t2, reviewer_model-t2, effective_temperature_and_dry-t2, revision_context_exclusions-t2, token_evidence_recorded-t2, token_budget_respected-t2, exact_selected_output-t2, semantic_reviewer_correction-t2, prior_results_stale_after_successor-t2, successor_validation_recorded-t2, successor_execution_succeeded-t2, successor_postgres_crosscheck-t2, successor_visible_under_three_minutes-t2, new_session_isolation | — |
catalyst-query-gemma-4-12b-qwen2.5-14b-checked · B1clarification | FAIL | no | — | 25/28 passedsession_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, base_writer_outcome, no_sql_after_non_ready_base, followup_http_created, followup_terminal_status, writer_outcome, base_classification, manual_version_classification, followup_profile, refresh_restored, timeline_current_turn, followup_evidence_available, invocation_duration_sum, invocation_digests, invocation_timestamp_reconciliation, writer_model, reviewer_model, effective_temperature_and_dry, revision_context_exclusions, token_evidence_recorded, token_budget_respected, exact_selected_output, semantic_reviewer_correction, successor_visible_under_three_minutes, new_session_isolation | judged · followup_terminal_status judged · writer_outcome judged · exact_selected_output |
catalyst-query-gemma-4-12b-qwen2.5-14b-checked · B2clarification | PASS | yes | 100 | 33/33 passedsession_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, token_budget_respected-base, base_writer_outcome, no_sql_after_non_ready_base, followup_http_created, followup_terminal_status, writer_outcome, base_classification, manual_version_classification, followup_profile, refresh_restored, timeline_current_turn, followup_evidence_available, invocation_duration_sum, invocation_digests, invocation_timestamp_reconciliation, writer_model, reviewer_model, effective_temperature_and_dry, revision_context_exclusions, token_evidence_recorded, token_budget_respected, exact_selected_output, semantic_reviewer_correction, successor_validation_recorded, successor_execution_succeeded, successor_postgres_crosscheck, successor_gold_execution_match, successor_visible_under_three_minutes, new_session_isolation | — |
catalyst-query-gemma-4-12b-qwen2.5-14b-checked · B3clarification | FAIL | no | 63 | 32/33 passedsession_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, token_budget_respected-base, base_writer_outcome, no_sql_after_non_ready_base, followup_http_created, followup_terminal_status, writer_outcome, base_classification, manual_version_classification, followup_profile, refresh_restored, timeline_current_turn, followup_evidence_available, invocation_duration_sum, invocation_digests, invocation_timestamp_reconciliation, writer_model, reviewer_model, effective_temperature_and_dry, revision_context_exclusions, token_evidence_recorded, token_budget_respected, exact_selected_output, semantic_reviewer_correction, successor_validation_recorded, successor_execution_succeeded, successor_postgres_crosscheck, successor_gold_execution_match, successor_visible_under_three_minutes, new_session_isolation | judged · successor_gold_execution_match scenarios/catalyst-query-gemma-4-12b-qwen2.5-14b-checked/B3/repetition-01/16-gold-execution-match-successor.json |
catalyst-query-gemma-4-12b-qwen2.5-14b-checked · U1unsupported | PASS | yes | — | 9/9 passedsession_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, token_budget_respected-base, base_writer_outcome, no_sql_after_non_ready_base, successor_visible_under_three_minutes, new_session_isolation | — |
catalyst-query-gemma-4-12b-qwen2.5-14b-checked · U2unsupported | FAIL | no | — | 8/9 passedsession_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, token_budget_respected-base, base_writer_outcome, base_version_available, successor_visible_under_three_minutes, new_session_isolation | judged · base_writer_outcome |
What each scenario does
catalyst-query-gemma-4-12b · A1
Question — List CD4 count results since 2026-02-01 with the patient, the value, the unit, and the observed date.
1 run — all passed · base query returned 11 rows
catalyst-query-gemma-4-12b · A2
Question — Count HIV visits by encounter type since 2025-01-01, highest count first.
1 run — all passed · base query returned 5 rows
catalyst-query-gemma-4-12b · A3
Question — Count medication requests for female patients by medication name, excluding do_not_perform, highest count first.
1 run — all passed · base query returned 30 rows
catalyst-query-gemma-4-12b · A4
Question — List each OpenMRS-native concept with no CIEL mapping, its name, and its total observation count, highest count first.
1 run — all passed · base query returned 2 rows
catalyst-query-gemma-4-12b · M1
Question — I need to get the list of medications prescribed for each patient, with the patient id and patient last name in the result.
Turn 2 — “The query needs to be updated - gender is labeled as last name, and the result should be collapsed per patient - so one row per patient, with comma-separated unique medications.” → ready
Turn 3 — “We need to join in the patient table by id to this query, so we can get patient last name” → ready
1 run — all passed · base query returned 100 rows · refined query returned 100 rows
catalyst-query-gemma-4-12b · M2
Question — Count medication requests by medication name.
Turn 2 — “Regroup that by patient gender as well as medication name.” → ready
Turn 3 — “Return only the ten highest medication-and-gender groups.” → ready
1 run — FAILURES present · base query returned 30 rows · refined query returned 59 rows
catalyst-query-gemma-4-12b · M3
Question — Count CD4 count results by patient.
Turn 2 — “Now count CD4 percentage results by patient instead.” → ready
Turn 3 — “Separately, count HIV visits by encounter type.” → ready
1 run — all passed · base query returned 100 rows · refined query returned 100 rows
catalyst-query-gemma-4-12b · B1
Question — Show recent HIV results.
Writer — 'HIV' is not one of the recorded result names. Did you mean Current WHO HIV stage, HIV viral load, or something else?
Turn 2 — “The last 90 days, and only CD4 count, CD4 percentage, and HIV viral load.” → rejected: The named analyte 'CD4 count' is not bound as a parameter. Add a predicate on concept_name using a named parameter and bind its string value exactly as 'CD4 count'. (1 more finding)
1 run — FAILURES present
catalyst-query-gemma-4-12b · B2
Question — Show patients with poor adherence.
Writer — What is the threshold or specific category for 'poor adherence' (e.g., a specific value for 'Overall drug adherence in last month' or 'Antiretroviral adherence in past week')?
Turn 2 — “Poor adherence means the patient's latest antiretroviral adherence result is anything other than 'All'.” → ready
1 run — FAILURES present · refined query returned 100 rows
catalyst-query-gemma-4-12b · B3
Question — Show patients overdue for follow-up.
Writer — What is the specific time window or number of days for a follow-up to be considered 'overdue'?
Turn 2 — “As of 2026-03-01, overdue means a recorded return visit date before that date with no later visit.” → ready
1 run — FAILURES present · refined query returned 100 rows
catalyst-query-gemma-4-12b · U1
Question — Show each patient's home address.
Writer — The catalog does not contain any fields or views that record a patient's home address.
1 run — all passed
catalyst-query-gemma-4-12b · U2
Question — Show the prescribing clinician's name for every medication request.
1 run — FAILURES present
catalyst-query-gemma-4-12b-q4-checked · A1
Question — List CD4 count results since 2026-02-01 with the patient, the value, the unit, and the observed date.
1 run — all passed · base query returned 11 rows
catalyst-query-gemma-4-12b-q4-checked · A2
Question — Count HIV visits by encounter type since 2025-01-01, highest count first.
1 run — all passed · base query returned 5 rows
catalyst-query-gemma-4-12b-q4-checked · A3
Question — Count medication requests for female patients by medication name, excluding do_not_perform, highest count first.
1 run — all passed · base query returned 30 rows
catalyst-query-gemma-4-12b-q4-checked · A4
Question — List each OpenMRS-native concept with no CIEL mapping, its name, and its total observation count, highest count first.
1 run — all passed · base query returned 2 rows
catalyst-query-gemma-4-12b-q4-checked · M1
Question — I need to get the list of medications prescribed for each patient, with the patient id and patient last name in the result.
Turn 2 — “The query needs to be updated - gender is labeled as last name, and the result should be collapsed per patient - so one row per patient, with comma-separated unique medications.” → ready
Turn 3 — “We need to join in the patient table by id to this query, so we can get patient last name” → ready
1 run — all passed · base query returned 100 rows · refined query returned 100 rows
catalyst-query-gemma-4-12b-q4-checked · M2
Question — Count medication requests by medication name.
Turn 2 — “Regroup that by patient gender as well as medication name.” → ready
Turn 3 — “Return only the ten highest medication-and-gender groups.” → ready
1 run — FAILURES present · base query returned 30 rows · refined query returned 59 rows
catalyst-query-gemma-4-12b-q4-checked · M3
Question — Count CD4 count results by patient.
Turn 2 — “Now count CD4 percentage results by patient instead.” → rejected: Query review failed: repaired query did not pass independent re-review
Turn 3 — “Separately, count HIV visits by encounter type.” → ready
1 run — FAILURES present · base query returned 100 rows
catalyst-query-gemma-4-12b-q4-checked · B1
Question — Show recent HIV results.
Writer — 'HIV' is not one of the recorded result names. Did you mean Current WHO HIV stage, HIV viral load, or something else?
Turn 2 — “The last 90 days, and only CD4 count, CD4 percentage, and HIV viral load.” → rejected: The named analyte 'CD4 count' is not bound as a parameter. Add a predicate on concept_name using a named parameter and bind its string value exactly as 'CD4 count'. (1 more finding)
1 run — FAILURES present
catalyst-query-gemma-4-12b-q4-checked · B2
Question — Show patients with poor adherence.
Writer — What is the threshold or specific category for 'poor adherence' (e.g., a specific value for 'Overall drug adherence in last month' or 'Antiretroviral adherence in past week')?
Turn 2 — “Poor adherence means the patient's latest antiretroviral adherence result is anything other than 'All'.” → ready
1 run — all passed · refined query returned 100 rows
catalyst-query-gemma-4-12b-q4-checked · B3
Question — Show patients overdue for follow-up.
Writer — What is the specific time window or number of days for a follow-up to be considered 'overdue'?
Turn 2 — “As of 2026-03-01, overdue means a recorded return visit date before that date with no later visit.” → ready
1 run — FAILURES present · refined query returned 100 rows
catalyst-query-gemma-4-12b-q4-checked · U1
Question — Show each patient's home address.
Writer — The catalog does not contain any fields or views that record a patient's home address.
1 run — all passed
catalyst-query-gemma-4-12b-q4-checked · U2
Question — Show the prescribing clinician's name for every medication request.
1 run — FAILURES present
catalyst-query-gemma-4-12b-qwen2.5-14b-checked · A1
Question — List CD4 count results since 2026-02-01 with the patient, the value, the unit, and the observed date.
1 run — all passed · base query returned 11 rows
catalyst-query-gemma-4-12b-qwen2.5-14b-checked · A2
Question — Count HIV visits by encounter type since 2025-01-01, highest count first.
1 run — all passed · base query returned 5 rows
catalyst-query-gemma-4-12b-qwen2.5-14b-checked · A3
Question — Count medication requests for female patients by medication name, excluding do_not_perform, highest count first.
1 run — all passed · base query returned 30 rows
catalyst-query-gemma-4-12b-qwen2.5-14b-checked · A4
Question — List each OpenMRS-native concept with no CIEL mapping, its name, and its total observation count, highest count first.
1 run — all passed · base query returned 2 rows
catalyst-query-gemma-4-12b-qwen2.5-14b-checked · M1
Question — I need to get the list of medications prescribed for each patient, with the patient id and patient last name in the result.
Turn 2 — “The query needs to be updated - gender is labeled as last name, and the result should be collapsed per patient - so one row per patient, with comma-separated unique medications.” → ready
Turn 3 — “We need to join in the patient table by id to this query, so we can get patient last name” → ready
1 run — all passed · base query returned 100 rows · refined query returned 100 rows
catalyst-query-gemma-4-12b-qwen2.5-14b-checked · M2
Question — Count medication requests by medication name.
Turn 2 — “Regroup that by patient gender as well as medication name.” → ready
Turn 3 — “Return only the ten highest medication-and-gender groups.” → ready
1 run — FAILURES present · base query returned 30 rows · refined query returned 59 rows
catalyst-query-gemma-4-12b-qwen2.5-14b-checked · M3
Question — Count CD4 count results by patient.
Turn 2 — “Now count CD4 percentage results by patient instead.” → ready
Turn 3 — “Separately, count HIV visits by encounter type.” → ready
1 run — all passed · base query returned 100 rows · refined query returned 100 rows
catalyst-query-gemma-4-12b-qwen2.5-14b-checked · B1
Question — Show recent HIV results.
Writer — 'HIV' is not one of the recorded result names. Did you mean Current WHO HIV stage, HIV viral load, or something else?
Turn 2 — “The last 90 days, and only CD4 count, CD4 percentage, and HIV viral load.” → rejected: The named analyte 'CD4 count' is not bound as a parameter. Add a predicate on concept_name using a named parameter and bind its string value exactly as 'CD4 count'. (1 more finding)
1 run — FAILURES present
catalyst-query-gemma-4-12b-qwen2.5-14b-checked · B2
Question — Show patients with poor adherence.
Writer — What is the threshold or category for 'poor adherence'? (e.g., a specific value for 'Overall drug adherence in last month' or a specific 'value_coded_name' like 'Poor' or 'Non-adherent'?)
Turn 2 — “Poor adherence means the patient's latest antiretroviral adherence result is anything other than 'All'.” → ready
1 run — all passed · refined query returned 100 rows
catalyst-query-gemma-4-12b-qwen2.5-14b-checked · B3
Question — Show patients overdue for follow-up.
Writer — What is the specific time window or number of days for a patient to be considered 'overdue' for a follow-up?
Turn 2 — “As of 2026-03-01, overdue means a recorded return visit date before that date with no later visit.” → ready
1 run — FAILURES present · refined query returned 100 rows
catalyst-query-gemma-4-12b-qwen2.5-14b-checked · U1
Question — Show each patient's home address.
Writer — The provided catalog does not contain any fields or views that record a patient's home address.
1 run — all passed
catalyst-query-gemma-4-12b-qwen2.5-14b-checked · U2
Question — Show the prescribing clinician's name for every medication request.
1 run — FAILURES present
SQL diffs
Turn / version / execution timeline
| scenario | session | versions | executions | gen wall |
|---|---|---|---|---|
catalyst-query-gemma-4-12b · A1 | 288232cd-5f6c-4086-a2b9-c6e6a3ff5201 | 588e0ccb-bf32-46aa-8e11-acb2844b78dd → | 59b2210e-81a6-466d-9745-3569d3edfbb7 / | 32106 ms |
catalyst-query-gemma-4-12b · A2 | 91fc6ef7-5e47-4c48-8255-95b0280d3652 | 0722ee75-1892-4c61-b5a3-b0979ad764dd → | e96b2f66-b63f-40fb-9d44-e35d858ea665 / | 31749 ms |
catalyst-query-gemma-4-12b · A3 | d71a5f6f-7aa6-4be0-89d3-015cfe650eb3 | b75f7703-2190-4b7a-b5f8-18a6ee441420 → | ac4a67e4-65af-490c-92cb-7becd57d8a65 / | 35570 ms |
catalyst-query-gemma-4-12b · A4 | f4784740-a2f2-417e-8771-7f485cbb6cc2 | 5e70b8c1-d20b-4443-b108-a78b27e27fac → | e82fcd93-2b13-493b-bd36-800e9d7fe415 / | 27342 ms |
catalyst-query-gemma-4-12b · M1 | 9f7ccc6f-e1f8-41b4-9152-4fc9d7aff772 | 18e942e1-ac8c-4562-90e4-f1bd459b3fbf → 18e942e1-ac8c-4562-90e4-f1bd459b3fbf | 28f6d9ed-e20b-443c-bf03-e973397d9b0a / 28f6d9ed-e20b-443c-bf03-e973397d9b0a | 61127 ms |
catalyst-query-gemma-4-12b · M2 | ea719da9-88ec-4e2b-a3df-ce358ba565df | c6a9dcdd-1e48-4eb2-a80d-7725cc3f6bbd → c6a9dcdd-1e48-4eb2-a80d-7725cc3f6bbd | 83dcc185-f456-4f27-adde-70781d5786d3 / 83dcc185-f456-4f27-adde-70781d5786d3 | 100817 ms |
catalyst-query-gemma-4-12b · M3 | 352ecad4-9ce9-414e-9687-5aaab198bd0c | d04ed859-7cb5-4726-8d23-8b76bcca6baf → d04ed859-7cb5-4726-8d23-8b76bcca6baf | e5818f98-1a7e-4209-8d86-da1558de6f90 / e5818f98-1a7e-4209-8d86-da1558de6f90 | 57705 ms |
catalyst-query-gemma-4-12b · B1 | e0674ad0-8f37-4992-bb4e-239648dd9e6a | → | / | 60830 ms |
catalyst-query-gemma-4-12b · B2 | 567f1ea8-7665-4e05-a13e-e5c00555aae7 | 4fd3e3b5-568c-4022-9171-ea665673225e → 4fd3e3b5-568c-4022-9171-ea665673225e | cf705f95-0274-4c3a-87a7-ea998788adca / cf705f95-0274-4c3a-87a7-ea998788adca | 55119 ms |
catalyst-query-gemma-4-12b · B3 | 0a977653-8130-4f08-823e-f939e5688f0a | 19cdbf25-ecae-47bc-bc67-51533e28ffee → 19cdbf25-ecae-47bc-bc67-51533e28ffee | 798427b4-0ef8-45bc-8454-20a634766d3a / 798427b4-0ef8-45bc-8454-20a634766d3a | 49920 ms |
catalyst-query-gemma-4-12b · U1 | 6e2734c8-a03a-4fd2-8fbc-e860c7882a79 | → | / | 14999 ms |
catalyst-query-gemma-4-12b · U2 | 0b6d3792-3130-45c8-93f7-929bd643144a | 4e4903e0-9228-445d-a492-6bec34d35f04 → | / | 28293 ms |
catalyst-query-gemma-4-12b-q4-checked · A1 | bb8147f2-e2e7-4c17-bf85-082337c87039 | bfd41602-5c23-4952-aad2-fb9a4fb4fce9 → | 4136a22c-1934-42f3-a89f-ab72308c0318 / | 67692 ms |
catalyst-query-gemma-4-12b-q4-checked · A2 | 1762baac-fdde-44f8-8adc-f7aadc3ce751 | a6ee49bc-82f7-410e-8132-9b8595b2bb78 → | ace13046-ddd2-492f-92a9-742fe974b0e7 / | 61097 ms |
catalyst-query-gemma-4-12b-q4-checked · A3 | d16b152d-f5fb-434b-9ab4-7e67f0e0bf95 | d53043f1-739a-4f5a-b13d-c37a567affa9 → | 9a210dd0-f4ab-4351-a04e-d525bd85d1ff / | 90627 ms |
catalyst-query-gemma-4-12b-q4-checked · A4 | 7e7f6e22-5346-40d0-8c7c-e172e089cf60 | 7a9523ff-0a32-4e44-8e45-4d69c28fc24b → | debfb9c4-c079-47ab-a8d0-bbb2d58414c0 / | 194136 ms |
catalyst-query-gemma-4-12b-q4-checked · M1 | e395100e-ae3a-49c1-96a0-97930d4d004c | 62c5bddd-7e4f-432c-893e-abb23b522574 → 62c5bddd-7e4f-432c-893e-abb23b522574 | e88cbb84-1da0-44bb-8311-2899a50b4a4a / e88cbb84-1da0-44bb-8311-2899a50b4a4a | 126016 ms |
catalyst-query-gemma-4-12b-q4-checked · M2 | dd8e7d90-777d-4abf-b777-d9482cd3de33 | 714f1fac-4635-4618-9b7c-474f51c94b84 → 714f1fac-4635-4618-9b7c-474f51c94b84 | 7b80ec24-7711-4819-adfd-6d3b393f7a4a / 7b80ec24-7711-4819-adfd-6d3b393f7a4a | 114305 ms |
catalyst-query-gemma-4-12b-q4-checked · M3 | bf85ac3a-7a0f-4966-8dfb-242b5e85018c | cc635f9b-7a9c-4b4f-a6a3-0bf2c742ba21 → cc635f9b-7a9c-4b4f-a6a3-0bf2c742ba21 | 6538d6bd-7b6b-4189-8fde-d030c4ecb811 / 6538d6bd-7b6b-4189-8fde-d030c4ecb811 | 117096 ms |
catalyst-query-gemma-4-12b-q4-checked · B1 | a908ec80-8e68-420a-ab96-b760050460a8 | → | / | 74302 ms |
catalyst-query-gemma-4-12b-q4-checked · B2 | 98248bb9-39ab-4a95-ac55-9cd7a0414c12 | f37ec15e-2f1e-4ae0-b316-ade6ed18d50c → f37ec15e-2f1e-4ae0-b316-ade6ed18d50c | e914e905-9b24-4370-82af-415d366a4afb / e914e905-9b24-4370-82af-415d366a4afb | 116127 ms |
catalyst-query-gemma-4-12b-q4-checked · B3 | b8f4433f-bd6e-4e2c-901a-a805c8095388 | d617fae2-4b11-4854-8305-77aee19377ec → d617fae2-4b11-4854-8305-77aee19377ec | 0132ea2c-8282-4159-80c9-b390397a92ea / 0132ea2c-8282-4159-80c9-b390397a92ea | 83994 ms |
catalyst-query-gemma-4-12b-q4-checked · U1 | ba9d9614-0b31-437b-bb64-345bbff60fee | → | / | 15544 ms |
catalyst-query-gemma-4-12b-q4-checked · U2 | e23cb339-6315-4750-9cf2-c77ca40724f6 | 97ffd2ee-e1d8-4748-84cb-dfa90fddd720 → | / | 51871 ms |
catalyst-query-gemma-4-12b-qwen2.5-14b-checked · A1 | 830ac3c2-3106-4d98-845f-c9fdbfd56847 | 3e3f2d48-09aa-4df1-b18a-03e4d7aa32c6 → | 34d776d3-3936-4e0e-90c4-b93a3cfa4e12 / | 98856 ms |
catalyst-query-gemma-4-12b-qwen2.5-14b-checked · A2 | 6c6f8313-d5b2-44dc-a8a5-b4f328021274 | 807c2af7-13e7-4753-a869-47e8f2e74a33 → | 6062472f-b97a-4e93-9649-10c97bdae4a1 / | 66445 ms |
catalyst-query-gemma-4-12b-qwen2.5-14b-checked · A3 | 0225c04b-57c4-4408-b56f-f37593cadd1e | bf90b391-375f-4136-88cd-f6bfd8cc6e48 → | 14606a37-a161-46f5-b7af-c3912a88cbc8 / | 66085 ms |
catalyst-query-gemma-4-12b-qwen2.5-14b-checked · A4 | 251a69f6-f30c-4def-aee5-80f1231b8a88 | 826f816a-1441-4aec-a250-67a7786f0f34 → | b080fa57-6057-48d6-8ab0-07ef49ad23ff / | 72961 ms |
catalyst-query-gemma-4-12b-qwen2.5-14b-checked · M1 | 3e67b9e5-e250-4acd-bafd-4cf0c2f4e577 | e4cb0a9c-f1f7-4a78-9048-609c20649963 → e4cb0a9c-f1f7-4a78-9048-609c20649963 | 21854c28-e88f-49e7-bccf-78850a62396d / 21854c28-e88f-49e7-bccf-78850a62396d | 145741 ms |
catalyst-query-gemma-4-12b-qwen2.5-14b-checked · M2 | bb7b83ed-722e-49e7-a2d6-9743c66eda90 | 4c0e9662-bd76-4c28-af52-1b482dc72a09 → 4c0e9662-bd76-4c28-af52-1b482dc72a09 | f424417d-9680-4108-b0a6-41aefee1de7a / f424417d-9680-4108-b0a6-41aefee1de7a | 137944 ms |
catalyst-query-gemma-4-12b-qwen2.5-14b-checked · M3 | bff945ac-4b7a-4703-8db5-ecd01562acc8 | 04831f8f-29ac-4d44-9c13-b7f8f98a026d → 04831f8f-29ac-4d44-9c13-b7f8f98a026d | a0cca333-e97a-470c-85b2-d67f31ff343a / a0cca333-e97a-470c-85b2-d67f31ff343a | 147088 ms |
catalyst-query-gemma-4-12b-qwen2.5-14b-checked · B1 | 1345d401-75f2-4067-bfbc-0b9a9512cbc8 | → | / | 75108 ms |
catalyst-query-gemma-4-12b-qwen2.5-14b-checked · B2 | 23f50b23-1d00-4ccf-a299-ca4ee43a24b9 | 39f97292-ab7e-4f4a-87b0-d912fbdeec8a → 39f97292-ab7e-4f4a-87b0-d912fbdeec8a | 43dbdc08-1cb9-433f-a884-a8a1ef70d38f / 43dbdc08-1cb9-433f-a884-a8a1ef70d38f | 100348 ms |
catalyst-query-gemma-4-12b-qwen2.5-14b-checked · B3 | 1946b3b9-e8fc-4ba8-9b76-8bf9ec61dc28 | 7d309e11-5df2-4168-a6d8-7b8f3ee2d462 → 7d309e11-5df2-4168-a6d8-7b8f3ee2d462 | c92fa964-b743-457f-8039-8ec3c570722d / c92fa964-b743-457f-8039-8ec3c570722d | 92414 ms |
catalyst-query-gemma-4-12b-qwen2.5-14b-checked · U1 | af1fbcf7-3f11-4695-841b-6cded3ebd824 | → | / | 15812 ms |
catalyst-query-gemma-4-12b-qwen2.5-14b-checked · U2 | b971bb0f-5aa5-471f-80c4-fd3bccfbacc3 | 866a1bf8-468d-42c3-8f00-743188b1547d → | / | 69258 ms |
Judge detailadvisory
writer-only catalyst-query-gemma-4-12b
B2 turn 1 · composite 55 — flagged
intent_fidelity=1 — The rule was 'latest antiretroviral adherence result is anything other than All', but the SQL ranks over concept_name IN ('Antiretroviral adherence in past week', 'Overall drug adherence in last month') plus a value_coded_name IS NOT NULL filter, changing which observation counts as latest; gold count 3002 vs reference 108 (scenarios/catalyst-query-gemma-4-12b/B2/repetition-01/16-gold-execution-match-successor.json) confirms the material mismatch.sql_quality=2 — The ROW_NUMBER() OVER (PARTITION BY patient_id ORDER BY observed_at DESC) subquery is structurally sound and executed, but ranking one 'latest' across a two-concept union with an extra IS NOT NULL guard misleads about what it computes - workable with issues.schema_discipline=3 — All tables/columns used (hiv_patient_dim_v1; hiv_observation_fact_v1.value_coded_name/observed_at/concept_name) are catalogued and the query executed cleanly; concept names are inline literals, which is acceptable.followup_coherence=1 — It applies the clarification's shape (latest-per-patient, <> 'All') to the original poor-adherence question but misapplies the rule by widening to a second adherence concept the user never gave, which is why the successor count diverges (3002 vs 108).
B3 turn 1 · composite 63 — flagged
intent_fidelity=1 — The rule requires 'a recorded return visit date before 2026-03-01 with no later visit', but the SQL never touches the 'Return visit date' observation - it selects patients whose hiv_visit_fact_v1.started_at is before :reference_date with no visit on/after it, the wrong entity for the rule; gold fails (5000 capped vs 4665, scenarios/catalyst-query-gemma-4-12b/B3/repetition-01/16-gold-execution-match-successor.json).sql_quality=3 — Readable IN-subquery plus correlated NOT EXISTS that executed fine, but it scans hiv_visit_fact_v1 twice where a single per-patient aggregate (MAX(started_at) < :reference_date) would do - workable with redundancy.schema_discipline=3 — Stays on catalogued hiv_patient_dim_v1/hiv_visit_fact_v1 columns and binds :reference_date as a date correctly (13-execute-successor.json).followup_coherence=1 — It picks up the as-of date 2026-03-01 from the clarification (bound as :reference_date) but misapplies the supplied overdue rule, substituting visit recency for the recorded return-visit-date condition.
M1 turn 0 · composite 84 — flagged
intent_fidelity=2 — The executed SQL (scenarios/catalyst-query-gemma-4-12b/M1/repetition-01/06-execute-base.json) is 'SELECT t1.patient_id, t2.family_name, t1.medication_name FROM analytics.hiv_medication_request_fact_v1 ... JOIN analytics.hiv_patient_dim_v1' - a per-prescription listing with exactly the requested patient id, last name and medication; the worklist sql field carries the later aggregated head query, so the per-version execution evidence governs.sql_quality=3 — Simple two-table equi-join projection, clean and executable; succeeded with 100 rows returned under the row cap.schema_discipline=3 — Catalogued hiv_medication_request_fact_v1 and hiv_patient_dim_v1 columns only; no parameters required.
M2 turn 0 · composite 84
intent_fidelity=2 — Grain matches the instruction - COUNT(*) by medication_name (scenarios/catalyst-query-gemma-4-12b/M2/repetition-01/06-execute-base.json) - but the query carries an unrequested 'WHERE do_not_perform = false' predicate; it is the reused catalog counting query and matches the guidance pinned later in the session, so this is minor filter drift rather than a wrong question, scored 2 not 3 because the opening instruction never asked for the exclusion.sql_quality=3 — Minimal correct single-table aggregate; no redundancy.schema_discipline=3 — medication_name and do_not_perform are catalogued hiv_medication_request_fact_v1 columns; no binding issues.
A1 turn 0 · composite 100
intent_fidelity=3 — Projects family_name/given_name, value_numeric, value_unit, observed_at and filters concept_name='CD4 count' AND observed_at >= 2026-02-01 via bound parameters - exactly the requested fields and window; gold count matches 11=11 (scenarios/catalyst-query-gemma-4-12b/A1/repetition-01/15-gold-execution-match-base.json).sql_quality=3 — Single clear join of hiv_patient_dim_v1 to hiv_observation_fact_v1 on patient_id with parameterized predicates; no dead branches, executed successfully (06-execute-base.json).schema_discipline=3 — Only catalogued tables/columns; :concept_name (string) and :start_date bound with correct types in scenarios/catalyst-query-gemma-4-12b/A1/repetition-01/06-execute-base.json.
A2 turn 0 · composite 100
intent_fidelity=3 — COUNT(*) grouped by encounter_type with started_at >= 2025-01-01 and ORDER BY visit_count DESC matches grain, filter and requested ordering; gold aggregate match is clean, 5=5 keys with no mismatches (scenarios/catalyst-query-gemma-4-12b/A2/repetition-01/15-gold-execution-match-base.json).sql_quality=3 — Minimal idiomatic aggregate 'SELECT encounter_type, COUNT(*) ... GROUP BY encounter_type ORDER BY visit_count DESC'; nothing extraneous.schema_discipline=3 — Uses only catalogued hiv_visit_fact_v1.encounter_type/started_at; :start_date bound with a correct temporal type (06-execute-base.json).
A3 turn 0 · composite 100
intent_fidelity=3 — Counts hiv_medication_request_fact_v1 by medication_name with the female filter and 'do_not_perform = false', ordered highest first - every requested constraint present; gold match 30=30 keys, zero valueMismatches (scenarios/catalyst-query-gemma-4-12b/A3/repetition-01/15-gold-execution-match-base.json).sql_quality=3 — Straightforward filtered aggregate with ORDER BY ... DESC; gender supplied via the bound :gender parameter.schema_discipline=3 — Catalogued columns only (medication_name, patient_gender, do_not_perform); execution evidence shows clean binding (06-execute-base.json).
A4 turn 0 · composite 100
intent_fidelity=3 — 'WHERE ciel_code IS NULL' projecting concept_name and observation_count ordered DESC is exactly the asked listing; it equals the gold reference relation with ordering added and the aggregate match passes 2=2 (scenarios/catalyst-query-gemma-4-12b/A4/repetition-01/15-gold-execution-match-base.json).sql_quality=3 — Minimal single-table select with the one needed predicate and sort; nothing extraneous.schema_discipline=3 — hiv_concept_mapping_v1.concept_name/ciel_code/observation_count are all catalogued; no parameters needed and none misbound.
M1 turn 1 · composite 100
intent_fidelity=3 — Collapses to one row per patient via STRING_AGG(DISTINCT medication_name, ', ') grouped by t1.patient_id, t2.family_name, and the last-name column is the true hiv_patient_dim_v1.family_name, so both requested fixes are reflected (scenarios/catalyst-query-gemma-4-12b/M1/repetition-01/13-execute-successor.json).sql_quality=3 — Appropriate aggregation with DISTINCT dedup inside STRING_AGG; grouping keys match the non-aggregated projection exactly.schema_discipline=3 — Same catalogued join; STRING_AGG over the text medication_name column is type-correct.followup_coherence=3 — Keeps the base join and projection (patient_id, family_name, medications) while applying the collapse-per-patient change; nothing required from the prior turn was dropped.
M1 turn 2 · composite 100
intent_fidelity=3 — The requested join to the patient table by id for the last name is already present ('JOIN analytics.hiv_patient_dim_v1 ... ON t1.patient_id = t2.patient_id' with family_name projected), so the unchanged SQL satisfies the cumulative request (scenarios/catalyst-query-gemma-4-12b/M1/repetition-01/13-execute-successor-t2.json).sql_quality=3 — Same sound STRING_AGG aggregate grouped by t1.patient_id, t2.family_name as the prior version; no degradation or dead branches introduced.schema_discipline=3 — Catalogued tables/columns only, unchanged from the validated prior version.followup_coherence=3 — Correctly recognizes the follow-up asks for a join the query already has and preserves every prior constraint (grouping, DISTINCT aggregation, family_name) instead of mangling the query.
M2 turn 1 · composite 100
intent_fidelity=3 — Adds patient_gender to both projection and GROUP BY exactly as asked; the gold FAIL (59 rows vs 10, scenarios/catalyst-query-gemma-4-12b/M2/repetition-01/16-gold-execution-match-successor.json) reflects the final top-ten reference being applied to this intermediate turn - all 10 reference groups match with zero valueMismatches, so the regrouped counts themselves are correct and the verdict is not inflated by ignoring it.sql_quality=3 — Clean regrouped aggregate; correctly leaves ordering/limit for the later turn that asks for them.schema_discipline=3 — Catalogued columns only; 'do_not_perform = false' keeps a valid boolean predicate form.followup_coherence=3 — Preserves the pinned 'do_not_perform = false' exclusion and the COUNT(*) grain from the base while applying only the requested regroup.
M2 turn 2 · composite 100
intent_fidelity=3 — ORDER BY ... DESC LIMIT 10 returns exactly the ten highest medication-and-gender groups; gold passes 10=10 with no value mismatches (scenarios/catalyst-query-gemma-4-12b/M2/repetition-01/16-gold-execution-match-successor-t2.json).sql_quality=3 — Idiomatic top-N aggregate; nothing extraneous.schema_discipline=3 — Catalogued columns only; no parameters needed.followup_coherence=3 — Keeps the gender+medication grouping and the pinned 'do_not_perform = false' exclusion from the prior turn, adding only the requested top-ten restriction.
M3 turn 0 · composite 100
intent_fidelity=3 — The executed SQL counts hiv_observation_fact_v1 rows per patient_id with :concept_name bound to 'CD4 count' (scenarios/catalyst-query-gemma-4-12b/M3/repetition-01/06-execute-base.json) - exactly 'count CD4 count results by patient'; the worklist sql field shows the later visits head query, so the per-version execution evidence governs.sql_quality=3 — Simple parameterized aggregate with the right GROUP BY grain.schema_discipline=3 — Catalogued observation-fact columns; concept bound as a string parameter.
M3 turn 1 · composite 100
intent_fidelity=3 — Same per-patient count with the concept parameter switched to 'CD4%' - the catalog's CD4 percentage concept (scenarios/catalyst-query-gemma-4-12b/M3/repetition-01/13-execute-successor.json shows :concept_name bound to 'CD4%').sql_quality=3 — Clean parameterized aggregate reuse; grain unchanged as requested.schema_discipline=3 — Catalogued columns; the concept parameter is re-bound correctly as a string.followup_coherence=3 — Keeps the verified per-patient counting shape and changes only the concept - exactly the 'instead' the follow-up asked for.
M3 turn 2 · composite 100
intent_fidelity=3 — Counts hiv_visit_fact_v1 rows by encounter_type - the fresh question asked; 11 groups returned (scenarios/catalyst-query-gemma-4-12b/M3/repetition-01/13-execute-successor-t2.json).sql_quality=3 — Minimal correct aggregate.schema_discipline=3 — Catalogued visit-fact columns only.followup_coherence=3 — 'Separately' is honored: it starts a fresh visits-by-encounter-type query and does not drag the prior CD4 concept filter or per-patient grain into it.
self-checked catalyst-query-gemma-4-12b-q4-checked
B3 turn 1 · composite 63 — flagged
intent_fidelity=1 — The rule requires 'a recorded return visit date before 2026-03-01 with no later visit', but the SQL never touches the 'Return visit date' observation - it selects patients whose hiv_visit_fact_v1.started_at is before :reference_date with no visit on/after it, the wrong entity for the rule; gold fails (5000 capped vs 4665, scenarios/catalyst-query-gemma-4-12b-q4-checked/B3/repetition-01/16-gold-execution-match-successor.json).sql_quality=3 — Readable IN-subquery plus correlated NOT EXISTS that executed fine, but it scans hiv_visit_fact_v1 twice where a single per-patient aggregate (MAX(started_at) < :reference_date) would do - workable with redundancy.schema_discipline=3 — Stays on catalogued hiv_patient_dim_v1/hiv_visit_fact_v1 columns and binds :reference_date as a date correctly (13-execute-successor.json).followup_coherence=1 — It picks up the as-of date 2026-03-01 from the clarification (bound as :reference_date) but misapplies the supplied overdue rule, substituting visit recency for the recorded return-visit-date condition.
M1 turn 0 · composite 84 — flagged
intent_fidelity=2 — The executed SQL (scenarios/catalyst-query-gemma-4-12b-q4-checked/M1/repetition-01/06-execute-base.json) is 'SELECT t1.patient_id, t2.family_name, t1.medication_name FROM analytics.hiv_medication_request_fact_v1 ... JOIN analytics.hiv_patient_dim_v1' - a per-prescription listing with exactly the requested patient id, last name and medication; the worklist sql field carries the later aggregated head query, so the per-version execution evidence governs.sql_quality=3 — Simple two-table equi-join projection, clean and executable; succeeded with 100 rows returned under the row cap.schema_discipline=3 — Catalogued hiv_medication_request_fact_v1 and hiv_patient_dim_v1 columns only; no parameters required.
M2 turn 0 · composite 84 — flagged
intent_fidelity=2 — Grain matches the instruction - COUNT(*) by medication_name (scenarios/catalyst-query-gemma-4-12b-q4-checked/M2/repetition-01/06-execute-base.json) - but the query carries an unrequested 'WHERE do_not_perform = false' predicate; it is the reused catalog counting query and matches the guidance pinned later in the session, so this is minor filter drift rather than a wrong question, scored 2 not 3 because the opening instruction never asked for the exclusion.sql_quality=3 — Minimal correct single-table aggregate; no redundancy.schema_discipline=3 — medication_name and do_not_perform are catalogued hiv_medication_request_fact_v1 columns; no binding issues.
B2 turn 1 · composite 87
intent_fidelity=2 — Implements 'latest antiretroviral adherence <> All' via DISTINCT ON (patient_id) ... ORDER BY patient_id, observed_at DESC over the single bound concept, and gold passes 108=108 (scenarios/catalyst-query-gemma-4-12b-q4-checked/B2/repetition-01/16-gold-execution-match-successor.json); scored 2 not 3 only for the unrequested "obs_status = 'final'" predicate - minor drift that provably did not change the result.sql_quality=3 — Idiomatic Postgres DISTINCT ON latest-per-patient pattern; compact, clear, and executed cleanly.schema_discipline=3 — hiv_observation_fact_v1.obs_status/value_coded_name/observed_at and the patient-dim columns are all catalogued (the query executed and returned rows); :concept_name bound as string (13-execute-successor.json).followup_coherence=3 — Applies the clarified rule to the original 'patients with poor adherence' question with the right concept and latest-per-patient semantics; the successor matches the reference count exactly.
A1 turn 0 · composite 100
intent_fidelity=3 — Projects family_name/given_name, value_numeric, value_unit, observed_at and filters concept_name='CD4 count' AND observed_at >= 2026-02-01 via bound parameters - exactly the requested fields and window; gold count matches 11=11 (scenarios/catalyst-query-gemma-4-12b-q4-checked/A1/repetition-01/15-gold-execution-match-base.json).sql_quality=3 — Single clear join of hiv_patient_dim_v1 to hiv_observation_fact_v1 on patient_id with parameterized predicates; no dead branches, executed successfully (06-execute-base.json).schema_discipline=3 — Only catalogued tables/columns; :concept_name (string) and :start_date bound with correct types in scenarios/catalyst-query-gemma-4-12b-q4-checked/A1/repetition-01/06-execute-base.json.
A2 turn 0 · composite 100
intent_fidelity=3 — COUNT(*) grouped by encounter_type with started_at >= 2025-01-01 and ORDER BY visit_count DESC matches grain, filter and requested ordering; gold aggregate match is clean, 5=5 keys with no mismatches (scenarios/catalyst-query-gemma-4-12b-q4-checked/A2/repetition-01/15-gold-execution-match-base.json).sql_quality=3 — Minimal idiomatic aggregate 'SELECT encounter_type, COUNT(*) ... GROUP BY encounter_type ORDER BY visit_count DESC'; nothing extraneous.schema_discipline=3 — Uses only catalogued hiv_visit_fact_v1.encounter_type/started_at; :start_date bound with a correct temporal type (06-execute-base.json).
A3 turn 0 · composite 100
intent_fidelity=3 — Counts hiv_medication_request_fact_v1 by medication_name with the female filter and 'do_not_perform = false', ordered highest first - every requested constraint present; gold match 30=30 keys, zero valueMismatches (scenarios/catalyst-query-gemma-4-12b-q4-checked/A3/repetition-01/15-gold-execution-match-base.json).sql_quality=3 — Straightforward filtered aggregate with ORDER BY ... DESC; gender appears as the inline literal 'female' rather than a bound parameter, a negligible style choice.schema_discipline=3 — Catalogued columns only (medication_name, patient_gender, do_not_perform); execution evidence shows clean binding (06-execute-base.json).
A4 turn 0 · composite 100
intent_fidelity=3 — 'WHERE ciel_code IS NULL' projecting concept_name and observation_count ordered DESC is exactly the asked listing; it equals the gold reference relation with ordering added and the aggregate match passes 2=2 (scenarios/catalyst-query-gemma-4-12b-q4-checked/A4/repetition-01/15-gold-execution-match-base.json).sql_quality=3 — Minimal single-table select with the one needed predicate and sort; nothing extraneous.schema_discipline=3 — hiv_concept_mapping_v1.concept_name/ciel_code/observation_count are all catalogued; no parameters needed and none misbound.
M1 turn 1 · composite 100
intent_fidelity=3 — Collapses to one row per patient via STRING_AGG(DISTINCT medication_name, ', ') grouped by t1.patient_id, t2.family_name, and the last-name column is the true hiv_patient_dim_v1.family_name, so both requested fixes are reflected (scenarios/catalyst-query-gemma-4-12b-q4-checked/M1/repetition-01/13-execute-successor.json).sql_quality=3 — Appropriate aggregation with DISTINCT dedup inside STRING_AGG; grouping keys match the non-aggregated projection exactly.schema_discipline=3 — Same catalogued join; STRING_AGG over the text medication_name column is type-correct.followup_coherence=3 — Keeps the base join and projection (patient_id, family_name, medications) while applying the collapse-per-patient change; nothing required from the prior turn was dropped.
M1 turn 2 · composite 100
intent_fidelity=3 — The requested join to the patient table by id for the last name is already present ('JOIN analytics.hiv_patient_dim_v1 ... ON t1.patient_id = t2.patient_id' with family_name projected), so the unchanged SQL satisfies the cumulative request (scenarios/catalyst-query-gemma-4-12b-q4-checked/M1/repetition-01/13-execute-successor-t2.json).sql_quality=3 — Same sound STRING_AGG aggregate grouped by t1.patient_id, t2.family_name as the prior version; no degradation or dead branches introduced.schema_discipline=3 — Catalogued tables/columns only, unchanged from the validated prior version.followup_coherence=3 — Correctly recognizes the follow-up asks for a join the query already has and preserves every prior constraint (grouping, DISTINCT aggregation, family_name) instead of mangling the query.
M2 turn 1 · composite 100
intent_fidelity=3 — Adds patient_gender to both projection and GROUP BY exactly as asked; the gold FAIL (59 rows vs 10, scenarios/catalyst-query-gemma-4-12b-q4-checked/M2/repetition-01/16-gold-execution-match-successor.json) reflects the final top-ten reference being applied to this intermediate turn - all 10 reference groups match with zero valueMismatches, so the regrouped counts themselves are correct and the verdict is not inflated by ignoring it.sql_quality=3 — Clean regrouped aggregate; correctly leaves ordering/limit for the later turn that asks for them.schema_discipline=3 — Catalogued columns only; 'do_not_perform = false' keeps a valid boolean predicate form.followup_coherence=3 — Preserves the pinned 'do_not_perform = false' exclusion and the COUNT(*) grain from the base while applying only the requested regroup.
M2 turn 2 · composite 100
intent_fidelity=3 — ORDER BY ... DESC LIMIT 10 returns exactly the ten highest medication-and-gender groups; gold passes 10=10 with no value mismatches (scenarios/catalyst-query-gemma-4-12b-q4-checked/M2/repetition-01/16-gold-execution-match-successor-t2.json).sql_quality=3 — Idiomatic top-N aggregate; nothing extraneous.schema_discipline=3 — Catalogued columns only; no parameters needed.followup_coherence=3 — Keeps the gender+medication grouping and the pinned 'do_not_perform = false' exclusion from the prior turn, adding only the requested top-ten restriction.
M3 turn 0 · composite 100
intent_fidelity=3 — The executed SQL counts hiv_observation_fact_v1 rows per patient_id with :concept_name bound to 'CD4 count' (scenarios/catalyst-query-gemma-4-12b-q4-checked/M3/repetition-01/06-execute-base.json) - exactly 'count CD4 count results by patient'; the worklist sql field shows the later visits head query, so the per-version execution evidence governs.sql_quality=3 — Simple parameterized aggregate with the right GROUP BY grain.schema_discipline=3 — Catalogued observation-fact columns; concept bound as a string parameter.
M3 turn 2 · composite 100
intent_fidelity=3 — Counts hiv_visit_fact_v1 rows by encounter_type - the fresh question asked; 11 groups returned (scenarios/catalyst-query-gemma-4-12b-q4-checked/M3/repetition-01/13-execute-successor-t2.json).sql_quality=3 — Minimal correct aggregate.schema_discipline=3 — Catalogued visit-fact columns only.followup_coherence=3 — 'Separately' is honored: it starts a fresh visits-by-encounter-type query and does not drag the prior CD4 concept filter or per-patient grain into it. (This team's turn-1 followup failed with no executed version - 10-final-turns.json shows ordinal 2 status 'failed' - so the prior selected version is still the turn-0 CD4-per-patient query, and none of its constraints belong here.)
qwen2.5-14b-checked catalyst-query-gemma-4-12b-qwen2.5-14b-checked
B3 turn 1 · composite 63 — flagged
intent_fidelity=1 — The rule requires 'a recorded return visit date before 2026-03-01 with no later visit', but the SQL never touches the 'Return visit date' observation - it selects patients whose hiv_visit_fact_v1.started_at is before :reference_date with no visit on/after it, the wrong entity for the rule; gold fails (5000 capped vs 4665, scenarios/catalyst-query-gemma-4-12b-qwen2.5-14b-checked/B3/repetition-01/16-gold-execution-match-successor.json).sql_quality=3 — Readable IN-subquery plus correlated NOT EXISTS that executed fine, but it scans hiv_visit_fact_v1 twice where a single per-patient aggregate (MAX(started_at) < :reference_date) would do - workable with redundancy.schema_discipline=3 — Stays on catalogued hiv_patient_dim_v1/hiv_visit_fact_v1 columns and binds :reference_date as a date correctly (13-execute-successor.json).followup_coherence=1 — It picks up the as-of date 2026-03-01 from the clarification (bound as :reference_date) but misapplies the supplied overdue rule, substituting visit recency for the recorded return-visit-date condition.
M1 turn 0 · composite 84 — flagged
intent_fidelity=2 — The executed SQL (scenarios/catalyst-query-gemma-4-12b-qwen2.5-14b-checked/M1/repetition-01/06-execute-base.json) is 'SELECT t1.patient_id, t2.family_name, t1.medication_name FROM analytics.hiv_medication_request_fact_v1 ... JOIN analytics.hiv_patient_dim_v1' - a per-prescription listing with exactly the requested patient id, last name and medication; the worklist sql field carries the later aggregated head query, so the per-version execution evidence governs.sql_quality=3 — Simple two-table equi-join projection, clean and executable; succeeded with 100 rows returned under the row cap.schema_discipline=3 — Catalogued hiv_medication_request_fact_v1 and hiv_patient_dim_v1 columns only; no parameters required.
M2 turn 0 · composite 84 — flagged
intent_fidelity=2 — Grain matches the instruction - COUNT(*) by medication_name (scenarios/catalyst-query-gemma-4-12b-qwen2.5-14b-checked/M2/repetition-01/06-execute-base.json) - but the query carries an unrequested 'WHERE do_not_perform IS FALSE' predicate; it is the reused catalog counting query and matches the guidance pinned later in the session, so this is minor filter drift rather than a wrong question, scored 2 not 3 because the opening instruction never asked for the exclusion.sql_quality=3 — Minimal correct single-table aggregate; no redundancy.schema_discipline=3 — medication_name and do_not_perform are catalogued hiv_medication_request_fact_v1 columns; no binding issues.
A4 turn 0 · composite 90
intent_fidelity=3 — Lists concepts with ciel_code IS NULL alongside their observation counts, highest first; the gold aggregate matches 2=2 with total_observation_count resolving to observation_count (scenarios/catalyst-query-gemma-4-12b-qwen2.5-14b-checked/A4/repetition-01/15-gold-execution-match-base.json).sql_quality=2 — Wraps the already per-concept observation_count in 'SUM(t1.observation_count) ... GROUP BY t1.concept_name', a redundant aggregation over a relation the gold reference reads directly - workable, minor redundancy, so 2.schema_discipline=3 — Catalogued hiv_concept_mapping_v1 columns only; no parameters needed.
A1 turn 0 · composite 100
intent_fidelity=3 — Projects family_name/given_name, value_numeric, value_unit, observed_at and filters concept_name='CD4 count' AND observed_at >= 2026-02-01 via bound parameters - exactly the requested fields and window; gold count matches 11=11 (scenarios/catalyst-query-gemma-4-12b-qwen2.5-14b-checked/A1/repetition-01/15-gold-execution-match-base.json).sql_quality=3 — Single clear join of hiv_patient_dim_v1 to hiv_observation_fact_v1 on patient_id with parameterized predicates; no dead branches, executed successfully (06-execute-base.json).schema_discipline=3 — Only catalogued tables/columns; :concept_name (string) and :start_date bound with correct types in scenarios/catalyst-query-gemma-4-12b-qwen2.5-14b-checked/A1/repetition-01/06-execute-base.json.
A2 turn 0 · composite 100
intent_fidelity=3 — COUNT(*) grouped by encounter_type with started_at >= 2025-01-01 and ORDER BY visit_count DESC matches grain, filter and requested ordering; gold aggregate match is clean, 5=5 keys with no mismatches (scenarios/catalyst-query-gemma-4-12b-qwen2.5-14b-checked/A2/repetition-01/15-gold-execution-match-base.json).sql_quality=3 — Minimal idiomatic aggregate 'SELECT encounter_type, COUNT(*) ... GROUP BY encounter_type ORDER BY visit_count DESC'; nothing extraneous.schema_discipline=3 — Uses only catalogued hiv_visit_fact_v1.encounter_type/started_at; :start_date bound with a correct temporal type (06-execute-base.json).
A3 turn 0 · composite 100
intent_fidelity=3 — Counts hiv_medication_request_fact_v1 by medication_name with the female filter and 'do_not_perform = false', ordered highest first - every requested constraint present; gold match 30=30 keys, zero valueMismatches (scenarios/catalyst-query-gemma-4-12b-qwen2.5-14b-checked/A3/repetition-01/15-gold-execution-match-base.json).sql_quality=3 — Straightforward filtered aggregate with ORDER BY ... DESC; gender supplied via the bound :gender parameter and a descriptive request_count alias.schema_discipline=3 — Catalogued columns only (medication_name, patient_gender, do_not_perform); execution evidence shows clean binding (06-execute-base.json).
B2 turn 1 · composite 100
intent_fidelity=3 — The CTE ranks only the bound 'Antiretroviral adherence in past week' observations per patient by observed_at DESC and keeps rank 1 with value_coded_name != 'All' - precisely the supplied rule; gold passes 108=108 (scenarios/catalyst-query-gemma-4-12b-qwen2.5-14b-checked/B2/repetition-01/16-gold-execution-match-successor.json).sql_quality=3 — Clear CTE + ROW_NUMBER latest-per-patient pattern with a tidy latest_adherence_result projection; no dead branches.schema_discipline=3 — Catalogued observation-fact and patient-dim columns only; :adherence_concept bound as string (13-execute-successor.json).followup_coherence=3 — Applies the clarified definition to the original poor-adherence question without adding or dropping constraints; the count matches the reference exactly.
M1 turn 1 · composite 100
intent_fidelity=3 — Collapses to one row per patient via STRING_AGG(DISTINCT medication_name, ', ') grouped by t2.patient_id, t2.family_name, and the last-name column is the true hiv_patient_dim_v1.family_name, so both requested fixes are reflected (scenarios/catalyst-query-gemma-4-12b-qwen2.5-14b-checked/M1/repetition-01/13-execute-successor.json).sql_quality=3 — Appropriate aggregation with DISTINCT dedup inside STRING_AGG; grouping keys match the non-aggregated projection exactly.schema_discipline=3 — Same catalogued join; STRING_AGG over the text medication_name column is type-correct.followup_coherence=3 — Keeps the base join and projection (patient_id, family_name, medications) while applying the collapse-per-patient change; nothing required from the prior turn was dropped.
M1 turn 2 · composite 100
intent_fidelity=3 — The requested join to the patient table by id for the last name is already present ('JOIN analytics.hiv_patient_dim_v1 ... ON t1.patient_id = t2.patient_id' with family_name projected), so the unchanged SQL satisfies the cumulative request (scenarios/catalyst-query-gemma-4-12b-qwen2.5-14b-checked/M1/repetition-01/13-execute-successor-t2.json).sql_quality=3 — Same sound STRING_AGG aggregate grouped by t2.patient_id, t2.family_name as the prior version; no degradation or dead branches introduced.schema_discipline=3 — Catalogued tables/columns only, unchanged from the validated prior version.followup_coherence=3 — Correctly recognizes the follow-up asks for a join the query already has and preserves every prior constraint (grouping, DISTINCT aggregation, family_name) instead of mangling the query.
M2 turn 1 · composite 100
intent_fidelity=3 — Adds patient_gender to both projection and GROUP BY exactly as asked; the gold FAIL (59 rows vs 10, scenarios/catalyst-query-gemma-4-12b-qwen2.5-14b-checked/M2/repetition-01/16-gold-execution-match-successor.json) reflects the final top-ten reference being applied to this intermediate turn - all 10 reference groups match with zero valueMismatches, so the regrouped counts themselves are correct and the verdict is not inflated by ignoring it.sql_quality=3 — Clean regrouped aggregate; correctly leaves ordering/limit for the later turn that asks for them.schema_discipline=3 — Catalogued columns only; 'do_not_perform IS FALSE' keeps a valid boolean predicate form.followup_coherence=3 — Preserves the pinned 'do_not_perform IS FALSE' exclusion and the COUNT(*) grain from the base while applying only the requested regroup.
M2 turn 2 · composite 100
intent_fidelity=3 — ORDER BY ... DESC LIMIT 10 returns exactly the ten highest medication-and-gender groups; gold passes 10=10 with no value mismatches (scenarios/catalyst-query-gemma-4-12b-qwen2.5-14b-checked/M2/repetition-01/16-gold-execution-match-successor-t2.json).sql_quality=3 — Idiomatic top-N aggregate; nothing extraneous.schema_discipline=3 — Catalogued columns only; no parameters needed.followup_coherence=3 — Keeps the gender+medication grouping and the pinned 'do_not_perform IS FALSE' exclusion from the prior turn, adding only the requested top-ten restriction.
M3 turn 0 · composite 100
intent_fidelity=3 — The executed SQL counts hiv_observation_fact_v1 rows per patient_id with :cd4_count_name bound to 'CD4 count' (scenarios/catalyst-query-gemma-4-12b-qwen2.5-14b-checked/M3/repetition-01/06-execute-base.json) - exactly 'count CD4 count results by patient'; the worklist sql field shows the later visits head query, so the per-version execution evidence governs.sql_quality=3 — Simple parameterized aggregate with the right GROUP BY grain.schema_discipline=3 — Catalogued observation-fact columns; concept bound as a string parameter.
M3 turn 1 · composite 100
intent_fidelity=3 — Same per-patient count with the concept parameter switched to 'CD4%' - the catalog's CD4 percentage concept (scenarios/catalyst-query-gemma-4-12b-qwen2.5-14b-checked/M3/repetition-01/13-execute-successor.json shows :cd4_percentage_name bound to 'CD4%').sql_quality=3 — Clean parameterized aggregate reuse; grain unchanged as requested.schema_discipline=3 — Catalogued columns; the concept parameter is re-bound correctly as a string.followup_coherence=3 — Keeps the verified per-patient counting shape and changes only the concept - exactly the 'instead' the follow-up asked for.
M3 turn 2 · composite 100
intent_fidelity=3 — Counts hiv_visit_fact_v1 rows by encounter_type - the fresh question asked; 11 groups returned (scenarios/catalyst-query-gemma-4-12b-qwen2.5-14b-checked/M3/repetition-01/13-execute-successor-t2.json).sql_quality=3 — Minimal correct aggregate.schema_discipline=3 — Catalogued visit-fact columns only.followup_coherence=3 — 'Separately' is honored: it starts a fresh visits-by-encounter-type query and does not drag the prior CD4 concept filter or per-patient grain into it.
Methods & provenance
How this works, the dataset, and the model lineup
Each conversation runs live: a question generates SQL (a writer model drafts it; reviewed profiles also invoke their configured reviewer), the query is validated and executed against PostgreSQL, then a follow-up instruction refines the exact current query and the successor is validated and executed again. Executed results are re-checked against an independently-authored gold query (byte-level row-set match) and an independent read-only PostgreSQL cross-check.
Dataset hiv-20260731T000741Z (openmrs-hiv-fhir-postgresql): 5285 patients · 427915 results · 143 test types · 2022-11-07 – 2026-08-11
| profile | writer model | reviewer model |
|---|---|---|
catalyst-query-gemma-4-12b | gemma-4-12b-q4 | — (writer only) |
catalyst-query-gemma-4-12b-q4-checked | gemma-4-12b-q4 | gemma-4-12b-q4 |
catalyst-query-gemma-4-12b-qwen2.5-14b-checked | gemma-4-12b | qwen2.5-14b |
run_id=9ae123db-8f40-4246-8769-d427a5551769 · evidence_status=development · dataset=hiv-20260731T000741Z · catalog=openmrs-hiv-catalog-v6 · provider=llama.cpp · profiles=catalyst-query-gemma-4-12b,catalyst-query-gemma-4-12b-q4-checked,catalyst-query-gemma-4-12b-qwen2.5-14b-checked