Catalyst notebook report · catalyst-phase1-comparison-v1

run_id=9ae123db-8f40-4246-8769-d427a5551769 · evidence_status=development · dataset=hiv-20260731T000741Z · catalog=openmrs-hiv-catalog-v6 · provider=llama.cpp · profiles=catalyst-query-gemma-4-12b,catalyst-query-gemma-4-12b-q4-checked,catalyst-query-gemma-4-12b-qwen2.5-14b-checked
Development fixture evidence — not a release claim.

In plain terms

We asked 3 AI “teams” the same 12 questions about a clinical database, in plain English, and let each team write and run the database queries itself — including asking us to clarify vague questions, and refusing ones the data can't answer. A team passes a question only when its final answer matches an independently written reference answer, row for row.

The teams finished within one question of each other — a practical tie at this sample size: qwen2.5-14b-checked 8 of 12 · writer-only 7 of 12 · self-checked 7 of 12.

The misses were largely shared, not team-specific: 4 of the 12 questions (B1, B3, M2, U2) were missed by every team; B2, M3 tripped only one team. A separate AI judge read every executed query. On the failed questions the SQL itself still scored at ceiling for construction and schema use — what dropped was fidelity to what the question actually asked. The errors are misreadings, not broken queries.

One conversation per question, on a demonstration dataset — read this as directional evidence, not a benchmark; differences of a question or two are within noise.

Result

Against the gates in force at publication (≥90% overall, ≥80% per scenario): no team qualified

8/12 qwen2.5-14b-checkedgemma-4-12b writer · qwen2.5-14b reviewer
7/12 writer-onlygemma-4-12b-q4, no reviewer
7/12 self-checkedgemma-4-12b-q4 writer · gemma-4-12b-q4 reviewer

22/36 conversations passed across all teams · 1038 assertions · pass = the final answer matches the independent reference answer row-for-row.

Judge summaryadvisory

A separate AI judge read every executed query against its recorded evidence and scored four axes from 0 to 3. The axes are what it judged, so they lead; each cell shows the median and, where the scores differ, the range behind it. Advisory means it never gates acceptance — the row-for-row reference check does. Where that check failed, the judge explains why; it cannot overrule it.

teamqueries judgedintentSQL craftschemafollow-upweakest queriescomposite
writer-only153 1–33 2–333 1–3B2·turn 1 — 55, B3·turn 1 — 63, M1·turn 0 — 84100 ↓55
self-checked143 1–3333 1–3B3·turn 1 — 63, M1·turn 0 — 84, M2·turn 0 — 84100 ↓63
qwen2.5-14b-checked153 1–33 2–333 1–3B3·turn 1 — 63, M1·turn 0 — 84, M2·turn 0 — 84100 ↓63
all teams443 1–33 2–333 1–3writer-only: B2·turn 1 — 55, writer-only: B3·turn 1 — 63, qwen2.5-14b-checked: B3·turn 1 — 63100 ↓55

Judge: claude-fable-5 (anthropic) · three independent passes finalized by per-axis medians · rubric 5dc94bbd3042…

intent
did the SQL answer the question as asked (0–3)
SQL craft
clean, minimal, executable construction — joins, predicates, no dead branches (0–3)
schema
only catalogued tables and columns, parameters bound with correct types (0–3)
follow-up
the revision honors the new instruction without breaking what worked (0–3, follow-up turns only)
weakest queries
every query the judge marked down on any axis, worst first (at most three shown) — the rubric's own anchors decide, not a composite cutoff. Blank when nothing was marked down.
composite
a weighted convenience score (0–100; opening queries 47/29/24, follow-ups 40/25/20/15) with the team's floor after ↓. Read the axes first: when they saturate, the composite is nearly a step function of whichever axis dropped.
One judge actor (anthropic/claude-fable-5/claude-fable-5), three passes finalized by median. That measures how stably this model scores, not whether it scores correctly — there is no cross-model agreement evidence in this run. No human adjudication. Nobody has confirmed or overruled a judged call in this run, so the scores carry no human anchor.

Scenario matrix

scenarioreportedgold okjudge by turn (advisory)assertionsgold FAIL detail
catalyst-query-gemma-4-12b · A1
single-ready
PASSyes100
13/13 passed
session_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, token_budget_respected-base, base_writer_outcome, base_version_available, base_validation_recorded, base_execution_succeeded, base_postgres_crosscheck, base_gold_execution_match, successor_visible_under_three_minutes, new_session_isolation
—
catalyst-query-gemma-4-12b · A2
single-ready
PASSyes100
13/13 passed
session_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, token_budget_respected-base, base_writer_outcome, base_version_available, base_validation_recorded, base_execution_succeeded, base_postgres_crosscheck, base_gold_execution_match, successor_visible_under_three_minutes, new_session_isolation
—
catalyst-query-gemma-4-12b · A3
single-ready
PASSyes100
13/13 passed
session_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, token_budget_respected-base, base_writer_outcome, base_version_available, base_validation_recorded, base_execution_succeeded, base_postgres_crosscheck, base_gold_execution_match, successor_visible_under_three_minutes, new_session_isolation
—
catalyst-query-gemma-4-12b · A4
single-ready
PASSyes100
13/13 passed
session_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, token_budget_respected-base, base_writer_outcome, base_version_available, base_validation_recorded, base_execution_succeeded, base_postgres_crosscheck, base_gold_execution_match, successor_visible_under_three_minutes, new_session_isolation
—
catalyst-query-gemma-4-12b · M1
multi-turn
PASSyes84 · 100 · 100
60/60 passed
session_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, token_budget_respected-base, base_writer_outcome, base_version_available, base_validation_recorded, base_execution_succeeded, base_postgres_crosscheck, followup_http_created, followup_terminal_status, writer_outcome, base_classification, manual_version_classification, followup_profile, refresh_restored, timeline_current_turn, followup_evidence_available, invocation_duration_sum, invocation_digests, invocation_timestamp_reconciliation, writer_model, reviewer_model, effective_temperature_and_dry, revision_context_exclusions, token_evidence_recorded, token_budget_respected, exact_selected_output, semantic_reviewer_correction, prior_results_stale_after_successor, successor_validation_recorded, successor_execution_succeeded, successor_postgres_crosscheck, followup_http_created-t2, followup_terminal_status-t2, writer_outcome-t2, base_classification-t2, manual_version_classification-t2, followup_profile-t2, refresh_restored-t2, timeline_current_turn-t2, followup_evidence_available-t2, invocation_duration_sum-t2, invocation_digests-t2, invocation_timestamp_reconciliation-t2, writer_model-t2, reviewer_model-t2, effective_temperature_and_dry-t2, revision_context_exclusions-t2, token_evidence_recorded-t2, token_budget_respected-t2, exact_selected_output-t2, semantic_reviewer_correction-t2, prior_results_stale_after_successor-t2, successor_validation_recorded-t2, successor_execution_succeeded-t2, successor_postgres_crosscheck-t2, successor_visible_under_three_minutes-t2, new_session_isolation
—
catalyst-query-gemma-4-12b · M2
retained-guidance
FAILno84 · 100 · 100
62/63 passed
session_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, token_budget_respected-base, base_writer_outcome, base_version_available, base_validation_recorded, base_execution_succeeded, base_postgres_crosscheck, guidance_pinned, followup_http_created, followup_terminal_status, writer_outcome, base_classification, manual_version_classification, followup_profile, refresh_restored, timeline_current_turn, followup_evidence_available, invocation_duration_sum, invocation_digests, invocation_timestamp_reconciliation, writer_model, reviewer_model, effective_temperature_and_dry, revision_context_exclusions, token_evidence_recorded, token_budget_respected, exact_selected_output, semantic_reviewer_correction, prior_results_stale_after_successor, successor_validation_recorded, successor_execution_succeeded, successor_postgres_crosscheck, successor_gold_execution_match, followup_http_created-t2, followup_terminal_status-t2, writer_outcome-t2, base_classification-t2, manual_version_classification-t2, followup_profile-t2, refresh_restored-t2, timeline_current_turn-t2, followup_evidence_available-t2, invocation_duration_sum-t2, invocation_digests-t2, invocation_timestamp_reconciliation-t2, writer_model-t2, reviewer_model-t2, effective_temperature_and_dry-t2, revision_context_exclusions-t2, token_evidence_recorded-t2, token_budget_respected-t2, exact_selected_output-t2, semantic_reviewer_correction-t2, prior_results_stale_after_successor-t2, successor_validation_recorded-t2, successor_execution_succeeded-t2, successor_postgres_crosscheck-t2, successor_gold_execution_match-t2, successor_visible_under_three_minutes-t2, new_session_isolation
catalyst-query-gemma-4-12b · M3
verified-examples
PASSyes100 · 100 · 100
60/60 passed
session_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, token_budget_respected-base, base_writer_outcome, base_version_available, base_validation_recorded, base_execution_succeeded, base_postgres_crosscheck, followup_http_created, followup_terminal_status, writer_outcome, base_classification, manual_version_classification, followup_profile, refresh_restored, timeline_current_turn, followup_evidence_available, invocation_duration_sum, invocation_digests, invocation_timestamp_reconciliation, writer_model, reviewer_model, effective_temperature_and_dry, revision_context_exclusions, token_evidence_recorded, token_budget_respected, exact_selected_output, semantic_reviewer_correction, prior_results_stale_after_successor, successor_validation_recorded, successor_execution_succeeded, successor_postgres_crosscheck, followup_http_created-t2, followup_terminal_status-t2, writer_outcome-t2, base_classification-t2, manual_version_classification-t2, followup_profile-t2, refresh_restored-t2, timeline_current_turn-t2, followup_evidence_available-t2, invocation_duration_sum-t2, invocation_digests-t2, invocation_timestamp_reconciliation-t2, writer_model-t2, reviewer_model-t2, effective_temperature_and_dry-t2, revision_context_exclusions-t2, token_evidence_recorded-t2, token_budget_respected-t2, exact_selected_output-t2, semantic_reviewer_correction-t2, prior_results_stale_after_successor-t2, successor_validation_recorded-t2, successor_execution_succeeded-t2, successor_postgres_crosscheck-t2, successor_visible_under_three_minutes-t2, new_session_isolation
—
catalyst-query-gemma-4-12b · B1
clarification
FAILno—
25/28 passed
session_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, base_writer_outcome, no_sql_after_non_ready_base, followup_http_created, followup_terminal_status, writer_outcome, base_classification, manual_version_classification, followup_profile, refresh_restored, timeline_current_turn, followup_evidence_available, invocation_duration_sum, invocation_digests, invocation_timestamp_reconciliation, writer_model, reviewer_model, effective_temperature_and_dry, revision_context_exclusions, token_evidence_recorded, token_budget_respected, exact_selected_output, semantic_reviewer_correction, successor_visible_under_three_minutes, new_session_isolation
judged · followup_terminal_status
judged · writer_outcome
judged · exact_selected_output
catalyst-query-gemma-4-12b · B2
clarification
FAILno55
32/33 passed
session_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, token_budget_respected-base, base_writer_outcome, no_sql_after_non_ready_base, followup_http_created, followup_terminal_status, writer_outcome, base_classification, manual_version_classification, followup_profile, refresh_restored, timeline_current_turn, followup_evidence_available, invocation_duration_sum, invocation_digests, invocation_timestamp_reconciliation, writer_model, reviewer_model, effective_temperature_and_dry, revision_context_exclusions, token_evidence_recorded, token_budget_respected, exact_selected_output, semantic_reviewer_correction, successor_validation_recorded, successor_execution_succeeded, successor_postgres_crosscheck, successor_gold_execution_match, successor_visible_under_three_minutes, new_session_isolation
catalyst-query-gemma-4-12b · B3
clarification
FAILno63
32/33 passed
session_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, token_budget_respected-base, base_writer_outcome, no_sql_after_non_ready_base, followup_http_created, followup_terminal_status, writer_outcome, base_classification, manual_version_classification, followup_profile, refresh_restored, timeline_current_turn, followup_evidence_available, invocation_duration_sum, invocation_digests, invocation_timestamp_reconciliation, writer_model, reviewer_model, effective_temperature_and_dry, revision_context_exclusions, token_evidence_recorded, token_budget_respected, exact_selected_output, semantic_reviewer_correction, successor_validation_recorded, successor_execution_succeeded, successor_postgres_crosscheck, successor_gold_execution_match, successor_visible_under_three_minutes, new_session_isolation
catalyst-query-gemma-4-12b · U1
unsupported
PASSyes—
9/9 passed
session_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, token_budget_respected-base, base_writer_outcome, no_sql_after_non_ready_base, successor_visible_under_three_minutes, new_session_isolation
—
catalyst-query-gemma-4-12b · U2
unsupported
FAILno—
8/9 passed
session_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, token_budget_respected-base, base_writer_outcome, base_version_available, successor_visible_under_three_minutes, new_session_isolation
judged · base_writer_outcome
catalyst-query-gemma-4-12b-q4-checked · A1
single-ready
PASSyes100
13/13 passed
session_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, token_budget_respected-base, base_writer_outcome, base_version_available, base_validation_recorded, base_execution_succeeded, base_postgres_crosscheck, base_gold_execution_match, successor_visible_under_three_minutes, new_session_isolation
—
catalyst-query-gemma-4-12b-q4-checked · A2
single-ready
PASSyes100
13/13 passed
session_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, token_budget_respected-base, base_writer_outcome, base_version_available, base_validation_recorded, base_execution_succeeded, base_postgres_crosscheck, base_gold_execution_match, successor_visible_under_three_minutes, new_session_isolation
—
catalyst-query-gemma-4-12b-q4-checked · A3
single-ready
PASSyes100
13/13 passed
session_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, token_budget_respected-base, base_writer_outcome, base_version_available, base_validation_recorded, base_execution_succeeded, base_postgres_crosscheck, base_gold_execution_match, successor_visible_under_three_minutes, new_session_isolation
—
catalyst-query-gemma-4-12b-q4-checked · A4
single-ready
PASSyes100
13/13 passed
session_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, token_budget_respected-base, base_writer_outcome, base_version_available, base_validation_recorded, base_execution_succeeded, base_postgres_crosscheck, base_gold_execution_match, successor_visible_under_three_minutes, new_session_isolation
—
catalyst-query-gemma-4-12b-q4-checked · M1
multi-turn
PASSyes84 · 100 · 100
60/60 passed
session_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, token_budget_respected-base, base_writer_outcome, base_version_available, base_validation_recorded, base_execution_succeeded, base_postgres_crosscheck, followup_http_created, followup_terminal_status, writer_outcome, base_classification, manual_version_classification, followup_profile, refresh_restored, timeline_current_turn, followup_evidence_available, invocation_duration_sum, invocation_digests, invocation_timestamp_reconciliation, writer_model, reviewer_model, effective_temperature_and_dry, revision_context_exclusions, token_evidence_recorded, token_budget_respected, exact_selected_output, semantic_reviewer_correction, prior_results_stale_after_successor, successor_validation_recorded, successor_execution_succeeded, successor_postgres_crosscheck, followup_http_created-t2, followup_terminal_status-t2, writer_outcome-t2, base_classification-t2, manual_version_classification-t2, followup_profile-t2, refresh_restored-t2, timeline_current_turn-t2, followup_evidence_available-t2, invocation_duration_sum-t2, invocation_digests-t2, invocation_timestamp_reconciliation-t2, writer_model-t2, reviewer_model-t2, effective_temperature_and_dry-t2, revision_context_exclusions-t2, token_evidence_recorded-t2, token_budget_respected-t2, exact_selected_output-t2, semantic_reviewer_correction-t2, prior_results_stale_after_successor-t2, successor_validation_recorded-t2, successor_execution_succeeded-t2, successor_postgres_crosscheck-t2, successor_visible_under_three_minutes-t2, new_session_isolation
—
catalyst-query-gemma-4-12b-q4-checked · M2
retained-guidance
FAILno84 · 100 · 100
62/63 passed
session_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, token_budget_respected-base, base_writer_outcome, base_version_available, base_validation_recorded, base_execution_succeeded, base_postgres_crosscheck, guidance_pinned, followup_http_created, followup_terminal_status, writer_outcome, base_classification, manual_version_classification, followup_profile, refresh_restored, timeline_current_turn, followup_evidence_available, invocation_duration_sum, invocation_digests, invocation_timestamp_reconciliation, writer_model, reviewer_model, effective_temperature_and_dry, revision_context_exclusions, token_evidence_recorded, token_budget_respected, exact_selected_output, semantic_reviewer_correction, prior_results_stale_after_successor, successor_validation_recorded, successor_execution_succeeded, successor_postgres_crosscheck, successor_gold_execution_match, followup_http_created-t2, followup_terminal_status-t2, writer_outcome-t2, base_classification-t2, manual_version_classification-t2, followup_profile-t2, refresh_restored-t2, timeline_current_turn-t2, followup_evidence_available-t2, invocation_duration_sum-t2, invocation_digests-t2, invocation_timestamp_reconciliation-t2, writer_model-t2, reviewer_model-t2, effective_temperature_and_dry-t2, revision_context_exclusions-t2, token_evidence_recorded-t2, token_budget_respected-t2, exact_selected_output-t2, semantic_reviewer_correction-t2, prior_results_stale_after_successor-t2, successor_validation_recorded-t2, successor_execution_succeeded-t2, successor_postgres_crosscheck-t2, successor_gold_execution_match-t2, successor_visible_under_three_minutes-t2, new_session_isolation
catalyst-query-gemma-4-12b-q4-checked · M3
verified-examples
FAILno100 · 100
53/57 passed
session_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, token_budget_respected-base, base_writer_outcome, base_version_available, base_validation_recorded, base_execution_succeeded, base_postgres_crosscheck, followup_http_created, followup_terminal_status, writer_outcome, base_classification, manual_version_classification, followup_profile, refresh_restored, timeline_current_turn, followup_evidence_available, invocation_duration_sum, invocation_digests, invocation_timestamp_reconciliation, writer_model, reviewer_model, effective_temperature_and_dry, revision_context_exclusions, token_evidence_recorded, token_budget_respected, exact_selected_output, semantic_reviewer_correction, prior_results_stale_after_successor, followup_http_created-t2, followup_terminal_status-t2, writer_outcome-t2, base_classification-t2, manual_version_classification-t2, followup_profile-t2, refresh_restored-t2, timeline_current_turn-t2, followup_evidence_available-t2, invocation_duration_sum-t2, invocation_digests-t2, invocation_timestamp_reconciliation-t2, writer_model-t2, reviewer_model-t2, effective_temperature_and_dry-t2, revision_context_exclusions-t2, token_evidence_recorded-t2, token_budget_respected-t2, exact_selected_output-t2, semantic_reviewer_correction-t2, prior_results_stale_after_successor-t2, successor_validation_recorded-t2, successor_execution_succeeded-t2, successor_postgres_crosscheck-t2, successor_visible_under_three_minutes-t2, new_session_isolation
judged · followup_terminal_status
judged · writer_outcome
judged · exact_selected_output
judged · prior_results_stale_after_successor
catalyst-query-gemma-4-12b-q4-checked · B1
clarification
FAILno—
25/28 passed
session_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, base_writer_outcome, no_sql_after_non_ready_base, followup_http_created, followup_terminal_status, writer_outcome, base_classification, manual_version_classification, followup_profile, refresh_restored, timeline_current_turn, followup_evidence_available, invocation_duration_sum, invocation_digests, invocation_timestamp_reconciliation, writer_model, reviewer_model, effective_temperature_and_dry, revision_context_exclusions, token_evidence_recorded, token_budget_respected, exact_selected_output, semantic_reviewer_correction, successor_visible_under_three_minutes, new_session_isolation
judged · followup_terminal_status
judged · writer_outcome
judged · exact_selected_output
catalyst-query-gemma-4-12b-q4-checked · B2
clarification
PASSyes87
33/33 passed
session_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, token_budget_respected-base, base_writer_outcome, no_sql_after_non_ready_base, followup_http_created, followup_terminal_status, writer_outcome, base_classification, manual_version_classification, followup_profile, refresh_restored, timeline_current_turn, followup_evidence_available, invocation_duration_sum, invocation_digests, invocation_timestamp_reconciliation, writer_model, reviewer_model, effective_temperature_and_dry, revision_context_exclusions, token_evidence_recorded, token_budget_respected, exact_selected_output, semantic_reviewer_correction, successor_validation_recorded, successor_execution_succeeded, successor_postgres_crosscheck, successor_gold_execution_match, successor_visible_under_three_minutes, new_session_isolation
—
catalyst-query-gemma-4-12b-q4-checked · B3
clarification
FAILno63
32/33 passed
session_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, token_budget_respected-base, base_writer_outcome, no_sql_after_non_ready_base, followup_http_created, followup_terminal_status, writer_outcome, base_classification, manual_version_classification, followup_profile, refresh_restored, timeline_current_turn, followup_evidence_available, invocation_duration_sum, invocation_digests, invocation_timestamp_reconciliation, writer_model, reviewer_model, effective_temperature_and_dry, revision_context_exclusions, token_evidence_recorded, token_budget_respected, exact_selected_output, semantic_reviewer_correction, successor_validation_recorded, successor_execution_succeeded, successor_postgres_crosscheck, successor_gold_execution_match, successor_visible_under_three_minutes, new_session_isolation
catalyst-query-gemma-4-12b-q4-checked · U1
unsupported
PASSyes—
9/9 passed
session_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, token_budget_respected-base, base_writer_outcome, no_sql_after_non_ready_base, successor_visible_under_three_minutes, new_session_isolation
—
catalyst-query-gemma-4-12b-q4-checked · U2
unsupported
FAILno—
8/9 passed
session_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, token_budget_respected-base, base_writer_outcome, base_version_available, successor_visible_under_three_minutes, new_session_isolation
judged · base_writer_outcome
catalyst-query-gemma-4-12b-qwen2.5-14b-checked · A1
single-ready
PASSyes100
13/13 passed
session_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, token_budget_respected-base, base_writer_outcome, base_version_available, base_validation_recorded, base_execution_succeeded, base_postgres_crosscheck, base_gold_execution_match, successor_visible_under_three_minutes, new_session_isolation
—
catalyst-query-gemma-4-12b-qwen2.5-14b-checked · A2
single-ready
PASSyes100
13/13 passed
session_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, token_budget_respected-base, base_writer_outcome, base_version_available, base_validation_recorded, base_execution_succeeded, base_postgres_crosscheck, base_gold_execution_match, successor_visible_under_three_minutes, new_session_isolation
—
catalyst-query-gemma-4-12b-qwen2.5-14b-checked · A3
single-ready
PASSyes100
13/13 passed
session_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, token_budget_respected-base, base_writer_outcome, base_version_available, base_validation_recorded, base_execution_succeeded, base_postgres_crosscheck, base_gold_execution_match, successor_visible_under_three_minutes, new_session_isolation
—
catalyst-query-gemma-4-12b-qwen2.5-14b-checked · A4
single-ready
PASSyes90
13/13 passed
session_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, token_budget_respected-base, base_writer_outcome, base_version_available, base_validation_recorded, base_execution_succeeded, base_postgres_crosscheck, base_gold_execution_match, successor_visible_under_three_minutes, new_session_isolation
—
catalyst-query-gemma-4-12b-qwen2.5-14b-checked · M1
multi-turn
PASSyes84 · 100 · 100
60/60 passed
session_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, token_budget_respected-base, base_writer_outcome, base_version_available, base_validation_recorded, base_execution_succeeded, base_postgres_crosscheck, followup_http_created, followup_terminal_status, writer_outcome, base_classification, manual_version_classification, followup_profile, refresh_restored, timeline_current_turn, followup_evidence_available, invocation_duration_sum, invocation_digests, invocation_timestamp_reconciliation, writer_model, reviewer_model, effective_temperature_and_dry, revision_context_exclusions, token_evidence_recorded, token_budget_respected, exact_selected_output, semantic_reviewer_correction, prior_results_stale_after_successor, successor_validation_recorded, successor_execution_succeeded, successor_postgres_crosscheck, followup_http_created-t2, followup_terminal_status-t2, writer_outcome-t2, base_classification-t2, manual_version_classification-t2, followup_profile-t2, refresh_restored-t2, timeline_current_turn-t2, followup_evidence_available-t2, invocation_duration_sum-t2, invocation_digests-t2, invocation_timestamp_reconciliation-t2, writer_model-t2, reviewer_model-t2, effective_temperature_and_dry-t2, revision_context_exclusions-t2, token_evidence_recorded-t2, token_budget_respected-t2, exact_selected_output-t2, semantic_reviewer_correction-t2, prior_results_stale_after_successor-t2, successor_validation_recorded-t2, successor_execution_succeeded-t2, successor_postgres_crosscheck-t2, successor_visible_under_three_minutes-t2, new_session_isolation
—
catalyst-query-gemma-4-12b-qwen2.5-14b-checked · M2
retained-guidance
FAILno84 · 100 · 100
62/63 passed
session_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, token_budget_respected-base, base_writer_outcome, base_version_available, base_validation_recorded, base_execution_succeeded, base_postgres_crosscheck, guidance_pinned, followup_http_created, followup_terminal_status, writer_outcome, base_classification, manual_version_classification, followup_profile, refresh_restored, timeline_current_turn, followup_evidence_available, invocation_duration_sum, invocation_digests, invocation_timestamp_reconciliation, writer_model, reviewer_model, effective_temperature_and_dry, revision_context_exclusions, token_evidence_recorded, token_budget_respected, exact_selected_output, semantic_reviewer_correction, prior_results_stale_after_successor, successor_validation_recorded, successor_execution_succeeded, successor_postgres_crosscheck, successor_gold_execution_match, followup_http_created-t2, followup_terminal_status-t2, writer_outcome-t2, base_classification-t2, manual_version_classification-t2, followup_profile-t2, refresh_restored-t2, timeline_current_turn-t2, followup_evidence_available-t2, invocation_duration_sum-t2, invocation_digests-t2, invocation_timestamp_reconciliation-t2, writer_model-t2, reviewer_model-t2, effective_temperature_and_dry-t2, revision_context_exclusions-t2, token_evidence_recorded-t2, token_budget_respected-t2, exact_selected_output-t2, semantic_reviewer_correction-t2, prior_results_stale_after_successor-t2, successor_validation_recorded-t2, successor_execution_succeeded-t2, successor_postgres_crosscheck-t2, successor_gold_execution_match-t2, successor_visible_under_three_minutes-t2, new_session_isolation
catalyst-query-gemma-4-12b-qwen2.5-14b-checked · M3
verified-examples
PASSyes100 · 100 · 100
60/60 passed
session_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, token_budget_respected-base, base_writer_outcome, base_version_available, base_validation_recorded, base_execution_succeeded, base_postgres_crosscheck, followup_http_created, followup_terminal_status, writer_outcome, base_classification, manual_version_classification, followup_profile, refresh_restored, timeline_current_turn, followup_evidence_available, invocation_duration_sum, invocation_digests, invocation_timestamp_reconciliation, writer_model, reviewer_model, effective_temperature_and_dry, revision_context_exclusions, token_evidence_recorded, token_budget_respected, exact_selected_output, semantic_reviewer_correction, prior_results_stale_after_successor, successor_validation_recorded, successor_execution_succeeded, successor_postgres_crosscheck, followup_http_created-t2, followup_terminal_status-t2, writer_outcome-t2, base_classification-t2, manual_version_classification-t2, followup_profile-t2, refresh_restored-t2, timeline_current_turn-t2, followup_evidence_available-t2, invocation_duration_sum-t2, invocation_digests-t2, invocation_timestamp_reconciliation-t2, writer_model-t2, reviewer_model-t2, effective_temperature_and_dry-t2, revision_context_exclusions-t2, token_evidence_recorded-t2, token_budget_respected-t2, exact_selected_output-t2, semantic_reviewer_correction-t2, prior_results_stale_after_successor-t2, successor_validation_recorded-t2, successor_execution_succeeded-t2, successor_postgres_crosscheck-t2, successor_visible_under_three_minutes-t2, new_session_isolation
—
catalyst-query-gemma-4-12b-qwen2.5-14b-checked · B1
clarification
FAILno—
25/28 passed
session_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, base_writer_outcome, no_sql_after_non_ready_base, followup_http_created, followup_terminal_status, writer_outcome, base_classification, manual_version_classification, followup_profile, refresh_restored, timeline_current_turn, followup_evidence_available, invocation_duration_sum, invocation_digests, invocation_timestamp_reconciliation, writer_model, reviewer_model, effective_temperature_and_dry, revision_context_exclusions, token_evidence_recorded, token_budget_respected, exact_selected_output, semantic_reviewer_correction, successor_visible_under_three_minutes, new_session_isolation
judged · followup_terminal_status
judged · writer_outcome
judged · exact_selected_output
catalyst-query-gemma-4-12b-qwen2.5-14b-checked · B2
clarification
PASSyes100
33/33 passed
session_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, token_budget_respected-base, base_writer_outcome, no_sql_after_non_ready_base, followup_http_created, followup_terminal_status, writer_outcome, base_classification, manual_version_classification, followup_profile, refresh_restored, timeline_current_turn, followup_evidence_available, invocation_duration_sum, invocation_digests, invocation_timestamp_reconciliation, writer_model, reviewer_model, effective_temperature_and_dry, revision_context_exclusions, token_evidence_recorded, token_budget_respected, exact_selected_output, semantic_reviewer_correction, successor_validation_recorded, successor_execution_succeeded, successor_postgres_crosscheck, successor_gold_execution_match, successor_visible_under_three_minutes, new_session_isolation
—
catalyst-query-gemma-4-12b-qwen2.5-14b-checked · B3
clarification
FAILno63
32/33 passed
session_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, token_budget_respected-base, base_writer_outcome, no_sql_after_non_ready_base, followup_http_created, followup_terminal_status, writer_outcome, base_classification, manual_version_classification, followup_profile, refresh_restored, timeline_current_turn, followup_evidence_available, invocation_duration_sum, invocation_digests, invocation_timestamp_reconciliation, writer_model, reviewer_model, effective_temperature_and_dry, revision_context_exclusions, token_evidence_recorded, token_budget_respected, exact_selected_output, semantic_reviewer_correction, successor_validation_recorded, successor_execution_succeeded, successor_postgres_crosscheck, successor_gold_execution_match, successor_visible_under_three_minutes, new_session_isolation
catalyst-query-gemma-4-12b-qwen2.5-14b-checked · U1
unsupported
PASSyes—
9/9 passed
session_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, token_budget_respected-base, base_writer_outcome, no_sql_after_non_ready_base, successor_visible_under_three_minutes, new_session_isolation
—
catalyst-query-gemma-4-12b-qwen2.5-14b-checked · U2
unsupported
FAILno—
8/9 passed
session_created, initial_turn_recorded, initial_evidence_available, token_evidence_recorded-base, token_budget_respected-base, base_writer_outcome, base_version_available, successor_visible_under_three_minutes, new_session_isolation
judged · base_writer_outcome

What each scenario does

catalyst-query-gemma-4-12b · A1

Question — List CD4 count results since 2026-02-01 with the patient, the value, the unit, and the observed date.

1 run — all passed · base query returned 11 rows

catalyst-query-gemma-4-12b · A2

Question — Count HIV visits by encounter type since 2025-01-01, highest count first.

1 run — all passed · base query returned 5 rows

catalyst-query-gemma-4-12b · A3

Question — Count medication requests for female patients by medication name, excluding do_not_perform, highest count first.

1 run — all passed · base query returned 30 rows

catalyst-query-gemma-4-12b · A4

Question — List each OpenMRS-native concept with no CIEL mapping, its name, and its total observation count, highest count first.

1 run — all passed · base query returned 2 rows

catalyst-query-gemma-4-12b · M1

Question — I need to get the list of medications prescribed for each patient, with the patient id and patient last name in the result.

Turn 2 — “The query needs to be updated - gender is labeled as last name, and the result should be collapsed per patient - so one row per patient, with comma-separated unique medications.” → ready

Turn 3 — “We need to join in the patient table by id to this query, so we can get patient last name” → ready

1 run — all passed · base query returned 100 rows · refined query returned 100 rows

catalyst-query-gemma-4-12b · M2

Question — Count medication requests by medication name.

Turn 2 — “Regroup that by patient gender as well as medication name.” → ready

Turn 3 — “Return only the ten highest medication-and-gender groups.” → ready

1 run — FAILURES present · base query returned 30 rows · refined query returned 59 rows

catalyst-query-gemma-4-12b · M3

Question — Count CD4 count results by patient.

Turn 2 — “Now count CD4 percentage results by patient instead.” → ready

Turn 3 — “Separately, count HIV visits by encounter type.” → ready

1 run — all passed · base query returned 100 rows · refined query returned 100 rows

catalyst-query-gemma-4-12b · B1

Question — Show recent HIV results.

Writer — 'HIV' is not one of the recorded result names. Did you mean Current WHO HIV stage, HIV viral load, or something else?

Turn 2 — “The last 90 days, and only CD4 count, CD4 percentage, and HIV viral load.” → rejected: The named analyte 'CD4 count' is not bound as a parameter. Add a predicate on concept_name using a named parameter and bind its string value exactly as 'CD4 count'. (1 more finding)

1 run — FAILURES present

catalyst-query-gemma-4-12b · B2

Question — Show patients with poor adherence.

Writer — What is the threshold or specific category for 'poor adherence' (e.g., a specific value for 'Overall drug adherence in last month' or 'Antiretroviral adherence in past week')?

Turn 2 — “Poor adherence means the patient's latest antiretroviral adherence result is anything other than 'All'.” → ready

1 run — FAILURES present · refined query returned 100 rows

catalyst-query-gemma-4-12b · B3

Question — Show patients overdue for follow-up.

Writer — What is the specific time window or number of days for a follow-up to be considered 'overdue'?

Turn 2 — “As of 2026-03-01, overdue means a recorded return visit date before that date with no later visit.” → ready

1 run — FAILURES present · refined query returned 100 rows

catalyst-query-gemma-4-12b · U1

Question — Show each patient's home address.

Writer — The catalog does not contain any fields or views that record a patient's home address.

1 run — all passed

catalyst-query-gemma-4-12b · U2

Question — Show the prescribing clinician's name for every medication request.

1 run — FAILURES present

catalyst-query-gemma-4-12b-q4-checked · A1

Question — List CD4 count results since 2026-02-01 with the patient, the value, the unit, and the observed date.

1 run — all passed · base query returned 11 rows

catalyst-query-gemma-4-12b-q4-checked · A2

Question — Count HIV visits by encounter type since 2025-01-01, highest count first.

1 run — all passed · base query returned 5 rows

catalyst-query-gemma-4-12b-q4-checked · A3

Question — Count medication requests for female patients by medication name, excluding do_not_perform, highest count first.

1 run — all passed · base query returned 30 rows

catalyst-query-gemma-4-12b-q4-checked · A4

Question — List each OpenMRS-native concept with no CIEL mapping, its name, and its total observation count, highest count first.

1 run — all passed · base query returned 2 rows

catalyst-query-gemma-4-12b-q4-checked · M1

Question — I need to get the list of medications prescribed for each patient, with the patient id and patient last name in the result.

Turn 2 — “The query needs to be updated - gender is labeled as last name, and the result should be collapsed per patient - so one row per patient, with comma-separated unique medications.” → ready

Turn 3 — “We need to join in the patient table by id to this query, so we can get patient last name” → ready

1 run — all passed · base query returned 100 rows · refined query returned 100 rows

catalyst-query-gemma-4-12b-q4-checked · M2

Question — Count medication requests by medication name.

Turn 2 — “Regroup that by patient gender as well as medication name.” → ready

Turn 3 — “Return only the ten highest medication-and-gender groups.” → ready

1 run — FAILURES present · base query returned 30 rows · refined query returned 59 rows

catalyst-query-gemma-4-12b-q4-checked · M3

Question — Count CD4 count results by patient.

Turn 2 — “Now count CD4 percentage results by patient instead.” → rejected: Query review failed: repaired query did not pass independent re-review

Turn 3 — “Separately, count HIV visits by encounter type.” → ready

1 run — FAILURES present · base query returned 100 rows

catalyst-query-gemma-4-12b-q4-checked · B1

Question — Show recent HIV results.

Writer — 'HIV' is not one of the recorded result names. Did you mean Current WHO HIV stage, HIV viral load, or something else?

Turn 2 — “The last 90 days, and only CD4 count, CD4 percentage, and HIV viral load.” → rejected: The named analyte 'CD4 count' is not bound as a parameter. Add a predicate on concept_name using a named parameter and bind its string value exactly as 'CD4 count'. (1 more finding)

1 run — FAILURES present

catalyst-query-gemma-4-12b-q4-checked · B2

Question — Show patients with poor adherence.

Writer — What is the threshold or specific category for 'poor adherence' (e.g., a specific value for 'Overall drug adherence in last month' or 'Antiretroviral adherence in past week')?

Turn 2 — “Poor adherence means the patient's latest antiretroviral adherence result is anything other than 'All'.” → ready

1 run — all passed · refined query returned 100 rows

catalyst-query-gemma-4-12b-q4-checked · B3

Question — Show patients overdue for follow-up.

Writer — What is the specific time window or number of days for a follow-up to be considered 'overdue'?

Turn 2 — “As of 2026-03-01, overdue means a recorded return visit date before that date with no later visit.” → ready

1 run — FAILURES present · refined query returned 100 rows

catalyst-query-gemma-4-12b-q4-checked · U1

Question — Show each patient's home address.

Writer — The catalog does not contain any fields or views that record a patient's home address.

1 run — all passed

catalyst-query-gemma-4-12b-q4-checked · U2

Question — Show the prescribing clinician's name for every medication request.

1 run — FAILURES present

catalyst-query-gemma-4-12b-qwen2.5-14b-checked · A1

Question — List CD4 count results since 2026-02-01 with the patient, the value, the unit, and the observed date.

1 run — all passed · base query returned 11 rows

catalyst-query-gemma-4-12b-qwen2.5-14b-checked · A2

Question — Count HIV visits by encounter type since 2025-01-01, highest count first.

1 run — all passed · base query returned 5 rows

catalyst-query-gemma-4-12b-qwen2.5-14b-checked · A3

Question — Count medication requests for female patients by medication name, excluding do_not_perform, highest count first.

1 run — all passed · base query returned 30 rows

catalyst-query-gemma-4-12b-qwen2.5-14b-checked · A4

Question — List each OpenMRS-native concept with no CIEL mapping, its name, and its total observation count, highest count first.

1 run — all passed · base query returned 2 rows

catalyst-query-gemma-4-12b-qwen2.5-14b-checked · M1

Question — I need to get the list of medications prescribed for each patient, with the patient id and patient last name in the result.

Turn 2 — “The query needs to be updated - gender is labeled as last name, and the result should be collapsed per patient - so one row per patient, with comma-separated unique medications.” → ready

Turn 3 — “We need to join in the patient table by id to this query, so we can get patient last name” → ready

1 run — all passed · base query returned 100 rows · refined query returned 100 rows

catalyst-query-gemma-4-12b-qwen2.5-14b-checked · M2

Question — Count medication requests by medication name.

Turn 2 — “Regroup that by patient gender as well as medication name.” → ready

Turn 3 — “Return only the ten highest medication-and-gender groups.” → ready

1 run — FAILURES present · base query returned 30 rows · refined query returned 59 rows

catalyst-query-gemma-4-12b-qwen2.5-14b-checked · M3

Question — Count CD4 count results by patient.

Turn 2 — “Now count CD4 percentage results by patient instead.” → ready

Turn 3 — “Separately, count HIV visits by encounter type.” → ready

1 run — all passed · base query returned 100 rows · refined query returned 100 rows

catalyst-query-gemma-4-12b-qwen2.5-14b-checked · B1

Question — Show recent HIV results.

Writer — 'HIV' is not one of the recorded result names. Did you mean Current WHO HIV stage, HIV viral load, or something else?

Turn 2 — “The last 90 days, and only CD4 count, CD4 percentage, and HIV viral load.” → rejected: The named analyte 'CD4 count' is not bound as a parameter. Add a predicate on concept_name using a named parameter and bind its string value exactly as 'CD4 count'. (1 more finding)

1 run — FAILURES present

catalyst-query-gemma-4-12b-qwen2.5-14b-checked · B2

Question — Show patients with poor adherence.

Writer — What is the threshold or category for 'poor adherence'? (e.g., a specific value for 'Overall drug adherence in last month' or a specific 'value_coded_name' like 'Poor' or 'Non-adherent'?)

Turn 2 — “Poor adherence means the patient's latest antiretroviral adherence result is anything other than 'All'.” → ready

1 run — all passed · refined query returned 100 rows

catalyst-query-gemma-4-12b-qwen2.5-14b-checked · B3

Question — Show patients overdue for follow-up.

Writer — What is the specific time window or number of days for a patient to be considered 'overdue' for a follow-up?

Turn 2 — “As of 2026-03-01, overdue means a recorded return visit date before that date with no later visit.” → ready

1 run — FAILURES present · refined query returned 100 rows

catalyst-query-gemma-4-12b-qwen2.5-14b-checked · U1

Question — Show each patient's home address.

Writer — The provided catalog does not contain any fields or views that record a patient's home address.

1 run — all passed

catalyst-query-gemma-4-12b-qwen2.5-14b-checked · U2

Question — Show the prescribing clinician's name for every medication request.

1 run — FAILURES present

SQL diffs

No multi-version SQL diffs found.

Turn / version / execution timeline

scenariosessionversionsexecutionsgen wall
catalyst-query-gemma-4-12b · A1288232cd-5f6c-4086-a2b9-c6e6a3ff5201588e0ccb-bf32-46aa-8e11-acb2844b78dd → 59b2210e-81a6-466d-9745-3569d3edfbb7 / 32106 ms
catalyst-query-gemma-4-12b · A291fc6ef7-5e47-4c48-8255-95b0280d36520722ee75-1892-4c61-b5a3-b0979ad764dd → e96b2f66-b63f-40fb-9d44-e35d858ea665 / 31749 ms
catalyst-query-gemma-4-12b · A3d71a5f6f-7aa6-4be0-89d3-015cfe650eb3b75f7703-2190-4b7a-b5f8-18a6ee441420 → ac4a67e4-65af-490c-92cb-7becd57d8a65 / 35570 ms
catalyst-query-gemma-4-12b · A4f4784740-a2f2-417e-8771-7f485cbb6cc25e70b8c1-d20b-4443-b108-a78b27e27fac → e82fcd93-2b13-493b-bd36-800e9d7fe415 / 27342 ms
catalyst-query-gemma-4-12b · M19f7ccc6f-e1f8-41b4-9152-4fc9d7aff77218e942e1-ac8c-4562-90e4-f1bd459b3fbf → 18e942e1-ac8c-4562-90e4-f1bd459b3fbf28f6d9ed-e20b-443c-bf03-e973397d9b0a / 28f6d9ed-e20b-443c-bf03-e973397d9b0a61127 ms
catalyst-query-gemma-4-12b · M2ea719da9-88ec-4e2b-a3df-ce358ba565dfc6a9dcdd-1e48-4eb2-a80d-7725cc3f6bbd → c6a9dcdd-1e48-4eb2-a80d-7725cc3f6bbd83dcc185-f456-4f27-adde-70781d5786d3 / 83dcc185-f456-4f27-adde-70781d5786d3100817 ms
catalyst-query-gemma-4-12b · M3352ecad4-9ce9-414e-9687-5aaab198bd0cd04ed859-7cb5-4726-8d23-8b76bcca6baf → d04ed859-7cb5-4726-8d23-8b76bcca6bafe5818f98-1a7e-4209-8d86-da1558de6f90 / e5818f98-1a7e-4209-8d86-da1558de6f9057705 ms
catalyst-query-gemma-4-12b · B1e0674ad0-8f37-4992-bb4e-239648dd9e6a → / 60830 ms
catalyst-query-gemma-4-12b · B2567f1ea8-7665-4e05-a13e-e5c00555aae74fd3e3b5-568c-4022-9171-ea665673225e → 4fd3e3b5-568c-4022-9171-ea665673225ecf705f95-0274-4c3a-87a7-ea998788adca / cf705f95-0274-4c3a-87a7-ea998788adca55119 ms
catalyst-query-gemma-4-12b · B30a977653-8130-4f08-823e-f939e5688f0a19cdbf25-ecae-47bc-bc67-51533e28ffee → 19cdbf25-ecae-47bc-bc67-51533e28ffee798427b4-0ef8-45bc-8454-20a634766d3a / 798427b4-0ef8-45bc-8454-20a634766d3a49920 ms
catalyst-query-gemma-4-12b · U16e2734c8-a03a-4fd2-8fbc-e860c7882a79 → / 14999 ms
catalyst-query-gemma-4-12b · U20b6d3792-3130-45c8-93f7-929bd643144a4e4903e0-9228-445d-a492-6bec34d35f04 → / 28293 ms
catalyst-query-gemma-4-12b-q4-checked · A1bb8147f2-e2e7-4c17-bf85-082337c87039bfd41602-5c23-4952-aad2-fb9a4fb4fce9 → 4136a22c-1934-42f3-a89f-ab72308c0318 / 67692 ms
catalyst-query-gemma-4-12b-q4-checked · A21762baac-fdde-44f8-8adc-f7aadc3ce751a6ee49bc-82f7-410e-8132-9b8595b2bb78 → ace13046-ddd2-492f-92a9-742fe974b0e7 / 61097 ms
catalyst-query-gemma-4-12b-q4-checked · A3d16b152d-f5fb-434b-9ab4-7e67f0e0bf95d53043f1-739a-4f5a-b13d-c37a567affa9 → 9a210dd0-f4ab-4351-a04e-d525bd85d1ff / 90627 ms
catalyst-query-gemma-4-12b-q4-checked · A47e7f6e22-5346-40d0-8c7c-e172e089cf607a9523ff-0a32-4e44-8e45-4d69c28fc24b → debfb9c4-c079-47ab-a8d0-bbb2d58414c0 / 194136 ms
catalyst-query-gemma-4-12b-q4-checked · M1e395100e-ae3a-49c1-96a0-97930d4d004c62c5bddd-7e4f-432c-893e-abb23b522574 → 62c5bddd-7e4f-432c-893e-abb23b522574e88cbb84-1da0-44bb-8311-2899a50b4a4a / e88cbb84-1da0-44bb-8311-2899a50b4a4a126016 ms
catalyst-query-gemma-4-12b-q4-checked · M2dd8e7d90-777d-4abf-b777-d9482cd3de33714f1fac-4635-4618-9b7c-474f51c94b84 → 714f1fac-4635-4618-9b7c-474f51c94b847b80ec24-7711-4819-adfd-6d3b393f7a4a / 7b80ec24-7711-4819-adfd-6d3b393f7a4a114305 ms
catalyst-query-gemma-4-12b-q4-checked · M3bf85ac3a-7a0f-4966-8dfb-242b5e85018ccc635f9b-7a9c-4b4f-a6a3-0bf2c742ba21 → cc635f9b-7a9c-4b4f-a6a3-0bf2c742ba216538d6bd-7b6b-4189-8fde-d030c4ecb811 / 6538d6bd-7b6b-4189-8fde-d030c4ecb811117096 ms
catalyst-query-gemma-4-12b-q4-checked · B1a908ec80-8e68-420a-ab96-b760050460a8 → / 74302 ms
catalyst-query-gemma-4-12b-q4-checked · B298248bb9-39ab-4a95-ac55-9cd7a0414c12f37ec15e-2f1e-4ae0-b316-ade6ed18d50c → f37ec15e-2f1e-4ae0-b316-ade6ed18d50ce914e905-9b24-4370-82af-415d366a4afb / e914e905-9b24-4370-82af-415d366a4afb116127 ms
catalyst-query-gemma-4-12b-q4-checked · B3b8f4433f-bd6e-4e2c-901a-a805c8095388d617fae2-4b11-4854-8305-77aee19377ec → d617fae2-4b11-4854-8305-77aee19377ec0132ea2c-8282-4159-80c9-b390397a92ea / 0132ea2c-8282-4159-80c9-b390397a92ea83994 ms
catalyst-query-gemma-4-12b-q4-checked · U1ba9d9614-0b31-437b-bb64-345bbff60fee → / 15544 ms
catalyst-query-gemma-4-12b-q4-checked · U2e23cb339-6315-4750-9cf2-c77ca40724f697ffd2ee-e1d8-4748-84cb-dfa90fddd720 → / 51871 ms
catalyst-query-gemma-4-12b-qwen2.5-14b-checked · A1830ac3c2-3106-4d98-845f-c9fdbfd568473e3f2d48-09aa-4df1-b18a-03e4d7aa32c6 → 34d776d3-3936-4e0e-90c4-b93a3cfa4e12 / 98856 ms
catalyst-query-gemma-4-12b-qwen2.5-14b-checked · A26c6f8313-d5b2-44dc-a8a5-b4f328021274807c2af7-13e7-4753-a869-47e8f2e74a33 → 6062472f-b97a-4e93-9649-10c97bdae4a1 / 66445 ms
catalyst-query-gemma-4-12b-qwen2.5-14b-checked · A30225c04b-57c4-4408-b56f-f37593cadd1ebf90b391-375f-4136-88cd-f6bfd8cc6e48 → 14606a37-a161-46f5-b7af-c3912a88cbc8 / 66085 ms
catalyst-query-gemma-4-12b-qwen2.5-14b-checked · A4251a69f6-f30c-4def-aee5-80f1231b8a88826f816a-1441-4aec-a250-67a7786f0f34 → b080fa57-6057-48d6-8ab0-07ef49ad23ff / 72961 ms
catalyst-query-gemma-4-12b-qwen2.5-14b-checked · M13e67b9e5-e250-4acd-bafd-4cf0c2f4e577e4cb0a9c-f1f7-4a78-9048-609c20649963 → e4cb0a9c-f1f7-4a78-9048-609c2064996321854c28-e88f-49e7-bccf-78850a62396d / 21854c28-e88f-49e7-bccf-78850a62396d145741 ms
catalyst-query-gemma-4-12b-qwen2.5-14b-checked · M2bb7b83ed-722e-49e7-a2d6-9743c66eda904c0e9662-bd76-4c28-af52-1b482dc72a09 → 4c0e9662-bd76-4c28-af52-1b482dc72a09f424417d-9680-4108-b0a6-41aefee1de7a / f424417d-9680-4108-b0a6-41aefee1de7a137944 ms
catalyst-query-gemma-4-12b-qwen2.5-14b-checked · M3bff945ac-4b7a-4703-8db5-ecd01562acc804831f8f-29ac-4d44-9c13-b7f8f98a026d → 04831f8f-29ac-4d44-9c13-b7f8f98a026da0cca333-e97a-470c-85b2-d67f31ff343a / a0cca333-e97a-470c-85b2-d67f31ff343a147088 ms
catalyst-query-gemma-4-12b-qwen2.5-14b-checked · B11345d401-75f2-4067-bfbc-0b9a9512cbc8 → / 75108 ms
catalyst-query-gemma-4-12b-qwen2.5-14b-checked · B223f50b23-1d00-4ccf-a299-ca4ee43a24b939f97292-ab7e-4f4a-87b0-d912fbdeec8a → 39f97292-ab7e-4f4a-87b0-d912fbdeec8a43dbdc08-1cb9-433f-a884-a8a1ef70d38f / 43dbdc08-1cb9-433f-a884-a8a1ef70d38f100348 ms
catalyst-query-gemma-4-12b-qwen2.5-14b-checked · B31946b3b9-e8fc-4ba8-9b76-8bf9ec61dc287d309e11-5df2-4168-a6d8-7b8f3ee2d462 → 7d309e11-5df2-4168-a6d8-7b8f3ee2d462c92fa964-b743-457f-8039-8ec3c570722d / c92fa964-b743-457f-8039-8ec3c570722d92414 ms
catalyst-query-gemma-4-12b-qwen2.5-14b-checked · U1af1fbcf7-3f11-4695-841b-6cded3ebd824 → / 15812 ms
catalyst-query-gemma-4-12b-qwen2.5-14b-checked · U2b971bb0f-5aa5-471f-80c4-fd3bccfbacc3866a1bf8-468d-42c3-8f00-743188b1547d → / 69258 ms

Judge detailadvisory

writer-only catalyst-query-gemma-4-12b

B2 turn 1 · composite 55 — flagged
version 4fd3e3b5-568c-4022-9171-ea665673225e
advisory composite: 55
  • intent_fidelity=1 — The rule was 'latest antiretroviral adherence result is anything other than All', but the SQL ranks over concept_name IN ('Antiretroviral adherence in past week', 'Overall drug adherence in last month') plus a value_coded_name IS NOT NULL filter, changing which observation counts as latest; gold count 3002 vs reference 108 (scenarios/catalyst-query-gemma-4-12b/B2/repetition-01/16-gold-execution-match-successor.json) confirms the material mismatch.
  • sql_quality=2 — The ROW_NUMBER() OVER (PARTITION BY patient_id ORDER BY observed_at DESC) subquery is structurally sound and executed, but ranking one 'latest' across a two-concept union with an extra IS NOT NULL guard misleads about what it computes - workable with issues.
  • schema_discipline=3 — All tables/columns used (hiv_patient_dim_v1; hiv_observation_fact_v1.value_coded_name/observed_at/concept_name) are catalogued and the query executed cleanly; concept names are inline literals, which is acceptable.
  • followup_coherence=1 — It applies the clarification's shape (latest-per-patient, <> 'All') to the original poor-adherence question but misapplies the rule by widening to a second adherence concept the user never gave, which is why the successor count diverges (3002 vs 108).
B3 turn 1 · composite 63 — flagged
version 19cdbf25-ecae-47bc-bc67-51533e28ffee
advisory composite: 63
  • intent_fidelity=1 — The rule requires 'a recorded return visit date before 2026-03-01 with no later visit', but the SQL never touches the 'Return visit date' observation - it selects patients whose hiv_visit_fact_v1.started_at is before :reference_date with no visit on/after it, the wrong entity for the rule; gold fails (5000 capped vs 4665, scenarios/catalyst-query-gemma-4-12b/B3/repetition-01/16-gold-execution-match-successor.json).
  • sql_quality=3 — Readable IN-subquery plus correlated NOT EXISTS that executed fine, but it scans hiv_visit_fact_v1 twice where a single per-patient aggregate (MAX(started_at) < :reference_date) would do - workable with redundancy.
  • schema_discipline=3 — Stays on catalogued hiv_patient_dim_v1/hiv_visit_fact_v1 columns and binds :reference_date as a date correctly (13-execute-successor.json).
  • followup_coherence=1 — It picks up the as-of date 2026-03-01 from the clarification (bound as :reference_date) but misapplies the supplied overdue rule, substituting visit recency for the recorded return-visit-date condition.
M1 turn 0 · composite 84 — flagged
version 01c8c87b-9782-4fb9-8c33-00968152598e
advisory composite: 84
  • intent_fidelity=2 — The executed SQL (scenarios/catalyst-query-gemma-4-12b/M1/repetition-01/06-execute-base.json) is 'SELECT t1.patient_id, t2.family_name, t1.medication_name FROM analytics.hiv_medication_request_fact_v1 ... JOIN analytics.hiv_patient_dim_v1' - a per-prescription listing with exactly the requested patient id, last name and medication; the worklist sql field carries the later aggregated head query, so the per-version execution evidence governs.
  • sql_quality=3 — Simple two-table equi-join projection, clean and executable; succeeded with 100 rows returned under the row cap.
  • schema_discipline=3 — Catalogued hiv_medication_request_fact_v1 and hiv_patient_dim_v1 columns only; no parameters required.
M2 turn 0 · composite 84
version 7e7c66b4-258e-4021-b1d2-f45a4089fbfe
advisory composite: 84
  • intent_fidelity=2 — Grain matches the instruction - COUNT(*) by medication_name (scenarios/catalyst-query-gemma-4-12b/M2/repetition-01/06-execute-base.json) - but the query carries an unrequested 'WHERE do_not_perform = false' predicate; it is the reused catalog counting query and matches the guidance pinned later in the session, so this is minor filter drift rather than a wrong question, scored 2 not 3 because the opening instruction never asked for the exclusion.
  • sql_quality=3 — Minimal correct single-table aggregate; no redundancy.
  • schema_discipline=3 — medication_name and do_not_perform are catalogued hiv_medication_request_fact_v1 columns; no binding issues.
A1 turn 0 · composite 100
version 588e0ccb-bf32-46aa-8e11-acb2844b78dd
advisory composite: 100
  • intent_fidelity=3 — Projects family_name/given_name, value_numeric, value_unit, observed_at and filters concept_name='CD4 count' AND observed_at >= 2026-02-01 via bound parameters - exactly the requested fields and window; gold count matches 11=11 (scenarios/catalyst-query-gemma-4-12b/A1/repetition-01/15-gold-execution-match-base.json).
  • sql_quality=3 — Single clear join of hiv_patient_dim_v1 to hiv_observation_fact_v1 on patient_id with parameterized predicates; no dead branches, executed successfully (06-execute-base.json).
  • schema_discipline=3 — Only catalogued tables/columns; :concept_name (string) and :start_date bound with correct types in scenarios/catalyst-query-gemma-4-12b/A1/repetition-01/06-execute-base.json.
A2 turn 0 · composite 100
version 0722ee75-1892-4c61-b5a3-b0979ad764dd
advisory composite: 100
  • intent_fidelity=3 — COUNT(*) grouped by encounter_type with started_at >= 2025-01-01 and ORDER BY visit_count DESC matches grain, filter and requested ordering; gold aggregate match is clean, 5=5 keys with no mismatches (scenarios/catalyst-query-gemma-4-12b/A2/repetition-01/15-gold-execution-match-base.json).
  • sql_quality=3 — Minimal idiomatic aggregate 'SELECT encounter_type, COUNT(*) ... GROUP BY encounter_type ORDER BY visit_count DESC'; nothing extraneous.
  • schema_discipline=3 — Uses only catalogued hiv_visit_fact_v1.encounter_type/started_at; :start_date bound with a correct temporal type (06-execute-base.json).
A3 turn 0 · composite 100
version b75f7703-2190-4b7a-b5f8-18a6ee441420
advisory composite: 100
  • intent_fidelity=3 — Counts hiv_medication_request_fact_v1 by medication_name with the female filter and 'do_not_perform = false', ordered highest first - every requested constraint present; gold match 30=30 keys, zero valueMismatches (scenarios/catalyst-query-gemma-4-12b/A3/repetition-01/15-gold-execution-match-base.json).
  • sql_quality=3 — Straightforward filtered aggregate with ORDER BY ... DESC; gender supplied via the bound :gender parameter.
  • schema_discipline=3 — Catalogued columns only (medication_name, patient_gender, do_not_perform); execution evidence shows clean binding (06-execute-base.json).
A4 turn 0 · composite 100
version 5e70b8c1-d20b-4443-b108-a78b27e27fac
advisory composite: 100
  • intent_fidelity=3 — 'WHERE ciel_code IS NULL' projecting concept_name and observation_count ordered DESC is exactly the asked listing; it equals the gold reference relation with ordering added and the aggregate match passes 2=2 (scenarios/catalyst-query-gemma-4-12b/A4/repetition-01/15-gold-execution-match-base.json).
  • sql_quality=3 — Minimal single-table select with the one needed predicate and sort; nothing extraneous.
  • schema_discipline=3 — hiv_concept_mapping_v1.concept_name/ciel_code/observation_count are all catalogued; no parameters needed and none misbound.
M1 turn 1 · composite 100
version c3dbec9f-1bbd-4a3f-bcf0-9e3ff68e11b4
advisory composite: 100
  • intent_fidelity=3 — Collapses to one row per patient via STRING_AGG(DISTINCT medication_name, ', ') grouped by t1.patient_id, t2.family_name, and the last-name column is the true hiv_patient_dim_v1.family_name, so both requested fixes are reflected (scenarios/catalyst-query-gemma-4-12b/M1/repetition-01/13-execute-successor.json).
  • sql_quality=3 — Appropriate aggregation with DISTINCT dedup inside STRING_AGG; grouping keys match the non-aggregated projection exactly.
  • schema_discipline=3 — Same catalogued join; STRING_AGG over the text medication_name column is type-correct.
  • followup_coherence=3 — Keeps the base join and projection (patient_id, family_name, medications) while applying the collapse-per-patient change; nothing required from the prior turn was dropped.
M1 turn 2 · composite 100
version 18e942e1-ac8c-4562-90e4-f1bd459b3fbf
advisory composite: 100
  • intent_fidelity=3 — The requested join to the patient table by id for the last name is already present ('JOIN analytics.hiv_patient_dim_v1 ... ON t1.patient_id = t2.patient_id' with family_name projected), so the unchanged SQL satisfies the cumulative request (scenarios/catalyst-query-gemma-4-12b/M1/repetition-01/13-execute-successor-t2.json).
  • sql_quality=3 — Same sound STRING_AGG aggregate grouped by t1.patient_id, t2.family_name as the prior version; no degradation or dead branches introduced.
  • schema_discipline=3 — Catalogued tables/columns only, unchanged from the validated prior version.
  • followup_coherence=3 — Correctly recognizes the follow-up asks for a join the query already has and preserves every prior constraint (grouping, DISTINCT aggregation, family_name) instead of mangling the query.
M2 turn 1 · composite 100
version 8cb08378-fca6-430b-86a4-271c4a670dda
advisory composite: 100
  • intent_fidelity=3 — Adds patient_gender to both projection and GROUP BY exactly as asked; the gold FAIL (59 rows vs 10, scenarios/catalyst-query-gemma-4-12b/M2/repetition-01/16-gold-execution-match-successor.json) reflects the final top-ten reference being applied to this intermediate turn - all 10 reference groups match with zero valueMismatches, so the regrouped counts themselves are correct and the verdict is not inflated by ignoring it.
  • sql_quality=3 — Clean regrouped aggregate; correctly leaves ordering/limit for the later turn that asks for them.
  • schema_discipline=3 — Catalogued columns only; 'do_not_perform = false' keeps a valid boolean predicate form.
  • followup_coherence=3 — Preserves the pinned 'do_not_perform = false' exclusion and the COUNT(*) grain from the base while applying only the requested regroup.
M2 turn 2 · composite 100
version c6a9dcdd-1e48-4eb2-a80d-7725cc3f6bbd
advisory composite: 100
  • intent_fidelity=3 — ORDER BY ... DESC LIMIT 10 returns exactly the ten highest medication-and-gender groups; gold passes 10=10 with no value mismatches (scenarios/catalyst-query-gemma-4-12b/M2/repetition-01/16-gold-execution-match-successor-t2.json).
  • sql_quality=3 — Idiomatic top-N aggregate; nothing extraneous.
  • schema_discipline=3 — Catalogued columns only; no parameters needed.
  • followup_coherence=3 — Keeps the gender+medication grouping and the pinned 'do_not_perform = false' exclusion from the prior turn, adding only the requested top-ten restriction.
M3 turn 0 · composite 100
version cee8779d-da8f-4d86-b0a4-c79e7b4438c7
advisory composite: 100
  • intent_fidelity=3 — The executed SQL counts hiv_observation_fact_v1 rows per patient_id with :concept_name bound to 'CD4 count' (scenarios/catalyst-query-gemma-4-12b/M3/repetition-01/06-execute-base.json) - exactly 'count CD4 count results by patient'; the worklist sql field shows the later visits head query, so the per-version execution evidence governs.
  • sql_quality=3 — Simple parameterized aggregate with the right GROUP BY grain.
  • schema_discipline=3 — Catalogued observation-fact columns; concept bound as a string parameter.
M3 turn 1 · composite 100
version 0c323a0f-6aca-4181-acd8-1482bb6c07d0
advisory composite: 100
  • intent_fidelity=3 — Same per-patient count with the concept parameter switched to 'CD4%' - the catalog's CD4 percentage concept (scenarios/catalyst-query-gemma-4-12b/M3/repetition-01/13-execute-successor.json shows :concept_name bound to 'CD4%').
  • sql_quality=3 — Clean parameterized aggregate reuse; grain unchanged as requested.
  • schema_discipline=3 — Catalogued columns; the concept parameter is re-bound correctly as a string.
  • followup_coherence=3 — Keeps the verified per-patient counting shape and changes only the concept - exactly the 'instead' the follow-up asked for.
M3 turn 2 · composite 100
version d04ed859-7cb5-4726-8d23-8b76bcca6baf
advisory composite: 100
  • intent_fidelity=3 — Counts hiv_visit_fact_v1 rows by encounter_type - the fresh question asked; 11 groups returned (scenarios/catalyst-query-gemma-4-12b/M3/repetition-01/13-execute-successor-t2.json).
  • sql_quality=3 — Minimal correct aggregate.
  • schema_discipline=3 — Catalogued visit-fact columns only.
  • followup_coherence=3 — 'Separately' is honored: it starts a fresh visits-by-encounter-type query and does not drag the prior CD4 concept filter or per-patient grain into it.

self-checked catalyst-query-gemma-4-12b-q4-checked

B3 turn 1 · composite 63 — flagged
version d617fae2-4b11-4854-8305-77aee19377ec
advisory composite: 63
  • intent_fidelity=1 — The rule requires 'a recorded return visit date before 2026-03-01 with no later visit', but the SQL never touches the 'Return visit date' observation - it selects patients whose hiv_visit_fact_v1.started_at is before :reference_date with no visit on/after it, the wrong entity for the rule; gold fails (5000 capped vs 4665, scenarios/catalyst-query-gemma-4-12b-q4-checked/B3/repetition-01/16-gold-execution-match-successor.json).
  • sql_quality=3 — Readable IN-subquery plus correlated NOT EXISTS that executed fine, but it scans hiv_visit_fact_v1 twice where a single per-patient aggregate (MAX(started_at) < :reference_date) would do - workable with redundancy.
  • schema_discipline=3 — Stays on catalogued hiv_patient_dim_v1/hiv_visit_fact_v1 columns and binds :reference_date as a date correctly (13-execute-successor.json).
  • followup_coherence=1 — It picks up the as-of date 2026-03-01 from the clarification (bound as :reference_date) but misapplies the supplied overdue rule, substituting visit recency for the recorded return-visit-date condition.
M1 turn 0 · composite 84 — flagged
version 63bc7eb0-ffa1-4385-b3e0-993211f7ae1f
advisory composite: 84
  • intent_fidelity=2 — The executed SQL (scenarios/catalyst-query-gemma-4-12b-q4-checked/M1/repetition-01/06-execute-base.json) is 'SELECT t1.patient_id, t2.family_name, t1.medication_name FROM analytics.hiv_medication_request_fact_v1 ... JOIN analytics.hiv_patient_dim_v1' - a per-prescription listing with exactly the requested patient id, last name and medication; the worklist sql field carries the later aggregated head query, so the per-version execution evidence governs.
  • sql_quality=3 — Simple two-table equi-join projection, clean and executable; succeeded with 100 rows returned under the row cap.
  • schema_discipline=3 — Catalogued hiv_medication_request_fact_v1 and hiv_patient_dim_v1 columns only; no parameters required.
M2 turn 0 · composite 84 — flagged
version 06305913-1820-4193-bb03-555fd5d96764
advisory composite: 84
  • intent_fidelity=2 — Grain matches the instruction - COUNT(*) by medication_name (scenarios/catalyst-query-gemma-4-12b-q4-checked/M2/repetition-01/06-execute-base.json) - but the query carries an unrequested 'WHERE do_not_perform = false' predicate; it is the reused catalog counting query and matches the guidance pinned later in the session, so this is minor filter drift rather than a wrong question, scored 2 not 3 because the opening instruction never asked for the exclusion.
  • sql_quality=3 — Minimal correct single-table aggregate; no redundancy.
  • schema_discipline=3 — medication_name and do_not_perform are catalogued hiv_medication_request_fact_v1 columns; no binding issues.
B2 turn 1 · composite 87
version f37ec15e-2f1e-4ae0-b316-ade6ed18d50c
advisory composite: 87
  • intent_fidelity=2 — Implements 'latest antiretroviral adherence <> All' via DISTINCT ON (patient_id) ... ORDER BY patient_id, observed_at DESC over the single bound concept, and gold passes 108=108 (scenarios/catalyst-query-gemma-4-12b-q4-checked/B2/repetition-01/16-gold-execution-match-successor.json); scored 2 not 3 only for the unrequested "obs_status = 'final'" predicate - minor drift that provably did not change the result.
  • sql_quality=3 — Idiomatic Postgres DISTINCT ON latest-per-patient pattern; compact, clear, and executed cleanly.
  • schema_discipline=3 — hiv_observation_fact_v1.obs_status/value_coded_name/observed_at and the patient-dim columns are all catalogued (the query executed and returned rows); :concept_name bound as string (13-execute-successor.json).
  • followup_coherence=3 — Applies the clarified rule to the original 'patients with poor adherence' question with the right concept and latest-per-patient semantics; the successor matches the reference count exactly.
A1 turn 0 · composite 100
version bfd41602-5c23-4952-aad2-fb9a4fb4fce9
advisory composite: 100
  • intent_fidelity=3 — Projects family_name/given_name, value_numeric, value_unit, observed_at and filters concept_name='CD4 count' AND observed_at >= 2026-02-01 via bound parameters - exactly the requested fields and window; gold count matches 11=11 (scenarios/catalyst-query-gemma-4-12b-q4-checked/A1/repetition-01/15-gold-execution-match-base.json).
  • sql_quality=3 — Single clear join of hiv_patient_dim_v1 to hiv_observation_fact_v1 on patient_id with parameterized predicates; no dead branches, executed successfully (06-execute-base.json).
  • schema_discipline=3 — Only catalogued tables/columns; :concept_name (string) and :start_date bound with correct types in scenarios/catalyst-query-gemma-4-12b-q4-checked/A1/repetition-01/06-execute-base.json.
A2 turn 0 · composite 100
version a6ee49bc-82f7-410e-8132-9b8595b2bb78
advisory composite: 100
  • intent_fidelity=3 — COUNT(*) grouped by encounter_type with started_at >= 2025-01-01 and ORDER BY visit_count DESC matches grain, filter and requested ordering; gold aggregate match is clean, 5=5 keys with no mismatches (scenarios/catalyst-query-gemma-4-12b-q4-checked/A2/repetition-01/15-gold-execution-match-base.json).
  • sql_quality=3 — Minimal idiomatic aggregate 'SELECT encounter_type, COUNT(*) ... GROUP BY encounter_type ORDER BY visit_count DESC'; nothing extraneous.
  • schema_discipline=3 — Uses only catalogued hiv_visit_fact_v1.encounter_type/started_at; :start_date bound with a correct temporal type (06-execute-base.json).
A3 turn 0 · composite 100
version d53043f1-739a-4f5a-b13d-c37a567affa9
advisory composite: 100
  • intent_fidelity=3 — Counts hiv_medication_request_fact_v1 by medication_name with the female filter and 'do_not_perform = false', ordered highest first - every requested constraint present; gold match 30=30 keys, zero valueMismatches (scenarios/catalyst-query-gemma-4-12b-q4-checked/A3/repetition-01/15-gold-execution-match-base.json).
  • sql_quality=3 — Straightforward filtered aggregate with ORDER BY ... DESC; gender appears as the inline literal 'female' rather than a bound parameter, a negligible style choice.
  • schema_discipline=3 — Catalogued columns only (medication_name, patient_gender, do_not_perform); execution evidence shows clean binding (06-execute-base.json).
A4 turn 0 · composite 100
version 7a9523ff-0a32-4e44-8e45-4d69c28fc24b
advisory composite: 100
  • intent_fidelity=3 — 'WHERE ciel_code IS NULL' projecting concept_name and observation_count ordered DESC is exactly the asked listing; it equals the gold reference relation with ordering added and the aggregate match passes 2=2 (scenarios/catalyst-query-gemma-4-12b-q4-checked/A4/repetition-01/15-gold-execution-match-base.json).
  • sql_quality=3 — Minimal single-table select with the one needed predicate and sort; nothing extraneous.
  • schema_discipline=3 — hiv_concept_mapping_v1.concept_name/ciel_code/observation_count are all catalogued; no parameters needed and none misbound.
M1 turn 1 · composite 100
version 2c9b2c73-9b7b-4114-a3a9-e1a1d1b232b2
advisory composite: 100
  • intent_fidelity=3 — Collapses to one row per patient via STRING_AGG(DISTINCT medication_name, ', ') grouped by t1.patient_id, t2.family_name, and the last-name column is the true hiv_patient_dim_v1.family_name, so both requested fixes are reflected (scenarios/catalyst-query-gemma-4-12b-q4-checked/M1/repetition-01/13-execute-successor.json).
  • sql_quality=3 — Appropriate aggregation with DISTINCT dedup inside STRING_AGG; grouping keys match the non-aggregated projection exactly.
  • schema_discipline=3 — Same catalogued join; STRING_AGG over the text medication_name column is type-correct.
  • followup_coherence=3 — Keeps the base join and projection (patient_id, family_name, medications) while applying the collapse-per-patient change; nothing required from the prior turn was dropped.
M1 turn 2 · composite 100
version 62c5bddd-7e4f-432c-893e-abb23b522574
advisory composite: 100
  • intent_fidelity=3 — The requested join to the patient table by id for the last name is already present ('JOIN analytics.hiv_patient_dim_v1 ... ON t1.patient_id = t2.patient_id' with family_name projected), so the unchanged SQL satisfies the cumulative request (scenarios/catalyst-query-gemma-4-12b-q4-checked/M1/repetition-01/13-execute-successor-t2.json).
  • sql_quality=3 — Same sound STRING_AGG aggregate grouped by t1.patient_id, t2.family_name as the prior version; no degradation or dead branches introduced.
  • schema_discipline=3 — Catalogued tables/columns only, unchanged from the validated prior version.
  • followup_coherence=3 — Correctly recognizes the follow-up asks for a join the query already has and preserves every prior constraint (grouping, DISTINCT aggregation, family_name) instead of mangling the query.
M2 turn 1 · composite 100
version 9c2e5eab-2c18-4aa2-b1df-ac7ae9f0379e
advisory composite: 100
  • intent_fidelity=3 — Adds patient_gender to both projection and GROUP BY exactly as asked; the gold FAIL (59 rows vs 10, scenarios/catalyst-query-gemma-4-12b-q4-checked/M2/repetition-01/16-gold-execution-match-successor.json) reflects the final top-ten reference being applied to this intermediate turn - all 10 reference groups match with zero valueMismatches, so the regrouped counts themselves are correct and the verdict is not inflated by ignoring it.
  • sql_quality=3 — Clean regrouped aggregate; correctly leaves ordering/limit for the later turn that asks for them.
  • schema_discipline=3 — Catalogued columns only; 'do_not_perform = false' keeps a valid boolean predicate form.
  • followup_coherence=3 — Preserves the pinned 'do_not_perform = false' exclusion and the COUNT(*) grain from the base while applying only the requested regroup.
M2 turn 2 · composite 100
version 714f1fac-4635-4618-9b7c-474f51c94b84
advisory composite: 100
  • intent_fidelity=3 — ORDER BY ... DESC LIMIT 10 returns exactly the ten highest medication-and-gender groups; gold passes 10=10 with no value mismatches (scenarios/catalyst-query-gemma-4-12b-q4-checked/M2/repetition-01/16-gold-execution-match-successor-t2.json).
  • sql_quality=3 — Idiomatic top-N aggregate; nothing extraneous.
  • schema_discipline=3 — Catalogued columns only; no parameters needed.
  • followup_coherence=3 — Keeps the gender+medication grouping and the pinned 'do_not_perform = false' exclusion from the prior turn, adding only the requested top-ten restriction.
M3 turn 0 · composite 100
version 75bd224d-4edd-48ed-89a3-700007352684
advisory composite: 100
  • intent_fidelity=3 — The executed SQL counts hiv_observation_fact_v1 rows per patient_id with :concept_name bound to 'CD4 count' (scenarios/catalyst-query-gemma-4-12b-q4-checked/M3/repetition-01/06-execute-base.json) - exactly 'count CD4 count results by patient'; the worklist sql field shows the later visits head query, so the per-version execution evidence governs.
  • sql_quality=3 — Simple parameterized aggregate with the right GROUP BY grain.
  • schema_discipline=3 — Catalogued observation-fact columns; concept bound as a string parameter.
M3 turn 2 · composite 100
version cc635f9b-7a9c-4b4f-a6a3-0bf2c742ba21
advisory composite: 100
  • intent_fidelity=3 — Counts hiv_visit_fact_v1 rows by encounter_type - the fresh question asked; 11 groups returned (scenarios/catalyst-query-gemma-4-12b-q4-checked/M3/repetition-01/13-execute-successor-t2.json).
  • sql_quality=3 — Minimal correct aggregate.
  • schema_discipline=3 — Catalogued visit-fact columns only.
  • followup_coherence=3 — 'Separately' is honored: it starts a fresh visits-by-encounter-type query and does not drag the prior CD4 concept filter or per-patient grain into it. (This team's turn-1 followup failed with no executed version - 10-final-turns.json shows ordinal 2 status 'failed' - so the prior selected version is still the turn-0 CD4-per-patient query, and none of its constraints belong here.)

qwen2.5-14b-checked catalyst-query-gemma-4-12b-qwen2.5-14b-checked

B3 turn 1 · composite 63 — flagged
version 7d309e11-5df2-4168-a6d8-7b8f3ee2d462
advisory composite: 63
  • intent_fidelity=1 — The rule requires 'a recorded return visit date before 2026-03-01 with no later visit', but the SQL never touches the 'Return visit date' observation - it selects patients whose hiv_visit_fact_v1.started_at is before :reference_date with no visit on/after it, the wrong entity for the rule; gold fails (5000 capped vs 4665, scenarios/catalyst-query-gemma-4-12b-qwen2.5-14b-checked/B3/repetition-01/16-gold-execution-match-successor.json).
  • sql_quality=3 — Readable IN-subquery plus correlated NOT EXISTS that executed fine, but it scans hiv_visit_fact_v1 twice where a single per-patient aggregate (MAX(started_at) < :reference_date) would do - workable with redundancy.
  • schema_discipline=3 — Stays on catalogued hiv_patient_dim_v1/hiv_visit_fact_v1 columns and binds :reference_date as a date correctly (13-execute-successor.json).
  • followup_coherence=1 — It picks up the as-of date 2026-03-01 from the clarification (bound as :reference_date) but misapplies the supplied overdue rule, substituting visit recency for the recorded return-visit-date condition.
M1 turn 0 · composite 84 — flagged
version 05b4d2f9-81c7-44ba-af5f-8686c90e6aec
advisory composite: 84
  • intent_fidelity=2 — The executed SQL (scenarios/catalyst-query-gemma-4-12b-qwen2.5-14b-checked/M1/repetition-01/06-execute-base.json) is 'SELECT t1.patient_id, t2.family_name, t1.medication_name FROM analytics.hiv_medication_request_fact_v1 ... JOIN analytics.hiv_patient_dim_v1' - a per-prescription listing with exactly the requested patient id, last name and medication; the worklist sql field carries the later aggregated head query, so the per-version execution evidence governs.
  • sql_quality=3 — Simple two-table equi-join projection, clean and executable; succeeded with 100 rows returned under the row cap.
  • schema_discipline=3 — Catalogued hiv_medication_request_fact_v1 and hiv_patient_dim_v1 columns only; no parameters required.
M2 turn 0 · composite 84 — flagged
version 1c624b17-699c-4f6a-82b1-a314bd83f263
advisory composite: 84
  • intent_fidelity=2 — Grain matches the instruction - COUNT(*) by medication_name (scenarios/catalyst-query-gemma-4-12b-qwen2.5-14b-checked/M2/repetition-01/06-execute-base.json) - but the query carries an unrequested 'WHERE do_not_perform IS FALSE' predicate; it is the reused catalog counting query and matches the guidance pinned later in the session, so this is minor filter drift rather than a wrong question, scored 2 not 3 because the opening instruction never asked for the exclusion.
  • sql_quality=3 — Minimal correct single-table aggregate; no redundancy.
  • schema_discipline=3 — medication_name and do_not_perform are catalogued hiv_medication_request_fact_v1 columns; no binding issues.
A4 turn 0 · composite 90
version 826f816a-1441-4aec-a250-67a7786f0f34
advisory composite: 90
  • intent_fidelity=3 — Lists concepts with ciel_code IS NULL alongside their observation counts, highest first; the gold aggregate matches 2=2 with total_observation_count resolving to observation_count (scenarios/catalyst-query-gemma-4-12b-qwen2.5-14b-checked/A4/repetition-01/15-gold-execution-match-base.json).
  • sql_quality=2 — Wraps the already per-concept observation_count in 'SUM(t1.observation_count) ... GROUP BY t1.concept_name', a redundant aggregation over a relation the gold reference reads directly - workable, minor redundancy, so 2.
  • schema_discipline=3 — Catalogued hiv_concept_mapping_v1 columns only; no parameters needed.
A1 turn 0 · composite 100
version 3e3f2d48-09aa-4df1-b18a-03e4d7aa32c6
advisory composite: 100
  • intent_fidelity=3 — Projects family_name/given_name, value_numeric, value_unit, observed_at and filters concept_name='CD4 count' AND observed_at >= 2026-02-01 via bound parameters - exactly the requested fields and window; gold count matches 11=11 (scenarios/catalyst-query-gemma-4-12b-qwen2.5-14b-checked/A1/repetition-01/15-gold-execution-match-base.json).
  • sql_quality=3 — Single clear join of hiv_patient_dim_v1 to hiv_observation_fact_v1 on patient_id with parameterized predicates; no dead branches, executed successfully (06-execute-base.json).
  • schema_discipline=3 — Only catalogued tables/columns; :concept_name (string) and :start_date bound with correct types in scenarios/catalyst-query-gemma-4-12b-qwen2.5-14b-checked/A1/repetition-01/06-execute-base.json.
A2 turn 0 · composite 100
version 807c2af7-13e7-4753-a869-47e8f2e74a33
advisory composite: 100
  • intent_fidelity=3 — COUNT(*) grouped by encounter_type with started_at >= 2025-01-01 and ORDER BY visit_count DESC matches grain, filter and requested ordering; gold aggregate match is clean, 5=5 keys with no mismatches (scenarios/catalyst-query-gemma-4-12b-qwen2.5-14b-checked/A2/repetition-01/15-gold-execution-match-base.json).
  • sql_quality=3 — Minimal idiomatic aggregate 'SELECT encounter_type, COUNT(*) ... GROUP BY encounter_type ORDER BY visit_count DESC'; nothing extraneous.
  • schema_discipline=3 — Uses only catalogued hiv_visit_fact_v1.encounter_type/started_at; :start_date bound with a correct temporal type (06-execute-base.json).
A3 turn 0 · composite 100
version bf90b391-375f-4136-88cd-f6bfd8cc6e48
advisory composite: 100
  • intent_fidelity=3 — Counts hiv_medication_request_fact_v1 by medication_name with the female filter and 'do_not_perform = false', ordered highest first - every requested constraint present; gold match 30=30 keys, zero valueMismatches (scenarios/catalyst-query-gemma-4-12b-qwen2.5-14b-checked/A3/repetition-01/15-gold-execution-match-base.json).
  • sql_quality=3 — Straightforward filtered aggregate with ORDER BY ... DESC; gender supplied via the bound :gender parameter and a descriptive request_count alias.
  • schema_discipline=3 — Catalogued columns only (medication_name, patient_gender, do_not_perform); execution evidence shows clean binding (06-execute-base.json).
B2 turn 1 · composite 100
version 39f97292-ab7e-4f4a-87b0-d912fbdeec8a
advisory composite: 100
  • intent_fidelity=3 — The CTE ranks only the bound 'Antiretroviral adherence in past week' observations per patient by observed_at DESC and keeps rank 1 with value_coded_name != 'All' - precisely the supplied rule; gold passes 108=108 (scenarios/catalyst-query-gemma-4-12b-qwen2.5-14b-checked/B2/repetition-01/16-gold-execution-match-successor.json).
  • sql_quality=3 — Clear CTE + ROW_NUMBER latest-per-patient pattern with a tidy latest_adherence_result projection; no dead branches.
  • schema_discipline=3 — Catalogued observation-fact and patient-dim columns only; :adherence_concept bound as string (13-execute-successor.json).
  • followup_coherence=3 — Applies the clarified definition to the original poor-adherence question without adding or dropping constraints; the count matches the reference exactly.
M1 turn 1 · composite 100
version 4febffe3-713c-4b4c-bc57-ed784ff39603
advisory composite: 100
  • intent_fidelity=3 — Collapses to one row per patient via STRING_AGG(DISTINCT medication_name, ', ') grouped by t2.patient_id, t2.family_name, and the last-name column is the true hiv_patient_dim_v1.family_name, so both requested fixes are reflected (scenarios/catalyst-query-gemma-4-12b-qwen2.5-14b-checked/M1/repetition-01/13-execute-successor.json).
  • sql_quality=3 — Appropriate aggregation with DISTINCT dedup inside STRING_AGG; grouping keys match the non-aggregated projection exactly.
  • schema_discipline=3 — Same catalogued join; STRING_AGG over the text medication_name column is type-correct.
  • followup_coherence=3 — Keeps the base join and projection (patient_id, family_name, medications) while applying the collapse-per-patient change; nothing required from the prior turn was dropped.
M1 turn 2 · composite 100
version e4cb0a9c-f1f7-4a78-9048-609c20649963
advisory composite: 100
  • intent_fidelity=3 — The requested join to the patient table by id for the last name is already present ('JOIN analytics.hiv_patient_dim_v1 ... ON t1.patient_id = t2.patient_id' with family_name projected), so the unchanged SQL satisfies the cumulative request (scenarios/catalyst-query-gemma-4-12b-qwen2.5-14b-checked/M1/repetition-01/13-execute-successor-t2.json).
  • sql_quality=3 — Same sound STRING_AGG aggregate grouped by t2.patient_id, t2.family_name as the prior version; no degradation or dead branches introduced.
  • schema_discipline=3 — Catalogued tables/columns only, unchanged from the validated prior version.
  • followup_coherence=3 — Correctly recognizes the follow-up asks for a join the query already has and preserves every prior constraint (grouping, DISTINCT aggregation, family_name) instead of mangling the query.
M2 turn 1 · composite 100
version 79fb851e-38eb-48c5-9df6-e259f68389b8
advisory composite: 100
  • intent_fidelity=3 — Adds patient_gender to both projection and GROUP BY exactly as asked; the gold FAIL (59 rows vs 10, scenarios/catalyst-query-gemma-4-12b-qwen2.5-14b-checked/M2/repetition-01/16-gold-execution-match-successor.json) reflects the final top-ten reference being applied to this intermediate turn - all 10 reference groups match with zero valueMismatches, so the regrouped counts themselves are correct and the verdict is not inflated by ignoring it.
  • sql_quality=3 — Clean regrouped aggregate; correctly leaves ordering/limit for the later turn that asks for them.
  • schema_discipline=3 — Catalogued columns only; 'do_not_perform IS FALSE' keeps a valid boolean predicate form.
  • followup_coherence=3 — Preserves the pinned 'do_not_perform IS FALSE' exclusion and the COUNT(*) grain from the base while applying only the requested regroup.
M2 turn 2 · composite 100
version 4c0e9662-bd76-4c28-af52-1b482dc72a09
advisory composite: 100
  • intent_fidelity=3 — ORDER BY ... DESC LIMIT 10 returns exactly the ten highest medication-and-gender groups; gold passes 10=10 with no value mismatches (scenarios/catalyst-query-gemma-4-12b-qwen2.5-14b-checked/M2/repetition-01/16-gold-execution-match-successor-t2.json).
  • sql_quality=3 — Idiomatic top-N aggregate; nothing extraneous.
  • schema_discipline=3 — Catalogued columns only; no parameters needed.
  • followup_coherence=3 — Keeps the gender+medication grouping and the pinned 'do_not_perform IS FALSE' exclusion from the prior turn, adding only the requested top-ten restriction.
M3 turn 0 · composite 100
version 07b3fdb0-1416-44f6-a6b8-579f6156970b
advisory composite: 100
  • intent_fidelity=3 — The executed SQL counts hiv_observation_fact_v1 rows per patient_id with :cd4_count_name bound to 'CD4 count' (scenarios/catalyst-query-gemma-4-12b-qwen2.5-14b-checked/M3/repetition-01/06-execute-base.json) - exactly 'count CD4 count results by patient'; the worklist sql field shows the later visits head query, so the per-version execution evidence governs.
  • sql_quality=3 — Simple parameterized aggregate with the right GROUP BY grain.
  • schema_discipline=3 — Catalogued observation-fact columns; concept bound as a string parameter.
M3 turn 1 · composite 100
version 9fecf9b9-70a4-4eb9-a83f-41310f621178
advisory composite: 100
  • intent_fidelity=3 — Same per-patient count with the concept parameter switched to 'CD4%' - the catalog's CD4 percentage concept (scenarios/catalyst-query-gemma-4-12b-qwen2.5-14b-checked/M3/repetition-01/13-execute-successor.json shows :cd4_percentage_name bound to 'CD4%').
  • sql_quality=3 — Clean parameterized aggregate reuse; grain unchanged as requested.
  • schema_discipline=3 — Catalogued columns; the concept parameter is re-bound correctly as a string.
  • followup_coherence=3 — Keeps the verified per-patient counting shape and changes only the concept - exactly the 'instead' the follow-up asked for.
M3 turn 2 · composite 100
version 04831f8f-29ac-4d44-9c13-b7f8f98a026d
advisory composite: 100
  • intent_fidelity=3 — Counts hiv_visit_fact_v1 rows by encounter_type - the fresh question asked; 11 groups returned (scenarios/catalyst-query-gemma-4-12b-qwen2.5-14b-checked/M3/repetition-01/13-execute-successor-t2.json).
  • sql_quality=3 — Minimal correct aggregate.
  • schema_discipline=3 — Catalogued visit-fact columns only.
  • followup_coherence=3 — 'Separately' is honored: it starts a fresh visits-by-encounter-type query and does not drag the prior CD4 concept filter or per-patient grain into it.

Methods & provenance

How this works, the dataset, and the model lineup

Each conversation runs live: a question generates SQL (a writer model drafts it; reviewed profiles also invoke their configured reviewer), the query is validated and executed against PostgreSQL, then a follow-up instruction refines the exact current query and the successor is validated and executed again. Executed results are re-checked against an independently-authored gold query (byte-level row-set match) and an independent read-only PostgreSQL cross-check.

Dataset hiv-20260731T000741Z (openmrs-hiv-fhir-postgresql): 5285 patients · 427915 results · 143 test types · 2022-11-07 – 2026-08-11

profilewriter modelreviewer model
catalyst-query-gemma-4-12bgemma-4-12b-q4— (writer only)
catalyst-query-gemma-4-12b-q4-checkedgemma-4-12b-q4gemma-4-12b-q4
catalyst-query-gemma-4-12b-qwen2.5-14b-checkedgemma-4-12bqwen2.5-14b

run_id=9ae123db-8f40-4246-8769-d427a5551769 · evidence_status=development · dataset=hiv-20260731T000741Z · catalog=openmrs-hiv-catalog-v6 · provider=llama.cpp · profiles=catalyst-query-gemma-4-12b,catalyst-query-gemma-4-12b-q4-checked,catalyst-query-gemma-4-12b-qwen2.5-14b-checked