Recorded ChartSearchAI answer evaluations and Catalyst query-workflow checks. Each report identifies its data, methods and limits. Historical findings do not establish current product or deployment acceptance.
Newest reports first.
How to read these reports. For reports with score tables, each answer is graded 0–10, higher is better — Accuracy: are the stated facts correct? Completeness: did it cover what mattered? Relevance: did it actually answer the question? Unsafe answers counts responses with a potential safety problem (lower is better). Questions is how many were scored per setup, and the best value in each column is highlighted. Benchmark /100 is the headline aggregate — it combines the dimensions (accuracy and completeness weighted most) minus soft penalties for safety and citation problems, so one number summarizes each setup. These come from a single automated judge on a small sample — they're directional, not a formal benchmark.
Catalyst SQL
Catalyst Phase 1: three model teams on the locked HIV suite
Historical development report. Its methods, acceptance rules and conclusions apply to that recorded run, not the current release.
Catalyst turns plain-English clinical questions into database queries it runs itself. Three AI team configurations each faced the same twelve questions about an HIV dataset — including vague ones they had to ask about, and impossible ones they had to refuse — with every answer checked row-for-row against an independently written reference.
Open the report for its recorded methods and results.
TakeawayNo team cleared the acceptance bar (≥90% overall, ≥80% per scenario). Best: the qwen-reviewed team at 8/12. An independent AI judge rated the SQL itself consistently well-built (median 100/100 across 44 queries) — the failures came from misreading ambiguous questions, the same two follow-ups for every team. Next lever: question understanding, not query craft.
Catalyst SQL
Catalyst notebook acceptance: can an AI-drafted SQL conversation survive real refinement?
Historical development report. Its methods, acceptance rules and conclusions apply to that recorded run, not the current release.
A different validation family from the chart-QA runs above: Catalyst turns a plain-language clinical question into governed SQL (a Gemma 4 12B writer drafted, a Qwen 2.5 14B reviewer checked, a deterministic policy enforcing read-only execution), then a follow-up instruction refines the exact current query. Four scenario families × 3 repetitions against the live stack and a real PostgreSQL: narrowing an unchanged base, aggregating from a dirty editor with a mid-session profile switch, correcting an unresolved parameter, and a semantic distinct-patient review. Every executed result is re-checked two ways — against an independently-run gold query (byte-level execution match) and by an independent read-only PostgreSQL cross-check.
TakeawayAll 12 scenario repetitions passed — 384/384 assertions, including every gold execution-match: the SQL the models produced returned exactly the rows an independently-authored gold query returns, on every repetition. This is an acceptance gate on one synthetic 96-patient laboratory cohort, not a benchmark — it demonstrates the governed notebook loop (draft → review → validate → execute → refine) is reliable end-to-end on real infrastructure, not that the models never err (the advisory validator and deterministic policy exist precisely because they do).
A 12-scenario product-path evaluation of the checked 12B profile. All cells completed, the implemented deterministic audit reported no blockers, and an independent Scout judge scored the answers 88.4/100 with no harm and 66/66 references resolved. The judge nevertheless found two strict six-month-window errors that the deterministic gate did not detect.
AI setup
Questions
Benchmark /100
Accuracy
Completeness
Relevance
Unsafe answers
product-12b-checked
12
88.4
9.2
8.7
9.0
0
In-Depth — scored separately on its own axes (its own parity Benchmark)
AI setup
In-Depth Benchmark
In-depth answers
Support
Added value
Unsafe
Padded
product-12b-checked
71.7
6
8.3
6.0
0
0
TakeawayThe checked 12B path is strong on exact dates, visit ordering, appointments, and weight trends, but it is not release-clean: one answer included an X-ray six days before the requested cutoff, while another included 10 out-of-window orders and omitted a qualifying order. This report demonstrates why deterministic checks and independent judging must remain separate, visible layers; the next correction is a generic dated-row window gate, not prompt tuning to these cells.
Twelve temporal and chart-grounding scenarios compare answer-only, deterministic temporal checking, and full checked profiles on E4B and 12B.
AI setup
Questions
Benchmark /100
Accuracy
Completeness
Relevance
Unsafe answers
Gemma 4 4B (E4B) · Q4 · fully checked
12
91.2
9.5
8.8
9.4
0
Gemma 4 4B (E4B) · Q4 · deterministic check
12
88.8
9.5
8.8
9.1
0
Gemma 4 12B · Q8 · fully checked
12
88.7
9.2
8.7
9.0
0
Gemma 4 12B · Q8 · answer only
12
88.2
9.2
9.1
9.2
0
Gemma 4 12B · Q8 · deterministic check
12
86.2
9.2
8.7
9.0
0
Gemma 4 4B (E4B) · Q4 · answer only
12
82.8
8.8
8.2
8.8
0
In-Depth — scored separately on its own axes (its own parity Benchmark)
AI setup
In-Depth Benchmark
In-depth answers
Support
Added value
Unsafe
Padded
Gemma 4 12B · Q8 · fully checked
74.2
6
8.8
6.0
0
0
Gemma 4 4B (E4B) · Q4 · fully checked
68.6
7
9.6
4.4
0
2
TakeawayDeterministic checking fixed E4B's past-appointment error, but did not fix strict six-month windows and degraded one 12B child-growth answer. The full checked E4B profile scored highest while adding reviewed, resolved evidence.
An 18-cell product-path diagnostic across the E4B single, 12B single, and medical-team profiles. All cells completed and stage timing was captured, but deterministic QA found answer-contract and grounding issues. This run is unjudged and is not a model-quality comparison.
Open the report for its recorded methods and results.
TakeawayThe consolidated hub path, readable team metadata, and stage observability worked end to end. The run exposed table-shape, exact-fact grounding, citation-set, temporal-fallback reporting, and team context-budget issues that are being remediated before the judged candidate run.
ChartSearchAI
Lever sweep (dev): coverage prompt + rewrite-validator across a dense 12B and a cheap MoE (A4B)
A fast 12-scenario dev sweep, holding two writers (Gemma-12B dense, Gemma-26B-A4B MoE ~4B active) against {bare, +coverage prompt, +rewrite-validator, +both}. Judged AFTER fixing a scoring bug where the judge couldn't see table answers (which had been deflating completeness). N=12/arm — directional, not definitive.
AI setup
Questions
Benchmark /100
Accuracy
Completeness
Relevance
Unsafe answers
wide-31b-contract-warn
12
86.5
9.0
8.2
9.2
0
wide-31b-current-warn
12
84.2
8.8
8.2
9.0
0
date-12b-contract-warn
12
83.5
8.9
7.9
9.0
1
date-12b-current-warn
12
80.2
8.5
7.6
8.7
0
date-a4b-contract-warn
12
77.3
8.3
7.3
8.8
1
wide-team-12b-contract-warn
12
73.1
7.6
7.8
7.8
1
wide-team-high-contract-warn
12
6.9
1.2
1.1
1.2
0
In-Depth — scored separately on its own axes (its own parity Benchmark)
AI setup
In-Depth Benchmark
In-depth answers
Support
Added value
Unsafe
Padded
wide-team-12b-contract-warn
71.4
11
7.6
6.9
1
0
TakeawayTwo findings. (1) A SCORING bug was capping the benchmark: the judge scored only the prose answer and never saw the structured table answers the prompts ask for, deflating completeness ~8pts; fixed. (2) The optimal config is WRITER-DEPENDENT: the dense 12B is best with coverage+rewrite-validator (83.5, 0 harm) vs bare 80.6, but the cheap A4B MoE is best BARE (83.7, 0 harm, highest completeness) and EVERY lever hurts it. Bare A4B edges the best scaffolded 12B — the cheap MoE alone is competitive, the on-device headline. The orchestrator hurts both (12B +rewrite: 72 w/ orch vs 78 w/o). Caveat: N=12, single judge, hard-weighted dev subset; needs full-run confirmation.
ChartSearchAI
Method scaffolding: validator-team vs the single writer (12B, A4B)
Does a small-orchestrator + Qwen-2.5-14B-validator team beat the bare single writer on the two strong writers (Gemma-12B, Gemma-26B-A4B)? Plus a cite-or-abstain prompt lever. 297 cells, judged.
AI setup
Questions
Benchmark /100
Accuracy
Completeness
Relevance
Unsafe answers
m-12b-rewriteval
62
80.7
8.3
8.0
8.8
2
m-12b-solo
62
80.1
8.4
7.7
8.7
0
m-a4b-team
62
80.0
8.4
7.7
8.6
0
m-a4b-solo
62
78.3
8.1
7.7
8.6
2
m-12b-team
62
77.1
8.2
7.4
8.2
1
m-12b-cite-abstain
62
76.4
8.2
7.3
8.5
3
In-Depth — scored separately on its own axes (its own parity Benchmark)
AI setup
In-Depth Benchmark
In-depth answers
Support
Added value
Unsafe
Padded
m-a4b-team
75.5
40
7.7
7.4
0
0
m-12b-solo
75.0
42
7.4
7.6
0
0
m-12b-team
74.8
43
7.6
7.4
0
2
m-a4b-solo
74.5
39
7.5
7.4
0
0
m-12b-rewriteval
73.7
42
7.5
7.3
0
1
m-12b-cite-abstain
73.5
42
7.3
7.5
1
1
TakeawayThe validator's value scales inversely with writer strength: it rescues the weaker A4B (2 harm to 0, +~3pts) but doesn't help the already-safe 12B, best as a bare single (~81, 0 harm). cite-or-abstain lowers the 12B via over-caution.
ChartSearchAI
Does 4-bit quantization break a clinical AI — and can a bigger model fix it?
Six single Google Gemma models reading the same patient charts, varied along two axes: precision (8-bit vs 4-bit) and size/architecture. Two are run at BOTH precisions — a 4-billion on-device model (E4B) and a 12B dense model — to isolate quantization; alongside them a 31B dense model and a 26B sparse Mixture-of-Experts (which fires only ~4B parameters per token), both at 4-bit. Each makes a fast Answer plus a separate In-Depth, both graded. 62 clinical questions across three patients (two adults with HIV/TB, a pediatric infant), judged against the full chart by an automated clinical rubric.
AI setup
Questions
Benchmark /100
Accuracy
Completeness
Relevance
Unsafe answers
q-12b-q8
62
83.6
8.6
7.9
9.5
0
q-26b-q4
62
79.6
8.2
7.5
9.1
1
q-31b-q4
62
77.7
8.2
7.2
9.1
1
q-12b-q4
62
76.4
7.9
7.3
8.8
2
q-e4b-q8
62
69.8
7.5
6.5
8.2
0
q-e4b-q4
62
68.1
7.3
6.4
8.4
3
In-Depth — scored separately on its own axes (its own parity Benchmark)
AI setup
In-Depth Benchmark
In-depth answers
Support
Added value
Unsafe
Padded
q-26b-q4
77.2
39
8.4
7.0
0
0
q-31b-q4
76.0
42
8.2
7.0
0
0
q-12b-q8
75.6
42
8.1
7.0
0
2
q-12b-q4
74.5
42
8.2
6.8
1
1
q-e4b-q8
66.5
46
7.6
5.9
1
7
q-e4b-q4
61.4
47
6.8
5.7
2
4
TakeawayCutting a model to 4-bit costs the 12B about 7 benchmark points AND its safety — at 8-bit it scored 84 with zero unsafe answers; at 4-bit, 76 with two. Recovering that loss by going BIGGER barely works: a 31B at 4-bit buys back only ~1 point, while simply restoring 8-bit on the 12B buys back all 7 — for this model class, spend memory on precision, not parameters. The efficiency standout was the sparse Mixture-of-Experts (26B-A4B): second-best quality (80) at HALF the latency of any other model, with just one unsafe answer. And 4-bit is most dangerous where there's least capacity to spare: the small 4B model at 4-bit was the floor (68), with the most unsafe answers (3) and the most invented facts for things not in the chart (6). Directional results from a single automated judge on a small sample, not a formal benchmark.
ChartSearchAI
Does the best AI model's lead come from quantization, or the model?
Eleven single-model setups, each making a fast Answer plus a separate In-Depth (both graded): five model classes at matched Q8 AND Q4 — Google Gemma 4 12B, Alibaba Qwen3-14B, Qwen2.5-14B, IBM Granite 3.3 8B, Mistral NeMo 12B — plus a reasoning model (Qwen3.6-35B). 32 clinical questions across three patients, judged against the full chart on two independent axes.
AI setup
Questions
Benchmark /100
Accuracy
Completeness
Relevance
Unsafe answers
gemma-12b-q4-2call
32
75.8
8.1
7.3
8.5
0
granite-3.3-8b-q8-2call
32
74.8
7.9
7.2
8.6
1
qwen3-14b-q8-2call
32
74.1
7.8
7.2
8.5
2
qwen3.6-35b-2call
32
73.5
7.8
7.2
8.3
0
12b-2call
32
72.9
7.9
7.3
8.6
4
qwen3-14b-2call
32
69.4
7.6
6.9
8.2
4
qwen2.5-14b-2call
32
67.7
7.3
6.8
7.9
3
qwen2.5-14b-q8-2call
32
66.7
7.1
6.7
7.8
2
mistral-nemo-12b-2call
32
65.2
7.0
6.7
7.9
4
granite-3.3-8b-2call
32
64.4
6.8
6.5
7.6
4
mistral-nemo-12b-q8-2call
32
55.7
6.3
5.8
7.2
7
In-Depth — scored separately on its own axes (its own parity Benchmark)
AI setup
In-Depth Benchmark
In-depth answers
Support
Added value
Unsafe
Padded
gemma-12b-q4-2call
65.7
27
6.4
6.7
0
0
12b-2call
64.6
28
6.2
6.8
1
1
qwen3-14b-q8-2call
61.2
32
6.1
6.5
2
5
qwen3.6-35b-2call
61.0
31
5.9
6.7
3
7
qwen2.5-14b-q8-2call
59.4
32
6.1
5.9
2
1
qwen3-14b-2call
57.7
32
5.8
6.3
5
5
granite-3.3-8b-q8-2call
55.2
32
5.6
6.0
5
2
qwen2.5-14b-2call
51.2
32
5.3
5.5
6
5
mistral-nemo-12b-2call
50.2
32
4.9
5.6
5
0
mistral-nemo-12b-q8-2call
47.7
32
4.6
5.6
8
1
granite-3.3-8b-2call
43.8
32
4.6
5.3
12
3
TakeawayGemma 4 12B tops BOTH axes (Answer 75.8, In-Depth 65.7) — and its lead is NOT a precision artifact: Gemma at Q4 matched (even edged) Gemma at Q8, while quantization moved the other models inconsistently and mostly within noise, so the earlier Gemma win was the model, not its higher precision. Mistral NeMo was the weakest clinical reader (Q8 worst, 55.8, 7 unsafe answers). The separate In-Depth axis earned its keep: it exposed safety failures the Answer hides — Granite's Q4 In-Depth produced 12 unsafe elaborations, Mistral-Q8's 8 — that an answer-only view would miss entirely. The reasoning model (Qwen3.6) was safe but not top. Directional: one judge, three patients, 32 questions.
ChartSearchAI
Does an AI team's deeper explanation beat a single model's?
The new two-call architecture: every setup returns a fast Answer and a separate, slower In-Depth, each independently timed and graded — for single models AND the multi-agent team, on equal footing, via the same shared in-depth prompt. Three patients, 32 clinical questions, five setups (Google Gemma 12B, Alibaba Qwen 14B, Google MedGemma 27B, Liquid LFM2 24B singles, plus a parity AI team).
AI setup
Questions
Benchmark /100
Accuracy
Completeness
Relevance
Unsafe answers
12b-2call
32
72.7
7.8
7.3
8.4
2
parity-team-2call
32
69.1
7.3
7.0
8.3
1
qwen2.5-14b-2call
32
68.5
7.4
6.9
7.9
3
medgemma-27b-2call
32
65.8
7.3
6.8
8.0
6
lfm2-24b-2call
32
51.7
5.8
5.4
6.9
5
In-Depth — scored separately on its own axes (its own parity Benchmark)
AI setup
In-Depth Benchmark
In-depth answers
Support
Added value
Unsafe
Padded
12b-2call
68.4
28
6.8
6.9
0
1
parity-team-2call
62.2
32
6.4
6.2
1
3
medgemma-27b-2call
57.8
32
5.8
6.2
5
1
qwen2.5-14b-2call
55.6
32
5.9
5.7
4
3
lfm2-24b-2call
29.7
31
3.5
3.5
10
9
TakeawayHolding the in-depth prompt fixed, the team's elaboration does NOT beat a single model's. The single Gemma 12B topped BOTH axes — best Answer (73) AND best In-Depth (added-value 6.9, zero unsafe elaborations) — while the parity team's in-depth was no better (6.2), and the fast on-device Liquid 24B's in-depth was actively dangerous (10 unsafe elaborations). Separate latencies confirm the design: a quick answer (Gemma ~20s, Liquid ~2s) followed by a slower in-depth (~12-24s). One strong general model gives the best answer AND the best background; the team scaffolding adds neither. Directional results from a single automated judge on a small sample, not a formal benchmark.
ChartSearchAI
Which single AI model reads a clinical chart best?
Three patients (two adults with HIV/TB and a pediatric infant), 32 clinical questions, seven setups: four single-model baselines across classes — Google Gemma 12B, Alibaba Qwen 14B, Google MedGemma 27B (a medical-specialist model), and the fast on-device Liquid LFM2 24B — plus a single model with an added elaboration and two multi-agent AI-team configurations. Every answer was graded against the full patient chart by an automated clinical rubric.
AI setup
Questions
Benchmark /100
Accuracy
Completeness
Relevance
Unsafe answers
12b-baseline
32
71.8
7.8
7.2
8.4
3
AI team — Standard + checker
32
71.8
7.8
7.4
8.3
4
qwen2.5-14b-baseline
32
70.5
7.5
7.2
8.2
1
Gemma 12B (single + in-depth)
32
70.3
7.8
6.5
7.9
1
AI team — matched to baseline + in-depth
32
66.3
7.0
6.8
8.0
1
medgemma-27b-baseline
32
64.5
7.1
6.8
8.1
5
lfm2-24b-baseline
32
48.9
5.6
5.3
6.6
8
In-Depth — scored separately on its own axes (its own parity Benchmark)
AI setup
In-Depth Benchmark
In-depth answers
Support
Added value
Unsafe
Padded
Gemma 12B (single + in-depth)
62.3
26
6.2
6.5
2
2
AI team — Standard + checker
61.4
32
6.3
6.2
3
0
AI team — matched to baseline + in-depth
58.6
32
6.1
6.0
4
1
TakeawayA strong GENERAL single model wins: Gemma 12B (72) tied the full multi-agent team, and Qwen 14B (71) was just behind AND the safest single arm (1 unsafe answer vs Gemma's 3). The surprise: the medical-specialist MedGemma 27B scored BELOW the general models (65) with more unsafe answers (5) — domain fine-tuning did not help here — and the fast on-device Liquid 24B was the weakest and least safe (49, 8 unsafe). Directional results from a single automated judge on a small sample, not a formal benchmark.
ChartSearchAI
Inside an AI team, which model decides if the answer is safe?
Three patients — two adults with HIV/TB and a pediatric perinatal-HIV infant — 62 clinical questions, six AI setups. A 2×2 ablation swaps the AI team's coordinator and answer-writer between a Google (Gemma) and a fast on-device (Liquid LFM2) model — every team cross-checked by a strong non-Liquid validator — alongside two single-model anchors (Gemma 12B and Liquid 24B). Every answer was graded against the full patient chart by an automated clinical rubric.
AI setup
Questions
Benchmark /100
Accuracy
Completeness
Relevance
Unsafe answers
12b-baseline
62
71.4
7.7
7.1
8.2
3
AI team — Standard
62
70.8
7.7
6.9
8.3
5
AI team — Standard + checker
62
70.4
7.8
6.9
8.3
5
lfm2-24b-baseline
62
50.8
5.6
5.5
6.7
11
AI team — Standard + checker
62
46.8
5.5
4.8
6.5
13
AI team — Standard + checker
62
43.3
5.3
4.4
6.2
17
In-Depth — scored separately on its own axes (its own parity Benchmark)
AI setup
In-Depth Benchmark
In-depth answers
Support
Added value
Unsafe
Padded
AI team — Standard + checker
62.9
62
6.7
6.1
3
3
AI team — Standard
60.0
62
6.3
5.8
2
6
AI team — Standard + checker
33.5
62
4.0
3.6
11
27
AI team — Standard + checker
33.3
62
3.9
3.7
10
29
TakeawayThe model that WRITES the answer decides clinical quality. Setups with a Qwen writer scored ~70; those with a Liquid writer ~43–51 — a ~25-point gap that tracks the writer alone, whichever model coordinates. A single strong general model (Gemma 12B, 71.4) matched or beat the full multi-agent team, and a strong non-Liquid validator did not rescue a Liquid writer (13–17 unsafe answers with it switched on). Pick a strong writer model before adding team scaffolding. Directional results from a single automated judge, not a formal benchmark.
ChartSearchAI
Can a careful AI team rescue fast on-device models?
Two patients, 8 clinical questions, five AI setups: a strong single model (Gemma 12B), our proper multi-agent AI team, a stripped-down 'parity' team, and two teams built around fast on-device 'Liquid' (LFM2) models — one all-Liquid, one that keeps a medical-specialist model. Every answer was graded against the full patient chart by an automated clinical rubric.
AI setup
Questions
Benchmark /100
Accuracy
Completeness
Relevance
Unsafe answers
AI team — Standard
8
87.0
9.0
8.6
9.2
0
AI team — matched to baseline
8
74.6
7.8
7.9
8.4
1
12b-baseline
8
74.4
7.8
8.1
8.5
1
AI team — Standard
8
44.6
5.4
5.0
7.0
4
AI team — team
8
39.8
4.9
4.9
6.8
4
In-Depth — scored separately on its own axes (its own parity Benchmark)
AI setup
In-Depth Benchmark
In-depth answers
Support
Added value
Unsafe
Padded
AI team — Standard
76.2
8
8.0
7.2
0
0
AI team — Standard
26.2
8
3.2
3.8
4
3
AI team — team
25.0
8
3.1
3.0
3
2
AI team — matched to baseline
0.0
7
0.0
0.0
0
0
TakeawayPutting fast on-device 'Liquid' models inside a careful multi-agent team improved their sourcing — they now cite real chart records instead of nothing — but did not make them safe: both Liquid teams still fabricated medications, reversed a weight trend, and hid an AIDS-defining CD4 count, landing far below both the proper model team and a single strong model. Team scaffolding helped grounding, not clinical accuracy. These are directional results from a single automated judge on a small sample, not a formal benchmark.
ChartSearchAI
Can a safety checker make a fast on-device AI team safe?
A follow-up to the Liquid-team run: same two patients, 16 clinical questions, four AI setups — our proper multi-agent team, a single strong model (Gemma 12B), and the fast on-device 'Liquid' team in two forms: as-is, and with a strong non-Liquid Gemma-12B 'validator' that independently re-checks each answer against the chart before it's returned. Every answer was graded against the full chart by an automated clinical rubric.
AI setup
Questions
Benchmark /100
Accuracy
Completeness
Relevance
Unsafe answers
AI team — Standard
16
81.1
8.5
7.9
9.0
0
12b-baseline
16
73.7
8.2
7.1
8.3
1
AI team — Standard + checker
15
51.5
6.3
4.9
6.7
4
AI team — Standard
16
43.2
5.0
4.4
6.3
5
In-Depth — scored separately on its own axes (its own parity Benchmark)
AI setup
In-Depth Benchmark
In-depth answers
Support
Added value
Unsafe
Padded
AI team — Standard
72.2
16
7.4
7.0
0
0
AI team — Standard + checker
39.3
15
4.5
4.3
3
4
AI team — Standard
29.7
16
3.3
3.8
5
7
TakeawayAdding a strong non-Liquid validator to the Liquid team helped — it eliminated the made-up 'I don't know' answers and removed one unsafe response, raising the headline score from 43 to 52 — but it did not make the team safe: the validated Liquid team still produced four unsafe answers (including a harmful regimen error the validator passed) and stayed far below both a proper model team (81) and a single strong model (74). A validator catches over-confident gaps, not the subtler clinical fabrications. These are directional results from a single automated judge on a small sample, not a formal benchmark.
Two patients, 16 questions that hinge on timing: the most-recent lab value, weight trends over months, what's changed recently, and whether a medication regimen is current. We compared plain single models (Gemma 4B → 12B → 26B) against our multi-agent AI team at four configurations.
AI setup
Questions
Benchmark /100
Accuracy
Completeness
Relevance
Unsafe answers
12b-baseline
16
80.2
8.7
7.4
8.8
0
AI team — Standard + checker
16
76.9
8.1
7.6
8.4
0
a4b-baseline
16
74.7
8.2
6.8
8.3
1
AI team — Advanced + checker
16
74.5
7.3
7.8
8.6
0
AI team — matched to baseline
16
70.7
7.6
6.9
8.2
0
gemma-e4b-llamacpp
16
62.5
6.9
6.4
7.6
1
AI team — Basic + checker
16
62.0
6.8
6.8
7.9
3
In-Depth — scored separately on its own axes (its own parity Benchmark)
AI setup
In-Depth Benchmark
In-depth answers
Support
Added value
Unsafe
Padded
AI team — Advanced + checker
71.3
15
7.4
7.0
0
2
AI team — Standard + checker
68.4
16
6.9
6.9
0
1
AI team — Basic + checker
57.5
16
6.3
6.0
2
7
12b-baseline
0.0
1
0.0
0.0
0
0
a4b-baseline
0.0
1
0.0
0.0
0
0
gemma-e4b-llamacpp
0.0
1
0.0
0.0
0
0
TakeawayRe-scored with our updated temporal rubric. The Gemma 12B single model led on the overall benchmark (80/100) and accuracy (8.7/10); the AI team's Advanced tier gave the most complete answers (7.8/10). The stricter date/trend scoring surfaced unsafe answers the first pass missed: a few setups fabricated dates or trends — the Basic AI team 3, and the Gemma 4B and 26B single models one each — while the Gemma 12B single and the Standard/Advanced/matched teams had none. Bigger isn't automatically better, and getting the timing right is where several setups slip.
ChartSearchAI
Fast on-device models or a careful AI team — which is safer?
Two patients, 8 clinical questions, five AI setups: Google's Gemma 12B and 4B single models, two brand-new on-device 'Liquid' (LFM2) single models (24B and 1.2B), and our multi-agent AI team. Every answer was graded against the full patient chart by an automated clinical rubric.
AI setup
Questions
Benchmark /100
Accuracy
Completeness
Relevance
Unsafe answers
12b-baseline
8
89.8
9.1
9.0
9.2
0
gemma-e4b-llamacpp
8
79.5
8.4
8.4
8.6
0
AI team — Basic + checker
8
74.1
7.6
7.8
8.5
1
lfm2-1.2b-baseline
8
46.4
5.1
5.9
6.8
3
lfm2-24b-baseline
8
31.0
3.9
4.1
5.2
4
In-Depth — scored separately on its own axes (its own parity Benchmark)
AI setup
In-Depth Benchmark
In-depth answers
Support
Added value
Unsafe
Padded
AI team — Basic + checker
60.0
8
6.8
5.9
1
2
TakeawayThe Gemma 12B single model was the most accurate and the safest — no unsafe answers, and it cited a chart record for every claim. The two new Liquid models were the fastest but the least safe: they often answered confidently without grounding their claims in the chart (almost never citing a record) and produced several unsafe answers — including a 'no allergies' reply for a patient whose chart documents a penicillin allergy. The AI team cited its sources and was far safer than the Liquid models, though it didn't beat the strong single model here. These are directional results from a single automated judge on a small sample, not a formal benchmark.
An earlier, broader sweep of 22 scenarios on the same patient, comparing the Gemma 12B single model with several AI-team configurations. This run predates our safety cross-checker and the temporal-reasoning fixes.
AI setup
Questions
Benchmark /100
Accuracy
Completeness
Relevance
Unsafe answers
AI team — Advanced
36
70.4
7.3
7.3
7.6
0
12b-baseline
36
67.0
7.2
6.1
7.5
0
AI team — Basic
36
62.0
6.6
6.3
7.1
5
AI team — 12B
36
58.8
6.3
6.3
7.1
4
AI team — Standard
36
54.8
5.8
6.0
6.6
2
TakeawayThe Advanced AI team led on both accuracy and completeness — but the smaller team tiers produced several unsafe answers (Basic tier: 5). That's exactly what motivated the cross-checking 'validator' added in the later runs.
ChartSearchAI
Single model or AI team — which answers priority questions better?
21 high-priority questions about one complex HIV/TB patient (Aloice), comparing two single models (Gemma 2B and 26B) with three AI-team configurations.
AI setup
Questions
Benchmark /100
Accuracy
Completeness
Relevance
Unsafe answers
a4b-baseline
21
84.0
9.1
7.6
9.2
0
e2b-baseline
21
79.9
8.5
7.6
8.6
0
AI team — Advanced
21
79.5
7.9
8.4
8.1
0
AI team — Standard
21
72.8
7.4
7.5
7.9
0
AI team — Basic
21
64.0
6.7
6.9
7.0
0
TakeawayThe 26B single model led on accuracy (9.1/10); the AI team's Advanced tier led on completeness (8.4/10). No unsafe answers in this run.
Catalyst SQL
Catalyst T094 validation
Report unavailable
Historical catalog entry. The linked report was unavailable when checked on 15 September 2026; no original summary was supplied.