Decision
Against the gates in force at publication (at least 90% of conversations overall and at least 80% on every scenario): No team qualified.
catalyst-query-gemma-4-12b— 7/12 conversations (58% overall, worst scenario 0%) — below the gatescatalyst-query-gemma-4-12b-q4-checked— 7/12 conversations (58% overall, worst scenario 0%) — below the gatescatalyst-query-gemma-4-12b-qwen2.5-14b-checked— 8/12 conversations (67% overall, worst scenario 0%) — below the gates
Summary
| Profile | Profile ID | Conversations passed | Assertions | Avg generation time |
|---|---|---|---|---|
| catalyst-query-gemma-4-12b | catalyst-query-gemma-4-12b | 7/12 | 347 | 46298 ms |
| catalyst-query-gemma-4-12b-q4-checked | catalyst-query-gemma-4-12b-q4-checked | 7/12 | 344 | 92734 ms |
| catalyst-query-gemma-4-12b-qwen2.5-14b-checked | catalyst-query-gemma-4-12b-qwen2.5-14b-checked | 8/12 | 347 | 90672 ms |
Per-scenario breakdown
PASS and FAIL are the judge's verdict on the answer. INVALID means the run broke its own contract there and measured nothing — it is not a score against the team.
| Scenario | catalyst-query-gemma-4-12b | catalyst-query-gemma-4-12b-q4-checked | catalyst-query-gemma-4-12b-qwen2.5-14b-checked |
|---|---|---|---|
| A1 single-turn | PASS | PASS | PASS |
| A2 single-turn | PASS | PASS | PASS |
| A3 single-turn | PASS | PASS | PASS |
| A4 single-turn | PASS | PASS | PASS |
| M1 multi-turn ×3 | PASS | PASS | PASS |
| M2 multi-turn ×3 | FAIL | FAIL | FAIL |
| M3 multi-turn ×3 | PASS | FAIL | PASS |
| B1 clarification | FAIL | FAIL | FAIL |
| B2 clarification | FAIL | PASS | PASS |
| B3 clarification | FAIL | FAIL | FAIL |
| U1 single-turn | PASS | PASS | PASS |
| U2 single-turn | FAIL | FAIL | FAIL |
What went wrong
| Scenario | Team | Verdict | What happened |
|---|---|---|---|
| M2 | catalyst-query-gemma-4-12b | FAIL | the answer has 49 groups the reference does not have (at turn 1 of 3) |
| M2 | catalyst-query-gemma-4-12b-q4-checked | FAIL | the answer has 49 groups the reference does not have (at turn 1 of 3) |
| M2 | catalyst-query-gemma-4-12b-qwen2.5-14b-checked | FAIL | the answer has 49 groups the reference does not have (at turn 1 of 3) |
| M3 | catalyst-query-gemma-4-12b-q4-checked | FAIL | answered 'rejected' where the scenario expects 'ready' (at turn 1 of 3) |
| B1 | catalyst-query-gemma-4-12b | FAIL | answered 'rejected' where the scenario expects 'ready' (at turn 1 of 2) |
| B1 | catalyst-query-gemma-4-12b-q4-checked | FAIL | answered 'rejected' where the scenario expects 'ready' (at turn 1 of 2) |
| B1 | catalyst-query-gemma-4-12b-qwen2.5-14b-checked | FAIL | answered 'rejected' where the scenario expects 'ready' (at turn 1 of 2) |
| B2 | catalyst-query-gemma-4-12b | FAIL | the answer returned 3002 rows; the independent reference returns 108 (at turn 1 of 2) |
| B3 | catalyst-query-gemma-4-12b | FAIL | the answer returned over 5000 rows; the independent reference returns 4665 (at turn 1 of 2) |
| B3 | catalyst-query-gemma-4-12b-q4-checked | FAIL | the answer returned over 5000 rows; the independent reference returns 4665 (at turn 1 of 2) |
| B3 | catalyst-query-gemma-4-12b-qwen2.5-14b-checked | FAIL | the answer returned over 5000 rows; the independent reference returns 4665 (at turn 1 of 2) |
| U2 | catalyst-query-gemma-4-12b | FAIL | answered 'ready' where the scenario expects 'unsupported' |
| U2 | catalyst-query-gemma-4-12b-q4-checked | FAIL | answered 'ready' where the scenario expects 'unsupported' |
| U2 | catalyst-query-gemma-4-12b-qwen2.5-14b-checked | FAIL | answered 'ready' where the scenario expects 'unsupported' |