MedGuard
ProductCompare with other models
- Boards
- 1 of 13
- Results
- 1
- Latest measurement
- Jul 2026
Results
Each line is placed on its own board. Dark tick: the board leader.
EHR and workflow agents
- 22.7Rank 11 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
Sources
Open a line for the quote and page.
11CHI-Bench All Domains pass@1; hermes harness; community submission 22.7
Printed as 22.7%Independent run, measured Jul 2026Configuration: All Domains pass@1; hermes harness; community submissionCHI-Bench leaderboard (actAVA) official leaderboard, actAVA, 12 Aug 2026. CHI-Bench v1.0.0, All Domains, rank 11, Agent/Model and Accuracy columns; PA/UM/CM follow; board last updated 2026-08-12; board Date column 2026-07-06; accessed 2026-09-30. Linked submission provenance: run 2026-07-01T06:59:37Z to 2026-07-01T08:17:44Z; submitted_at 2026-07-06T09:10:18Z.11 | hermes submitted by cuilinke | MedGuard | Open-source | 22.7% | 4.0% | 4.0% | 60.0% | 2026-07-06
Every result from this document- Also printed in CHI-Bench submission manifest: hermes + MedGuard, submission.model; provenance.started_at / finished_at; results.overall.pass_at_1; retrieved 2026-09-30