PRISM Leaderboard

2026-Q4

Automated tests read the DOM; people read the screen. PRISM measures the gap across the same 50 scenarios, run each quarter, scoring reliability, perception and intent as the RPS-Index. Three counts sit beside it: False Heals, a green result over a broken flow; intent violations; and integrity violations. How this is measured →

RankEntryRPS-IndexR̄ / S̄ / ĪFalse HealsViolations
1Kane CLI
Kanev0.8.17Vendor submission
0.76740.8871 / 0.8667 / 0.77879 of 503 intent
0 integrity
2Momentic
Momenticv3.58.4Vendor submission
0.76470.8652 / 0.8433 / 0.77475 of 503 intent
0 integrity
3Passmark
Passmarkv1.0.16Operator baseline
0.67100.7853 / 0.7667 / 0.69008 of 503 intent
0 integrity
4Shiplight CLI
Shiplightv0.1.105Operator baseline
0.62270.7033 / 0.6800 / 0.63533 of 503 intent
0 integrity
5Magnitude
Magnitudev0.3.13Operator baseline
0.62140.7903 / 0.8000 / 0.641315 of 504 intent
0 integrity
6Hercules (open source, AGPL v3)
TestZeusv1.0.2Operator baseline
0.53230.6439 / 0.6167 / 0.537326 of 503 intent
0 integrity

Hercules is TestZeus's free, open-source agent (AGPL v3), not our commercial platform. We welcome PRISM. This row runs Hercules 1.0.2 on one operator-chosen model. Even so, it matched the board's best score on hydration and timing races, and scored 0.91 on React state races with zero integrity violations. We treat every false pass as a bug. We've asked for the False Heal telemetry and will ship fixes in public. Canvas, shadow DOM and multi-step intent are now open roadmap issues. Pick a gap and ship a fix: github.com/test-zeus-ai/testzeus-hercules. Closed tools ask for trust. Open ones earn it.

Hercules maintainers, TestZeus
7Playwright
PRISM Benchmarkv1.60.0Operator baseline
0.45990.6409 / 0.6900 / 0.52008 of 502 intent
0 integrity