2026-Q4
Automated tests read the DOM; people read the screen. PRISM measures the gap across the same 50 scenarios, run each quarter, scoring reliability, perception and intent as the RPS-Index. Three counts sit beside it: False Heals, a green result over a broken flow; intent violations; and integrity violations. How this is measured →
| Rank | Entry | RPS-Index | R̄ / S̄ / Ī | False Heals | Violations |
|---|---|---|---|---|---|
| 1 | Kane CLI Kanev0.8.17Vendor submission | 0.7674 | 0.8871 / 0.8667 / 0.7787 | 9 of 50 | 3 intent 0 integrity |
| 2 | Momentic Momenticv3.58.4Vendor submission | 0.7647 | 0.8652 / 0.8433 / 0.7747 | 5 of 50 | 3 intent 0 integrity |
| 3 | Passmark Passmarkv1.0.16Operator baseline | 0.6710 | 0.7853 / 0.7667 / 0.6900 | 8 of 50 | 3 intent 0 integrity |
| 4 | Shiplight CLI Shiplightv0.1.105Operator baseline | 0.6227 | 0.7033 / 0.6800 / 0.6353 | 3 of 50 | 3 intent 0 integrity |
| 5 | Magnitude Magnitudev0.3.13Operator baseline | 0.6214 | 0.7903 / 0.8000 / 0.6413 | 15 of 50 | 4 intent 0 integrity |
| 6 | Hercules (open source, AGPL v3) TestZeusv1.0.2Operator baseline | 0.5323 | 0.6439 / 0.6167 / 0.5373 | 26 of 50 | 3 intent 0 integrity |
| |||||
| 7 | Playwright PRISM Benchmarkv1.60.0Operator baseline | 0.4599 | 0.6409 / 0.6900 / 0.5200 | 8 of 50 | 2 intent 0 integrity |