PRISM Leaderboard

How this is measured

The three dimensions

Every scenario is scored on three dimensions, each in [0,1]: R, deterministic reliability; S, semantic synchronisation, a binary perception gate; and I, intent alignment. R and I are graded; S is pass or fail.

RPS-Index

A run reports the vector (R̄, S̄, Ī), the mean of each dimension across scenarios, and the RPS-Index, which is R̄·S̄·Ī. It is report-only: there are no tiers, no pass mark, and no threshold anyone is certified against.

False Heals

A False Heal is the framework reporting a pass while the sandbox independently confirms no genuine conclusion occurred: a green test over a broken flow. It is the least gameable number the benchmark produces: a system cannot earn a low count by being cautious, because caution yields honest failures, nor by perceiving the page better. The only way to register one is to claim a success that can be disproved. That is why it is a first-class column and not a footnote.

Intent and integrity violations

Two further counts sit beside False Heals, and the three are mutually exclusive: a scenario produces at most one of them.

An intent violation is the system engaging a control it should not have: acting on the wrong target, or taking an action the task never asked for. It is scored separately from failing to finish, because a system that does the wrong thing confidently is not in the same position as one that does nothing.

An integrity violation is a run detected gaming the measurement rather than performing the task: arriving at the answer by a route a person using the page could not take. It voids that scenario outright rather than scoring it, because a score derived from it would not describe the system’s perception at all.

Incomplete entries

A run that did not execute all 50 scenarios is reported as incomplete and carries no number at all. This is not leniency. Some scenarios cannot be completed by design, and they pay full credit to a harness that never attempted them, so a run that died early still aggregates to a value. Printing that value would read as a poor result rather than an absent one.

Ranked and not ranked

A quarter is ranked only when it ran that quarter’s held-out seed set. A quarter run against a calibration set is a real measurement of a real run, but it is not a ranking: it is shown without a rank column and says so above the table.

Right of reply

Every entry may carry a reply of up to 100 words, published verbatim and attributed, at the same weight as the row it answers. The operator never edits it, may append a factual correction as a clearly marked operator note, and may decline a statement that is false or defamatory, in which case the row says that a statement was declined, without repeating it.

Where the code is

The scenario specifications, the methodology and the scoring code are public: a number on a public board computed by private code is not auditable.

What is not public is the sandbox the scenarios run in. Publishing it would let a system be tuned against the exact pages it is measured on, which is the one thing that would make every number here meaningless.