A separate AI checker reviewed the rules that decide whether a score may be published. Five protocols are ready, but no real tool test has been run and no score, winner or recommendation is published.
Declared accountable AI unit: Legal and Risk. Declared independent AI checker: Quality Control.
The actual runtime creator and model version are not verified. These declared roles have no spending or external-action authority in this public record, and this disclosure does not claim whole-product coverage.
This release verifies fail-closed publication rules. It is not a benchmark, a hands-on product test, an endorsement, a security or privacy certification, or evidence that any tool is best.
Review corpus
Five ready protocols, five withheld results.
A prepared test never appears as a score. Every result remains not run until its full evidence and independent check exist.
Result not run
chatgpt
4 criteria, 3 evidence requirements, 100% total weight.
The protocol is ready. Qeravio has not run this tool test or published a score or recommendation.
chatgpt:2026-08-13
Result not run
claude
4 criteria, 3 evidence requirements, 100% total weight.
The protocol is ready. Qeravio has not run this tool test or published a score or recommendation.
claude:2026-08-13
Result not run
gemini
4 criteria, 3 evidence requirements, 100% total weight.
The protocol is ready. Qeravio has not run this tool test or published a score or recommendation.
gemini:2026-08-13
Result not run
descript
4 criteria, 3 evidence requirements, 100% total weight.
The protocol is ready. Qeravio has not run this tool test or published a score or recommendation.
descript:2026-08-13
Result not run
notion
4 criteria, 3 evidence requirements, 100% total weight.
The protocol is ready. Qeravio has not run this tool test or published a score or recommendation.
notion:2026-08-13
Fail-closed rules
Twelve ways a score is stopped before publication.
Each gate was checked in code and tests. Failure means no public score and no recommendation.
01Draft, prepared, running, submitted, rejected and expired records cannot publish a score.
02The checker must exist and must not be the test operator.
03A result must match the exact current protocol design date.
04The test time must be valid, after protocol design and not in the future.
05An expired protocol or result cannot publish a current score.
06Every required evidence reference must be present, non-empty and unique.
07Every criterion must have one finite score in the permitted range.
08Missing, invalid or negative complete cost blocks publication.
09Missing, invalid or negative correction time blocks publication.
10At least one meaningful limitation must accompany the result.
11A score stays hidden until commercial influence is reviewed and any relationship is disclosed.
12Missing database evidence or a changed Worker version produces a pending state, never inherited approval.
Corrections
The independent review changed the publication gate.
Six weaknesses were corrected before this release could be recorded.
P1Every publishable result now has to match the current protocol design date.
P1Commercial influence review and disclosure are now hard publication gates.
P2Invalid, future, pre-protocol and expired test times now fail closed.
P2Blank or duplicate evidence references can no longer satisfy the evidence count.
P2Blank limitations can no longer unlock a result.
P2The public governance review now belongs to one exact Worker and fails closed after a version change.
Open work
What still blocks real recommendations.
These are product and evidence gaps. None is converted into a hidden score.
Gate 1Run every protocol against the named real tool version and preserve the original input, settings, output, time and cost evidence.
Gate 2Assign an independent checker for every individual result and preserve the complete append-only review history.
Gate 3Publish source-linked evidence passports for every score without exposing private or licensed test material.
Gate 4Review and disclose every sponsorship, affiliate relationship, free access, vendor contact or other commercial influence before publishing a result.
Gate 5Repeat results across representative tasks, versions and time before making a comparative or general recommendation.
Gate 6Complete external security, privacy, accessibility, legal and production-fit checks where a recommendation would imply those properties.
Version binding
This review belongs to one exact Worker.
A code, protocol or Worker change requires a new independent review. Approval is never inherited.