M-0101 · The examTrap-case benchmarks
Our analyzer is tested against corpora of deliberately borderline commits: hope-based parameter twiddling that soundsdeliberate, routine work dressed in experiment vocabulary, genuine investigation described in flat language, bugfixes on the eligibility boundary. The classes come from how real claims fail review — not from what is easy to score.
M-0202 · The judgeA 26-check deterministic reviewer
Every draft faces checks encoding CRA T4088 doctrine before a human sees it: word limits, citation coverage for every assertion, presence of negative results, chronology and cross-section date consistency, banned-language and boilerplate-variance screens. Deterministic checks — the same draft always gets the same verdict.
M-0303 · ReproducibilitySame inputs, same claim
Generation is seeded and pinned: identical evidence produces an identical draft, bit for bit. When our numbers move, it is because the evidence or the method changed — never because a model rolled dice. Your auditor will appreciate that property more than any demo.
Failure is a feature of the record
Abandoned branches. Reverted approaches. Measured dead ends. We find them and keep them — because documented failure is what systematic investigation looks like to a reviewer.
The interviewer asks for outcomes and measurements. It is explicitly forbidden from coaching your engineers into eligibility vocabulary.
What we refuse to do
- No invented metrics
- No synthetic failures
- No “AI-optimized” claim inflation
- No assertion without a citable artifact behind it
The industry has been burned by tools that maximize first and verify never. We built the opposite, on purpose.