Everything on this page is public on purpose. A verifier should be able to judge whether our signal means anything without ever seeing a candidate's transcript — and a candidate should know the criteria in advance. If knowing the anchors let you fake them, the anchors would be the problem.
Every anchor score above zero cites the exact thing you said, or a measurement taken from what you submitted. No quote, no points.
For judgement tasks there is no answer key. The evaluator scores your process, whether your reasoning is consistent with what you actually discovered, and how you defend it — never whether we agree with your conclusion.
The model proposes anchor scores. The pass threshold and the universal-fail rules are fixed in code and applied afterwards. A model cannot decide to pass you.
Alongside the scored anchors runs something computed rather than judged — a client trust curve, or geometric measurements of what you built. You can talk your way past a rubric; you cannot talk your way past arithmetic.
Assume you will look things up. The assessments are built so that shortcuts still require understanding — the situation changes underneath you, and you have to explain what you did.
No item bank, no fixed questions. The scenario instance is private and the shift is the point, so last year’s answers are worth nothing.
Levels differ by how much the situation misleads or moves, not by how hard the task is. The ladder doubles as a trust ladder: each level states the assurance it can honestly claim, so a low-assurance result is legible as one.
These are the live rubrics, printed from the same modules the evaluator loads. Scores are 0 (absent), 1 (present), 2 (present with distinction).
Listing the gaps is part of the method. A verification product that overstates itself is worse than none.
We know a session happened and what was said in it. We do not know who was sitting there. Proctoring and photo checks belong to L3 and above, which do not exist yet.
No appeals process, no expert audit, no random sampling for evaluator drift. All of that is designed and none of it is built.
Results accumulate on a profile, but there is no transfer between skills, no gap analysis, and no decay. A verification here does not yet expire, which means recency is not being priced in.
The instance should be generated fresh per candidate from a private template. Today there is one hand-authored instance, so it is repeatable in a way production would not be.
The bar should be set as a delta above what current frontier models produce unaided. We have not run those baselines, so we do not claim it.
Skill-level claims need an integrative assessment. We verify individual competencies, and the credential says so explicitly.
Hiring managers included. Both live assessments are open, and take about twenty minutes.