Methodology · public

How we decide
you can.

Everything on this page is public on purpose. A verifier should be able to judge whether our signal means anything without ever seeing a candidate's transcript — and a candidate should know the criteria in advance. If knowing the anchors let you fake them, the anchors would be the problem.

Evidence-linked

Every anchor score above zero cites the exact thing you said, or a measurement taken from what you submitted. No quote, no points.

Conclusion-blind

For judgement tasks there is no answer key. The evaluator scores your process, whether your reasoning is consistent with what you actually discovered, and how you defend it — never whether we agree with your conclusion.

Policy, not opinion

The model proposes anchor scores. The pass threshold and the universal-fail rules are fixed in code and applied afterwards. A model cannot decide to pass you.

Two streams

Alongside the scored anchors runs something computed rather than judged — a client trust curve, or geometric measurements of what you built. You can talk your way past a rubric; you cannot talk your way past arithmetic.

Open-book

Assume you will look things up. The assessments are built so that shortcuts still require understanding — the situation changes underneath you, and you have to explain what you did.

Nothing memorisable

No item bank, no fixed questions. The scenario instance is private and the shift is the point, so last year’s answers are worth nothing.

Levels

Levels differ by how much the situation misleads or moves, not by how hard the task is. The ladder doubles as a trust ladder: each level states the assurance it can honestly claim, so a low-assurance result is legible as one.

L1 Structured Performs under cooperative conditions, where what the client states is essentially what they need. Self-serve · AI-assessed · light identity checks
L2 Shifting The situation misleads or moves: what you are given at the start is not what you end up having to handle. Self-serve · AI-assessed
L3 Integrative A whole-skill capstone: orchestrating the competencies together under real conditions. Adds photo verification (checked, never stored) and evidence review
L4 Contested Adverse conditions — competing stakeholders, active resistance, incomplete information. Proctored, or accepted proof of work
L5 Real-world Attested delivery on real engagements, evidenced by timestamped builder logs. Attested project evidence plus proctored defense
Only L2 is available today, at one assurance tier: Self-serve · AI-assessed. Assessed by AI against published anchors, with every score linked to evidence from the session. No identity verification, no proctoring, no human review at this tier.

The anchors we score against

These are the live rubrics, printed from the same modules the evaluator loads. Scores are 0 (absent), 1 (present), 2 (present with distinction).

problem_framing · L2 Pass at 8 of 12 points, and no universal fail
PF2-1 Treats the stated problem as a hypothesis: at least one question probes whether the stated problem is actually the problem.
PF2-2 Picks up at least two buried cues AND connects them to each other (connection, not mere mention).
PF2-3 Proposes a reframe explicitly ("you came in about X, but this suggests Y because...").
PF2-4 Grounds the reframe ONLY in information from this session (no phantom facts).
PF2-5 Handles client resistance to the reframe without capitulating instantly or bulldozing — walks through the evidence.
PF2-6 Secures explicit agreement on the reframe, or explicitly logs the disagreement and its consequence.
Universal fails — any one of these fails the level regardless of points
UF-1 Solution-jumping: proposes solutions before any problem statement exists.
UF-2 Frame absorption: adopts the stated problem verbatim with no testing — including competently solving the WRONG problem.
UF-3 Contrarian reframing: rejects the stated frame without citing discovered cues.
UF-4 Phantom facts: the frame relies on information that appears nowhere in the session.
UF-5 No agreement loop: never confirms the frame with the client.
UF-6 Frozen frame: holds the initial frame after clearly contradicting information.
Composition under a changing brief · L2 Pass at 8 of 12 points, and no universal fail
CC2-1 The first composition establishes a clear focal point: one element carries dominant visual weight rather than competing evenly with the rest.
CC2-2 The first composition is structurally sound — elements sit inside the margins, essential text does not collide, and positions align to a consistent axis or grid.
CC2-3 The revised composition satisfies the constraint introduced by the changed brief.
CC2-4 The revision adapts the existing composition rather than restarting: parts that still worked are recognisably preserved.
CC2-5 Structural quality survives the change — the revision introduces no new margin breaches, collisions or alignment regressions.
CC2-6 The written rationale explains the trade-off made, naming specific elements and the new constraint, and is consistent with what actually changed.
Universal fails — any one of these fails the level regardless of points
UF-C1 Restart: the revision discards the first composition and rebuilds from scratch instead of adapting it.
UF-C2 No adaptation: the revision is effectively unchanged while the new constraint remains unmet.
UF-C3 Constraint ignored: the requirement introduced by the changed brief is not satisfied at all.
UF-C4 Broken artefact: elements fall outside the canvas, or essential text is obscured by another element.
UF-C5 Unsupported rationale: the explanation describes changes that the submitted layouts show were never made.
Secondary evidence — recorded, and can accrue even when the primary competency does not pass
AL-1 References and builds on information the client offered unprompted.
AL-2 Summarizes accurately without inserting assumptions.
Q-1 Uses open questions to explore before narrowing with closed ones.
Q-2 Follows up on answers rather than running a fixed script.

What this version does not do

Listing the gaps is part of the method. A verification product that overstates itself is worse than none.

No identity verification

We know a session happened and what was said in it. We do not know who was sitting there. Proctoring and photo checks belong to L3 and above, which do not exist yet.

No human review

No appeals process, no expert audit, no random sampling for evaluator drift. All of that is designed and none of it is built.

No evidence ledger yet

Results accumulate on a profile, but there is no transfer between skills, no gap analysis, and no decay. A verification here does not yet expire, which means recency is not being priced in.

One scenario per competency

The instance should be generated fresh per candidate from a private template. Today there is one hand-authored instance, so it is repeatable in a way production would not be.

No baseline-relative scoring

The bar should be set as a delta above what current frontier models produce unaided. We have not run those baselines, so we do not claim it.

No capstone

Skill-level claims need an integrative assessment. We verify individual competencies, and the credential says so explicitly.

Try it yourself

The fastest way to judge
an assessment is to take it.

Hiring managers included. Both live assessments are open, and take about twenty minutes.

See the 2 live assessments