Reviewer summary

The candidate report. After submit, hiring teams get evidence, not impressions. Every claim points back to a prompt, a commit, or a failing test.

Fit for the role

Scored against the role requirements, one at a time.

Hire signals and concerns

Each one points to a specific prompt, commit, or failing test in the session.

Prompts, phase-tagged

The full transcript, tagged as planning, building, or debugging, with file references highlighted.

Failing tests, explained

Which line failed, what it returned, and what it should have returned.

Tokens used

Planning, building, and debugging shown separately, against the budget.

Active time

Time to first edit, how long they read the brief, and when the sample suite went green.

Hire or No Hire

An overall score, a recommendation, and a confidence level. Fit is scored against the job description, requirement by requirement, with a reviewer summary written for the debrief.

The scorecard

Four scores hiring managers can compare across candidates: coding test, code quality, interview, and AI fluency. AI fluency covers whether they planned first, pushed back on a bad suggestion, edited by hand, and wrote their own tests.

Integrity flags

Proctoring covers the whole session, from the first edit to the last interview answer. A separate check flags code written only to pass the visible tests.