TestUtopia
SolutionsPricingAboutContact
Help Center
  • Getting started

Assessments

  • Creating assessments
  • Inviting candidates
  • Proctoring
  • Reviewing results

Questions

  • Question types
  • Multiple-choice questions
  • Essay questions
  • Code questions
  • Test case format

Library

  • The Library
  • Importing and publishing

Interviews

  • Managing interviews
  • Live interviews
  • Interview questions
  • Scorecards
  • Interview templates

AI Assessments

  • AI Assessments
  • Templates
  • The question pool
  • Reading the report

Streams

  • Streams
  • Enrolling students
  • Stream tests
  • Stream results

Training

  • Training & Certification
  • Participants
  • Training tasks
  • Certificates and evidence

Account & billing

  • Team and roles
  • Plans and billing
  • Account and security

For candidates

  • Before you start
  • What is recorded
  • Coding questions
  • If something goes wrong
  • After you submit
  • Your data

Developers

  • REST API
  • Webhooks
  • Greenhouse
  • SSO and SCIM

Cannot find what you need?

Contact support
  1. Help Center
  2. /
  3. AI Assessments

Reading the report

A completed session produces a report. It goes to you, never to the candidate.

What is in it

FieldWhat it is
Overall scoreA single number for the session
Assessed levelThe model's judgement of where the candidate sits
Skill breakdownThe same score per focus area, with correct, partial and incorrect counts
Strengths and weaknessesShort lists, in prose
SummaryA narrative account of the session
RecommendationA suggestion about proceeding

Underneath sits every turn: the question, the answer, the evaluation and the model's confidence in it.

How much to trust it

The report is generated by a model reading a transcript. It is evidence, and it is a summary of better evidence that is one click away.

Three specific cautions:

The skill breakdown is thinner than it looks. Fifteen questions across four areas is three or four questions each. "Testing: 40%" often means one wrong answer out of three, and presented as a percentage it invites a confidence the sample does not support.

Evaluation confidence varies per turn, and it is recorded. An answer marked incorrect with low confidence is worth reading yourself — free-text answers that are right but unusually phrased are the common case.

The assessed level is a comparison to a norm nobody wrote down. It is useful as a rough sort, not as a title. Whether someone is "senior" depends on the job, and the model does not know the job.

The recommendation

This is the field to be most careful with, because it is the one that looks like a decision.

It is a suggestion produced by a model, and it has no authority. Nothing in the platform acts on it: no status changes, no candidate is filtered, and the candidate is never told it exists.

Under GDPR Article 22, a decision produced by automated processing with no meaningful human involvement is exactly what a candidate can object to. Reading a model's recommendation and forwarding it is not meaningful involvement — it is the automated decision, with a person's name on it. If you reject someone, the reasons need to be reasons you hold, from evidence you looked at.

This is also why the candidate never sees the report. If they did, an AI-generated verdict would be reaching them directly, and no amount of internal process would change what happened.

What the candidate sees

During the session: whether each answer was right, and an explanation. Afterwards: nothing.

No score, no level, no breakdown, no strengths, no weaknesses, no recommendation — through any endpoint. That is enforced by an automated contract test over the candidate handlers, so it cannot be reintroduced by an accidental field addition.

If you want to give a candidate feedback, that is yours to write and yours to stand behind — see Reviewing results.

Before you act on a report

For any candidate you are about to reject or advance, open the turns. It takes a couple of minutes and answers the question the report cannot: were these the right questions, and was the answer really wrong?

If a question was bad, flag it — and re-read the score before using it.

Was this helpful?
PreviousThe question pool
TestUtopia

Advanced technical assessment engine designed for high-precision engineering teams. Curating talent through rigorous data-driven evaluation.

Solutions

  • Solutions
  • Pricing
  • Features

Company

  • About

Support

  • Help Center
  • Contact support

Legal

  • Privacy Policy
  • Terms of Service
  • Cookie Policy
  • Security
  • Trust Center
Test Utopia Ltd · Razsadnika-Konyiovitsa, Bl. 22, fl. 6, ap. 38, Sofia, 1330, Bulgaria
Reg. No.: 207409973|VAT: BG207409973
[email protected]|+359 886 363 248

© 2026 Test Utopia Ltd. All rights reserved.