The question pool
Questions come from a pool of pre-written, human-verified questions rather than being invented from nothing each time. Selection is deterministic: given the technology, the skill area furthest behind its coverage target, and the current difficulty, the pool decides what comes next.
That matters for defensibility. A question drawn from a reviewed pool is one you can show someone; a question generated live is one you first see when a candidate complains about it.
What is in the pool
Questions are either platform-wide — available to everyone — or your company's own. Each carries its technology, one or more skill areas, and a difficulty.
Skill areas work the same way: a shared set per technology, plus any your company adds. They are the vocabulary your template focus areas draw on, so adding an area is what lets you weight it.
Reviewing what candidates were asked
Every session records each turn: the question, the difficulty, the skill area, the candidate's answer, how it was evaluated and how confident the evaluation was.
Read the turns before you read the report on a candidate you care about. The report is a summary of these turns; if a question was bad, the summary inherits the problem without showing it.
Flagging a bad question
Any question in a session can be flagged, with a reason:
| Reason | Use when |
|---|---|
| Hallucinated | The question asserts something untrue about the technology |
| Incorrect answer | The question is fine; the expected answer is wrong |
| Inappropriate | It should not have been asked |
| Ambiguous | More than one answer is defensible |
| Too easy / Too hard | Wrong for the difficulty it was assigned |
| Other | With a note |
A flagged question becomes a blocked question for your company: it is kept out of future sessions, with the reason and reviewer note attached. Blocking is company-scoped, so your judgement about a question does not silently change what other companies see.
Flag as you review, not later. The cost of a bad question compounds — every candidate who gets it is measured on it, and the report never says so.
A candidate assessed on a question you later flag deserves a second look. Blocking fixes the future; it does nothing for the person who already answered it. If the flag was incorrect_answer or hallucinated and the question sat in an area you weighted heavily, the score is unsafe.
Reusing a good question
A question that works can be imported into your own question bank, where it becomes an ordinary question you can put in a written assessment. The import is recorded against the turn, so you can see which questions have already been promoted.
This is the payoff of running AI assessments early in a funnel: over time the ones that discriminate well become material you own and control.