Skip to main content

How we mark your AI mock interview

After your AI mock interview, every station you answered is marked against that question’s own mark scheme. You can check each mark: it quotes your own words, it tells you how far two independent AI examiners agreed, and we test the AI’s marking against human markers who cannot see the AI’s marks.

Last updated: 30 September 2026

The mark scheme

Each question in our bank has its own mark scheme: a few skills, such as ethical reasoning or communication, each with a maximum mark and a note on what earns it. The AI examiner gives a mark for every skill. Your station mark out of 10 is then worked out from those skill marks by our own code (your total as a share of the available marks, rounded), not picked separately by the AI. On a panel, each panellist’s questions are marked one by one, each against its own scheme.

These are NextGen MedPrep’s mark schemes, not the university’s own. Universities do not publish their full interview marking schemes, so treat a mark as practice feedback on your answer, not a prediction of your offer.

What the examiner reads

The AI examiner reads the whole conversation for that station, not just your first answer. It is told to:

  • credit only reasoning you produced yourself. A hint or framework the interviewer gave you earns nothing on its own, only what you then did with it;
  • mark you on the questions you were actually asked, never on a point no question invited;
  • ignore the ums, repeats and mis-heard words that any speech transcript contains;
  • mark on its merits an answer our own turn-taking or the station clock may have cut short, without deducting for the ending you did not get to give;
  • take into account how hard the interviewer was pushing. The interviewer adapts during the interview, and an answer that holds up under tough follow-ups is credited as such, while the mark scheme still decides the mark;
  • judge Australian, New Zealand and US answers against their own health systems, not the NHS.

Every mark shows its evidence

For each skill mark, the examiner must quote up to 2 short passages of what you said that the mark rests on: the words that earned a strong mark, or the words that show the gap for a weak one. We then check every quote word for word against your side of the transcript. A quote that is not your exact words (the interviewer’s words, a paraphrase, two fragments stitched together) is thrown away.

In your debrief, “Why this mark” shows those quotes, each with a button that jumps to that moment in your session replay while the video is kept. If no quote survived the check, the mark is shown as provisional, never as a confident verdict.

Two examiners, not one

An AI marking the same answer twice does not always land in the same place, so one opinion is not enough. Each station is marked by two independent AI examiner passes that take the skills in a different order. If they disagree by more than 1 mark on any skill, or by more than 2 points on the mark out of 10, a third pass marks it too, and the middle mark stands.

Your debrief shows how far they agreed. High confidence means the examiners were within a point of each other. When they could not agree, the station says “Examiners could reasonably disagree on this one” and shows the range, so you can weigh the comments and your own words more than the number.

How we check the AI against people

Agreement between two AI passes shows the AI is consistent. It does not show the AI is right. So NextGen MedPrep markers re-mark real stations the AI has marked, using the same mark scheme and the same transcript, without seeing the AI’s marks until they have submitted their own. Stations are picked across universities, types of station and the AI’s mark bands (low, middle and high), so the check covers the whole range and not just the most common case. Some stations are given to a second marker as well, which tells us how closely human markers agree with each other.

We compare the marks skill by skill: how often the AI is within one mark of the person, the average gap, agreement beyond chance, and whether the AI tends to mark higher or lower. Staff test sessions are left out.

Calibration in progress

0 of the 30 stations we need have been re-marked by a person so far. We will publish how closely the AI agrees with human markers once 30 different stations have been re-marked, because fewer than that could make the AI look better or worse than it is.

If a mark looks wrong

Your marks are produced by AI (Anthropic’s Claude) and are not checked by a person before you see them. If you think a mark is unfair or wrong, flag the debrief from its page and a NextGen MedPrep coach will re-review your session within 48 hours; their notes appear on your debrief.

Read more about how we produce and check our content in our editorial policy, or go back to AI mock interviews.