Skip to main content
The council chamber at Halifax Town Hall, ranks of wooden benches curving towards a raised seat
Stage 4 · Performance

Data and calculation stations

Graphs, drug doses, risk figures and screening results. What the numerical stations are really marking, the handful of concepts that keep recurring, and how to keep thinking when the arithmetic wobbles.

In 30 seconds

  • Numerical stations are marking reasoning and communication at least as much as arithmetic — a wrong figure reasoned aloud usually outscores a right figure produced in silence.
  • A relative risk reduction means nothing without the baseline risk it applies to, so always ask for the absolute numbers.
  • When a disease is rare, most positive results from even a good test are false positives — that is arithmetic, not test failure.
  • Read the axes, units, scale and sample size before you conclude anything, and say out loud what the graph does not show.
  • If you get stuck, describe what you would do next; freezing costs far more than the error itself.

What the numerical stations are actually for

Many MMI circuits include a station with a sheet of paper on it, and some panel interviews do the same. A graph, a short table, a drug chart, a paragraph describing a trial result. You get the same few minutes as every other station — the circuit mechanics are in the MMI guide — and the number you arrive at is worth far less than you think.

These stations exist because clinical medicine runs on evidence nobody can see directly. Whether to start a statin, whether a screening letter means anything, whether the deteriorating patient in bay four is seen before the stable one in bay one: each is a judgement made from imperfect numbers, at speed, usually in front of somebody who is frightened.

So this is not a maths test. It asks whether you can reason with evidence calmly, notice when a figure does not make sense, and stay honest about the limits of what you were handed. The candidate who says "before I answer, I want to check what the denominator is" is giving an examiner far more to mark than the one who silently produces the right decimal.

The station types recur across schools, and they fall into a handful of recognisable shapes.

  • A graph or table to read, with a conclusion to draw and then defend.
  • A drug dose or fluid calculation, often weight-based and often paediatric.
  • A risk statistic to interpret, usually phrased the way a newspaper would phrase it.
  • A screening or diagnostic result to explain, sometimes to a role-played patient.
  • A prioritisation exercise, where the marks sit in the justification rather than the order.

None of this needs mathematics beyond GCSE. What it needs is that you keep thinking out loud while you do it, which is a genuinely separate skill from doing it. The signposting habits in structuring answers under pressure transfer here almost unchanged.

The mark is on the reasoning, not the answer

Marking schemes are rarely published, but stations of this kind are generally built to credit method as well as the final figure. A candidate who states their assumptions, works aloud, reaches a wrong figure and then notices it looks implausible tends to fare better than one who writes the correct number without narrating. Silence is unmarkable, however good the thinking behind it was.

Reading a graph properly before you say anything about it

One of the commonest failures at a data station is answering the question you expected rather than the one the sheet supports. The fix is a fixed sweep of the page before you offer any conclusion, done aloud so the examiner can hear you doing it.

  1. Read the title and establish the population: which patients, which country, which years.
  2. Read both axes, including the units. Deaths per 100,000 and total deaths tell different stories.
  3. Check the scale. A vertical axis starting at 40 rather than 0 makes a modest change look dramatic; a logarithmic axis flattens a steep one.
  4. Find the sample size. The same difference across 60 patients and across 60,000 deserves very different confidence.
  5. Look for error bars or confidence intervals, and whether they overlap.
  6. Ask what is missing: no control group, no denominator, a truncated time period, or an outcome that is only a proxy for the one that matters.

Then describe what you see, keeping description separate from interpretation on purpose. "What the graph shows is that admissions rose steadily across the four years. What that might mean is a separate question, and there are at least three candidate explanations." That pair of sentences tells the examiner you know the difference between data and inference.

Two ideas do most of the work in the interpretation half. Correlation is not causation: things moving together may share a cause, may run in the opposite direction to the one you assumed, or may simply coincide. Confounding is the mechanism behind most of it, a third variable driving both. Ice cream sales and drowning both rise in summer, and the confounder is the weather. In clinical datasets the usual suspects are age, sex, deprivation and smoking, and naming them specifically reads far better than the phrase "there could be confounders".

Do not let a graph become a hot-topic monologue

A graph about waiting lists is not an invitation to deliver your prepared answer on NHS pressures. The examiner is scoring the station in front of you: what these data support, what they do not, and what you would want to know next. Bring wider context in afterwards, in a sentence or two, and only if it bears on the reading.

Absolute risk, relative risk and number needed to treat

If you learn one piece of statistics for interview, make it this one. Relative figures are the ones that get quoted, in press releases and in clinics, and on their own they are close to uninterpretable.

Suppose a drug halves the risk of a heart attack over five years. That is a relative risk reduction of 50 per cent. Whether it matters depends entirely on where you started. If your five-year risk was 4 per cent, the drug takes you to 2 per cent, an absolute risk reduction of two percentage points. If your five-year risk was 0.04 per cent, the same drug takes you to 0.02 per cent, and the benefit is close to invisible next to the side effects, the cost and the burden of a daily tablet.

Number needed to treat turns that into something a patient can picture. It is one divided by the absolute risk reduction expressed as a decimal. In the first case, 1 divided by 0.02 gives an NNT of 50: fifty people take the drug for five years so that one extra heart attack is avoided. In the second case the NNT is 5,000. Identical headline, entirely different conversation.

Absolute risk
The chance of an event in a given group over a stated period, such as 4 in 100 over five years.
Relative risk reduction
The proportion by which a treatment lowers that risk. Halving 4 per cent and halving 0.04 per cent are both a 50 per cent reduction.
Absolute risk reduction
The difference between the two absolute risks, in percentage points. The figure a patient can actually feel.
Number needed to treat
One divided by the absolute risk reduction as a decimal: how many people must be treated, over the stated period, for one extra person to benefit.
Number needed to harm
The same arithmetic applied to a side effect. Quoting both is what balanced consent looks like.
Incidence and prevalence
Incidence is new cases arising in a period; prevalence is the proportion who have the condition at a point in time.
Confidence interval
The range of values compatible with the data, and so a measure of how precise the estimate is. A wide interval, or one that includes no difference, means the study cannot settle the question.
In the room

A patient has read that a new drug "cuts your risk of a heart attack by half". They ask whether they should take it. What do you say?

Name the gap before answering the question as asked: half of what matters more than the half. You would want their own baseline risk, over what period, and the absolute numbers. Then translate. If their five-year risk falls from 4 per cent to 2 per cent, that is two people in every hundred avoiding an event, or one benefiting for every fifty treated. Then put the other column up honestly, covering side effects, monitoring and the daily commitment, and finish by saying this is a decision the two of you make together once they can see both sides. That last move turns a statistics answer into a medical one.

Sensitivity, specificity and why a positive result can mean very little

Sensitivity is how good a test is at catching the people who do have the disease. At 90 per cent sensitivity it finds 90 of every 100 people who genuinely have it and misses ten. Specificity is how good it is at correctly clearing the people who do not: 90 per cent specificity means that of every 100 healthy people tested, 90 are correctly reassured and ten receive a false positive.

Those are properties of the test. The number a patient actually cares about is different: given that my result came back positive, what is the chance I really have this? That is the positive predictive value, and it depends not only on the test but on how common the disease is in the population being tested. It is the most counterintuitive idea in the topic, and far better worked through with numbers than asserted.

A worked screening example

Take that same test, 90 per cent sensitivity and 90 per cent specificity, applied to 10,000 people in a population where 1 in 100 has the disease.

  • Of the 10,000, 100 people have the disease and 9,900 do not.
  • Of the 100 who have it, the test correctly identifies 90 and misses 10.
  • Of the 9,900 who do not, the test correctly clears 8,910 and wrongly flags 990.
  • Total positive results: 90 true positives plus 990 false positives, which is 1,080.
  • Positive predictive value: 90 out of 1,080, or roughly 8 per cent.

Read that last line again. The test is good on both measures, and yet more than nine in ten people who receive a positive result do not have the disease. Nothing has malfunctioned. The healthy group is simply so much larger that a small false-positive rate applied to it swamps the true positives.

Now make the disease rarer, 1 in 1,000 rather than 1 in 100. Ten people have it and the test finds nine. Of the 9,990 who do not, it wrongly flags 999. The positive predictive value collapses to 9 in 1,008, under 1 per cent. This is one reason programmes are targeted at groups whose prevalence is high enough to make the arithmetic work, and why the debate about genomic screening turns on prevalence at least as much as on the technology.

The clinical consequence follows directly: a positive screening result is a reason for a further, more specific test, not a diagnosis. That one sentence will carry you a long way in a station where you have to explain a result to a worried role-played patient.

Should a screening programme be widened?

The case for widening
  • Earlier detection usually means earlier, less invasive and cheaper treatment.
  • People generally want to know, and withholding an available test can feel paternalistic.
  • A narrow programme can entrench inequality if uptake already differs by deprivation.
The case for caution
  • At low prevalence most positives are false, and each one means anxiety plus a follow-up investigation carrying its own risks.
  • Overdiagnosis: finding disease that would never have harmed that person, then treating it anyway.
  • Screening budgets are not new money, and a negative result can falsely reassure someone out of reporting real symptoms later.
In the room

A patient has had a positive result from a screening test and is very distressed. How would you explain what it means?

Deal with the feeling before the arithmetic: acknowledge that a letter like that is frightening, and find out what they already understand. Then give the headline in plain words, that a positive screening result is not a diagnosis but a flag that something now needs a more accurate look. If they want numbers, offer them concretely rather than as percentages: of every hundred people who get this result, roughly this many turn out to have the condition. Say what happens next and when, admit what is still unknown rather than over-reassuring, and check back what they have understood. The examiner is marking honesty about uncertainty and whether the person was treated as a person, not your recall of the term positive predictive value.

Doses, fluids and doing arithmetic while somebody watches

Dose calculations are the station candidates dread and the one with the clearest technique. The arithmetic is GCSE. The difficulty is that you are doing it standing up, at speed, with a stranger holding a clipboard.

  1. Say the units out loud as you go. Most errors here are unit errors rather than arithmetic errors: milligrams for micrograms, millilitres for milligrams.
  2. State your assumptions first. If a weight or a concentration is missing, say what you are assuming and why.
  3. Work in visible steps rather than one leap, so a slip in step two does not erase the credit for steps one and three.
  4. Sanity-check the magnitude. Ask aloud whether the answer is a plausible thing to give a human being.
  5. Say so if the figure looks wrong. Flagging an implausible number is a patient-safety behaviour and is marked as one.

A worked shape. A child weighs 18 kg, the dose is 15 mg per kg, and the suspension is 250 mg in 5 mL. Eighteen times fifteen is 270 mg. The suspension is 50 mg per mL, so 270 divided by 50 is 5.4 mL. Then the check: a little over a teaspoon of a paediatric suspension is plausible, so nothing has gone badly wrong. Had the answer come out at 54 mL, the correct move is to say out loud that it cannot be right and go back a step.

It is worth adding that this is exactly why high-risk doses are commonly second-checked in practice, and why electronic prescribing systems flag doses outside the expected range. Showing that you understand the system a calculation sits inside, rather than treating it as a private test of your accuracy, is worth a mark by itself.

Prioritisation stations, and what to do when it goes wrong

Prioritisation exercises hand you four or five tasks or patients and ask for an order. The order is rarely where the marks are. What gets scored is whether you sort by clinical urgency rather than by who asked first, whether you delegate or run tasks in parallel, whether you escalate what is beyond your competence, and whether the people you deprioritise stay in view. Telling a waiting relative when you will come back is part of the answer, not an afterthought.

State your principle before your list: immediate risk to life, then time-critical treatment, then everything else, delegating where you safely can. Then apply it out loud. If the examiner pushes back, revising your order for a stated reason reads as judgement; defending an indefensible order to look decisive does not.

Now the part guides tend to leave out. Arithmetic under a timer is genuinely harder than arithmetic at a desk, and it is harder for everybody. Anxiety tends to eat into working memory, which is precisely the faculty a two-step calculation depends on. If your mind empties at a data station, that is a response to the room, not evidence about your ability.

Recovery matters more than perfection. Two sentences will rescue almost any station: "I have lost my thread, let me take that from the start", or "I cannot finish this in the time, so let me tell you how I would approach it and what I would check." Both are markable. Freezing, apologising repeatedly and going quiet are not. There is more on holding a station that starts to slide in nerves, delivery and presence.

Practise the way the station runs rather than the way revision runs. Set a timer for however long your stations are — five minutes is a fair default if you have not been told — take a graph from a published health report, and talk to an empty room until you can describe, interpret, qualify and stop inside the time. If you sat the UCAT recently, the decision-making and quantitative reasoning habits transfer almost directly; the interview simply adds the requirement that your reasoning is audible, and that it is kind.

Sources

  1. Professional standards for doctors, including Good Medical Practice General Medical Council
  2. NHS screening programmes and what a screening result means NHS
  3. UK National Screening Committee — how screening programmes are appraised GOV.UK
  4. NICE guidance and the evidence base behind it National Institute for Health and Care Excellence

Common questions

Not in any depth. The recurring set is small: absolute versus relative risk, number needed to treat, sensitivity, specificity, positive predictive value, correlation versus causation and confounding. If you can define each in plain English and work one screening example with round numbers, you are ahead of most candidates. Beyond that, examiners are assessing reasoning and communication rather than statistical technique.

Reaching the end of an article ticks it off automatically.

Knowing it and saying it are different skills

A mock interview is the only way to find out which parts of this you can actually deliver under a timer, with someone scoring you.