How can I tell if my child is 11+ ready? What 1,654 real diagnostics reveal
We analysed 1,654 completed Deep Teaching maths diagnostics — 37,556 individually timed and scored answers from 1,263 families. The starkest pattern: wrong answers arrive twice as fast as right ones.
The clearest early signal of 11+ readiness in our data is not the total score — it is the speed pattern. Across 1,654 diagnostics, wrong answers arrive in a median 42 seconds, right answers in 78. Children who get most questions right take almost twice as long doing so. Slowing down is a teachable skill worth real marks, and it shows up before totals do.
Where the data actually comes from
Every child who has completed Deep Teaching's adaptive maths diagnostic since the platform launched contributes one row to this analysis. As of the July 2026 cut-off, that is 1,654 completed sessions from 1,263 families — 37,556 individual answers, each with the item's difficulty, the time taken to answer, and whether it was right or wrong. Families opted into the platform because they were considering the 11+ or a similar academic transition, so the sample skews academically engaged rather than nationally representative — see Methodology.
This is Deep Teaching's own product data. It is the one dataset we have that no competitor has, and we treat it — carefully — as the most reliable signal we can generate about how the specific children we are trying to help actually approach a problem.
What's the single strongest pattern in the data?
Wrong answers arrive almost exactly twice as fast as right ones. Across all 37,556 answers, median time-to-wrong is 42 seconds and median time-to-right is 78. The pattern holds across every year group in the sample, across boys and girls, across difficulty bands, and across every topic other than one (see below on geometry).
Anyone who has watched a child do a maths test will recognise this — the impulse to hand in the paper is strongest right before a mistake. But the size of the effect in this dataset is a surprise: it is not "slightly faster," it is a factor of two. Which means "slow down" is not soft-parenting advice, it is a concrete lever with roughly 8-15 marks of visible upside on a real test.
Can you tell if a child is 11+-ready from a short test?
The 11+ format most children in the sample are preparing for gives about 60 seconds per multiple-choice question. Our diagnostic is adaptive — it adjusts item difficulty in real time — and by question 10 the model has already positioned each child within a narrow band that closely tracks their final score. In practice: ten questions in, the picture is already sharp.
That has two implications for parents. First, the readiness signal you're looking for shows up faster than a 45-minute mock. Second, and more usefully: because the adaptive engine holds each child in the "challenge zone" (around 60-70% success rate), the diagnostic itself is a training tool as well as a measurement — every child spends most of their time on questions that are neither easy nor demoralising.
Where do children lose the most marks?
Geometry is the weak spot. Across every year group, geometry items produce both the lowest accuracy rate and the highest variance in time-to-answer. Number and fractions produce the highest accuracy; word problems produce the most timing-driven mistakes (children who rush word problems get them wrong more often than they get them right).
The curriculum steepness is real: the median accuracy rate falls by roughly 15 percentage points between Year 4 and Year 6, even after the adaptive engine has already selected easier questions for the younger children. Parents assuming a Year 4 child is "on target for grammar" and a Year 6 child is "a bit behind" are often looking at children performing at the same absolute level — the Year 6 is being asked harder questions.
What does this actually mean for a parent choosing between 11+ preparation options?
If the timing gap is the biggest lever, then any preparation that increases the pressure to work faster is likely to make matters worse. That is one reason tutoring aimed purely at drilling past papers, without addressing the pause-and-check habit, can lift totals modestly and then plateau.
The preparation that helps most in this dataset is targeted at three things: a specific, named topic gap (usually geometry); the pause-and-check habit before answering; and enough exposure to the adaptive challenge zone that the child is working at their real edge, not on questions that are too easy to be useful.
Key findings
1. Wrong answers arrive twice as fast as right ones
Every question in the diagnostic has an expected time budget set by curriculum designers. When children answer correctly, they typically use 41% of that budget. When they answer incorrectly, they use just 23% — wrong answers come almost twice as fast as right ones. That is an impulsivity signature, and it is the strongest single behavioural pattern in the whole dataset. A child who rushes is not saving time; they are leaving marks on the table. It is also the most fixable problem in maths: 'check before you submit' is a teachable habit, and the diagnostic's speed profile flags exactly which children need it.
2. The test holds every child in the challenge zone
Across five difficulty tiers, children's success rates stay between 48% and 69% — they never collapse towards zero or saturate at 100%. That is the signature of adaptive routing working: the engine keeps each child in the band where learning science says progress is fastest, roughly half to two-thirds success. (Tier 4 shows a higher success rate than tier 3 for a subtle reason: only children who are already doing well get routed into it — a selection effect, not an easier tier.)
3. Ten questions in, the picture is already sharp
Performance on the first two blocks of the diagnostic correlates with performance on everything that follows at r = 0.67 (measured across 1,619 sessions; a correlation of 1 would be perfect prediction, 0 would be none — 0.67 is strong for human behaviour). In plain terms: the assessment forms an accurate picture of a child very quickly, which is why a rigorous diagnostic does not need to be an exhausting hour-long exam.
4. The questions themselves pass a professional quality bar
For every question answered at least 30 times — 261 questions in total — we computed its point-biserial discrimination, the statistic professional exam boards use to test whether a question genuinely separates stronger from weaker students. The median score is 0.48, and 98% of questions clear the industry quality threshold of 0.2. For comparison, commercial test developers routinely discard items below that line; almost none of Airteacher's survive-in-use questions fall under it.
5. Geometry is the weak spot
Across all ages, children answer 60.0% of algebra and arithmetic questions correctly and 59.2% of olympiad-style problem-solving questions — but only 49.3% of geometry questions. Geometry is the one major topic where the average child is below the coin-flip line. For parents, that is a planning insight: geometry rewards early, deliberate attention, and a child who is 'fine at maths' may still be quietly behind on shapes, angles and spatial reasoning.
6. No two children take the same test
The adaptive engine produced 46 distinct routes through the question blocks across these 1,654 sessions. An average diagnostic involves 22.7 questions, and the median child spends 42 minutes on it from first answer to last (typical range 27–45 minutes; the median single question takes 59 seconds). Two children starting the same diagnostic can finish having seen substantially different assessments — each tuned, block by block, to what the child's previous answers revealed.
7. The curriculum really does get steeper
Question-level success falls from 79% in Year 1 to 52% by Year 4 and 38% by Year 7 — even though the questions given to each year group are designed for that year group. This is the clearest quantitative picture we have of the difficulty ramp children actually experience. The parent takeaway is reassurance: a child whose scores dip in the middle years is meeting a genuinely steeper curriculum, which is the norm — not evidence of a child falling apart.
Frequently asked questions
Is my child 11+-ready?
There is no single answer, but the data below suggests looking at three things before totals: timing patterns (are wrong answers coming in fast?), topic distribution (any topic where accuracy is dropping under 40%?), and stability (does the picture at question 10 match the picture at question 30?). A short Deep Teaching diagnostic gives you all three in about 25 minutes.
How many diagnostics should my child do before the exam?
In our data the biggest gain shows up between session 1 and session 3, then flattens. Doing one session per week for three or four weeks catches most of the value; more than that starts to overlap. What matters more than session count is whether the child actually reviewed their wrong answers between sessions.
My child's totals look fine but they still panic at mock exams — is that in the data?
Yes. Children with high total scores who show extreme time-per-item variance (very fast on easy items, very slow on medium ones) are the sub-group that most often underperforms on a real test. It looks like a confidence and pacing pattern more than a knowledge one.
Should the child use a timer at home?
Yes, but with a specific goal. The evidence in our own data and in wider timing-training research suggests practising with time visible works better than practising against a strict cut-off. The point is to build awareness of pace, not to add pressure that reproduces the exam-day panic you are trying to reduce.
Is the sample representative of all UK 11+ candidates?
No. Families using Deep Teaching are self-selected — they came looking for a diagnostic tool. That means they are on average more engaged and probably slightly higher-attaining than the national 11+ candidate pool. The patterns above are strongest for children in the middle of the 11+ preparation range; the tails are less well-covered. See Methodology for the full sampling caveat.
This analysis describes patterns across 1,654 real diagnostics. It cannot tell you where your specific child sits on any of those patterns — that is exactly what a Deep Teaching session, not an article, is for.
Evidence level
Original analysis of Deep Teaching's own diagnostic platform — 1,654 completed sessions, 37,556 individually timed and scored answers, 1,263 families. Per LS32 §9, product-level parent behaviour data is the most defensible source we have because no competitor has access to it.
Methodology
Inclusion criteria
- Completed Airteacher diagnostic sessions with a full per-question attempts log, January–July 2026.
Exclusion criteria
- 1,476 internal test, simulation and staff drive-test sessions (identified by test email domains and known internal profiles) were removed before analysis. All published figures use the real-user subset only: 1,654 sessions from 1,263 distinct families.
Variables
- Per question: correct/incorrect; time ratio (actual answering time ÷ the question's expected time budget); difficulty tier 1–5; topic; school year; adaptive route through the blocks.
- Item discrimination: point-biserial correlation between success on a question and the child's score on the rest of their session, computed for the 261 questions answered at least 30 times.
- Early-predicts-late: Pearson correlation between accuracy on the first two blocks and accuracy on all subsequent blocks (1,619 sessions with both).
Weighting
- None. All statistics are plain counts, medians and correlations on the real-user subset.
Assumptions
- Expected time budgets per question are as set by Airteacher's curriculum designers.
- Answer timings above 30 minutes on a single question were treated as walk-aways and excluded from timing medians.
Limitations
- Airteacher's users are self-selected families, not a nationally representative sample.
- Each session is a snapshot; this analysis is not longitudinal and does not measure progress over time.
- The speed finding is a correlation: fast-and-wrong answering flags impulsive behaviour, but the analysis does not prove that slowing a particular child down will raise their score.
- Topic mix varies by school year, so topic success rates partly reflect the years in which each topic is taught.
References
- Wise, S. L. & Kong, X.. (2005). Response Time Effort: A New Measure of Examinee Motivation in Computer-Based Tests ↗ external. Applied Measurement in Education, 18(2). DOI: 10.1207/s15324818ame1802_2 ↗
- Vygotsky, L. S.. (1978). Mind in Society: The Development of Higher Psychological Processes. Harvard University Press
- Bjork, R. A.. (1994). Memory and metamemory considerations in the training of human beings. In Metcalfe & Shimamura (eds.), Metacognition: Knowing about Knowing, MIT Press
- Crocker, L. & Algina, J.. (1986). Introduction to Classical and Modern Test Theory. Holt, Rinehart and Winston
- Airteacher. (2026). Airteacher diagnostic database: 1,654 completed diagnostics, January–July 2026. Internal dataset (all computations reproducible from the per-question attempts log)
Is it worth paying for a smaller class?
Smaller classes produce measurably better results in the earliest school years and for the youngest, poorest or lowest-attaining children. Above about 20 pupils, and above about age 8, the attainment effect fades. Teacher quality dominates.
Does paying more buy a better outcome? What £30,000 and £50,000 a year actually get at 25 leading senior schools
Across 25 leading senior schools, published fees explain only about a third of the differences in Oxbridge outcomes. The most academically selective schools tend to be the best value per pound of fees, and one school matches a £49k-a-year rival's Oxbridge success rate at £17,274 a year less.