Question generation across all four papers
Listening, reading, writing and speaking, generated rather than drawn from a fixed pool so that practice does not become memorisation of the answer key.
Candidates do not need more practice questions. They need to know why they got a 6.5. Scoring is the entire product, which means a model that marks generously is pleasant, popular and worthless.
Preparation meant a book and a guess. Where feedback existed it arrived from a tutor days later, described the answer rather than the marking, and rarely explained why a piece of writing landed on one band rather than the one above it. Candidates were practising without a signal.
Listening, reading, writing and speaking, generated rather than drawn from a fixed pool so that practice does not become memorisation of the answer key.
A band, plus what pushed it there: task response, coherence, lexical range, grammatical accuracy. The reasoning is the part candidates act on; the number alone teaches nothing.
The scoring was measured against real examiner marks and adjusted until it agreed, because a scorer that is consistently half a band optimistic is worse than no scorer.
What the candidate keeps getting wrong drives what they see next, rather than a fixed syllabus that ignores the evidence sitting in their own attempt history.
Progress across a group, where a cohort is collectively weak, and which candidates have stopped. Institutions buy this half; candidates use the other half.
Everything in an assessment product depends on the marking being trustworthy, and every incentive pushes the other way. Generous marking demos better, retains better and reviews better. The engineering discipline was refusing to tune towards the pleasant answer, and the way you enforce that is to fix the calibration set first and treat any drift away from it as a regression rather than an improvement.
Most of our contracts ask us not to name the client, and we would rather keep the work than the logo. We have also not put a percentage improvement on this page, because we would be estimating it and you would be right to discount it.
What we will do is arrange a reference conversation with the client, with their agreement, for work that resembles yours. Ask on the first call.
Thirty minutes. If your problem rhymes with this one we will say what we would do differently the second time, which is usually the more useful half.
Either one reaches Umer directly. No forms sitting in a queue.