Piloting Properly
A pilot that produces the three numbers determining viability, rather than a demonstration that produces enthusiasm.
In practice · Procedure
Most pilots in this area establish that the software runs. The useful pilot establishes what it will cost to operate. A trial of time tracking software shows the right principle: test a real workflow with representative users, then keep only the fields and controls that answer the stated question.
The three numbers
Flag rate: flags per hundred sessions, by type.
Review load: minutes actually spent per flag by a reviewer doing the job properly.
Support volume: contacts per hundred candidates, and when they arrive.
Multiply the first two by cohort size and compare against available staff. That calculation is the decision.
Designing the pilot
A real assessment with real stakes, not a mock. Behaviour differs and so does the flag rate.
A cohort large enough for the rates to mean something — a few hundred rather than twenty.
Diverse enough to see disparity, which a single small module will not show.
And run for the full process, including reviewing every flag and handling any case, because the review step is what is being tested.
What to measure beyond the three
Flag rate by student group, which is the fairness baseline.
Proportion of flags resolved as nothing.
Check-in failure rate and time to resolve.
Session dropout rate.
And student experience, asked directly and anonymously, which will be the most uncomfortable and most useful output.
Asking students
A short anonymous survey after the pilot exam.
What was difficult, what was unclear, what would they change, would they prefer a supervised alternative.
Include an open field and read the responses rather than the averages, because the specific incidents are where the problems are.
What a pilot usually reveals
That the flag rate is several times higher than expected.
That reviewing them properly is impossible at full scale with current staff.
That check-in takes longer than anybody budgeted, and fails for a predictable subset.
And that a meaningful proportion of students had an experience the institution would not want described publicly.
All four are better discovered in a pilot.
The decision after
Three outcomes are legitimate: proceed with a changed configuration, proceed with a narrower scope, or do not proceed.
The third is a real option and pilots are supposed to be able to produce it.
A pilot that could only ever have concluded "proceed" was a rollout with a different name.
Writing it up
The three numbers, the fairness figures, the student feedback, and the recommended configuration.
Circulated to whoever signs the contract, before signature.
And kept, because it is the evidence that the decision was made on something.
What to check
Did your pilot produce a flag rate and a review time?
Was every flag actually reviewed?
Were students asked, anonymously?
And was "do not proceed" ever a possible outcome?
The point
A pilot that could only ever have concluded proceed was a rollout with a different name.
Three outcomes are legitimate, and not proceeding is one of them.
Worth stating
Ask students directly and anonymously after the pilot, with an open field, and read the responses rather than the averages.
The specific incidents are where the problems are, and they will be the most uncomfortable output.
Also worth knowing
Run the pilot on a real assessment with real stakes, large enough and diverse enough for the rates to mean something, and review every flag properly.
A mock exam produces different behaviour and a different flag rate.
And finally
Measure beyond the three numbers: flag rate by group, check-in failure rate and time to resolve, session dropout rate, and student experience asked anonymously.
The last will be the most uncomfortable and most useful output.
Summary
Make not proceeding a possible outcome. A pilot that could only ever conclude proceed was a rollout with a different name.
In summary
A useful pilot produces the flag rate, the review load and the support volume.
Multiply the first two by cohort size, and that calculation is the decision. For wider institutional context, consult the Information Commissioner's Office.