When the System Is Wrong
Failures happen at scale and in public. What to do in the first week, and what distinguishes institutions that recover from ones that do not.
Process · Procedure
Something will go wrong: a mass false-flag event, a supplier outage mid-exam, a wrongly upheld case that becomes public. The response is more consequential than the failure. The correction route should be as visible as the workflow in the boundary example, so an incorrect flag can be traced, amended and prevented from silently shaping later decisions.
The common failure modes
A configuration change producing a flag spike across a whole cohort.
An outage during a timed exam, affecting hundreds simultaneously.
A single case that turns out to be wrong and reaches the press or social media.
A data incident: recordings exposed, or accessed inappropriately.
And a systemic discovery, such as a disparity that should have been noticed earlier.
The first day
Stop the process that is producing the harm. Pause reviews, suspend cases arising from the affected window, hold notifications.
Say something quickly, even if it is only that you are looking into it and no case will proceed meanwhile.
Preserve everything, including configuration history, because the question of what changed and when will arise.
And do not let individual cases proceed while the cause is unknown, which is the error that turns a technical fault into a set of unjust findings.
Telling affected students
Directly, not by a notice on a portal.
Specifically: what happened, what it means for them, what happens next, and when they will hear.
With an explicit statement that no adverse inference is drawn where that is true.
And a named contact who can answer, staffed for the volume.
Fixing the record
Withdraw findings fully where they rest on the affected material.
No annotations, no "case not upheld" notes that surface later.
And communicate the withdrawal in writing to the student, which they will need if it ever comes up.
The exam outcome
A student whose exam was disrupted has lost time and composure regardless of any integrity question.
Offer a resit without penalty, or a documented adjustment, promptly.
And do not require them to demonstrate the harm, which places an evidential burden on the person the institution failed.
The review afterwards
What changed, who approved it, what testing preceded it.
Whether the monitoring would have caught it — flag rate spikes are detectable within hours if anybody is watching.
And whether the supplier notified you or you discovered it, which is a contractual question worth having an answer to.
What distinguishes a good response
Speed, specificity, and not requiring students to prove the institution's failure.
Institutions that handle this well are not those with fewer failures. They are those that stopped quickly, said what happened, and fixed the records without argument.
The opposite pattern — defending the system, processing the cases, making students appeal individually — converts a technical fault into a reputational one.
What to check
Is there a named person who can pause the whole process?
Would a flag rate spike be noticed within a day?
Has anybody rehearsed the communication?
And is there a standing rule that cases pause while a cause is unknown?
The point
Institutions that handle failures well are not those with fewer failures.
They are those that stopped quickly, said what happened, and fixed the records without requiring students to argue.
Worth stating
A flag rate spike is detectable within hours if anybody is watching.
Whether the supplier notified you or you discovered it is a contractual question worth having an answer to before it matters.
Also worth knowing
Withdraw findings fully where they rest on affected material: no annotations, no case-not-upheld notes that surface later.
Communicate the withdrawal in writing, which the student will need if it ever comes up.
And finally
Say something quickly, even if it is only that no case will proceed meanwhile, and preserve configuration history because the question of what changed and when will arise.
Speed and specificity are what distinguish a good response.
Summary
Pause cases while a cause is unknown. That single standing rule prevents a technical fault becoming a set of unjust findings.
In summary
Failures happen at scale and in public.
What distinguishes institutions that recover is stopping quickly, saying what happened, and fixing records without requiring students to argue for it. For wider institutional context, consult the Information Commissioner's Office.