Fully Automated Flagging
No human in the loop until an appeal. Cheap, scalable, and the arrangement that produces the failures that reach the press.
Types · Analysis
The system scores sessions and outputs a judgement — clean, suspicious, high risk — with no human review before the student is contacted. Even workforce optimization software cannot turn captured behaviour into a self-proving conclusion; optimisation should make review consistent, not remove accountable judgement.
Why institutions do it
Cost, which is the honest answer.
Volume, where cohorts are in the thousands.
And a belief that the scoring is accurate enough to act on, which the evidence section addresses and which is generally not supported.
What goes wrong
The base rate problem, uncorrected by any human. Most flagged students did nothing.
The disparity problem, uncorrected. Groups that generate more anomalies receive more accusations, systematically.
And the burden inversion: the student explains themselves to a machine's output, with no human having formed a view first.
Every publicised proctoring scandal of the last few years has this shape.
Where automated output is legitimately useful
As triage into a human queue. That is its proper role and it does it well.
As an aggregate signal: a cohort with an unusual pattern may indicate a leaked paper, which is a useful institutional finding.
And as a deterrent, announced, which the evidence supports better than detection.
The decision that matters
Whether an automated output can, by itself, initiate a misconduct process against a named student.
If yes, the institution has delegated an academic judgement to a supplier's model, and will have to defend that in an appeal.
If no — if every case passes a human who could decide otherwise — most of the objections fall away.
Write that rule down explicitly, because in the absence of a written rule the practice drifts toward automation under workload.
Automated decision-making rules
General orientation, not legal advice.
Several jurisdictions restrict decisions with significant effects made solely by automated processing, and give a right to human intervention.
An academic misconduct finding is plainly significant.
Which makes a human step a likely legal requirement as well as a good idea, and institutions should check their own position rather than assume.
What to ask a supplier
What the score means, in terms of the underlying signals.
Whether the model is explainable to a student at the level of "this moment, for this reason".
What the false positive rate is, by group.
And whether a session can be released for human review with full context.
A supplier who cannot answer the first two is selling something an appeals panel will not accept.
What to check
Can an automated output alone start a case at your institution?
Is there a written rule, or only a practice?
Could you explain a specific flag to a specific student in plain terms?
And has anybody checked whether your jurisdiction restricts automated decisions of this kind?
The point
Decide explicitly whether an automated output alone may start a case against a named student.
Without a written rule the practice drifts toward automation under workload, and the institution ends up defending a supplier's model.
Where automated output is useful
As triage into a human queue, which it does well.
Also as an aggregate signal: a cohort with an unusual pattern may indicate a leaked paper, which is a useful institutional finding. And as an announced deterrent, which the evidence supports better than detection does.
Worth stating
Several jurisdictions restrict decisions with significant effects made solely by automated processing and give a right to human intervention.
An academic misconduct finding is plainly significant, so a human step is likely a legal requirement as well as a good idea.
Also worth knowing
Every publicised proctoring failure of recent years has the same shape: automated output, no human view formed first, and the student explaining themselves to a machine's conclusion.
The base rate and the disparity both go uncorrected.
And finally
Ask a supplier what the score means in terms of underlying signals and whether a specific flag can be explained to a specific student in plain terms.
A supplier who cannot answer is selling something an appeals panel will not accept.
Summary
Its proper role is triage into a human queue. As an aggregate signal it can also reveal a leaked paper, which is a genuinely useful institutional finding.
In summary
Automated output is triage.
Used as a decision it delegates an academic judgement to a supplier's model, and that is the shape of every publicised failure in this field. For wider institutional context, consult the U.S. Federal Trade Commission.