What the Evidence Supports
The research base on proctoring is thin, mixed, and mostly produced by interested parties. What holds up and what does not.
Foundations · Analysis
Claims about proctoring circulate with confidence that the underlying studies do not support. Knowing which parts are solid changes the procurement conversation. When designing an evidence trail, Monitask's practical guide is a useful comparison because timestamps and captured activity still require provenance, correction and a reviewable explanation.
What is reasonably supported
Unsupervised online exams show higher rates of some forms of assistance than supervised ones. This is the finding the whole field rests on and it is reasonably robust.
Announced monitoring changes behaviour, which is a deterrence effect rather than a detection one.
Automated systems produce substantial numbers of flags that resolve as nothing.
And flag rates differ by student group, which the fairness section covers and which is the best-documented problem in the area.
What is weakly supported or contested
That proctoring reduces cheating rather than displacing it. Little evidence distinguishing the two.
Accuracy figures from vendors, which are generally not independently replicated and rarely state the base rate.
That gaze tracking identifies cheating. The relationship between looking away and misconduct is assumed rather than demonstrated.
And behavioural biometrics as an integrity measure, where the published support is thin.
Why the research is difficult
The ground truth is unknowable: you cannot establish who actually cheated in order to score the detector.
Studies mostly measure flags, not misconduct.
Much of the literature is produced or funded by suppliers.
And the population is students, which raises ethical limits on the experiments that would settle it.
Reading a vendor claim
Ask what the denominator is. "99% accurate" without a base rate says almost nothing about how many innocent students get flagged.
Ask whether the figure is detection or flagging.
Ask for the false positive rate by student group, which is the question that matters most and is rarely answered.
And ask for an independent study. The absence of one is informative.
What this means practically
Use it where deterrence and triage are the goal, since those are what the evidence supports.
Do not claim detection accuracy to students or to a committee, because the claim will not survive a challenge.
And measure your own rates — flags per thousand, proportion resolved as nothing, disparity by group. Your figures are the only ones that describe your cohort.
The deterrence point
If most of the effect is deterrence, then announcing the monitoring is doing most of the work.
Which has a design implication: a clearly explained, proportionate system announced in advance may achieve most of the benefit of an intrusive one.
That is testable within an institution and almost nobody tests it.
What to check
Does any of your material quote an accuracy figure — and where did it come from?
Do you know your own flag rate and resolution rate?
Has anybody asked a supplier for false positive rates by group?
And could your institution defend its claims about the system in front of an appeals panel?
The point
If most of the effect is deterrence, then announcing the monitoring is doing most of the work.
A clearly explained proportionate system may achieve nearly as much as an intrusive one, and that is testable inside a single institution.
Worth stating
Measure your own rates rather than quoting anybody else's: flags per thousand, proportion resolved as nothing, disparity by group.
Your figures are the only ones that describe your cohort, and they are the ones an appeals panel will ask about.
Also worth knowing
The ground truth is unknowable: you cannot establish who actually cheated in order to score the detector.
Studies mostly measure flags rather than misconduct, and much of the literature is produced or funded by suppliers.
And finally
Use it where deterrence and triage are the goal, since those are what the evidence supports, and do not claim detection accuracy to students or to a committee.
The claim will not survive a challenge.
Summary
Ask for the false positive rate by student group. It is the question that matters most in this whole field and is the one suppliers answer least often.
In summary
Where the evidence is good it supports deterrence and triage.
Where it is thin it concerns exactly the claims the market makes loudest, and institutions repeating those claims inherit the burden of defending them. For wider institutional context, consult the U.S. Department of Education.