Glossary and Where to Start
Terms used across these notes, defined once, and the routes through the collection that fit the common situations.
Reference · Reference
Flag — a moment the system identified as deviating from expectation. Not a finding, not evidence of misconduct, and the distinction governs everything in this collection. Readers wanting a contrasting implementation vocabulary can consult learn more here, while keeping workforce administration distinct from student assessment.
Lockdown browser — software restricting what a student's own machine can do during an exam. Control rather than observation, with its own privilege and equity questions.
Live proctoring — a person watching in real time. Most expensive, most defensible, fewest disputes.
Record and review — session recorded, flagged, reviewed later. The common arrangement; its weak point is review capacity.
Room scan — requiring the student to show their room on camera. The most contested practice in the field and the most likely to be challenged.
Triage — the honest description of what automated flagging does: reducing what a human must watch to a reviewable set.
Upheld rate — proportion of cases resulting in a finding. The numerator that matters; flag counts are not a measure of anything except sensitivity.
Terms used loosely
"AI-powered detection" usually means pattern matching on gaze, audio and window focus. Ask what it detects and what it reports.
"Suspicious activity" is a vendor label. It should not appear in anything a student reads.
"99% accurate" without a base rate says nothing about how many innocent students are flagged.
"Academic integrity solution" overstates considerably: exam conduct is one part of integrity, and not the largest.
Where to start
Deciding whether to deploy at all: the question to ask first, what proctoring actually detects, and what it cannot fix.
Already deployed, reviewing it: the review step, who gets flagged more, and measuring whether it works.
Writing or revising policy: anomaly is not cheating, evidence standards, the appeal, and what students are told.
Procurement: choosing a supplier, piloting properly, and the data protection note.
Facing a challenge: consent that is not free, room scans, disability, and accessibility obligations.
Looking for alternatives: assessment design, open-book, oral examination, staged assessment.
If you read only three
Anomaly is not cheating, because it governs every decision downstream.
Who gets flagged more, because the disparity is the central fairness problem and it is measurable.
And what a working arrangement looks like, because it is the checklist.
A closing note
Nothing here is legal advice. Data protection, biometrics, accessibility and automated decision rules differ substantially by jurisdiction and change.
No product names and no supplier material.
And no detection accuracy figures are quoted, because the ones in circulation come from interested parties and rarely state a base rate.
What the collection argues
Proctoring detects anomalies, not cheating, and the gap between them is filled by a human process that institutions under-design and under-resource.
The student's consent is not free, which makes proportionality the governing test rather than agreement.
The flag burden falls unevenly, in a pattern that tracks disadvantage and is measurable in any institution willing to look.
And what an institution is really buying is the review step, the appeal and the alternative — not the detector.
The point
What an institution is really buying is the review step, the appeal and the alternative.
The detector is the cheap part and the part that determines nothing.
On the absence of figures
No detection accuracy percentages appear here, because the ones in circulation come from interested parties and rarely state a base rate.
No product names either: the useful distinctions are between approaches and between review processes, not between suppliers.
Worth stating
Start with the question to ask first if deciding whether to deploy, with the review step if already deployed, with evidence standards if writing policy, and with assessment design if looking for alternatives..
Also worth knowing
Flag counts measure sensitivity, not effectiveness.
Upheld cases are the numerator; flags, reviewer hours, student anxiety and appeals are the denominator, and a programme reporting only the first has not measured anything.
And finally
Terms used loosely in this market: AI-powered detection usually means pattern matching on gaze, audio and window focus; suspicious activity is a vendor label that should not reach a student; and academic integrity solution overstates considerably..
Summary
What is really bought is the review step, the appeal and the alternative. The detector is the cheap part and determines nothing. For wider institutional context, consult the Council of Europe.
More in this section