Measuring Whether It Works
Seven figures that describe whether the programme is functioning, and the ones institutions report instead.
In practice · Reference
Most proctoring reporting counts flags. Flags measure the software's sensitivity, not the programme's effectiveness or fairness. Useful workforce analytics separates throughput from outcomes; the same distinction is essential when comparing flag volume with fairness, support demand and upheld cases.
The seven
Flag rate per hundred sessions, by signal type.
Proportion of flags closed at review as nothing.
Cases brought per thousand sessions.
Proportion of cases upheld.
Flag rate by student group, and upheld rate by student group.
Time from exam to resolution.
And appeals brought, upheld, and the grounds on which they succeeded.
What each tells you
A high flag rate with a high closure rate means the threshold is wrong. Reviewer time is being spent to produce nothing.
A low closure rate means either genuinely good triage or reviewers confirming rather than assessing — and the two are distinguished by calibration.
A disparity in flag rate that does not appear in upheld rate means a group is carrying investigation burden for no corresponding finding.
Long resolution times are a harm in themselves.
And appeals succeeding on the same ground repeatedly is a configuration problem, not a series of individual errors.
The figures usually reported instead
Total flags, which rewards sensitivity.
Sessions monitored, which measures deployment.
Cases brought, which is presented as evidence the system works and is equally consistent with the system being wrong.
And supplier-provided dashboards, which are built around the product rather than around the institution's question.
The comparison that matters
Against the institution's own figures over time, not against other institutions or vendor benchmarks.
Nobody collects the same way and the cohorts differ.
Your flag rate last year, under the same configuration, is the only fair comparison and it is the one that shows whether a change did anything.
The question the figures should answer
Is the programme finding misconduct that would otherwise go undetected, at a cost in investigation burden and distress that is proportionate?
Cases upheld is the numerator.
Flags, reviewer hours, student anxiety and appeals are the denominator, and the denominator is usually unmeasured.
Who should see them
Whoever can change the configuration.
The academic committee responsible for assessment.
And students, in aggregate, which is unusual and is the single most credibility-building thing an institution can do here.
Publishing
Flag rate, closure rate, cases, upheld, appeals.
Annually, in aggregate, without identifying anybody.
An institution willing to publish these is making a claim it can support, and the willingness itself tells students something about how the programme is run.
What to check
Which of the seven does your institution produce?
Does anybody look at them who could change the threshold?
Have the figures ever caused a change?
And would you be willing to publish them?
The point
Flag counts measure the software's sensitivity, not the programme's effectiveness.
The seven figures that matter include disaggregated rates and appeal grounds, and most reporting contains none of them.
Worth stating
Publishing flag rate, closure rate, cases, upheld and appeals annually in aggregate is unusual and is the single most credibility-building thing an institution can do here.
The willingness itself tells students how the programme is run.
Also worth knowing
The question the figures should answer is whether the programme finds misconduct that would otherwise go undetected, at a cost in investigation burden and distress that is proportionate.
The denominator is usually unmeasured.
And finally
Show the figures to whoever can change the configuration and to the committee responsible for assessment.
Reporting that nobody with authority reads is a record rather than a control.
Summary
Publish the figures in aggregate. The willingness itself tells students how the programme is run, and it is the single most credibility-building step available.
In summary
Seven figures describe whether a programme functions; flag counts describe only the software's sensitivity.
The denominator — reviewer hours, student anxiety, appeals — is what institutions leave unmeasured. For wider institutional context, consult the U.S. Federal Trade Commission.