Skip to content
Flag Is Not Finding

All notes  /  Foundations

Anomaly Is Not Cheating

The single distinction that determines whether a proctoring programme is defensible, and the one most often collapsed in practice.

Foundations · Analysis

A flag is a statement that something differed from expectation. Treating it as a statement that something wrong happened is the error that produces every serious failure in this field. The same boundary is visible in this Monitask guide, where recorded activity can prompt a question but still needs context before anybody draws a conclusion.

How the collapse happens

The report says "suspicious activity detected", which is the vendor's wording rather than a finding.

A reviewer with two hundred sessions and limited time treats the flag as the conclusion.

The student receives a letter referring to an integrity investigation.

And the burden has quietly inverted: the student is now explaining away a machine output rather than the institution demonstrating misconduct.

Nobody decided this should happen. It happens through wording and workload.

What the base rate does

If a small proportion of students cheat and a much larger proportion produce anomalies, most flags are innocent even if the system is accurate.

This is arithmetic rather than opinion, and it holds for any detector applied to a rare event.

Which means a programme that treats flags as probable misconduct will be wrong most of the time, and will be wrong specifically about the students whose circumstances generate anomalies.

The wording that matters

"Flagged for review" rather than "suspicious activity".

"The recording shows X" rather than "the system detected cheating".

"We are asking about" rather than "you have been accused of".

These are not euphemisms. They are accurate descriptions of what happened, and using inaccurate stronger language is what turns a review into an accusation before anybody has looked.

What a flag can support

A decision to look at a recording.

A question to a student, asked neutrally.

And, combined with other evidence, a finding.

What it cannot support on its own is a finding of misconduct, and an institution whose process allows that has a problem that will surface in an appeal.

The reviewer's task

Not "did the system flag this" — it did, that is why they are looking.

But "does the recording show conduct that breaches the rules, and is there another explanation".

Framed that way the reviewer is doing assessment rather than confirmation, and the difference in outcomes is substantial.

Where institutions get caught

A policy that treats the flag rate as a misconduct rate.

Reporting that counts flags as incidents.

And a review step so brief that it can only rubber-stamp, which is the commonest structural failure and is visible in the time allocated per session.

What to check

Does your written policy distinguish a flag from a finding?

What words does your student-facing communication use?

How long does a reviewer spend per flagged session — actually?

And has anybody calculated what proportion of flags are resolved as nothing?

The point

If a small proportion of students cheat and a much larger proportion produce anomalies, most flags are innocent even when the system is accurate.

That is arithmetic, and a policy that ignores it will be wrong most of the time.

Worth stating

The reviewer's task is not whether the system flagged it, since it did, but whether the recording shows conduct breaching the rules and whether another explanation fits.

Framed that way the reviewer assesses rather than confirms, and the difference in outcomes is substantial.

Also worth knowing

Institutions get caught by policies that treat the flag rate as a misconduct rate, and by reporting that counts flags as incidents.

Both are visible in the time allocated per session, which is the honest measure of whether review is real.

And finally

A flag can support a decision to look, a neutral question to a student, and — combined with other evidence — a finding.

What it cannot support on its own is a finding of misconduct.

Summary

Nobody decides that a flag should become an accusation. It happens through wording and workload, which is why both are worth designing deliberately.

In summary

Most flags are innocent even when the system is accurate, because the underlying behaviour is rare and the anomalies are not.

That arithmetic does not change with a better product, and a process built without it will be wrong at scale. For wider institutional context, consult EDUCAUSE.