Skip to content
Flag Is Not Finding

All notes  /  Types

Recorded and Reviewed Later

The session is recorded, software flags moments, and a human looks at the flags afterwards. The most common arrangement and the one whose weak point is review capacity.

Types · Analysis

Nobody watches at the time. The recording plus a flag list goes to a reviewer, who decides what to do with each. A recorded-proctoring queue is partly an employee productivity reporting problem because institutions must see the volume, age and ownership of cases without confusing reviewer throughput with student guilt.

Why it is the common choice

It scales. Students sit when they like and the review happens later.

It is cheaper than live watching by a large factor.

And it produces evidence — a recording that can be re-examined during an appeal, which live watching does not.

Where it fails

Review capacity. Two thousand exams producing four thousand flags requires somebody to look at four thousand moments, and institutions consistently under-resource this.

The failure mode is not that flags are ignored — it is that they are confirmed quickly, which is worse.

And the gap in time: a student asked about a moment from three weeks ago cannot remember it, which disadvantages them in a way nobody intended.

Sizing the review

Estimate flags per session from a pilot, multiply by cohort, divide by the time a reviewer needs per flag.

Two to five minutes per flagged moment is realistic if the reviewer is actually watching the surrounding context.

If the resulting figure is impossible, the programme is not viable as designed and the answer is a higher flag threshold, a smaller scope, or a different approach entirely.

Do this arithmetic before purchase. Almost nobody does.

Who reviews

Not the module leader who wrote the exam, ideally: they have a view about the cohort and the assessment.

A trained reviewer with no stake in the outcome, working to written criteria.

And a second reviewer for anything proceeding to a case, which is cheap insurance and is standard in other evidence-handling contexts.

Time limits

Set a maximum interval between exam and review.

Beyond a few weeks the student's ability to explain collapses, and the fairness of the process with it.

A programme that reviews at the end of term is running an unfair process, however good its detection.

What the recording should contain

Enough context around each flag to interpret it — thirty seconds either side at minimum, and the ability to view the whole session.

A reviewer who can only see the flagged clip cannot assess an explanation, which is a common and avoidable limitation.

Retention

Keep recordings long enough for the appeal period and no longer.

Which means the retention period follows the appeal deadline rather than a default in the supplier's settings.

And deleted sessions should actually be deleted, including from any supplier backup, which is worth confirming in writing.

What to check

How many flags per thousand sessions does your system produce?

How many reviewer hours are budgeted against that?

How long between exam and review, in practice?

And can a reviewer see the whole session, or only the flagged clip?

The point

Estimate flags per session from a pilot, multiply by cohort, divide by the minutes a reviewer needs.

If the answer is impossible, the programme is not viable as designed, and that arithmetic belongs before purchase.

Worth stating

A reviewer who can only see the flagged clip cannot assess an explanation.

Thirty seconds either side at minimum, with the whole session available, or the review step cannot do what it exists to do.

Also worth knowing

The failure mode of an under-resourced review is not that flags are ignored — it is that they are confirmed quickly, which is worse.

And a student asked about a moment from three weeks ago cannot remember it.

And finally

Set a maximum interval between exam and review, because beyond a few weeks the student's ability to explain collapses and the fairness of the process goes with it.

Reviewing at the end of term is an unfair process however good the detection.

Summary

Give the reviewer the whole session, not only the flagged clip, or they cannot assess an explanation against what surrounds it.

In summary

Record and review scales and produces evidence, which live watching does not.

Its weak point is entirely the human step, and institutions consistently under-resource exactly the part they are actually buying. For wider institutional context, consult the National Institute of Standards and Technology.