Skip to content
Flag Is Not Finding

All notes  /  In practice

Training the Reviewers

The people who turn flags into decisions usually receive an hour on the interface and nothing on judgement. That gap produces most unfair outcomes.

In practice · Procedure

Reviewer training in most institutions covers how to operate the software. What it needs to cover is what the flags mean and what they do not. A weekly timesheet template can make calibration and case-review time visible, helping managers budget the human work that automated flags do not remove.

What the training should contain

That a flag is a moment for review, not a finding — stated first and repeated.

What each signal actually measures and how weak the inference is, signal by signal.

What disability-related behaviour looks like on camera: stimming, gaze aversion, vocalisation, medication breaks.

What a child entering a room looks like, and how it differs from a collaborator.

What a failed connection looks like, and that it is never an integrity matter.

And the evidence hierarchy, with worked examples.

Worked examples

Twenty real sessions, anonymised, with the correct dispositions agreed in advance by experienced staff.

Reviewers work through them and compare.

This is the part that changes behaviour; abstract principles do not transfer to a screen at four in the afternoon with sixty sessions left.

Rebuild the set annually as the cohort and configuration change.

Calibration

The same twenty sessions, several reviewers, compared quarterly.

Disagreement is expected at first and is the point.

Track whether it narrows. A team whose disagreement does not narrow has criteria that are too vague to apply.

What to tell them about their own position

That closing a flag as nothing is a correct and expected outcome, and the great majority should be closed.

That they will not be judged on how many they escalate.

And that they may mark something inconclusive — which sounds obvious and, without permission, does not happen: reviewers under pressure resolve ambiguity toward escalation because it feels safer.

Saying this explicitly changes the distribution of outcomes measurably.

Workload and fatigue

Review quality degrades over a long session, as with any inspection task.

Cap the number per sitting, build in breaks, and avoid scheduling the whole cohort's review in two days.

A reviewer on their ninetieth session is not doing the same job as on their tenth, and the students at the end of the queue should not be disadvantaged by their position in it.

Who should not review

Anybody who taught the module, in most cases.

Anybody with a view about the student.

And anybody whose performance is measured on cases brought, which is an incentive that should not exist anywhere in this process.

Support for reviewers

This is unpleasant work: watching students' homes, making decisions that affect their futures.

Reviewers see distressing material occasionally — illness, domestic incidents, distress during an exam.

There should be a route for that, and it should be mentioned in training rather than discovered.

What to check

How long is your reviewer training, and how much of it is about judgement?

Is there a worked example set?

Has calibration been run, and did disagreement narrow?

And has anybody told reviewers that closing flags as nothing is the expected outcome?

The point

Reviewers under pressure resolve ambiguity toward escalation because it feels safer.

Telling them explicitly that closing a flag as nothing is the expected outcome changes the distribution measurably.

Worth stating

This is unpleasant work — watching students' homes and making decisions that affect their futures — and reviewers occasionally see distressing material.

There should be a route for that, mentioned in training rather than discovered.

Also worth knowing

Worked examples are the part that changes behaviour: twenty real sessions, anonymised, with dispositions agreed in advance.

Abstract principles do not transfer to a screen at four in the afternoon with sixty sessions left.

And finally

Cap the number of reviews per sitting and avoid scheduling a whole cohort in two days, because inspection quality degrades over long sessions.

Students at the end of the queue should not be disadvantaged by their position in it.

Summary

Train on judgement rather than on the interface, with worked examples and quarterly calibration. Abstract principles do not survive the sixtieth session of an afternoon.

In summary

Reviewers are usually trained on the interface and not on judgement.

That gap produces most of the unfair outcomes, and a worked example set plus quarterly calibration closes most of it. For wider institutional context, consult the National Institute of Standards and Technology.