Skip to content
Flag Is Not Finding

All notes  /  Fairness

What to Do About the Disparity

Measuring the gap is the easy part. Four responses that actually reduce it, and two that only appear to.

Fairness · Procedure

An institution that measures flag rates by group will find a disparity. What follows determines whether the measurement was worth making. Institutions can use the connected-work guide to structure the corrective work—owners, deadlines and follow-up—without treating a productivity record as evidence about individual students.

Measure first, properly

Flag rate by group.

Proportion of flags proceeding to a case, by group.

Proportion upheld, by group.

And time from exam to resolution, by group, because a slow process is a heavier burden on students with less slack.

Four figures. Quarterly. Seen by somebody with authority to change the configuration.

Response one: lower the sensitivity

Most disparity comes from signals that generate high volumes of weak flags — gaze, movement, brief absence.

Raising the threshold reduces the total and reduces the disparity disproportionately, because the excess flags on affected groups are concentrated in exactly those weak signals.

This is the single most effective change and it is a configuration setting.

Response two: remove signals that do not work

Behavioural and gaze-only flags that never survive review are pure cost.

Count how many of each type led to an upheld case in the last year.

Any signal with a near-zero conversion rate should be switched off, and this is an easy internal argument because it also reduces reviewer workload.

Response three: get the context to the reviewer

Access arrangements, declared circumstances, known connection problems — available before review rather than discovered during a case.

A reviewer with the relevant fact resolves in seconds what otherwise becomes a letter.

This does not reduce flags and substantially reduces harm.

Response four: make the alternative real

A supervised on-campus option, offered openly, no reason required.

Students most affected by the disparity are the ones most likely to take it, which removes them from the flagged population entirely.

This is the expensive response and the most complete one.

What only appears to help

Telling students to prepare their environment better, which places the burden on the affected and changes nothing structural.

Adding a sentence to the policy about fairness.

And reviewer discretion without training or data, which produces inconsistency rather than fairness.

The uncomfortable finding

Sometimes the measurement shows that a particular assessment cannot be proctored fairly for a particular cohort.

The correct response is to change the assessment, not to accept the disparity as a cost of the method.

An institution unwilling to reach that conclusion should think carefully about whether to measure at all — because knowing and not acting is a worse position than not knowing.

What to check

Are the four figures produced, and by whom?

Who sees them, and can that person change the configuration?

When the disparity was last found, what changed?

And is there any assessment where the honest answer is that the method does not fit?

The point

Raising the threshold is the single most effective change and it is a configuration setting.

The excess flags on affected groups are concentrated in the weak signals, so reducing those reduces the disparity disproportionately.

Worth stating

Four figures, quarterly, seen by somebody with authority to change the configuration: flag rate by group, proportion proceeding, proportion upheld, and time to resolution.

Without the last one, a slow process hides as a fair one.

Also worth knowing

What only appears to help: telling students to prepare their environment better, adding a sentence about fairness to the policy, and reviewer discretion without training or data.

The first shifts the burden to the affected.

And finally

Remove signals that never survive review: count how many of each type led to an upheld case last year.

Any with a near-zero conversion rate is pure cost, and switching them off also reduces reviewer workload.

Summary

Lower the sensitivity first: the excess flags on affected groups sit in the weak signals, so raising the threshold reduces the disparity disproportionately.

In summary

Measure the disparity, then lower the sensitivity, remove signals that never convert, get context to the reviewer, and make the alternative real.

Four responses, in that order of cost. For wider institutional context, consult the Council of Europe.