Skip to content
Flag Is Not Finding

All notes  /  In practice

Choosing a Supplier

The questions that separate suppliers, which are mostly about the human process and the disaggregated figures rather than about detection features.

In practice · Reference

Procurement in this market compares feature lists. The differences that matter are elsewhere. A comparison of workforce monitoring software pricing belongs beside licence quotes because review effort, support, data export and exceptions often cost more than the headline plan.

The fairness questions

What are your false positive rates, disaggregated by skin tone, and against which benchmark?

What is the failure-to-enrol rate for identity verification, by group?

Has any independent party evaluated the system?

Can specific signals be disabled per candidate?

A supplier who has not measured these has not tested for them, and the answer to the first question sorts the market faster than anything else.

The process questions

Can a session be released in full to the institution, and to the student?

Can a reviewer see context around a flag, and the whole session?

Can access arrangements be attached to a session before review?

And can flag thresholds be configured by assessment rather than globally?

The data questions

Where is data processed and stored, and which subprocessors are involved?

Who at the supplier can view a session, and is that access logged and visible to us?

What is deleted, when, and how is deletion confirmed including backups?

What happens to everything at the end of the contract?

The accessibility questions

Which reading software and assistive technologies has this been tested with?

Does the identity flow ever request removal of head coverings, and can that be disabled?

Is there a conformance statement, and who wrote it?

The claims to interrogate

Any accuracy percentage: ask for the denominator and the base rate.

"AI-powered detection": ask what it detects and what it reports.

"Reduces cheating by N percent": ask for the study.

And anything describing flags as detected misconduct, which is a wording problem that will propagate into your student communications if you let it.

The pilot requirement

Do not buy on a demonstration.

A pilot on a real assessment with a real cohort produces the flag rate, the review load and the support volume — three numbers that determine whether the programme is viable and which no demonstration provides.

Its own note covers how to run one.

Contract terms worth insisting on

Release of sessions to students.

Configurable thresholds and per-candidate signal control.

Deletion confirmation covering backups.

Disaggregated performance reporting, annually.

And exit terms: export or deletion, in a defined format, at a defined time.

What to check

Does your specification mention the review process at all?

Have you asked for disaggregated performance figures?

Is session release to students in the contract?

And did anybody pilot before signing, or did the pilot follow the purchase?

The point

Ask for false positive rates disaggregated by skin tone, and for the evidence behind any accuracy claim.

A supplier who has not measured this has not tested for it, and the answer sorts the market quickly.

Worth stating

Do not buy on a demonstration.

A pilot on a real assessment produces the flag rate, the review load and the support volume — three numbers that decide viability and that no demonstration provides.

Also worth knowing

Insist on contract terms for session release to students, configurable thresholds, deletion confirmation covering backups, annual disaggregated reporting, and defined exit terms.

Those five are worth more than any feature comparison.

And finally

Interrogate any accuracy percentage for its denominator and base rate, and any claim to reduce cheating for the study behind it.

Anything describing flags as detected misconduct is a wording problem that will propagate into your own communications.

Summary

Pilot before signing rather than after. A demonstration establishes that the software runs; only a pilot produces the flag rate and the review load.

In summary

The differences between suppliers that matter are the disaggregated figures and the human process, not the feature list.

Two questions — false positive rates by group, and session release to students — sort the market quickly. For wider institutional context, consult the U.S. Department of Education.