Skip to content
examina.io

ExamsProctoringComputer-Based Testing

Your exam is already open book

Almost every student now uses AI on work that counts. The useful question is no longer how to keep it out, but what your assessment is actually measuring now that it is in the room.

An exam window open on top of a large open book, showing that the reference material is already there
The book was open before anyone decided to allow it

In 2024, two-thirds of UK undergraduates told researchers they used generative AI in their studies. By 2026, that figure is 95%, and 94% say they use it on work that counts towards a grade.

Those are not cheating statistics. They describe the room your exam is already sitting in.

Which means the interesting question stopped being "how do we keep AI out" some time ago. The question now is what your assessment actually measures once it is in there, and whether you would still be happy with the answer if a student told you exactly what they did.

The detector arms race is one you have already lost

Detection tools promise a percentage. What they deliver is a number with no defensible meaning attached, and a false-positive rate that lands hardest on students writing in a second language or in a plain, tidy style.

Consider the practical end of this. You accuse a student. They ask what the evidence is. You say a tool scored their essay at 78 per cent. They ask what 78 per cent means. You discover, in front of a disciplinary panel, that nobody in the room can say. Meanwhile, a student who genuinely likes the word "delve" is now defending their own vocabulary as though it were a criminal association.

Even where detection works today, it works against a specific generation of models. Every model released since has been trained by people who read the detection research too. This is a treadmill that bills you monthly.

What a monitored session actually proves

Proctoring is worth having, but it is worth being precise about what you get for it.

Three things a monitored session establishes, identity, presence and a reviewable record, set against the one thing it cannot establish, which is intent

A monitored session gives you a strong claim about identity, a strong claim about presence, and a record a human can review afterwards. Those are real, and they are exactly what a take-home essay cannot give you at any price.

What it does not give you is intent. A session can show that a second window opened. It cannot show whether the student opened it, whether a housemate walked past with a phone, or whether the tab had been sitting there since breakfast. Anyone selling you certainty about the last one is selling you something they do not have.

Assessment security tells you what happened in the room. It has never told you what happened in someone's head, and it was never supposed to.

That gap is not a flaw in the technology. It is the space where professional judgement lives, and the honest design goal is to make that judgement easy rather than to pretend it is unnecessary.

Write questions a model finds boring

The cheapest defence available to you costs nothing and ships this term: ask things a general model has no purchase on.

A model is excellent at the recallable and the generic. It is weak on the local, the recent and the personal. So:

  • Ask about data the cohort generated themselves in week four.
  • Ask students to critique a specific flawed answer you wrote, rather than produce a good one.
  • Ask them to apply the idea to something from their own placement, and to say what did not transfer.
  • Ask for the reasoning behind a choice, then ask what would have changed their mind.

Questions like these are not AI-proof, and anyone promising you AI-proof is guessing. They are AI-resistant, which is achievable, and they have a pleasant side effect: they are better questions. The generic essay prompt was never measuring much anyway. It just took a machine writing a competent one in nine seconds for everybody to notice.

There is a bonus. When a model does attempt these, it is confidently, fluently, gloriously wrong in ways that are very easy to spot, because it is filling gaps it does not know exist.

Two lanes beat a traffic light

Many institutions responded to all this with a traffic light scheme. Red means no AI, amber means some, green means go ahead. It looked tidy on a policy page and then met actual coursework.

Two assessment lanes side by side, one with assistance switched off and observed conditions, one with assistance allowed and the work assessed on process

The trouble is that amber has to be interpreted, and it gets interpreted differently by every module leader, every marker and every student, all of whom are confident they read it correctly. Several universities are now moving away from these schemes for 2026 and 2027 entry, having discovered that a policy needing its own explanatory policy is not a policy.

What replaces it is simpler. Some assessments are observed, with assistance switched off, because you genuinely need to know what a person can do unaided. Those get supervised properly. Everything else allows assistance openly and moves the marks onto process: the drafts, the prompts, the rejected approaches, the reasoning about which suggestion was wrong and why.

Process is much harder to fabricate than a polished output, and considerably more interesting to read. One university has rebuilt assessments specifically around evidence of thinking rather than evidence of typing, on exactly this reasoning.

What to actually do this term

You do not need to redesign a degree. Try four things.

  1. Pick your genuinely unaided assessments deliberately, and make them fewer, shorter and properly supervised.
  2. On everything else, say plainly what assistance is allowed, in the assessment brief rather than a policy document nobody opens.
  3. Ask for one process artefact alongside the output. A short reflection on what the student tried and discarded does more work than any detector.
  4. Write down what you would accept as evidence before you need it. Deciding that during a disciplinary hearing goes badly for everyone.

We have all done this before

In the 1970s, mathematics departments argued seriously about whether the pocket calculator would end the discipline. Students would never learn arithmetic. Standards would collapse. Something had to be done.

What happened instead is that the questions changed. Arithmetic stopped being the thing worth testing, because a four-pound device could do it perfectly, and the assessment shifted to whether a student could set up a problem, choose a method, and tell when an answer was nonsense. Mathematics education survived, and got slightly more honest about what it valued.

Your exam is already open book. It has been for a while. The useful move is to decide what you are measuring now that the book is open, and to write the question you would still want to ask if every student told you the truth about how they answered it.

Share LinkedIn X

See how examina.io runs proctored exams

Author once, deliver securely, and watch every session live — exam design, delivery, proctoring and reporting in one place.