Since late 2022, every educator setting a written assessment has had to reckon with a new reality: a free tool can produce a competent answer to most exam questions in seconds. The instinct is to reach for stricter monitoring — lockdown browsers, cameras, proctors. Monitoring has its place, and we have written about it elsewhere. But there is a quieter, more durable line of defence that gets less attention than it deserves: writing questions that a language model cannot easily answer in the first place.
This is not about outsmarting the technology in an arms race you will lose. Models improve every few months, and any question defined by "the current model can't do this" has a short shelf life. It is about something older and sturdier — the difference between assessing recall and assessing thinking. Questions that test whether a student can reproduce information were always the weakest form of assessment. AI has simply made their weakness impossible to ignore.
Why "just ban AI" is not a strategy
Prohibition has two problems. First, it is largely unenforceable for any work done outside a supervised room, and AI-detection tools are unreliable enough that acting on their output risks falsely accusing honest students — a serious harm in its own right. Second, and more importantly, a blanket ban trains students for a world that no longer exists. The graduates we are assessing will spend their working lives with these tools at their fingertips. An assessment strategy built entirely on pretending otherwise is preparing them for nothing.
The more defensible goal is to design assessment where using AI either doesn't help much, or where the AI-assisted version still requires — and therefore demonstrates — genuine understanding. That reframing turns out to be liberating, because the techniques that achieve it are simply the techniques of good assessment.
The core principle: move up the ladder of thinking
The revised version of Bloom's taxonomy (Anderson & Krathwohl, 2001) arranges cognitive tasks from remembering and understanding at the base, through applying and analysing, up to evaluating and creating. Generative AI is extraordinarily good at the bottom two rungs and progressively shakier as you climb. Every technique below is, at heart, a way of moving your questions up that ladder.
1. Anchor the question in specific, local context
A model trained on the public internet has never seen the dataset you handed out in week six, the guest lecture your class attended, or the particular case study you built the unit around. Tie the question to material that exists only inside your course.
Weak
Explain the causes of the 2008 financial crisis.
Stronger
Using the three primary-source memos we analysed in the week-seven seminar, explain how the incentives described in Memo B contributed to the risk build-up. Refer to specific passages.
2. Ask for the reasoning, not just the answer
Require students to show the path, not the destination. A model will happily produce a correct final answer; a well-designed prompt asks for the working, the assumptions, the discarded alternatives, and the justification for choices — the parts where understanding actually lives.
3. Demand application to a genuinely novel scenario
Take a principle from the course and drop it into a situation the student has never encountered. The novelty is the point: retrieval and paraphrase won't carry a student who doesn't understand the principle well enough to transfer it.
Weak
What is the difference between correlation and causation?
Stronger
Here is a finding from a study you have not seen [supplied]. Identify the strongest causal claim the authors make, explain one plausible confounder they did not control for, and describe an experiment that would test it.
4. Build multi-step problems where each part depends on the last
When a later part of a question uses the student's own earlier answer, a generic AI response cannot simply be pasted in — it has no access to what this student concluded in part (a). Chained questions reward coherent reasoning across a whole problem rather than a series of disconnected lookups.
5. Ask students to critique, not produce
One of the most AI-resistant question types flips the task: give students a flawed answer — one an AI might plausibly generate — and ask them to find and explain its errors. Evaluating a specious argument requires deeper understanding than producing a fluent one, and fluent-but-wrong is exactly where current models are most dangerous.
Example
The answer below was submitted for this problem and contains at least two significant errors. Identify them, explain why each is wrong, and provide the corrected reasoning.
6. Require personal or process-based evidence
For coursework, ask for artefacts of the student's own process: annotated drafts, a reflective log of how their thinking changed, screenshots of their working, a short recorded walkthrough of their reasoning. A polished final product is easy to fake; a documented trail of how it came to exist is not.
7. Use current, post-cutoff, or hyper-specific material
Questions built around events, data, or readings more recent than a model's training data force genuine engagement. This is a temporary advantage for any given model, but combined with the other techniques it raises the effort required.
The honest limits
None of this is a silver bullet, and it would be dishonest to pretend otherwise. A determined student with a capable model and enough time can still get help on applied, reasoned questions — feeding in the supplied materials, describing the scenario, and prompting for the analysis. What good question design does is change the economics. It removes the trivial wins, forces the student to understand the problem well enough to even prompt effectively, and — for anything truly high-stakes — makes a supervised or oral component a reasonable final check rather than the whole strategy.
This is the same layered logic that applies to exam security generally. Assessment design is the first and cheapest layer: it does the most work for the least intrusion. Proctoring is a later layer, reserved for the assessments where the stakes justify it. A programme that leans entirely on surveillance while still asking recall questions has its priorities inverted.
A short checklist
- Could a student answer this by pasting it into a chatbot? If yes, revise.
- Does it require course-specific material a model has never seen?
- Does it ask for reasoning, process, or evaluation rather than a retrievable fact?
- Does any later part depend on the student's own earlier answer?
- Is the scenario novel enough that understanding, not paraphrase, is required?
- Is the monitoring you're applying proportionate to what the question already protects?
References & Further Reading
- Anderson, L. W., & Krathwohl, D. R. (Eds.) (2001). A Taxonomy for Learning, Teaching, and Assessing: A Revision of Bloom's Taxonomy of Educational Objectives. Longman.
- Wiggins, G. (1998). Educative Assessment: Designing Assessments to Inform and Improve Student Performance. Jossey-Bass.
- QAA (Quality Assurance Agency for Higher Education) — guidance on academic integrity and assessment design: qaa.ac.uk.