The one rule
Automate the logistics. Never automate the retrieval. Anything that is a chore standing between you and thinking — finding material, splitting notes into cards, reformatting, rephrasing, translating jargon — is fair game. The moment the AI produces an answer you were capable of producing yourself, you have spent the part of the session that was going to teach you something.
This sounds like a slogan. It's actually a fairly precise operational test, and you can apply it in real time. Before sending a message, ask: am I removing an obstacle to thinking, or am I removing the thinking?
"Split these three paragraphs of notes into atomic question-and-answer pairs" removes an obstacle. "What are the three causes of X?" — when you have just read the section on the three causes of X and could have recalled them with effort — removes the thinking. The messages look equally productive. Only one of them leaves you knowing more.
Which difficulties are worth keeping
The reason the rule works comes from Robert and Elizabeth Bjork's idea of desirable difficulties: some kinds of struggle during learning worsen your performance now and improve your retention later. Others are just friction.
Learners are consistently bad at telling the two apart, because both feel like difficulty. AI removes friction indiscriminately, which is exactly why it needs a rule rather than a vibe.
| Difficulty | Keep or remove? | Why |
|---|---|---|
| Retrieving an answer from memory | Keep | This is the learning event |
| Working out why two ideas differ | Keep | Builds the distinction you'll be tested on |
| Attempting a problem before seeing the method | Keep | Failed attempts prime you to encode the solution |
| Spacing and interleaving your practice | Keep | Feels worse, retains better |
| Finding a clear explanation of a concept | Remove | Pure search cost, no learning content |
| Splitting notes into individual facts | Remove | Mechanical and rule-driven |
| Rephrasing a card for consistency | Remove | Formatting, not thinking |
| Decoding unnecessarily dense academic prose | Remove | The difficulty is the writing, not the idea |
| Transcribing a photo of a whiteboard | Remove | Genuinely just typing |
Nearly every good use of AI in studying is in the bottom half of that table. Nearly every bad one is a case of quietly reaching into the top half.
What happens when students get unrestricted access
A 2024 field experiment with around a thousand Turkish high-school maths students provides the clearest evidence available. Students given unrestricted access to a general chatbot did substantially better on practice problems — and then performed roughly 17% worse than the control group on a later exam without it. A restricted version that gave hints instead of answers eliminated the harm.
This is the study to know, because it isolates exactly the mechanism this article is about.
Bastani and colleagues ran three conditions across a set of practice sessions: no AI, an unrestricted GPT-4 assistant, and a "tutor" version with pedagogical guardrails — instructed to give hints, withhold final answers, and prompt the student to attempt the step themselves.
- During practice, both AI conditions helped a lot. The unrestricted version improved practice performance markedly; the tutor version improved it even more.
- On the subsequent unassisted exam, the unrestricted group did meaningfully worse than students who'd had no AI at all. The tutor group performed about the same as the control.
Read those two bullets together and you have the entire problem in miniature. The unrestricted assistant produced better work and less learning. The students in that condition were not lazy or dishonest — they were using an available tool in the obvious way, and the obvious way happened to remove the retrieval.
The optimistic reading is equally important: the guardrails worked. The difference between the two AI conditions was not the model. It was the instruction to withhold answers. That is a design choice, and — crucially — it is one you can impose on yourself with a sentence of prompting.
One subject, one age group, one country, one model generation. The tutor condition prevented harm rather than exceeding the control, so it is not evidence that AI tutoring beats ordinary study. Research on AI and learning is a few years old and mostly short-term. Treat the direction as informative and any specific percentage as provisional.
The more encouraging side of the literature: a 2024 study in a Harvard physics course reported that students working with a purpose-built AI tutor learned more, in less time, than a matched group in an active-learning classroom. Note the qualifier — purpose-built, with deliberate pedagogical scaffolding. It is not evidence about chatting with a general-purpose assistant, which is what most students actually do.
Both results point the same way. What matters is not whether you use AI; it's whether the way you use it preserves the difficulty that does the teaching.
The fluency trap
AI explanations are unusually well-written: clear, structured, appropriately paced. That fluency produces a strong sense of understanding — and the feeling of understanding is a notoriously bad predictor of what you'll remember. A brilliant explanation you read once and never retrieve is close to worthless a week later.
Roediger and Karpicke showed in 2006 that students who reread material predicted they would outperform students who tested themselves, and were wrong — badly wrong, at a one-week delay. The mechanism is processing fluency: material that goes down easily feels learned.
An AI explanation is essentially fluency maximised. It is tailored to your question, written at your level, and free of the digressions that make textbooks hard. Everything that makes it a good explanation also makes it a more convincing illusion.
The fix is cheap and takes about ninety seconds:
- Read the explanation until it makes sense.
- Close the chat.
- Say it back, out loud, in your own words, without looking.
- Note the exact point where you go vague. That's the bit you didn't get.
- Reopen, ask about that specific point, and repeat.
Step 3 is non-negotiable and is the whole intervention. It converts a reading session into a retrieval session, which is the difference the research keeps finding. If you want the full argument for why, see active recall.
Is this cheating?
Two different questions get bundled here. Academic cheating is about submitted work, and your institution's policy is the authority — passing off generated text as your own is generally a violation. Cheating yourself is about whether you traded a skill for a shortcut. You can be entirely within the rules and still hollow out your own learning.
On the institutional question, three sensible defaults regardless of what the rules say: don't submit text you didn't write, disclose assistance where a policy asks for it, and assume anything an assessment measures is something someone wanted you to be able to do.
On the second question, the test is simple. Would you be able to do this in an exam room, alone? If the honest answer is no, and the AI is what closed the gap, you have a problem that will surface later — under worse conditions, with less time to fix it.
What is clearly fine: asking for explanations, asking for examples, asking it to check your reasoning after you've reasoned, asking for practice questions, having it do formatting and clerical work, asking it to tell you what you've misunderstood. None of that substitutes for your thinking; all of it removes obstacles to it.
Eight prompts that hold up
Adapt the wording; the structure is what matters.
1. Explain at three levels
"Explain [concept] three times: once for someone with no background, once at undergraduate level, and once at the level of a specialist. Make clear what the simpler versions leave out."
The third part is the valuable one. Simplified explanations are lies of omission, and knowing which lies were told is often the actual learning.
2. Make it ask you
"I'm studying [topic]. Ask me one question at a time. Don't tell me the answer — if I'm wrong, give me a hint and ask again. Only explain fully after I've attempted it twice."
This is the guardrail from the study above, imposed by hand. It is the single highest-value prompt on this list, and it works with any assistant.
3. Practice questions, answers withheld
"Generate 10 exam-style questions on [topic] at [level]. Put all answers at the end under a heading, and don't reference them in the questions."
Then answer all ten before scrolling. The failure mode this avoids is reading each answer immediately after each question, which is rereading with extra steps.
4. Find the hole in my explanation
"Here's my explanation of [concept], written from memory: […]. Don't rewrite it. Tell me what's wrong, what's missing, and what I've stated more confidently than the evidence supports."
The Feynman technique with a marker. Writing the explanation from memory is retrieval; the critique is feedback. Both halves are doing work, which is rare for a single prompt.
5. Give me the edge case
"Give me a case where [rule] appears to apply but doesn't, and explain what distinguishes it."
Exams live at boundaries. Most study material lives in the middle. This prompt manufactures the boundary cases you'd otherwise only meet in the exam.
6. Separate two things I keep confusing
"I keep mixing up [A] and [B]. Give me a single distinguishing feature for each that I could put on a flashcard, plus one example where the distinction matters."
Interference between similar items is one of the biggest causes of persistent forgetting, and the fix is a distinguishing cue rather than more repetitions. See how to make flashcards that actually work.
7. Turn this into atomic cards
"Turn these notes into flashcards. One fact per card. Each front must be a question with exactly one correct answer, scoped with the subject. Keep answers under 15 words. Flag anything you're not confident is correct."
Card generation is the clearest case of legitimate automation — it's mechanical, and the constraints that make cards good are exactly the sort of rules a model follows well. The final clause matters more than it looks.
8. Predict the exam
"Here's my syllabus and the topics covered: […]. What are the five most likely exam questions, and what does each one actually test?"
Useful for direction rather than truth. Treat the output as a prompt to think about coverage, not as a prediction.
Prompts 2, 4 and 7 work best when the explanation and the cards live in the same place, so a conversation you've just worked through can become scheduled review rather than a chat log you never open again. That's the loop Memori is built around: talk a concept through with the tutor, tap Save Memori on the explanation to turn it into draft cards, edit them, and let FSRS decide when they come back.
Three habits that replace learning
1. "Summarise this so I don't have to read it"
A summary is a compressed representation of someone else's understanding. Reading one gives you a map without the territory — you can now discuss the chapter and cannot use it.
The inversion works far better: read the material, write your own summary from memory, then ask the AI what you left out. Same tool, opposite direction, completely different outcome.
2. Generating cards you never read
Generate 200 cards from a PDF, import them, start reviewing. Two weeks later you're failing cards you don't recognise, on facts you never verified, phrased in someone else's words.
Generated cards are drafts. Read every one before it enters your deck. This takes a fraction of the time writing them would have, which is the actual saving — and it's where you catch the confidently-wrong card before you memorise it for a decade.
3. Asking rather than attempting
The habit that does the most damage, and the hardest to notice, because each individual instance is defensible. You're stuck for eight seconds, so you ask. Each time, you skip a retrieval attempt.
Impose a rule: attempt first, always. Write down your best answer, however poor, before asking. Kornell and colleagues found that failed retrieval attempts followed by feedback produced better learning than simply studying the answer — so your wrong guess isn't wasted time, it's the thing that makes the correction stick.
A protocol for confident wrong answers
Language models state incorrect things in the same confident register as correct ones, and there's no reliable tell. Calibrate verification to stakes: anything entering a flashcard deck deserves a check against your course material, because a memorised error is expensive to undo. Explanations used to build intuition are lower risk.
A workable triage:
| What you're getting | Risk | Verification |
|---|---|---|
| An explanation of a concept | Low–medium | Cross-check against your course material once; errors usually surface as inconsistency |
| An analogy or intuition pump | Low | None needed — it's a thinking aid, not a claim |
| Specific numbers, dates, dosages, cut-offs | High | Always verify against an authoritative source |
| Citations and references | High | Always verify the source exists and says what's claimed |
| Flashcards you'll memorise | High | Read every card; check anything you can't independently confirm |
| A worked solution to a problem | Medium | Follow the reasoning step by step rather than checking the final answer |
One further trick: your own knowledge is the cheapest detector you have. If an explanation contradicts something you already firmly believe, stop and resolve it. Roughly half the time you've found an error, and the other half you've found a gap in your understanding. Both are worth the two minutes.
An end-to-end workflow
Putting it together, for one topic:
- Read the primary source first. Lecture, chapter, paper. Don't outsource the first encounter — you need to know what the material actually says before you can tell whether an explanation of it is right.
- Close it and write what you remember. Five to ten minutes, blank page. This is diagnosis, and it is much faster than rereading.
- Take your gaps to the AI, one at a time. Use prompt 2 so it questions you rather than lecturing you. Use prompt 4 on the concepts you thought you understood.
- Say each explanation back before moving on. Ninety seconds. Non-negotiable.
- Generate cards from the gaps only. Not from the chapter — from what you failed to recall in step 2. Then read and edit every card.
- Schedule them and stop thinking about it. Spaced repetition handles the rest; see the complete guide.
- Do the problem sets yourself. Unassisted, closed-book. This is the step that tells you whether any of the above worked, and it's the one people skip.
Steps 1, 2, 4 and 7 are you doing the work. Steps 3, 5 and 6 are logistics. That ratio is roughly what a healthy AI-assisted study session looks like.
Frequently asked questions
Is using AI to study cheating?
For learning, no. For submitted work, follow your institution's policy — and don't submit text you didn't write. The more useful question is whether you're cheating yourself: if the AI is doing the thinking the assessment exists to measure, you've traded a skill for a grade.
Does AI actually make people learn less?
It can. In the 2024 Bastani study, students with unrestricted chatbot access did better on practice and roughly 17% worse on a later unassisted exam. A version constrained to hints rather than answers removed the effect. The tool isn't the variable; the interaction pattern is.
Can AI replace a tutor?
It replaces unlimited patient explanation on demand, which is a large part of what tutoring is and was previously very expensive. It doesn't replace a tutor's diagnostic judgement about your specific misconceptions, or the simple accountability of someone expecting work from you next week. Early research on purpose-built AI tutors is promising; it isn't settled.
How do I stop explanations from feeling like learning?
Close the chat and say it back from memory before moving on. Ninety seconds. Fluent explanations create a strong feeling of understanding that doesn't survive to tomorrow; the retrieval attempt is what converts one into the other.
Should I trust AI-generated flashcards?
As drafts, yes; as final cards, no. Generation handles the mechanical part well — atomising, phrasing, formatting — but it can't know what you already find easy, and a confidently wrong card is worse than no card because you'll drill it for years. Read every one before saving.
What about using AI for maths and problem-solving?
This is the highest-risk case, and the Bastani study was a maths study for a reason. Watching a correct solution appear is nearly worthless; it feels like understanding and transfers poorly. Attempt the problem, get stuck, ask for a hint at the specific step — never the full solution — and finish it yourself.
Where this comes from
- Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2024). Generative AI can harm learning. A field experiment with approximately one thousand high-school mathematics students comparing unrestricted and guardrailed AI assistance.
- Kestin, G., Miller, K., Klales, A., Milbourne, T., & Ponti, G. (2024). AI tutoring outperforms active learning. A controlled comparison in a Harvard undergraduate physics course.
- Bjork, R. A., & Bjork, E. L. (2011). Making things hard on yourself, but in a good way: Creating desirable difficulties to enhance learning. In Psychology and the Real World.
- Roediger, H. L., & Karpicke, J. D. (2006). Test-enhanced learning: Taking memory tests improves long-term retention. Psychological Science, 17(3), 249–255.
- Kornell, N., Hays, M. J., & Bjork, R. A. (2009). Unsuccessful retrieval attempts enhance subsequent learning. Journal of Experimental Psychology: Learning, Memory, and Cognition, 35(4), 989–998.
- Bloom, B. S. (1984). The 2 sigma problem. Educational Researcher, 13(6), 4–16. Frequently cited for the promise of one-to-one tutoring; the two-standard-deviation figure has not replicated robustly, and later estimates of tutoring effects are smaller.