The short answer
Understand the material first, decide what genuinely needs to be memorised, then have the AI write cards for only that — with an explicit maximum. Review every card for about ten seconds before saving. A one-hour lecture should produce 10–20 cards, not 80. Cutting half the generated output is normal and makes the deck better.
The failure this is designed to prevent is specific and very common: paste 40 pages of notes into a model, ask for flashcards, receive 300, save them all, and acquire roughly 3,000 reviews of obligation for material you mostly didn't need to memorise. That deck gets abandoned within a month, and the conclusion drawn is usually "AI flashcards don't work" rather than "I skipped the selection step".
What you give up by automating
Writing a card yourself is itself a learning act. You have to decide what the question is, what counts as the answer, and where the boundary of the idea sits — and producing information yourself creates a stronger memory than reading the same information, which is the generation effect. Automating card-writing gives some of that away. It's a real cost, and worth being honest about.
Slamecka and Graf demonstrated the generation effect in 1978 and it has held up well: material you generate is better remembered than material you read. So a guide claiming AI card generation is strictly superior is selling something.
Three things make the trade worthwhile anyway:
- The realistic alternative is no cards. Not hand-written cards — no cards. Card-writing is a chore that happens at the end of a study session when you're already tired, and it is the single most commonly skipped step in the whole method. A deck that exists beats a better deck that doesn't.
- The generation effect applies to the retrieval, too, and that part stays yours. You still have to produce the answer at every review, thousands of times, which is where most of the learning lives.
- You can keep the valuable half. The genuinely cognitive part of card-writing is deciding what matters. The mechanical part is phrasing it into a question and answer. Automate the second, keep the first, and you lose much less than you'd think.
That split is the principle behind this whole workflow, and it's the same rule set out in how to use AI to study: automate the logistics, never automate the retrieval.
Step 1: what deserves a card
Before generating anything, go through your notes and mark what's cardable. Three questions, and a point needs all three:
- Will I need to retrieve this without context? Facts you'll need to produce cold — a drug's mechanism, a date, a conjugation, a threshold value — are cardable. Things you'll always have in front of you are not.
- Would I be annoyed to have forgotten this in six months? If not, don't card it. Most of a lecture is scaffolding that helped you understand something; the scaffolding doesn't need to be retained.
- Is it something I can't re-derive? If it follows from a principle you already know, card the principle instead. Ten cards derived from one rule are nine cards of waste and one card of value.
What almost never deserves a card:
- Anything the lecturer said was non-examinable or "just for interest"
- Long lists, unless the list itself is the examinable object
- Definitions of terms you already use fluently
- Anything you can look up in two seconds and rarely need
- Worked examples — card the method or the step you keep getting wrong, not the example
- Anything you didn't understand (fix that first; see below)
If there's something in the notes you can't explain, resolve it before making cards. Flashcards retain understanding; they don't create it. A card made from a sentence you didn't follow produces a memorised string you can recite and can't use — and it will be one of the cards you fail repeatedly, because there's nothing underneath it to recall from.
The card budget
Decide the maximum before you generate, not after. Without a number, models will happily produce a card per sentence, and you will accept them because deleting feels wasteful.
| Source | Sensible budget | Review cost at ~10 reviews/card |
|---|---|---|
| One 50-minute lecture | 10–20 cards | 100–200 reviews |
| One textbook chapter | 15–30 cards | 150–300 reviews |
| A week of one course | 40–60 cards | 400–600 reviews |
| A dense vocabulary list | 1 card per item | Genuinely the exception |
| A full semester course | 300–600 cards | 3,000–6,000 reviews |
That third column is what makes the budget feel real. Each card costs roughly ten reviews over its life — the arithmetic is in how many cards to review per day. So 80 cards from one lecture isn't 80 units of work, it's about 800 retrievals, spread across a year, on material you sat through once.
Pricing cards in reviews rather than in cards changes the selection decision entirely, and it's the mental shift that makes people ruthless enough.
Step 2: the prompt
"Make flashcards from these notes" produces bad cards reliably. The prompt has to constrain: one fact per card, a hard maximum, short answers, no cards answerable from their own phrasing, and an instruction to skip anything not clearly stated in the source rather than filling gaps from general knowledge.
A prompt that works:
Turn the notes below into flashcards for spaced repetition.
Rules:
— Maximum 15 cards. Fewer is better. Choose only what genuinely needs memorising.
— One fact per card. Never combine two facts with "and".
— Answers under 15 words.
— The question must not contain its own answer, and must not be answerable by elimination.
— Ask for specifics, not for summaries. No "describe the process of…" cards.
— Use only what's in the notes. If something is unclear or absent, skip it — do not fill it in from general knowledge.
— Skip anything marked non-examinable, and skip long lists.
— After the cards, list anything in the notes that seemed important but was too vague to card.
Notes: [paste]
Three of those rules are doing most of the work:
- The maximum. Without it you get one card per sentence. With it, the model is forced to prioritise, and its prioritisation is decent.
- "Use only what's in the notes." This is the anti-hallucination clause, and it does something more valuable than reducing errors — it makes the errors checkable. If every card must trace to your notes, you can verify cards in a subject you don't yet know by checking against the source rather than against your own knowledge.
- The final list of unclear points. The most useful output on the page. It tells you exactly where your notes are thin, which is where to go back to the lecture or ask a follow-up question. Almost nobody asks for this.
Useful variations
- Two-way cards: "For vocabulary, generate both directions as separate cards." Worth it for production-heavy learning; skip for recognition-only material, since it doubles the review cost.
- Cloze from prose: "Where a fact only makes sense in context, use a cloze deletion with exactly one blank." Good for definitions and statute-like text; bad when the surrounding sentence gives the answer away.
- Difficulty control: "Prefer questions that require producing an answer over questions that require recognising one." Small change in wording, large change in output quality.
- Exam-shaped: "Phrase questions the way an examiner would ask them." Useful for professional exams with a house style.
What bad output looks like, and the fix
Four failure patterns account for most of what needs deleting. Each is recognisable in about two seconds once you know the shape.
The summary card
Describe the process of glycolysis.
Glucose is phosphorylated to glucose-6-phosphate, isomerised to fructose-6-phosphate, phosphorylated again… (10 steps, 90 words)
Which enzyme catalyses the rate-limiting step of glycolysis?
Phosphofructokinase-1.
The left card can't be graded — you'll recall six of ten steps and have no idea whether that's a pass. Anything starting "describe", "explain" or "list all" is an essay prompt wearing a flashcard's clothes.
The self-answering card
What is the maternally inherited organelle that produces ATP?
The mitochondrion.
Which organelle is inherited exclusively from the mother?
The mitochondrion.
Models love to over-specify questions. Every extra qualifier is a clue, and after twenty reviews you're retrieving from the phrasing rather than from the fact — which is the failure mode active recall exists to avoid.
The compound card
What are the causes and treatment of hyperkalaemia?
Renal failure, acidosis, tissue breakdown, potassium-sparing diuretics; treat with calcium gluconate, insulin with dextrose, salbutamol…
Which drug is given first in hyperkalaemia with ECG changes, and why?
Calcium gluconate — it stabilises the myocardium without lowering potassium.
"And" in a question is the reliable signal. The compound card fails as a unit, so you re-review the half you knew every time the other half slips — the fastest route to a leech.
The trivia card
In what year was the enzyme first isolated?
1926.
Nothing depends on this. The model carded it because it was a fact in the notes, not because it was worth remembering.
Models can't tell importance from mere presence. This is exactly the judgement the selection step reserves for you, and the one you shouldn't delegate.
Step 3: the ten-minute review
Ten seconds per card, seven checks. For a batch of 20 cards that's under four minutes, and it is the step that separates a deck you'll keep from one you'll abandon. Skipping it is the most consequential shortcut available in this workflow.
For each card, ask:
- Is it true? Check against your notes, not your memory.
- Is it one thing? If there's an "and", split or cut.
- Could I answer it in six months, cold? If it depends on remembering which lecture it came from, add the context to the question.
- Does the question give the answer away? Strip the extra qualifiers.
- Is the answer short enough to grade? Over about 20 words, you can't say cleanly whether you got it.
- Do I actually want to know this? The trivia filter. Be harsh.
- Do I already know it? Cards for things you know cold are pure cost. Delete them.
Deleting half the batch is a normal outcome and a good one. The instinct to keep cards because generating them was cheap is the same instinct that fills a deck with material you'll resent — and the cards were cheap to make, but they aren't cheap to own.
This is the part Memori compresses. Cards are generated from an explanation you were already reading, shown to you selectable and editable before anything is saved, so the review step happens at the moment of creation rather than as a separate chore you'll skip. Magic Import does the same for photos of notes, pasted text, PDFs and existing Anki decks, and Magic Reformat fixes a whole batch at once when the problem is systematic — every answer too long, say — rather than card by card.
Handwritten, slides, and recordings
Most real notes aren't clean text. The workflow is the same; the input step differs.
- Handwritten notes. Photograph a page at a time, in good light, straight on. Multi-page dumps degrade extraction badly. Check that any numbers, units and drug doses survived transcription — this is where OCR errors are both most likely and most dangerous.
- Lecture slides. Deceptively bad input: slides are compressed prose, so the model expands the bullets by guessing at what the lecturer said. If you use slides, tighten the "use only what's stated" rule and expect to delete more. Slides plus your own notes work far better than slides alone.
- Recordings and transcripts. The opposite problem — enormously verbose, mostly filler. Ask for a structured summary first, check it, then generate cards from the summary rather than from the transcript.
- Textbook PDFs. Work section by section with a per-section budget. Feeding a whole chapter in and accepting the output is how people end up with 200 cards from 30 pages.
- Diagrams and images. Text cards about a diagram are usually weaker than the diagram itself. If the spatial relationship is the point, use image occlusion rather than prose.
Step 4: getting them into rotation
You now have 15 good cards. Don't dump all 15 into today.
- Let your daily new-card limit meter them in. If your limit is 10, today's batch spreads over two days, which is what keeps your review load predictable.
- Don't binge before an exam. Adding 300 cards in the final fortnight gives each one two reviews at most — not enough to be reliable, and enough to wreck your daily count. Cards introduced under three weeks before an exam are mostly cramming with extra steps.
- Tag by source (lecture, week, chapter). Not for organisation's sake — so that when you fail a cluster of cards from one lecture, you know to go back and re-learn that topic rather than drilling the cards harder.
- Expect the first week to be rough. Fresh cards fail more. That's the schedule doing its job, not a sign the cards are bad.
A weekly rhythm that works
The whole workflow, as a habit rather than a project:
| When | What | Time |
|---|---|---|
| Same day as the lecture | Read the notes, resolve anything you can't explain, mark what's cardable | 15 min |
| Same day | Generate with a budget, review, save | 10 min |
| Every day | Reviews at your usual anchor time | 15–25 min |
| Weekly | Look at what's failing; rewrite the worst 3–5 cards | 10 min |
| Monthly | Prune cards you no longer need | 10 min |
Roughly half an hour on the day of a lecture, plus the daily review you'd be doing anyway. The version that fails is the one where card-making becomes a separate project undertaken three weeks before the exam on eleven lectures at once — which is precisely when the selection step gets skipped and the 400-card deck appears.
Frequently asked questions
How many flashcards should one lecture produce?
Ten to twenty for a typical hour. If you're generating eighty from one lecture you've transcribed it rather than studied it — and created roughly 800 reviews of obligation for material that mostly didn't need memorising.
Is it bad to let AI write your flashcards?
It costs you the generation effect — producing something yourself creates a stronger memory than reading it. The loss is real but small next to the benefit of having cards at all, since the realistic alternative is usually no cards. Keeping the selection and review steps recovers most of it.
What's the best prompt for generating flashcards?
One that constrains rather than requests: one fact per card, a hard maximum, answers under 15 words, no card answerable from its own phrasing, and skip anything not clearly stated in the source rather than inferring it.
Should I make cards during the lecture or afterwards?
Afterwards, ideally the same day. Making them live splits your attention between understanding and formatting, and you can't yet tell which points mattered. Within 24 hours is the sweet spot — recent enough to remember context, late enough to see structure.
How do I check cards for a subject I don't know yet?
Check them against your source rather than your knowledge. If a card states something your notes don't contain, delete it. That catches most errors without requiring expertise — and it only works because you restricted generation to the supplied material.
Can AI turn a PDF or textbook chapter into flashcards?
Yes, same workflow — but selection matters more, because a chapter contains far more cardable material than you need. Work section by section with an explicit budget rather than feeding in the whole thing.
What about cards for subjects that aren't fact-based?
Card the components rather than the whole. For maths, the definition, the condition a theorem requires, the step you always forget. For essay subjects, the specific evidence, dates and names you'd otherwise have to look up. The reasoning itself is trained by doing problems and writing essays, not by reviewing cards about them.
Should I keep the original notes after making cards?
Yes. Cards are retrieval prompts, not a record — they deliberately strip context, which makes them poor for relearning something you've lost entirely. When a card stops making sense, you'll want the source it came from.
Where this comes from
- Slamecka, N. J., & Graf, P. (1978). The generation effect: Delineation of a phenomenon. Journal of Experimental Psychology: Human Learning and Memory, 4(6), 592–604. The basis for the trade-off described in "what you give up by automating".
- Woźniak, P. Effective learning: Twenty rules of formulating knowledge. The minimum information principle underlying the one-fact-per-card constraint.
- Bjork, R. A., & Bjork, E. L. (1992). A new theory of disuse and an old theory of stimulus fluctuation. Desirable difficulties, and the rule about which difficulties are safe to automate away.
- Roediger, H. L., & Karpicke, J. D. (2006). Test-enhanced learning. Psychological Science, 17(3), 249–255. On why the retrieval step has to remain effortful.
- Dunlosky, J., et al. (2013). Improving students' learning with effective learning techniques. Psychological Science in the Public Interest, 14(1), 4–58. On practice testing versus summarising and highlighting.
- Nielsen, M. (2018). Augmenting Long-term Memory. On card design as an ongoing editorial practice rather than a one-time act.