When I was building Sahayak — an AI-powered study companion for Indian CBSE students built during ByteVerse 2026 — one of the core features was adaptive quiz generation. You pick a subject and topic, the app generates a five-question quiz using Gemini Flash-lite, you answer it, and the app grades and explains each response.
The problem showed up in testing almost immediately: take the same quiz three times in a row and you'll see the same photosynthesis light reaction question, the same Ohm's Law resistor problem, the same neutralization equation. Not identical phrasing, but the same underlying concept, the same difficulty, the same example every single time.
Why this happens
Foundation models are next-token predictors trained on text corpora. When you ask for a "Class 10 Biology question on Life Processes," the model generates the statistically most likely output given that prompt — which is whatever appeared most frequently in textbooks, study guides, and online resources in its training data. That's going to be:
- Heredity → Mendel's pea plant monohybrid cross, 3:1 phenotypic ratio
- Electricity → Ohm's Law V = IR, series resistor calculation
- Acids and Bases → HCl + NaOH neutralization reaction
- Life Processes → light-dependent reactions of photosynthesis
Every time. The model isn't being lazy — it's doing exactly what it's designed to do. Dominant examples in the training data produce dominant outputs.
Adjusting temperature helps marginally. Telling the model "don't repeat previous questions" in the system prompt does nothing, because it has no memory of what it generated in previous API calls.
The sub-topic cooldown system
The fix is to give the system memory that the model doesn't have. In src/db/schema.ts, every generated question gets logged with a granular academic sub-topic label — not just "Biology" but something like "Xylem Transpiration", "Mendelian Dihybrid Ratio", or "Refraction through Glass Prism":
export const askedQuestions = pgTable('asked_questions', {
id: uuid('id').defaultRandom().primaryKey(),
userId: text('user_id').notNull(),
subject: text('subject').notNull(),
subTopic: text('sub_topic').notNull(),
questionHash: text('question_hash').notNull(),
createdAt: timestamp('created_at').defaultNow().notNull(),
});
When a new quiz request arrives, src/app/api/quiz/generate/route.ts queries the user's recent history:
const recentAsked = await db
.select({ subTopic: askedQuestions.subTopic })
.from(askedQuestions)
.where(and(
eq(askedQuestions.userId, userId),
eq(askedQuestions.subject, primarySubject)
))
.orderBy(desc(askedQuestions.createdAt))
.limit(50);
const cooldownSubTopics = Array.from(new Set(recentAsked.map(q => q.subTopic)))
.filter(Boolean)
.slice(0, 15);
const cooldownInstruction = cooldownSubTopics.length > 0
? `\n\nSUB-TOPIC COOLDOWN (AVOID THESE SUB-TOPICS): ${cooldownSubTopics.join(', ')}. Focus on other areas of the curriculum.`
: '';
The last 50 questions are deduplicated by sub-topic, the 15 most recent unique topics are extracted, and they get appended to the system prompt as an explicit negative constraint: "avoid these, cover something else."
This forces the model to traverse corners of the syllabus it would never organically reach. Anaerobic fermentation instead of aerobic. Sex determination chromosomes in reptiles instead of humans. Resistivity temperature coefficients instead of Ohm's Law. The full CBSE curriculum has hundreds of testable sub-topics — the cooldown system makes the quiz generator actually use them.
The sub-topic label sourcing
The sub-topic labels themselves are generated by Gemini as part of the quiz JSON response. Each question object includes a subTopic field that the model fills in. This creates a slight chicken-and-egg dependency (the model that generates questions also labels them), but in practice the labels are consistent enough to be useful as cooldown keys. A question about \frac{1}{2}mv^2 labeled "Kinetic Energy" and another labeled "Work-Energy Theorem" are distinct enough that the cooldown system treats them separately.
The questionHash field is an MD5 or SHA-based hash of the question text — a backup deduplication layer that catches exact question repeats even if sub-topic labeling is inconsistent.
What it felt like after the fix
Before: five consecutive Biology quizzes, same five concepts, different numbers.
After: ten consecutive Biology quizzes covering photosynthesis, then transpiration, then enzyme action, then excretion in amoeba, then double fertilization, then sex-linked inheritance, then anaerobic respiration, then nephron structure, then stomatal regulation, then seed germination. Actual curriculum coverage.
The model wasn't holding back before — it just didn't know what it had already said. Give it a memory of what's been covered and it naturally explores outward.
Takeaway
Don't rely on stochastic model parameters (temperature, Top-P) to generate content variety in generative AI applications. They help slightly, but they can't overcome the gravitational pull of dominant training examples. Build a relational history ledger of past outputs, extract the recency-weighted distribution of covered concepts, and inject algorithmic negative constraints at runtime. Let the database remember what the model can't.