BACK TO PROJECTS
2026-07-03

Sahayak (Byteverse AI)

AI study companion for Grades 5–12 matching Indian curricula (CBSE/ICSE), featuring LaTeX math parsing, curriculum-aligned textbook prompting, Neon PostgreSQL schemas, and Gemini exponential retry backoff.

Next.jsReactTypeScriptGoogle GeminiClerkDrizzle ORMNeon PostgresKaTeXTailwind CSS
[ Source Code ]✓ VERIFIED STABLE
[ Target Hardware ]Microcontroller / Edge Node

I built Sahayak during the 48-hour Byteverse Hackathon in July 2026 because every single "AI tutor" on the internet felt like a lazy ChatGPT wrapper that completely falls apart the moment an Indian high school student asks a real physics derivation or calculus question. Most models either choke on mathematical notation, spit out broken markdown tables, or speak in this generic American college tone that has zero relevance to what CBSE or ICSE actually tests in board exams.

The stack was Next.js 16.2.9 App Router, React 19.2.4, Tailwind CSS v4, Clerk for auth, and Drizzle ORM over Neon serverless PostgreSQL. For the AI backbone, I used Google's Gemini 2.5 Flash-lite (gemini-3.1-flash-lite via @google/generative-ai). But connecting to Gemini at 3:00 AM during a hackathon is an exercise in pure pain. The API was constantly throwing HTTP 503 "Service Unavailable" and HTTP 429 rate limit spikes. If you don't handle that gracefully, the student's entire active chat session dies. I wrote an exponential backoff wrapper in src/lib/gemini.ts (generateContentWithRetry()) that catches 503 and 429 exceptions, pauses execution for a strict 3000ms delay, and retries the prompt before ever letting an error surface to the client.

The biggest technical disaster of the entire build was what I now call the "Great Backslash Disaster." In our quiz generator (/api/quiz/generate), I instructed Gemini to output rigorous multi-choice and numerical questions with formulas in LaTeX (like \frac{a}{b} or \sqrt{x}) inside a strict JSON payload. But Gemini emitted raw, single backslashes. In standard JSON, \f is a form-feed character and \s is an invalid escape sequence. The result? JSON.parse() crashed violently every single time. Simple regex find-and-replaces were even worse because they blindly doubled valid escapes like \n and \", corrupting legitimate strings. At 4:00 AM, running on zero sleep, I ended up writing a custom character-by-character lookahead tokenizer (sanitizeJsonString()) that inspects the character immediately following every backslash—if it isn't a valid JSON control sequence, it safely escapes it to \\.

Beyond raw chat, Sahayak has a full relational pedagogical engine mapped across 14 tables in src/db/schema.ts. When a student asks for help, buildSystemPrompt.ts inspects their board, class, and enrolled subjects, then pulls their syllabus state from studentSyllabus (Completed, In Progress, Not Started) so the AI never assumes mastery over chapters the student hasn't touched yet. It even injects specific board pedagogy: CBSE Class 9 gets Ganita Manjari worked derivations, Class 10–12 gets strict NCERT step-by-step layouts, and ICSE gets formal Selina / ML Aggarwal proofs. To prevent the model from repeating identical questions during test generation, the backend runs a rolling 50-item query on askedQuestions, extracts the top 15 recent subTopic tags, and injects them as negative prompt constraints.

The git log from that sprint tells the whole story: cd94069 (initial scaffold on July 1), efb1973 ("middleware to proxy and so much more" after fighting Clerk edge middleware routing conflicts), cc3feaf ("ulta latka diya lol"), and commit 11e09b4 at 06:17 AM on July 3 titled "jaishriallahchrist" where I finally stitched together the curriculum picker, multi-language support (18 Indian languages), and auto-grading logic across 12 files just hours before our final deployment (60dd592). It was chaotic, exhausting, and easily one of the most rewarding full-stack systems I've ever built.