It's 3 AM during ByteVerse 2026. My teammate is asleep on the bean bags behind me. The quiz generation endpoint I've been building for the past six hours — Sahayak, an AI study companion for CBSE students — is nearly complete. I fire off a test query: "Generate a 5-question quiz on Class 10 Algebra."
Gemini responds. The stream closes. My code calls JSON.parse().
SyntaxError: Unexpected token f in JSON at position 142.
I stare at it. Position 142. I count the characters manually in the raw response. There it is, sitting in the middle of a beautifully written math question: $\frac{x}{2} = 4$. A LaTeX fraction. And \f is, apparently, a JSON form feed escape character that nobody ever uses in real life — but JSON.parse() absolutely does not accept when it appears raw inside a string.
Welcome to the great backslash disaster.
What's actually happening
LaTeX uses backslashes for everything. \frac, \sqrt, \theta, \Delta, \times. They're how you write math. JSON also uses backslashes — but exclusively as escape delimiters. \" for embedded quotes, \\ for a literal backslash, \n for newlines. Every other \x combination is a syntax error.
When Gemini generates a JSON payload containing LaTeX math, it's doing something the JSON RFC never anticipated: embedding a language that structurally conflicts with JSON's own escape grammar. The model writes {"question": "Solve $\frac{x}{2} = 4$ for x."} and the parser reads \f as an attempted form feed escape and immediately throws.
Why "just tell the model to escape it" doesn't work
My first instinct was to fix it in the prompt. I added this to the system prompt, in all caps:
CRITICAL: You MUST double-escape ALL LaTeX backslashes as \\\\ in your JSON output.
That reduced crashes from 100% of math questions to about 20%. Still completely unacceptable in a live product. Foundation models don't understand JSON RFC. They're predicting plausible character sequences based on training data, and sometimes the training data had single backslashes, sometimes double. There's no way to make this reliable through prompt pressure alone.
The model also selectively obeys it. It'll get \frac right and then casually write a raw \sqrt in the explanation field three keys later because that's what felt natural to generate in that context.
Why naive regex makes it worse
My next attempt: globally replace \ with \\ before calling JSON.parse().
const fixed = raw.replace(/\\/g, '\\\\'); JSON.parse(fixed);
This spectacularly broke everything. Real JSON escape sequences like \" (embedded quote inside a string) became \\", which the parser reads as an escaped backslash followed by a bare double-quote — which immediately terminates the string, corrupting the entire document structure. Newlines \n became \\n, turning them from actual newline characters into the literal two-character sequence backslash-n.
Global replacement of escape characters in a JSON document is a document-destruction tool disguised as a fix.
The actual solution: a lookahead character parser
What I needed was something that could distinguish between "this backslash is a valid JSON escape I should leave alone" and "this backslash is a rogue LaTeX prefix I need to double up."
Valid JSON escape continuations: ", \, /, b, f, n, r, t, u. Anything else is structurally illegal and needs to be escaped.
I wrote a streaming lookahead tokenizer in src/app/api/quiz/generate/route.ts:
function sanitizeJsonString(rawJson: string): string {
let result = '';
for (let i = 0; i < rawJson.length; i++) {
const char = rawJson[i];
if (char === '\\') {
const nextChar = rawJson[i + 1];
if (nextChar === undefined) {
result += '\\\\';
continue;
}
if (nextChar === '"' || nextChar === 'n' || nextChar === '\\') {
result += '\\';
if (nextChar === '\\') {
result += '\\';
i++;
}
} else {
result += '\\\\';
}
} else {
result += char;
}
}
return result;
}
Before calling JSON.parse(), the raw model output goes through sanitizeJsonString(). The function inspects every backslash and checks what follows it. Structural JSON escapes (\", \n, \\) are preserved untouched. Everything else — \frac, \sqrt, \theta, \Delta, \times — gets its backslash doubled so JSON sees a properly escaped literal backslash instead of a malformed escape sequence.
The result is clean. JSON.parse(sanitizeJsonString(cleaned)) stopped crashing. Not "crashed less" — stopped crashing.
The commit message says it all
The commit that landed this fix was 11e09b4, timestamped 06:17 AM +0530. The message: "jaishriallahchrist".
Make of that what you will.
Takeaway
Never assume an LLM will correctly honor JSON escape grammar, no matter how much you yell at it in the system prompt. Tokenizers work in semantic chunks, not RFC specifications. When you're ingesting structured payloads that contain rich syntax (LaTeX, Markdown, regex, code) from AI models, build a defensive micro-tokenizer at your network boundary. The grammar mismatch is fundamental — it won't be fixed by a better prompt.