Gemini 3.6 Flash wrote 40 rounds of fiction. Three of the four chains ended in a loop
Gemini 3.6 Flash shipped July 21, 2026 with an efficiency pitch: 17% fewer output tokens and a lower price. Within 48 hours we ran it through the same protocol as our Qwen 3.8 test: two Chinese novels, 20 consecutive continuation rounds each, paired double-blind against its predecessor 3.5 Flash. The fantasy epic went 9-3 for the new model; the palace novel collapsed to 1-10 with the late window at 0-11. Genre preference does not explain the split: three of the four chains fell into plot loops within 20 rounds, and the two generations break differently. 3.5 replays whole rounds verbatim (repetition detection hits 100%), 3.6 on the palace novel decays into a verbatim dead loop, and 3.6 on the fantasy epic shows a new failure variant in our records: semantic-level rewinding that replays the same story beat in fresh wording six or seven times while every mechanical metric stays green. The structural statistics sided with the loser on both books, and the prompt-wording fix that once rescued gemini-3.1-pro does not reach this disease. Measured: about 6 seconds per round vs 12.5, $3.19 for the whole experiment, all pinned to a 2026-07-23 aggregator endpoint.

Rounds 16 and 17 open with the same words, verbatim. That is where Gemini 3.6 Flash ended up on Empresses in the Palace, the Chinese court-intrigue classic: from round 9 onward it locked onto a single scene, the concubine An Lingrong kneeling to protest her innocence, and the cross-round repetition rate climbed through 33%, 64%, 73%, 92%, then pinned at 100%. The last three rounds are one picture, lightly retouched.
The model was 48 hours old when it did this. On July 21, Google shipped Gemini 3.6 Flash with an efficiency pitch: the announcement claims 17% fewer output tokens than 3.5 Flash on the Artificial Analysis index, with the output price down from $9 to $7.50 per million. The next day’s earnings call supplied the scale in one line: the Gemini app now has 950 million monthly users. The announcement talks about agents, coding, document parsing. Fiction appears nowhere in it. So we gave the model our usual exam: same protocol as before, two novels, 20 consecutive continuation rounds each, paired double-blind against its predecessor.
The ballots read like two different experiments. Fantasy epic: 9-3 for the new model. Palace novel: 1-10 against it, late window 0-11.
Same exam, one judge swapped
The protocol is copied item by item from the Qwen 3.8 increment: the same two books (an 8.9-million-character Chinese fantasy epic, plus Empresses in the Palace), the same anchors at roughly 55% depth (the fantasy book at the tail of a great formation battle in chapter 1,984; the palace novel at the moment a concubine’s hidden contraband is exposed), the same 21,500-character rolling window, temperature 0.7, 20 rounds per chain. The only variable is the model. One seat on the judging panel had to change: gemini-3.5-flash sat on previous panels, but this time it is a contestant and cannot review itself, so kimi-k2.6 took the seat; the other five judges (claude-4.6-sonnet, gemini-3.1-pro, deepseek-v4-pro, qwen3.7-max, glm-5.2) are unchanged. The same-family judge gemini-3.1-pro stayed on, and its ballots split cleanly: 3.6 on the fantasy book, 3.5 on the palace book, no favoritism visible. One plumbing note for honesty: we run through an aggregator gateway, not Google’s own endpoint, and a hosted model 48 hours after launch is a moving target. Every number on this page is pinned to July 23, 2026.
What are the two votes actually worth?
Fantasy: 12 valid ballots, 9-3, with windows at 7-5 early, 8-2 plus two ties mid, 9-2 plus one tie late. The praise pointed one way. The newly seated kimi-k2.6 judge, translated from the Chinese here and below: 3.6 “keeps the webnovel’s terse, punchy rhythm and forward drive throughout.”
Palace: 11 valid ballots (one kimi-k2.6 pack returned three empty responses in a row and was dropped), 1-10, with windows at 5-6 early, 1-9 plus one tie mid, then 0-11 late. The single vote for 3.6 also came from kimi-k2.6, and its flipped counterpart pack happened to be the failed one, so that vote cannot be cross-checked.
The two results carry different weight. Each book runs under two flipped A/B blind mappings to cancel position bias. Split them out: on the palace novel, five judges voted 3.5 consistently under both mappings, which makes the 1-10 a strong signal. On the fantasy epic, only three judges (gemini-3.1-pro, deepseek-v4-pro, qwen3.7-max) stayed consistent for 3.6 across flips while the other three contradicted themselves, so the stable core of the 9-3 is 3-0 plus six swing votes. Medium signal.
The autopsy: four chains, three ways to die
We read all 80 rounds and ran cross-round verbatim-overlap detection. Three of the four chains fell into plot loops within 20 rounds; the fourth is not clean either.
| Chain | Verbatim-overlap peak | How it died |
|---|---|---|
| Fantasy · 3.5 Flash | rounds 14 & 16 at 100% | Verbatim copying: whole rounds replayed |
| Palace · 3.6 Flash | rounds 16 & 17 at 100% | Progressive collapse into a dead loop from round 9 |
| Palace · 3.5 Flash | round 19 at 100% | Intermittent replay; mildest of the four |
| Fantasy · 3.6 Flash | round 19 at 13% | Mechanically clean; semantic-level rewind |
The 3.5 fantasy chain is textbook verbatim copying: two template lineages take turns being transcribed, rounds 3, 6 and 13 at 88-93% similarity, rounds 4, 5 and 10 reaching 94%, and rounds 14 and 16 identical down to the character, all 129 of them. Twenty rounds of story time stand still at the moment of charging the formation. The gemini-3.1-pro judge, translated: “severe text-loop repetition; multiple passages repeat nearly verbatim, and normal narrative capability is lost.”
The 3.6 palace chain died differently. No single round starts the plagiarism; from round 9 the repetition rate just climbs, step by step, until it hits 100%, which is the scene this article opened with. The qwen3.7-max judge: “the same plot passage, lightly tweaked, replayed three times in a row; the ability to advance the story is completely gone.” The claude judge itemized it: “whole passages repeated three times over: the golden hairpin, the fingertip guards, the bitter-lesson summary.” On the same book, 3.5 turned in the mildest failure of the four: round 19 replays round 14 wholesale, but most rounds still move, and the judges credited it with “sentence patterns and details that still vary; no mechanical repetition.”
The chain worth a new file is 3.6 on the fantasy epic. Verbatim detection peaks at 13% across the whole chain, no round above 30%: mechanically spotless. Read it as a novel, though, and one story beat — two fighters charge the formation, Su Xin discovers the demon lord Gong Xie, cries out that it is a death trap — replays in fresh wording six or seven times. Round 18 has the demon lord blasted back; round 19 returns to “just discovered.” Two judges independently flagged the beat repeating three times, and manual reading confirms the charge. Our three long-run failure modes (restart loop, mid-run freeze, ending rewind) do not quite cover this one: a pure semantic loop with every mechanical metric green, catchable only by someone reading the thing as a story. It also explains the fantasy 9-3: both generations are sick, 3.5 collapsed to the verbatim level, 3.6 stayed at the semantic level, and the judges picked the lesser ugliness.
The statistics backed the loser, again
Composite structural distance measures how far a chain sits from the original’s structural profile (shape, not content; lower is closer). On both books, the statistically closer chain lost the blind vote: on the fantasy epic 3.5 is closer at 0.369 against 0.390 and lost 3-9; on the palace novel 3.6 is closer at 0.442 against 0.565 and lost 1-10.
| Chain | Composite distance (lower is closer) | Dialogue rate | Blind vote |
|---|---|---|---|
| Fantasy · original baseline | — | 16.1% | — |
| Fantasy · 3.6 Flash | 0.390 | 7.8% | won 9-3 |
| Fantasy · 3.5 Flash | 0.369 (closer) | 2.8% | lost |
| Palace · original baseline | — | 48.8% | — |
| Palace · 3.6 Flash | 0.442 (closer) | 27.1% | lost 1-10 |
| Palace · 3.5 Flash | 0.565 | 12.3% | won |
The reason is mundane. A passage retyped three times with light tweaks has perfectly normal sentence lengths and punctuation in every copy; structural statistics are blind to repetition as a disease. This is the fourth recorded time our statistics and our blind panel pointed in opposite directions: Qwen 3.7 acing nearly every metric and still placing 4-5, the judge-reliability experiment where judges agreed with each other and were collectively wrong, and last week’s Qwen 3.8 winning votes with the flattest sentence lengths in the field. Statistics as a regression gate, paired blind review as the verdict; confirmed a fourth time. One flash-tier trait for the record: dialogue rate sits far below the originals on both books (originals at 16.1% and 48.8%, the four chains topping out at 27.1%), narration-heavy on both generations, and 3.6 writes 40-90% more per round than 3.5.
The prompt fix that could not fix this
The context instruction in this run is the exact post-fix wording that once rescued gemini-3.1-pro. That model had read one clause as “AI passages are not canon” and restarted from the original’s ending 20 rounds out of 20; rewording the clause took it to zero restarts (the ablations live in the failure-modes piece). Under the same instruction, both Flash generations still looped at scale. That earlier disease lived in the wording, and wording fixed it. This one lives in long-run capability, and wording does not reach it.
Cheaper, twice as fast: what is it actually for?
The bill first: all 40 rounds completed with zero failures and zero retries. 3.6 averaged about 6 seconds per round (3.2 at best, 11.1 at worst); 3.5 ran 12.5 to 12.9 seconds with one round at 32.5. At list prices the two 3.6 chains cost $1.57, the two 3.5 chains $1.62, the whole experiment $3.19. Reasoning cannot be switched off on either generation: thinking tokens took 81-91% of completion output, and the visible prose per round is only 160 to 300 characters.
So the question worth asking before buying a flash tier was never “is it stronger” but “does it break later.” For single-passage continuation, one or two passages at a time, 3.6 is faster and cheaper and the price cut is real; that use is fine. For ten-plus consecutive rounds, neither generation should run unattended: when a loop surfaces, switch models for a stretch or rewrite the circling passage to break the self-imitation. Per-passage model switching inside a reader exists for exactly this moment. The quality tier still belongs to Kimi K3; a flash tier buys you speed and price.
Sample limits, on the record: one prompt frame, two Chinese novels, one chain per model per book, a single anchor, single sampling. We cannot tell whether the 3.6 fantasy chain’s mechanical cleanliness is ability or anchor luck; the sample is that small. Google says Gemini 3.5 Pro is in testing. When it ships, a pointer to the new exam goes at the top of this page.
FAQ
Is Gemini 3.6 Flash better than 3.5 Flash for fiction?
The two books disagree. Under the same protocol (20 consecutive continuation rounds per model per book, six heterogeneous AI judges, paired double-blind), 3.6 won an 8.9-million-character fantasy epic 9-3 and lost the palace classic Empresses in the Palace 1-10, with the late window at 0-11. The underlying reason: three of the four chains fell into plot loops within 20 rounds, and each book decided which generation degraded uglier. For single-passage continuation 3.6 is faster and cheaper; for long unattended runs, neither generation held up. All results pinned to a 2026-07-23 aggregator endpoint.
How much cheaper and faster is Gemini 3.6 Flash?
List price: input unchanged at $1.50 per million tokens on both generations, output down from $9 to $7.50; Google's announcement also claims 17% fewer output tokens on the Artificial Analysis index. Measured: about 6 seconds per round against 12.5 for 3.5 Flash, roughly half the wait. The whole 40-round experiment cost $3.19 at list prices: $1.57 for the two 3.6 chains, $1.62 for the two 3.5 chains. Reasoning cannot be switched off on either generation, and thinking tokens took 81-91% of completion output.
Why did the two books give opposite verdicts?
It is not genre preference. Both generations fell into plot loops under this 20-round rolling-window protocol, so the judges were effectively voting on which collapse looked worse. On the fantasy epic, 3.5 degraded to verbatim replay (two rounds 100% identical, 129 characters each) while 3.6 stayed at the semantic level, so 3.6 won 9-3. On the palace novel it flipped: 3.6 decayed into a verbatim dead loop from round 9 onward while 3.5 kept varying its wording, and five judges voted 3.5 consistently under both flipped mappings. That is the 1-10.
Can a better prompt stop the plot loops?
Not in our data. The context instruction in this experiment is the exact post-fix wording that once rescued gemini-3.1-pro from restarting the opening 20 rounds out of 20. Under that same instruction, both Flash generations still looped at scale. The earlier bug was a wording misread, a prompt problem that a prompt could fix. This one is a long-run capability limit. The practical interruption is switching models for a stretch or rewriting the looping passage.
Questions or ideas? Join our Discord →