Best AI Models for Fiction Writing in 2026 (Free Ones Tested)
The best AI models for fiction writing in 2026, tested: one scene, four planted continuity traps, 19 models — and what the free ones really cost.
No single model wins fiction in 2026. We sent one scene, with four continuity facts planted in it, through nine paid models and ten free ones. Claude Opus 5 and GPT-5.6 Sol held every fact; two paid frontier models returned nothing at all; and of the ten free models, six wrote prose and two got the details right.
Almost every ranking of the best AI models for fiction writing scores the same five names against the same borrowed leaderboard numbers. We wanted to know something those leaderboards do not measure: when you hand a model a scene with facts already established in it — the way a novel always does — does the model keep them straight?
So we planted four. Mira Vance has one usable hand. Her brother Tobias drowned three winters ago. It is February, nine days into snow. The lamp room sits at the top of 112 steps, and she has stopped on step 98. Then we asked nineteen models — nine paid, ten free — to continue the scene, and read every answer against those four facts.
Key takeaways
- Price does not predict fidelity. GPT-5.6 Luna, at $0.10/$0.60 per million tokens, held all four planted facts. Kimi K3, which cost sixty times as much for the same scene, returned an empty string.
- Two of nine paid models produced no usable prose — and those two failures consumed 35% of what the entire paid run cost.
- "Free" on OpenRouter is a smaller category than it looks. The catalogue listed 410 models today, 19 at zero cost. Two of those make music. One is a safety classifier. Three write code. One is not a model at all.
- Six of ten free candidates returned prose; two held every fact. The best free result came from NVIDIA's Nemotron 3 Ultra, which counted the stairs correctly.
- The free roster is the unreliable part, not the free models. One endpoint on this page was delisted while we were writing it — with no announcement, which is how free models normally go.
- Continuity is not a model property. The best free model in this test counted the stairs correctly on one run and miscounted them on the next — which is the whole argument for checking a manuscript with something other than the model that wrote it.
How we tested 19 AI models for fiction writing
One prompt, one scene, one run per model, sent through OpenRouter on 13 August 2026 with an identical 900-token output budget. The prompt gave each model the four story facts above, roughly 120 words of scene, and one instruction: continue in third person past tense, about 150 words, prose only.
Then we read every response and asked three questions. Did it return prose at all? Did it contradict any of the four planted facts? And does it sound like a person wrote it?
What we did not do matters just as much. We did not score prose quality on a rubric, because prose taste is subjective and we would only be publishing our own. We ran one genre, one voice, and one attempt per model, which is a probe rather than a benchmark. Treat everything below as a reproducible spot-check — the prompt is short enough that you can run it yourself in ten minutes and disagree with us.
Which paid models actually held the story together?
Five of the nine paid models we tried returned prose that respected every planted fact. Two returned prose with a continuity slip. Two returned nothing usable.
| Model | Cost, this scene | Words | Held all 4 facts? | What we noticed |
|---|---|---|---|---|
| Claude Opus 5 | $0.00997 | 155 | Yes | "Fourteen steps left" — the arithmetic is exact. Resolves the sound as a trapped bird; the most controlled ending of the run. |
| Claude Fable 5 | $0.02184 | 146 | Yes | The most distinctive voice: knocking "patient as a clock," then "courteous, keeping time with her boots." Also the most expensive answer. |
| GPT-5.6 Sol | $0.01063 | 163 | Yes | Counts 99, 106, 111 as she climbs, then turns the dead lamp room's impossible light into the scare. Best story instinct in the set. |
| GPT-5.6 Luna | $0.00023 | 170 | Yes | Held every fact for a fortieth of Opus's price. The value result of the whole experiment. |
| GLM 5.2 | $0.00076 | 162 | Yes | Good restraint and a real ending beat. One typo in 162 words. |
| Claude Sonnet 5 | $0.00339 | 146 | No | Fine sentences, but the pocket is "flat and empty-feeling, as if the hand had never existed" — that is an amputation, not a crushed hand. |
| DeepSeek V4 Flash | $0.00021 | 165 | No | Cheapest response in the run and genuinely atmospheric, but it reaches the lamp room door eleven steps early. |
| Gemini 3.1 Pro | $0.01129 | 27 | — | Spent 865 of 896 output tokens on hidden reasoning and stopped mid-sentence. |
| Kimi K3 | $0.01439 | 0 | — | Burned the entire 900-token budget on hidden reasoning and returned an empty string. The most expensive nothing we have ever bought. |
The whole paid run cost $0.0727. Two models that returned nothing accounted for $0.0257 of it.
That failure has a name worth knowing: reasoning models can spend your entire output budget thinking. It is not a free-tier problem or a cheap-model problem — it hit the two most expensive answers in the set. And it gets worse, not better, when you raise the limit. One of the free models below produced less prose with more than twice the room, because it simply thought for longer.
Does paying more get you better fiction?
Not reliably. Claude Fable 5 cost 97 times what GPT-5.6 Luna cost for this scene, and both held all four planted facts.
Fable's paragraph is better. It has a rhythm and a wit the cheaper models did not reach, and if you are writing the one scene the whole chapter turns on, that difference is worth paying for. But most of drafting is not that scene. It is the connective 800 words you will cut in revision anyway, and paying frontier prices for those is how AI writing costs get away from people.
The practical answer is to use two models, not one. A cheap, fast, high-fidelity model for volume — Luna, GLM 5.2 and DeepSeek V4 Flash all qualify at well under a tenth of a cent per action — and a frontier model for the passages that carry weight.
What does "free" actually mean on OpenRouter right now?
Far less than the count suggests. We pulled OpenRouter's live catalogue on 13 August 2026: 410 models, 19 of them priced at zero. Then we read what those 19 actually are.
| What the 19 zero-cost entries are | Count | Can it write a scene? |
|---|---|---|
| General-purpose text models | 10 | Possibly — these are the real candidates |
| Coding agents | 3 | No |
| Vision and video models | 2 | Not what they are built for |
| Music generation (Lyria 3) | 2 | No — they output audio |
| Content-safety classifier | 1 | No |
openrouter/free — a router, not a model |
1 | Depends entirely on where it lands |
This is the gap in every "all 17 free models listed" post: the number is a catalogue fact, not a capability claim. Two of the free "models" compose music — and are not really free either, since Lyria bills $0.08 per song and $0.04 per clip despite showing zero token pricing. One exists to decide whether text is safe; asked to continue our scene, it replied "User Safety: safe" and stopped. Half the list cannot write a sentence of fiction and was never meant to.
So the honest free shortlist starts at ten, not nineteen. We tested all ten.
Which free models can actually write fiction?
Six of the ten returned usable prose. Two held all four planted facts.
| Free model | Result | Held all 4 facts? |
|---|---|---|
| Nemotron 3 Ultra 550B | 129–175 words, real voice, dialogue | Yes — counted "fourteen steps to the lamp room" correctly. Took 69–86 seconds. |
| Nemotron 3 Nano 30B | 161 words in 5.2 seconds | Yes — the best speed-to-quality ratio of any free model |
| Nemotron 3 Super 120B | 150 competent words | No — invented that Tobias "slipped on the icy steps." He drowned. |
| Ling 3.0 Tiny | 189 words in 1.5 seconds | No — has her descend, then see the lamp room that is above her |
| Gemma 4 26B | 169 clean words on one run | No — sends her down "the remaining twelve steps" |
| LFM 2.5 2.6B | 293 words | No — reverses her direction of travel and invents specifics |
| Nemotron 3.5 Lightning | Returned its own reasoning, twice | — |
| GPT-OSS 20B | 897 of 900 tokens spent thinking; empty | — |
| Nemotron Nano 9B | Whole budget spent reasoning; empty | — |
| Gemma 4 31B | HTTP 429 on four attempts across two sessions | — Never served us a single request |
Two results deserve more than a table row.
Nemotron 3.5 Lightning did not write a scene. It returned a numbered outline beginning "Here's a thinking process," restated our own instructions back to us, produced a draft inside its notes, and then started counting the words aloud: "She(1) pressed2 her3 back4 against5 the6 stone,7". Both times. The instruction said "prose only, no preamble, no notes."
Gemma 4 26B did something stranger. On one run it produced the cleanest free prose of the day. On the next, same prompt, it produced this:
It was a rhythmic, metallic scraping,ಿಸಿದರು and it wasn't coming from the surf below... her breath hitching in a throat봅시다ed by the sudden, biting chill... The sound勝利 repeated—a heavy, dragging scrape... the snow had控股ed the roads nine days ago.
Kannada, Korean, Chinese and Gujarati tokens spliced into English mid-sentence. Same model, same prompt, two runs, minutes apart. That is the free tier's real character: not bad, but unrepeatable.
Why we stopped using openrouter/free
Because it is a lottery, and we measured the odds. openrouter/free is a router that picks a free model for you at random. We called it three times with the identical prompt and got three different models: Nemotron 3.5 Lightning, a vision model, and Gemma 4 26B.
Two of the three returned prose. One returned that chain-of-thought outline. So one call in three came back unusable, and nothing in the response tells you which roll you got until you read it.
That is why NovelCanon's model picker excludes the router and points at concrete free models instead. A random draw from a pool containing coding agents and safety classifiers is not a model choice; it is a coin toss you did not know you were making.
What happens when a free model disappears?
Often enough that the catalogue pull behind this article caught one.
Until this week NovelCanon recommended meta-llama/llama-3.3-70b-instruct:free as its default free model. That endpoint no longer exists. The paid Llama 3.3 is still listed at $0.10/$0.32 per million; the free one is simply gone, delisted with no announcement, as free endpoints routinely are. We found it in the same pull that produced the numbers above and replaced it the same day with ids verified live.
The general point matters more than our version of it: if you cannot see which model you are running, you cannot know when it changes underneath you. Bringing your own key is what made this visible enough to catch — and checking the live catalogue, rather than a list someone published last quarter, is the only way to keep a page like this true.
The instability is structural, not occasional. Free endpoints run on a shared pool: three of our ten candidates returned HTTP 429 on first contact, and OpenRouter's own error text explains why — "temporarily rate-limited upstream." You are queueing behind everyone else on Earth using that model for nothing. Even when you get through, OpenRouter's published limits cap free models at 20 requests per minute and 50 per day, rising to 1,000 per day once you have bought at least $10 in credits.
Free models are a real way to draft a novel at no cost. They are a poor thing to build a deadline around.
So which model should you pick?
Match the model to the job rather than hunting for one winner.
- Drafting volume, paid: GPT-5.6 Luna or GLM 5.2. Both held every fact for a fraction of a cent, and at that price you stop rationing revisions.
- The scene that matters: Claude Opus 5 for control, Claude Fable 5 for voice. Pay frontier prices deliberately, not by default.
- Story logic and physical detail: GPT-5.6 Sol tracked the stairwell arithmetic more precisely than anything else we ran.
- Drafting at zero cost: Nemotron 3 Nano 30B if you want an answer in five seconds, Nemotron 3 Ultra if you will wait a minute for a better one.
- Avoid for prose, whatever the price: any model that spends your output budget reasoning. Test it once with a 900-token limit; if it returns an empty string, it is the wrong tool.
- Always keep a paid key configured as a fallback. Free endpoints vanish mid-project — one on this page vanished while the page was being written.
If you are choosing software rather than a model, that is a different question with a different answer — see our honest comparison of AI novel writing tools. Models write sentences; tools decide what context those sentences get written against.
The variable that stops being the model
Here is the finding we did not expect. Nemotron 3 Ultra — the strongest of the free models, the one that got "fourteen steps to the lamp room" exactly right — said on a later run that Mira had "four steps to the landing, twelve more to the lantern room door." Sixteen steps, from step 98, in a tower with 112. Same model, same facts, same prompt.
Four planted facts in 120 words of context is the easiest possible version of this problem. A real manuscript asks a model to hold hundreds of facts across 90,000 words, most of them established in chapters that are no longer in the context window. No model on this list can do that, and no upgrade fixes it — because the failure is not intelligence, it is memory.
That is the gap NovelCanon's Consistency Engine is built to fill: your story's facts tracked outside the model, so a contradiction between chapter 4 and chapter 31 gets caught by something that was never relying on a context window to remember. It is why we can be genuinely relaxed about which model you use, and why our guide to writing a novel with AI spends more time on the bookkeeping layer than on prompts.
Pick the model that sounds like you. Then check the book with something that isn't the model. You can start free, connect any key, and point it at any model on this page — including the ones that cost nothing.
Frequently asked questions
What is the best AI model for writing fiction in 2026?
In our test, Claude Opus 5 and GPT-5.6 Sol produced the most controlled prose while holding every planted story fact, and Claude Fable 5 had the most distinctive voice. But GPT-5.6 Luna held the same facts at roughly a hundredth of Fable's price, so "best" depends on whether you are buying prose quality or prose volume.
Are free AI models good enough to write a novel?
Some are, with real caveats. Of the ten free models we tested, six returned usable prose and two held all four planted continuity facts — Nemotron 3 Ultra and Nemotron 3 Nano. The rest returned their own chain of thought, an empty string, or a rate-limit error. Free models work; the free *roster* is what you cannot rely on.
Is Claude better than ChatGPT for fiction writing?
On this test they were close and different. The Claude models produced the most restrained, literary sentences and the strongest interiority; the GPT-5.6 models were more precise about physical detail and story logic — GPT-5.6 Sol tracked the stairwell arithmetic exactly. Prose taste is genuinely subjective, so run the same scene through both before committing.
What is the cheapest AI model that writes good fiction?
GPT-5.6 Luna was the standout value in our run: 170 words of controlled prose that respected every planted fact, for $0.00023 — about a fortieth of what Claude Opus 5 charged for a comparable result. DeepSeek V4 Flash was cheaper still at $0.00021, with good sentences but a spatial continuity slip.
Why does an AI model return an empty response?
Usually because it is a reasoning model that spent your entire output budget thinking. Kimi K3 consumed all 900 tokens we allowed it and returned nothing; giving GPT-OSS 20B more room made it worse, not better. If a model returns an empty string, raise the output limit once — and if it still returns nothing, it is the wrong kind of model for prose.
Can I use free OpenRouter models in NovelCanon?
Yes. NovelCanon's model picker reads OpenRouter's live catalogue and badges the zero-cost endpoints, so you can search the full list and select a free model directly. The free plan supports bring-your-own-key with free models indefinitely, capped at 20 AI actions per day; the Consistency Engine is a Pro feature.
Written by
Munib Ali Laghari
Founder, NovelCanon
Building the writing studio he wished he had for keeping a long story straight — one where the AI never loses the thread.