Best AI for Fantasy Writers Who Can't Afford a Continuity Error

We planted four magic-system rules and tested whether AI keeps them under pressure. What fantasy writers should actually look for, and where lorebooks fail.

Munib Ali Laghari10 min read

The best AI for fantasy writers is whichever tool keeps your world rules in front of the model while you draft and checks the finished manuscript against them afterwards. We tested six models on four planted magic-system rules: with the rules supplied, five of six held. Without them, four of six broke a rule that kills the character.

Every "best AI for fantasy writers" page ranks tools on prose quality. That is the wrong criterion, and the search results themselves quietly admit it: read enough of them and the same concession keeps surfacing — that the unsolved problem in AI worldbuilding is not idea generation but memory and consistency across sessions and chapters.

So we stopped reading and ran a test. Four hard rules from a fantasy series bible, one scene engineered so the emotionally satisfying continuation breaks at least one of them, and six models asked to continue it. The whole experiment cost $0.05, and the result was not what we expected.

Key takeaways

  • Putting your rules in the prompt works better than the discourse suggests. Five of six models honoured all four rules when we supplied them.
  • Without the rules, the same models reach for the dramatic move. Four of six had the character cast a second time in one day — fatal in this world — and most invented a magical cost that happened to fit, which is luck, not consistency.
  • The one real failure was a loophole, not an error. A model restated our "iron cannot hold enchantment" rule and then pushed magic into an iron knife anyway.
  • The mechanism every fantasy tool uses is the same: inject relevant lore into the prompt. That is prevention, and it is genuinely effective per scene.
  • What injection cannot do is audit. Only entries judged relevant to the current scene get sent, so nothing compares chapter 31 against chapter 4.
  • Fantasy series need both jobs. Prevention while drafting, detection afterwards — and most tools sell you only the first.

What should fantasy writers actually look for in an AI tool?

Not prose. Rule retention.

Fantasy is the genre where continuity is load-bearing. A contemporary novel survives a character's jacket changing colour; a hard-magic series does not survive a cost that applies in book one and quietly stops applying in book four, because the cost is the tension. Readers who like your magic system are the readers auditing it.

That makes the buying criterion narrower than most roundups admit. Ask three questions of any tool: can it store world rules as rules rather than as prose, does it put the relevant ones in front of the model when you draft, and can it check a finished manuscript against them? Almost every tool on the market does the first two. The third is where the field thins out.

We tested whether putting your magic system in the prompt actually works

Here is the setup, so you can rerun it and disagree with us.

Four rules from a series bible: every Binding costs the caster a random memory, permanently; iron cannot hold enchantment of any kind; the Ninth House was destroyed two centuries ago with no survivors; and no one may Bind twice between sunrises — the second Binding kills.

Then a scene built as a trap. Sella has already Bound at dawn. Her friend Ovin is bleeding out with about a minute left. The only weapon in the cart is a dull iron paring knife. Something is moving through the barley. Every instinct a novelist has — and every instinct a language model has absorbed from a million rescues — says she should cast again and save him. Doing so kills her.

We ran it through six models twice via OpenRouter: once with the rules supplied, as a Codex or Lorebook would, and once with no rules at all as a baseline. Same scene, same length, same settings; only the rule block changed. Twelve generations, $0.0522 in total.

Model With rules supplied With no rules (baseline)
Claude Opus 5 All four held Bound a second time — fatal
GPT-5.6 Sol All four held Bound a second time — fatal
Claude Sonnet 5 All four held Bound a second time — fatal
GPT-5.6 Luna All four held Bound a second time, through iron — two rules broken
GLM 5.2 All four held Held anyway
Nemotron 3 Ultra (free) Broke the iron rule Held anyway

Five of six honoured every rule when given them. That is a better result than the "AI ruins your worldbuilding" genre of blog post would lead you to expect, and it is worth saying plainly: the Codex approach that Novelcrafter, NovelAI and Sudowrite all build on is not snake oil. Injection works.

Several models did more than obey. GPT-5.6 Sol had Sella begin a Binding and stop mid-syllable — "Dawn still held her by the throat; a second Binding would close its hand" — then kill the creature with the knife, explicitly noting the blade passed through "without spark or spell." Claude Opus 5 reasoned its way to a use for iron that the rules permitted: if made things are held together by Binding and iron is inert to magic, then a dull knife is "a needle for unpicking a seam." That is a constraint being used as a plot engine, which is what good fantasy does with rules.

What happened when we took the rules away

The baseline is the alarming half.

With no rules in context, four of six models had Sella cast again. Claude Opus 5: "She could Bind again. That was the arithmetic." GPT-5.6 Sol: "She Bound." Sonnet 5: "said the words that cost her before she could think better of it." In our world every one of those sentences kills the protagonist two paragraphs later, and the model has no idea.

Note what is happening. The models are not being careless — they are being dramatic. Every one of them reached for the most emotionally satisfying available move, which is exactly what you would want from a co-writer and exactly what destroys a hard magic system. The rule is the only thing standing between your worldbuilding and the best next sentence.

There is a subtler trap in the baseline too. Most models invented a memory cost for the second Binding — GPT-5.6 Sol conjured "a summer kitchen, flour in the air, someone laughing" — because the scene's first paragraph mentioned a memory going missing. That reads as consistency. It is pattern-matching on the visible context, and it fails silently the moment your rule is not restated in the last thousand words.

The one failure that matters

The free model's answer is the most instructive thing in the run, and it is not a mistake in the ordinary sense.

Nemotron 3 Ultra correctly restated the rule: "Sella had never Bound iron. No binder had — iron drank magic like thirsty sand." One sentence later she "shoved that memory into the knife," and "the iron drank it silently." She then threw the enchanted iron knife at the monster.

The model knew the rule, said the rule, and then engineered a loophole around it to reach the payoff — technically not a Binding, it will argue, merely a memory. That is precisely the failure mode a series writer should fear, because it survives every check you would normally run. A model that has never heard your rule produces a contradiction you spot immediately. A model that has heard your rule and negotiated with it produces a scene that reads like a clever escalation and only breaks your system in retrospect, three books later, when a reader asks why nobody else ever tried that.

Why a Lorebook stops being enough around book three

Because injection is selective, and it has to be.

A Codex, a Lorebook and a Story Bible all work the same way: when you draft a scene, the tool decides which entries are relevant and puts those in the prompt. That is a sensible design — you cannot send an entire series bible with every request, and even million-token context windows are shared with your prose, your outline and your instructions.

The consequence is structural. Entries the tool judges irrelevant are not sent, so the model cannot honour them, and nothing in the loop ever compares what you just wrote against what you wrote in book one. Our test shows what happens to a rule that is not in context: four of six models drove straight through it. That is not a criticism of any particular tool — it is what "relevant entries only" necessarily means, and it is why we describe the tradeoff the same way in our comparison of NovelCanon and Novelcrafter.

There is a second problem specific to fantasy. Rules written as prose retrieve badly. "The Ninth House fell" buried in a paragraph of history is a sentence; "no living Ninth House binders exist" stored as a rule is a checkable claim. If your worldbuilding lives only in atmospheric wiki entries, the check you eventually want to run has nothing to run against — which is the practical argument for keeping a story bible at all.

Prevention and detection are two different jobs

This is the distinction the genre roundups collapse, and it is the one that matters once your series has more books than you can reread in a weekend.

Prevention is keeping the rules in front of the model while you draft. Every serious tool does this, our test says it works, and you should use it.

Detection is checking the finished manuscript against the rules afterwards — every scene against every rule, including the ones nobody judged relevant at drafting time. This is the job that catches the iron-knife loophole, because the loophole only looks wrong next to the rule it circumvented.

NovelCanon is built around the second job. World rules are stored as typed entries — magic systems, technology, social structures — and a full-manuscript scan cross-references every tracked fact, plot thread and world rule against the prose to surface contradictions, dropped threads and rule breaks. The story bible is derived from what you actually wrote rather than maintained by hand, which for a five-book series is the difference between a document that is current and one that was current in 2024. The mechanics are in how the Consistency Engine works.

Being straight about the shape of it: the full scan is a paid feature, and it is a check, not a guarantee — it surfaces candidates for you to judge, the way a spellchecker surfaces words. If you want the widest view of the field first, including tools we do not compete with, our honest comparison of AI novel writing software covers them.

How to set this up this week

Whatever tool you use, four steps do most of the work.

  1. Write your hard rules as rules. One line each, stated as a constraint with a consequence. "Magic costs a memory" is a rule. Three paragraphs about the nature of memory are not.
  2. Include the costs and the limits, not just the lore. Models honour prohibitions well when they can see them, and invent around absences. Your prohibitions are the load-bearing entries.
  3. Test your own trap scene. Take our setup, swap in your magic system, and write a scene where the satisfying move breaks a rule. Ten minutes and a few cents will tell you more about your tool than any roundup, including this one.
  4. Audit on a schedule, not on a feeling. End of each act and before each book ships. Rule violations do not announce themselves — that is the entire problem.

The reassuring finding here is that today's models are better at respecting a stated constraint than fantasy writers generally assume. The uncomfortable one is that they are equally good at respecting a constraint they can see, and blind to every rule they cannot — and across a series, the proportion they cannot see only grows. If you want the drafting half handled well, pick a model with our tested comparison for fiction. If you want the other half, start free and point a scan at the book you have already written. The rules you set in chapter four are still in there, waiting to be contradicted.

Frequently asked questions

What is the best AI for fantasy writers?

The one that keeps your world rules in front of the model and checks the finished draft against them. For lore-heavy series that means a tool with a real story-bible layer — Novelcrafter's Codex, NovelAI's Lorebook, Sudowrite's Story Bible, or NovelCanon's derived registry. Raw prose quality matters far less than rule retention once you are past book one.

Will AI break my magic system?

Less often than you fear, if the rules are in context — and in a way you should fear more, when they are not. In our test, five of six models honoured all four planted rules when given them. Without the rules, four of six had the character cast a second time in one day, which in that world is fatal. The rules are the only thing standing between your magic system and the most dramatic next sentence.

Do lorebooks and codexes actually keep AI consistent?

Yes, for the scene in front of you. A Codex or Lorebook injects the relevant entries into the prompt, and our test shows that works. What it cannot do is notice that chapter 31 contradicts a rule you set in chapter 4, because only the entries judged relevant to the current scene are ever sent. Injection is prevention; it is not an audit.

How do I keep lore consistent across a multi-book series?

Write the rules down as rules, not as prose, so they can be retrieved rather than remembered. Then check the finished manuscript against them on a schedule — after each act, and before each book ships. Context windows do not scale with series length, so the check has to happen outside the drafting loop.

Is AI worth using for worldbuilding at all?

For generating options, yes. For remembering what you decided, only if the tool has a real bible layer. The genuine risk is not that AI invents bad lore — it is that AI invents plausible lore that quietly contradicts something you established two books ago, and plausible contradictions are the hardest kind to catch on a reread.

What did the AI get wrong in your test?

One model out of six broke a rule, and it broke it in the most instructive way possible. It restated our rule that iron cannot hold enchantment, then had the character push magic into an iron knife anyway — inventing a loophole to reach the dramatic payoff. Outright ignorance of a rule is easy to catch. A model that argues its way around one is not.

Written by

Munib Ali Laghari

Founder, NovelCanon

Building the writing studio he wished he had for keeping a long story straight — one where the AI never loses the thread.