Sherlock Holmes Canon Inconsistencies: What Our AI Scan Missed
We scanned all 60 Sherlock Holmes stories — 657,553 words, 23,074 facts. It found none of Doyle's four famous continuity errors. Here is exactly why.
We ran our Consistency Engine over the complete Sherlock Holmes canon — 4 novels, 56 stories, 657,553 words — and extracted 23,074 facts. It raised 21 flags, none of them a mistake Conan Doyle made, and caught none of the four Sherlock Holmes canon inconsistencies scholars have documented. Why it missed them is the useful part.
The last time we did this, we scanned Dracula and learned that a naive checker flags a great novel 164 times for the crime of having a plot. That post was about precision — not crying wolf. This one is about the bill that precision runs up.
Sherlock Holmes is the right book to send that bill to, because it is the one classic that comes with an answer key.
Key takeaways
- We scanned the complete canon — 4 novels, 56 stories, 657,553 words — and extracted 23,074 facts. No chunk failed.
- A naive "any changed value is a contradiction" check flags 1,144 conflicts. The engine's own policy cuts that to 21.
- None of the 21 is an error Conan Doyle made. Ten are the plot, seven are unnamed subjects colliding, four are extraction slips.
- It caught none of the four documented canon errors — and each miss failed for a different reason, which is the actual finding.
- Watson's famous wandering wound is read correctly every time and then deliberately not flagged:
injurycounts as a changeable state, so the later value supersedes the earlier one. We re-ran it four times to confirm. - That same rule is what makes the engine trustworthy. Precision has a price, and this is what the price looks like.
Which Sherlock Holmes canon inconsistencies are actually documented?
Holmesians have already published the list of Doyle's mistakes, so we could score the engine against ground truth instead of grading our own homework. For over a century, readers playing what they call the Grand Game have catalogued every slip in the canon and invented elaborate in-universe explanations for them. Doyle wrote fast and rarely looked anything up, so there is plenty to catalogue.
That gave us something a data story usually cannot have: a pre-registered test. We wrote the four best-documented errors down before looking at the results, and verified each one directly in the Project Gutenberg text rather than trusting a secondary source.
| # | The documented error | The two halves, verbatim |
|---|---|---|
| 1 | Watson's war wound moves | "struck on the shoulder by a Jezail bullet" (A Study in Scarlet) → "my wounded leg. I had a Jezail bullet through it" (The Sign of the Four) |
| 2 | Watson's own first name | Titled "JOHN H. WATSON, M.D." → his wife asks "should you rather that I sent James off to bed?" (The Man with the Twisted Lip) |
| 3 | Holmes's literary knowledge | "Upon my quoting Thomas Carlyle, he inquired in the naivest way who he might be" (A Study in Scarlet) → one book later Holmes asks "Are you well up in your Jean Paul?" (The Sign of the Four) |
| 4 | Red-Headed League arithmetic | "The Morning Chronicle of April 27, 1890. Just two months ago" → "THE RED-HEADED LEAGUE IS DISSOLVED. October 9, 1890" |
There is a nice detail in the first one. By The Noble Bachelor, Doyle appears to have noticed: the wound becomes "the jezail bullet which I had brought back in one of my limbs." Having contradicted himself, he simply stopped being specific.
What did the scan actually find?
23,074 facts, 1,144 differences a naive check would call contradictions, and 21 flags after the engine's own policy — of which not one is an error Conan Doyle made.
The funnel matters more than any single number, so here it is in full:
| Stage | Count |
|---|---|
| Words scanned (9 books, publication order) | 657,553 |
| Scene-sized chunks analysed (0 failed) | 344 |
| Facts extracted | 23,074 |
| Facts kept after the junk filter | 20,622 |
| Distinct subjects tracked | 1,821 |
Distinct subject · attribute keys |
11,611 |
| A naive "any changed value is a contradiction" check | 1,144 |
| After the engine's derivation policy | 21 |

Every number here comes from the actual scan run described in this post — 4 novels, 56 stories, 657,553 words, no failed chunks.
That 1,144 → 21 is the same collapse the Dracula scan produced, at four times the scale. And the 21 that survived break down like this:
| What the flag really was | Count |
|---|---|
| The plot: someone reported dead is alive elsewhere in the narrative | 10 |
| Unnamed subjects colliding — "man", "woman" | 7 |
| Extraction or attribution slips | 4 |
The ten "plot" flags are funny in aggregate, because they expose Doyle's favourite trick. Drebber and Stangerson are murdered in A Study in Scarlet and alive again in the Utah flashback. Blessington has committed suicide and is also showing people into a room. Professor Coram is dead and is also turning his white mane towards us.
And twice — twice — the engine flags Sherlock Holmes himself. "It was the last that I was ever destined to see of him in this world," says The Final Problem; then The Empty House answers, "Holmes! Is it really you? Can it indeed be that you are alive?" The most famous resurrection in detective fiction, filed as a continuity error, exactly as a machine should file it and exactly as no reader ever would.
Did it catch Watson's wandering war wound?
No — and because this is the centrepiece, we ran it four times to be certain of why. It is not that the engine missed the sentences. It reads both, every run, and in three runs of four it files them against the same character under the same attribute:
| Book | Attribute | Value | Confidence |
|---|---|---|---|
| A Study in Scarlet | injury |
"shattered bone in shoulder" | high |
| The Sign of the Four | injury |
"Jezail bullet wound in leg" | high |
Two values, one key, both high confidence — a textbook contradiction sitting in the registry. Flags raised: zero, in every single run.
Here is why. Every attribute the engine tracks is classified. eye_color and name are identity — they should never change, so a difference is an error. location and mood are volatile — they change constantly and are never flagged. Everything unfamiliar defaults to state: attributes that legitimately move over time, where a later value simply supersedes the earlier one instead of contradicting it.
injury is on no list, so it defaults to state. And as a general rule that is correct. People get hurt and recover; a character wounded in chapter two and wounded elsewhere in chapter twenty has not created a plot hole. The engine reads "shoulder, then leg" as Watson being injured again later, which is a sensible reading of almost any novel.
It is only wrong here because Watson's wound is not really a state at all. It is a fixed biographical fact wearing the costume of a changeable one — one bullet, at Maiwand, in 1880, that either hit his shoulder or his leg. No attribute classification can tell you that, because the distinction lives in the story, not the vocabulary.
The repeats turned up a second, quieter fragility worth knowing about. In one run of four, the extractor filed the leg wound under an invented key — has_wounded_leg instead of injury — so the two halves never met in the registry at all. The extractor writes its own attribute names, and when it picks a different one for the same idea, a contradiction dissolves silently. Same model, same text, same code; different vocabulary.
That is the trade, stated plainly: the default-to-state rule is what turned 1,144 flags into 21 and made the Dracula scan trustworthy. The same rule is why the most famous continuity error in English literature went by without a word.
What happened to the other three?
Three misses, three completely different mechanisms — which is why "did the AI find the errors" is the wrong question.
- Watson called "James" was never extracted at all. The line is "should you rather that I sent James off to bed?", spoken by Watson's wife about Watson. Catching it means inferring that a name in someone else's dialogue refers to the narrator, who is not named in the sentence. Our extractor never made that leap, so no competing value for
nameever entered the registry. A genuine detection failure, not a policy choice. - The Carlyle reversal failed for a subtler reason. The engine did record Holmes's famous self-assessment —
knowledge_literature = "Nil"— straight from the list Watson draws up. But the other half is never stated; it is demonstrated. Holmes discussing Jean Paul is behaviour, and nothing in that scene says "Holmes is well read." A fact registry compares claims to claims. It cannot compare a claim to a performance. - The Red-Headed League dates are out of scope by design, and we said so last time. "April 27 to October 9 is not two months" is arithmetic over a calendar, not two different values for one property. It is the same boundary that kept the Dracula scan from catching that book's tangled 1893 chronology — we named that limit then, and it has not moved.
One of our predictions was wrong, and it belongs in the record. Before reading the results we expected the Carlyle miss to come from the engine's deliberate exclusion of knowledge facts from contradiction checking. It didn't. The real reason is that the two halves were never the same kind of statement — a more interesting failure, and one no policy change would fix.
So what does a 0-for-4 scan tell you about your own manuscript?
That the scan is strong exactly where your memory is weak, and weak exactly where a reader is strong — which is why they are not substitutes.
Look at what it did do. It read 657,553 words without tiring, held 1,821 subjects and 11,611 distinct claims in mind at once, and never got bored on page four hundred. No human editor does that. What it surfaces are fixed, restated facts in conflict — colours, names, stated physical traits — precisely the details that drift across a long draft and that nobody catches by re-reading.
What it could not do was infer. Every miss here needed a leap: that a name in dialogue means the narrator, that quoting a philosopher implies knowing him, that two dates five months apart contradict a character's "two months ago," that one attribute called injury was really a permanent biographical fact. Those leaps are what a reader is for.
One honest caveat on the numbers. We scanned the canon as raw text with no Characters board, so the engine had no alias list — which is why "man" and "woman" were tracked as if they were characters, and why seven of the 21 flags are unrelated people colliding under one label. In a real project you name your cast, and that whole category largely disappears. The flags a writer would see are fewer and cleaner than the ones we got.
So the practical read is this. Run a full-manuscript continuity check to catch the drifting details you have no chance of holding in your head, and expect it to raise questions rather than deliver verdicts. Then read your own book for the things that require knowing what a sentence implies. The engine is a tireless first pass over twenty thousand facts, not a replacement for the one reader who understands the story — and how it decides what counts as a contradiction is worth understanding before you trust either of you.
We could have written "AI finds 21 errors in Sherlock Holmes." It would have been a better headline and a lie — there are no errors of Doyle's in that 21, and the four he actually made are still sitting in the canon, unflagged. A consistency tool that invents findings on a book you can go and read is worth less than nothing. What it is worth is knowing exactly where it can see and where it cannot. If you want the craft side of that judgement — telling a real slip from a character who is simply allowed to change — the complete guide to story consistency covers it, and you can point the same scan at your own draft with NovelCanon.
Frequently asked questions
What are the most famous inconsistencies in the Sherlock Holmes canon?
Four are cited most often. Watson's war wound is in his shoulder in "A Study in Scarlet" and his leg in "The Sign of the Four". His wife calls him James in "The Man with the Twisted Lip" though he is John H. Watson. Holmes has never heard of Thomas Carlyle in the first novel and is discussing Jean Paul by the second. And "The Red-Headed League" calls an April 1890 newspaper "just two months ago" while dating its own ending to October 1890.
Did the AI scan actually find Watson's wandering war wound?
No, and we tested this four separate times to be sure. The engine reads both halves every run, and in three runs out of four it files them on the same character under the same attribute — "shattered bone in shoulder" and "Jezail bullet wound in leg", both at high confidence. It then raises no flag at all, because it treats "injury" as a state that legitimately changes over time, so the later value quietly supersedes the earlier one instead of contradicting it.
Why would a consistency checker deliberately ignore a change like that?
Because most changes in a novel are the story, not a mistake. Characters get hurt, recover, move, age and die, and a checker that flags every one of those trains you to ignore it. Treating unfamiliar attributes as changeable state is what keeps the false-positive rate low — and the cost of that choice is exactly the kind of miss this scan demonstrates.
Does this mean AI consistency checking does not work?
It means it has a shape you should understand before you rely on it. The scan is strong on fixed, restated facts — eye colour, a name, a stated physical trait — and weak where a contradiction is implicit in behaviour, buried in date arithmetic, or hidden in an attribute that usually does change. It is a first pass that reads all 657,000 words without tiring, not a replacement for a reader who knows the book.
Did the scan accuse Conan Doyle of any mistakes?
No, and that is worth stating plainly. It raised 21 flags across the whole canon and not one is an error Doyle made. Ten are the plot — a character reported dead who is alive elsewhere in the narrative, including Holmes's own death at the Reichenbach Falls. Seven are unnamed subjects like "man" or "woman" colliding into a single phantom character. Four are extraction slips.
Can I run this kind of scan on my own manuscript?
Yes — it is the same full-manuscript deep scan NovelCanon runs on your own book, using your own AI key, with a cost estimate shown before it starts. The difference is that your project has a Characters board, so the engine can resolve aliases and stop treating "man" and "woman" as characters, which is where a third of these flags came from.
Written by
Munib Ali Laghari
Founder, NovelCanon
Building the writing studio he wished he had for keeping a long story straight — one where the AI never loses the thread.