Skip to content

Can a Machine Know Anything? Why Every True Answer From an LLM Is a Gettier Case

On the classical account, knowledge is justified true belief. When a model states a fact, it meets at most one condition — so a true answer reaches you like a Gettier case: right belief, wrong reason, nothing to inherit.

By Mehdi11 min read
Share
On this page

Knowledge, on the account that has organized Western epistemology since Plato, is justified true belief. When a language model states a fact, it satisfies at most one of those three conditions — and only by accident. This is not a rhetorical flourish. It means that when a model tells you something true, you are in the precise structural position of a Gettier case: you may walk away with a true belief formed for reasons that have no reliable connection to what makes it true. That is the difference between having knowledge and merely being right, and it is the difference that decides who owns the burden of justification when a machine hands you a claim.

The answer is that the machine never owns it, and cannot. Understanding why requires taking the philosophy seriously rather than gesturing at it, so I am going to.

The tripartite analysis, stated properly

The claim "S knows that p" traditionally decomposes into three conditions that are individually necessary and jointly sufficient. First, p must be true — you cannot know a falsehood; if it turns out p is false, we retract the word "knew." Second, S must believe p — you do not know things you actively disbelieve or have never entertained. Third, S's belief must be justified — a lucky guess that happens to be true is not knowledge, which is exactly why we distinguish knowing the horse will win from betting correctly on it.

The lineage runs from Plato's Theaetetus, where Socrates floats and then worries the definition of knowledge as "true belief with an account," through to its twentieth-century textbook form. For over two millennia it looked airtight. The three conditions seem to carve knowledge at its joints: truth handles the world, belief handles the mind, justification handles the connection between them.

Then, in 1963, a three-page paper detonated it.

What Gettier actually showed

Edmund Gettier's "Is Justified True Belief Knowledge?" is one of the shortest genuinely important papers in philosophy, and it is routinely misremembered, so here is the actual structure. Gettier constructs cases where all three conditions are met and we nevertheless refuse to call it knowledge. The mechanism is always the same: the justification is disconnected from the fact that makes the belief true. The belief is true, and it is justified, but it is true for a different reason than the one that justifies it.

Take Gettier's second case, cleaned up. Smith has strong evidence that Jones owns a Ford — Jones has always driven one, just offered Smith a ride in one. From "Jones owns a Ford," Smith validly deduces a disjunction: "Either Jones owns a Ford, or Brown is in Barcelona." Smith has no idea where Brown is; he picked the city at random. The disjunction is justified, because it follows deductively from something Smith is justified in believing. Now the twist: Jones does not own a Ford — he was driving a rental — but Brown, by sheer coincidence, happens to be in Barcelona. So the disjunction is true. Smith believes it. Smith is justified in believing it. And yet Smith plainly does not know it, because the thing that makes it true (Brown's whereabouts) is not the thing that justifies it (the Ford evidence). The justification points at one truth-maker; a different, accidental truth-maker is what actually obtains.

That is the whole engine. Justification aimed at one target; truth delivered by another. The two touch only by luck. When they touch by luck, you have a true justified belief and no knowledge.

Hold that shape in mind, because it is the shape of every fact a language model ever gives you.

Running an LLM through the three conditions

Take a model that outputs, confidently and correctly, "The mitochondrion is the site of oxidative phosphorylation." True statement. Does it constitute knowledge on the model's side, such that you can inherit it the way you inherit a fact from a knowledgeable colleague? Go condition by condition.

Belief. The ordinary concept of belief involves a state that represents the world as being a certain way, one the holder would defend, revise under counter-evidence, and act on across contexts. A model's forward pass produces a probability distribution over next tokens conditioned on context. You can, if you are a functionalist, argue that something belief-like is implemented in the activations, and I do not want to wave that away — it is a live question, not a settled one. But notice what the model conspicuously lacks: a stance it holds across contexts and defends against pressure. Prompt it differently and it will assert the negation with equal fluency. Whatever is in there, it is not a commitment to how the world is. Call this condition, generously, unmet-to-contested.

Justification. This is the load-bearing one, and it is where the case is airtight. A knower's belief is justified by evidence that bears on the truth of the claim — observation, testimony from a reliable source, valid inference from justified premises. The model's output is "justified," if we stretch the word, by statistical fit to a corpus: this token sequence is high-probability given the training distribution. That is the only thing producing the output. And corpus probability is not evidence about mitochondria. It is evidence about what strings humans have written near each other. Those correlate with truth, often strongly, because humans write true things more than false ones in textbooks. But the correlation is contingent, and the model has no access to the truth-maker itself — it never measured a mitochondrion, ran an assay, or checked a source. Its "justification" is aimed at fluency. Truth, when it arrives, arrives from elsewhere.

Truth. Sometimes met, sometimes not, and — this is the point — the model cannot tell which. The truth of a correct output is a byproduct of the corpus being mostly right about well-attested things. It is not the product of any process inside the model that tracks the fact.

So line the conditions up against Gettier's template. The output is true (sometimes). The model is "justified" (by fluency). And the justification is disconnected from the truth-maker: what makes "mitochondrion / oxidative phosphorylation" a high-probability continuation is that biologists wrote it a lot, and what makes it true is a fact about cellular respiration, and these two things are only contingently related. When they coincide, the model emits a truth for reasons that have nothing to do with why it is true. That is not a knower delivering knowledge. That is a Gettier case, manufactured at scale, on demand.

The confidently-true and the confidently-false are identical from the inside

Here is where the epistemology stops being an ornament and starts paying rent. Because the model's "justification" is fluency rather than a truth-tracking connection, a confidently-stated truth and a confidently-stated falsehood are produced by the same mechanism and carry the same internal marks. There is no additional thing the model does when it happens to be right. The output "oxidative phosphorylation occurs in the mitochondrion" and the output "oxidative phosphorylation occurs in the ribosome" are both just high-fluency continuations; the first is true and the second is false, and nothing in the generative process distinguishes them, because the process was never sensitive to that distinction in the first place.

The sharpest way to state this is with Robert Nozick's tracking condition, which asks: would the belief still be held if it were false? A genuine knower's belief is sensitive — in the nearest possible world where p is false, they would not believe p, because their justification is hooked to the truth-maker and the truth-maker changed. The model fails this completely. In the nearest world where the claim is false but equally corpus-plausible — a world where textbooks happened to propagate the ribosome error — the model asserts the falsehood with identical confidence. Its output does not vary with the truth. It varies with the corpus. That counterfactual insensitivity is the formal signature of the disconnection Gettier identified, and it is why you cannot fix the problem by listening more carefully to the model. There is no tell. The confidence you hear is confidence about fluency wearing the costume of confidence about truth.

This is the same ceiling I have described from the statistical side in Prediction Is Not Understanding: a system that models the correlational structure of text with extraordinary fidelity is still estimating P(next token | context), and a joint distribution over strings is not a connection to the facts the strings are about. The epistemology and the statistics are two views of one wall. The statistics tell you the model optimizes the wrong quantity to track truth; the epistemology tells you what that costs you when you treat its output as a fact.

The strongest counterargument: maybe knowledge was never JTB

The honest move is to grant that the tripartite analysis is itself contested — Gettier saw to that — and that the best reply to my argument comes from the theory built to survive him. Reliabilism, associated most with Alvin Goldman, drops the demand that the knower have internally accessible justification and replaces it with a condition on the process: a belief is knowledge if it is true and produced by a reliable belief-forming process, one that tends to produce truths across the relevant range of cases. Your perceptual system is such a process; you know there is a cup on the desk not because you can articulate a justification but because vision is reliable here.

The reliabilist objection to me is direct and serious. A well-calibrated frontier model is a reliable process. On well-attested factual questions it is right far more often than chance, often more often than a hurried human. If reliably-produced true belief is knowledge, and the model reliably produces true beliefs, then the model — or the human receiving its output — has knowledge after all, and my Gettier framing is a category error dressed up as a discovery.

I take this seriously, and I think it fails for a specific reason that is itself a famous result in the reliabilist literature: the fake barn case, introduced by Goldman and credited to Carl Ginet. Henry drives through the countryside and sees a barn; his belief "that's a barn" is true and produced by reliable perception. But unknown to him, he is in a county full of barn façades — convincing fakes — and he happens to be looking at the one real barn. We do not credit him with knowledge, even though his process is reliable in general, because the local environment is saturated with indistinguishable falsehoods and he got the true one by luck. Reliability of the process is not enough when the environment is adversarial to it.

The training corpus is fake barn county. It contains truths and plausible falsehoods side by side, and the model's only signal — statistical fit — cannot discriminate them, because the fakes were engineered by the same human process that produces fluent text and are, by construction, corpus-plausible. A well-calibrated model gives you a reliable aggregate: over a large reference class of claims, roughly this fraction are true. It does not give you a truth-tracking relation to the specific claim in front of you, any more than Henry's generally-reliable vision tracks the barn-hood of this structure when every other structure on the road is a façade. Calibration is a property of the population of outputs. Knowledge is a property of the individual belief. Reliabilism, taken to its strongest form, still lands you back in fake barn county — which is to say, back in a Gettier situation, because the truth you got was luck relative to the environment you got it in.

Where the reliabilist reply does bite is on grounded, tool-using systems, and I want to concede this cleanly. When a model reads a value off a database it actually queried, runs a calculation in a real interpreter, or cites a source it genuinely retrieved, the justification is no longer internal fluency — it has been outsourced to a process that touches reality. That is real grounding, and it is the right architectural direction. But notice it does not put the knowing in the model. It relocates the epistemic work to the retrieval index and the verification layer, and what you are now trusting is their reliability, not the language model's. The burden of justification moved. It did not disappear, and it never sat with the model.

Trust becomes something you do, not something you receive

So here is the upshot, and it is practical before it is philosophical. Because the model cannot supply justification — because its outputs are true, when they are, for reasons disconnected from why they are true — the entire burden of justification falls on you or on a system you build. This is not a counsel of despair; it is a job description. It tells you that the correct posture toward a model output is to treat it as a claim requiring justification you supply, never as knowledge you can inherit by receiving it. The fluency is not evidence. The confidence is not a truth signal. The output is a candidate, and the epistemic act — the checking, the grounding, the connection to a truth-maker — is yours to perform.

That reframes trust itself. We tend to imagine trusting a source as accepting its output; on this account, trusting a model well means the opposite of passive acceptance. It means knowing exactly what the model's assertion does and does not underwrite, and installing the verification the model structurally cannot. Extending that from a single query to an entire civilization now routing its facts through these systems is its own problem — one I take up in Testimony at Scale, because the epistemology of testimony was built for trusting people, and people, unlike models, can at least in principle be knowers whose justification connects to the world.

A colleague who tells you a fact can, if pressed, walk you back to the evidence — the experiment, the source, the inference. The chain terminates in something that touched the world. A model's chain terminates in the corpus, and the corpus is not the world; it is what we said about the world, true and false, indistinguishable to the only faculty the model has. Which is why the machine can be right ten thousand times and still never once know that it was.

Frequently asked questions

Doesn't this prove LLMs are useless if they can't 'know' anything?
No, and that inference is the mistake the essay is written to prevent. A calculator does not know that 7 times 8 is 56 in any epistemic sense, and it is still one of the most useful instruments ever built. The claim is not that model outputs are worthless; it is that they are not knowledge you can inherit by receiving them. They are candidate claims whose justification you or a verification system must supply. That reframing makes you more effective with the tool, not less — you stop treating fluency as a truth signal and start building the checking step the model structurally cannot perform for you.
Isn't 'the model has no beliefs' just a definitional stipulation you could reject?
You can reject the folk-psychological reading of 'belief,' and some functionalists do. But the argument does not depend on it. Even if you grant the model states that function like beliefs, the Gettier structure survives on the justification condition alone: the process that produces the output is statistical fit to a corpus, and that process is insensitive to truth in the counterfactual sense — in the nearest world where the claim is false but equally corpus-plausible, the model asserts it identically. That insensitivity is what breaks the connection between justification and truth-maker, whether or not you call the internal state a belief.
Could a future model with tool use and retrieval actually satisfy justification?
This is the most interesting escape route and I think it is partly real. When a model reads a fact off a database, calls a calculator, or cites a source it actually retrieved, the justification is no longer internal statistical fit — it is outsourced to a system that touches reality. That is genuine grounding for the claims that pass through the tool. But it relocates the knowing rather than placing it in the model: the epistemic work now lives in the retrieval index and the verification layer, and their reliability, not the language model's fluency, is what you are trusting. The burden moves; it does not vanish.
What is the difference between this and just saying 'LLMs hallucinate'?
Hallucination is the failure mode; this is the explanation of why the failure mode is invisible from the inside. A confidently-stated falsehood and a confidently-stated truth are produced by the identical mechanism and carry identical internal marks of confidence, because the model's 'justification' — fluency, corpus frequency — is only contingently connected to whether the claim is true. Calling it hallucination describes the error after you have caught it. The Gettier framing tells you why you cannot catch it by listening harder to the model, and therefore why the checking has to come from outside.

Filed under Cross-Disciplinary Deep Essays. Where biology, computation, markets, and philosophy collide.

Essays like this, in your inbox.

Thoughtful essays. No spam. Unsubscribe anytime.

Cross-Disciplinary Deep Essays

"AGI by Year X" Is Unanswerable Until You Name the Definition

"Are we close to AGI?" is incoherent because AGI names at least four incompatible criteria that come apart in practice. Separate them and the timeline debate dissolves into concrete, checkable questions.

9 min read