By Kareem Farid, founder of Kunkafa
Why ChatGPT can't predict the stock market
Language models learn words about markets, not the price patterns themselves - and fluent writing reads as certainty even when nothing sits underneath it.
Ask ChatGPT where a stock goes next and the answer arrives in the voice of an analyst: measured, balanced, quietly authoritative. It reads like knowledge, and wanting it to be knowledge is greed talking - the shortcut we would all take if it existed. The opposite reflex, that AI has nothing to offer a market, is just as wrong and costs just as much. The useful question sits between them: what is this kind of AI actually built to learn, and what would a market require instead?
A language model learns words, not price patterns
What a large language model has absorbed about a stock is what people have written about it. Analyst notes, news stories, earnings write-ups, forum arguments, a decade of hindsight explaining why a move that already happened was obvious. That corpus is genuinely enormous and it is genuinely about markets. It is not the market.
The market itself is a long sequence of numbers: for every period, an opening price, a high, a low, a close and a volume, repeating for years across thousands of instruments. Whatever structure exists in a market lives in the relationships between those numbers - how far price ran before it turned, how often a move of a given size followed a stretch of quiet, how the same shape resolved across thousands of earlier occurrences. None of that is written down in prose anywhere, so none of it reaches the training data.
The film critic
Asking a language model to forecast a market is like asking a film critic to direct. The critic has read and written more about film than most directors ever will. They can tell you why a scene works, place a picture in its tradition, describe exactly what a great performance does to an audience.
They have still never lit a set, never cut a sequence at three in the morning, never had to get a performance out of a tired actor on the ninth take. Fluency about a craft and command of it are separate skills, and no amount of the first turns into the second. The critic's authority is real, and it is authority about the writing, not about the doing.
Fluent is not the same as grounded
The second problem is what fluency does to a reader. A language model is trained to produce text that flows, and the financial writing it learned from is full of hedged, sober, confident-sounding sentences. So it produces hedged, sober, confident-sounding sentences, and they land with the weight of a considered view.
Three things are missing behind that voice. There is no chance attached to the claim, so nothing tells you how strongly to hold it. There is no duration, so the claim can never quite be wrong - a stock that falls for a month and recovers in the thirteenth can be said to have vindicated it. And there is no record: the model cannot tell you how often sentences of that shape turned out right, because nothing in it was ever counted against an outcome.
Something that cannot be scored is not a cautious forecast. It is a paragraph.
Fine-tuning does not fix the representation
Training a language model further on financial text teaches it more financial text, which is not the missing ingredient. The deeper obstacle is how numbers survive the trip into the model at all: a price becomes pieces of text, a run of characters that happens to look numeric, rather than a value sitting on a line at a fixed distance from the values above and below it.
Almost everything a market question depends on lives in that distance. How far is this from where it opened, how does today's range compare with the last thirty, is this move large for this instrument or unremarkable. A system that handles prices as text can imitate sentences about distance without ever computing one, and imitation is what comes back.
What working on numbers looks like instead
Seventy experts, one question each. Each expert judges how far a price could move and which way, on its own, and each is trained on the numbers rather than on writing about the numbers - roughly 10 billion data points of price history across stocks, indices, currencies, commodities and crypto.
Seventy independent views of the same market. When they agree you see it; when they disagree, that tells you something too. Disagreement is information that a single confident paragraph structurally cannot carry, and on most days it is the more useful half of the answer.
The part that matters most is the checking. Every expert is tested on millions of past moves it never saw while training, and the chances it states are compared with what actually happened afterwards and corrected where they drift. That loop is the whole difference between a number that means something and a number that sounds like it does, and no amount of eloquence substitutes for it.
It also produces a far less exciting product. Markets are efficient - whatever is known is already in the price, so most forecasts sit near 50/50. A system that has actually been scored will say so most of the time, where a system built to sound insightful will always find something to say.
What language models genuinely are good at
Language. They summarise a filing faster than you can open it, explain an unfamiliar term without condescension, draft the note you were dreading, turn a dense disclosure into something a person can act on, and hold a conversation about what a company said and did not say. Those are real capabilities and they are useful in and around investing.
The failure appears only when a tool built for words is asked to do arithmetic on the future, and answers anyway, in the same steady voice it uses for everything else. The trouble is not that the answer is wrong. It is that the answer is indistinguishable, in tone, from one with evidence behind it.
What to ask of anything that claims to see ahead
Four questions separate a forecast from a fluent paragraph, and they work on any tool, this one included.
- Is a chance attached, and is it labelled with what it measures?
- Is there a distance and a duration, so the claim can be scored rather than reinterpreted later?
- Are both directions shown, or only the one you were hoping for?
- Is there a public record of how the same kind of claim performed, and does it admit when the sample is too thin to judge?
Kunkafa's answers to those four are on the Kunkafa results page, updated daily, with how the forecasts are produced and scored on the Kunkafa methodology page and the common questions on the Kunkafa FAQ. Two related pieces go further into the numbers side: tabular data, the AI problem nobody talks about, and how we trained an AI on 10 billion data points.
Next time someone says AI can call the market, the question is not how clever the AI is. It is what the AI was trained on, and who checked it afterwards. Forecast, not advice. For the rational investor: emotion out, scenarios in.
Continue Reading
Tabular data: the AI problem nobody talks about
Prediction from structured, noisy numbers is harder than it looks and language models do not solve it. Why markets are the proving ground.
Probability, Not Prophecy: what an honest forecast looks like
The tagline as a design rule: the other side always a glance away, a chance for every level, a record beside it, and a 50/50 said out loud.