What happens when a language model turns out to be a better financial analyst than most investors — and what that means for price discovery, market efficiency, and the advisors caught in between.
A working paper from the University of Florida is quietly rewriting what the industry thought it knew about both artificial intelligence and market efficiency. In "Can ChatGPT Forecast Stock Price Movements? Return Predictability and Large Language Models"1 (first version April 6, 2023; updated April 9, 2026), Alejandro Lopez-Lira and Yuehua Tang present a finding that is, by their own measure, not obvious: a general-purpose language model, trained without a single day of explicit financial supervision, can correctly identify the direction of the stock market's immediate reaction to a news headline roughly 90% of the time. The fact that it does so — reliably, at scale, across 4,123 U.S. common stocks from October 2021 to May 2024 — is not a party trick. It is a window into how markets actually process information, and where they fall short.
The Finding That Changes the Question
The research does not ask whether ChatGPT is a good trading tool. The more important question is what GPT-4's performance reveals about the market itself. Lopez-Lira and Tang put it plainly: "this high accuracy provides a unique window into market information processing." By comparing what GPT-4 identifies as the economic implications of a news headline with how markets actually respond over time, the authors can pinpoint exactly where prices underreact, and why.
The setup is elegant. Headlines for each company are fed into GPT-4 with a prompt asking the model to assess whether the news is good, bad, or neutral for the stock price in the short term. Scored as +1, -1, or 0, these signals are then tested against actual market returns. The out-of-sample integrity is unimpeachable: the model version used (gpt-4-0314) had a training cutoff of September 2021, and the sample begins in October 2021.
The results are striking. GPT-4 achieves daily portfolio hit rates of 93.3% for overnight news and 88.8% for intraday news in predicting the direction of the initial market reaction. If those initial reactions could be traded — by an insider, hypothetically — the returns would be 3.06% for overnight news and 4.44% for intraday. That is not the tradable result. What is tradable, for investors with sufficiently low transaction costs, is the subsequent drift.
Underreaction Is Not Random — It Has an Address
After the initial price move, stocks continue drifting in the direction GPT-4 predicted for one to two additional trading days. A long-short strategy based on GPT-4 assessments of overnight news generates average returns of 34 basis points per day before transaction costs, with an annualized Sharpe ratio of 2.97. The drift is more pronounced for negative news and for smaller stocks — both patterns that Lopez-Lira and Tang trace directly to limits to arbitrage. Short positions are costly and constrained; small-cap markets have fewer sophisticated participants to correct mispricings quickly. "For negative news," the authors observe, "shorting faces higher costs and regulatory restrictions, which limit the ability of attentive arbitrageurs to correct mispricings."
Importantly, not all news is created equal when it comes to underreaction. Using topic modeling to cluster headlines, the authors find substantial heterogeneity in how markets process different information types. Earnings reports, strategic partnership announcements, and clinical trial results are processed efficiently: GPT-4's initial alignment is strong, but there is minimal subsequent drift. The market gets those right, fast. In contrast, insider stock transactions, dividend announcements, and healthcare conference presentations show the underreaction pattern: strong initial alignment combined with significant subsequent drift. "Markets are slow to fully process and incorporate the information content for these news types," the authors write, "given that GPT-4 correctly identifies their implications from the outset."
Size Matters: The Threshold Effect
Perhaps the most consequential theoretical contribution is what the authors call the quality threshold. Not all language models perform equally. GPT-1, GPT-2, and BERT produce hit rates below 65% on initial reactions and negative Sharpe ratios on drift strategies. GPT-3.5 improves meaningfully, but still trails GPT-4 by a significant margin on every dimension. The annualized Sharpe ratios tell the story: 2.97 for GPT-4, 1.66 for GPT-3.5, 1.26 for DistilBART-MNLI, and negative for most basic models.
Lopez-Lira and Tang formalize this in their theoretical model: "there is a critical threshold in LLM sophistication, below which models cannot generate profitable predictions, and only sufficiently advanced LLMs demonstrate robust predictive power." The threshold exists because markets contain inherent frictions, including noise trader risk, transaction costs, and limits to information-processing capacity, that any predictive signal must overcome. Below the quality threshold, a model may extract some information from news, but its signal-to-noise ratio is too low to be actionable.
The topic-level comparison between GPT-4 and GPT-3.5 is revealing. The largest performance gaps appear precisely in the categories requiring the most analytical synthesis: insider stock transactions (25.2 basis point advantage for initial reactions), healthcare conference presentations (20.4 basis point drift advantage), and complex earnings and strategic announcements. Dividend announcements, by contrast, show virtually no difference between the two models. The pattern is consistent: "topics requiring deeper analytical reasoning, medical or scientific knowledge, and complex financial interpretation show the largest benefits from increased model sophistication."
The Market Is Getting Smarter — Because of Itself
The most unsettling finding, from a trading perspective, is the one with the clearest long-run implications. The strategy is losing its edge — and the authors believe LLM adoption is the reason. The annualized Sharpe ratio of the GPT-4 overnight strategy drops from 6.54 in 2021Q4 to 3.68 in 2022, 2.33 in 2023, and 1.22 over January through May 2024. "A key corollary concerns the potential erosion of LLMs' predictability," the authors note. "As high-quality LLMs become widely adopted, the very predictability they initially exploited may diminish or disappear."
This is markets working as theorized. As more participants use the same signal, their collective demand drives prices closer to what the model would have predicted, closing the gap before it can be harvested. The LLM's information advantage vanishes when everyone is looking at the same output. "Thus, the LLM's predictive success inherently contains the seeds of its own obsolescence as a trading signal." In a structurally important sense, the paper argues that AI adoption is actively improving price efficiency in real time — not in some hypothetical future, but right now, quarter by quarter, in the data.
5 Key Takeaways for Advisors and Investors
1. GPT-4 reliably identifies the economic content of news headlines, but the tradable edge lies in the subsequent drift, not the initial reaction. The initial reaction is not accessible; the one-to-two-day drift is — for participants with very low transaction costs. At 20 basis points round-trip, the strategy becomes unprofitable.
2. Markets process transparent, quantifiable information efficiently but systematically underreact to information requiring complex synthesis. Earnings announcements and clinical trials get priced in fast. Insider transactions, dividend signals, and specialized conference disclosures do not. Advisors should treat these categories differently in their information-processing frameworks.
3 .The edge belongs to the most sophisticated models. Basic LLMs and domain-specific models like FinBERT fail where GPT-4 succeeds. The threshold effect implies that below a certain level of model complexity, the analytical sophistication simply is not there. This mirrors the reasoning capacity available to most market participants.
4. Small-cap stocks exhibit meaningfully more underreaction than large-caps. The asymmetry is largest for negative news in small-cap names where short-selling constraints are binding. Advisors managing small-cap exposure should be especially attentive to how and when that information gets priced.
5. The return predictability from LLMs is declining as adoption rises, which is the expected outcome of improved market efficiency. The structural implication is not that AI creates persistent alpha — it is that AI accelerates price discovery. For long-term investors and advisors, this is constructive: markets are getting better at reflecting fundamental information, faster.
Footnote:
1 Lopez-Lira, Alejandro, and Yuehua Tang. "Can ChatGPT Forecast Stock Price Movements? Return Predictability and Large Language Models." Social Science Research Network, 6 Apr. 2023, updated 9 Apr. 2026, https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4412788.