AI Projects · Technical writing and AI-assisted research
Research to Readable
FinanceQA, a January 2025 benchmark of how AI models handle financial analysis, turned into a plain-language explainer, a LinkedIn post and a five-post X thread.
Every figure links to a numbered source, so you can check what the study found and where its evidence stops. Researched and drafted with AI assistance; How I Build With AI explains the process.
When the Number Isn't in the Filing: What FinanceQA Tested
Based on the FinanceQA paper (arXiv:2501.18062v1, January 2025) and its public dataset as observed October 4, 2026. Not commissioned by or affiliated with AfterQuery. Not investment advice.
The question
A company's annual report can hold nearly every number an analyst needs and still leave one out. Costco's 10-K, for example, reports variable lease costs but not the variable lease assets that go with them [1].
FinanceQA, a benchmark published by AfterQuery researchers in January 2025, turns that kind of gap into a test [1]. One of its tasks asks for an estimate of those lease assets. The study's question was practical: can large language models do the everyday analysis expected of junior investment professionals, including the parts where the documents run out?
What the study tested
The public release contains 148 questions [2][3]. The paper doesn't state the size of its evaluation set, though its reported results are consistent with this release. According to the authors, the questions were written by annotators with experience at hedge funds, private equity firms or investment banks, without help from language models, and every answer was checked by hand [1]. The release groups them into three types [3]:
- 38 basic tactical questions: calculations from excerpts of Costco's annual report.
- 46 assumption-based tactical questions: the same filing, but the answer needs a figure it doesn't state, so an assumption has to be supplied and labeled.
- 64 conceptual questions with no document at all, such as working out an EBITDA margin from two valuation multiples.
Each model had one attempt, with a prompt asking for a concise answer. Scoring was strict: an answer counted only if it matched exactly and reflected the assumptions and reasoning an analyst would use. There was no partial credit [1].
What the results showed
Four models were tested: OpenAI's GPT-4o and o1, Anthropic's Claude 3.5 Sonnet, and Meta's Llama 3.3 70B Instruct. Overall accuracy ranged from 31.1% for Llama to 48.7% for o1 [1]. The abstract rounds this to models "failing approximately 60% of realistic tasks"; the table behind it runs from about 51% wrong for o1 to about 69% for Llama.
The pattern by question type says more than the totals. Conceptual questions went best, with scores from about 52% to about 86%. Basic tactical questions landed between roughly 32% and 45%. On assumption questions, every model scored between 2.2% and 4.3% [1].
So the models did best where a question resembled a textbook problem, and worst where the work depended on noticing that something was missing.
Why the assumption questions are hard
Go back to the lease example. The method the paper describes is to assume that variable lease assets relate to operating lease assets in the same proportion that variable lease costs relate to operating lease costs, then note that assumption [1]. None of those steps is exotic. The hard part is that the question asks for an estimate without saying which figure is missing or how to fill it. The analyst has to recognize the gap, pick a defensible method and make the choice visible to whoever relies on the number.
The authors describe a related failure: every model missed a question about adjusted EBITDA for 2024. In their account, the models tended to apply a standard formula and left out adjustments an analyst would make, such as adding back operating lease costs under current lease accounting [1]. One missed adjustment is enough to change the final figure.
What fine-tuning changed
The authors also fine-tuned GPT-4o on 9,078 training examples, mostly synthetic variations of separately written human questions [1]. On the same question set, its overall accuracy rose from 39.2% to 56.8%, and its assumption score went from 2.2% to 15.2% [1].
That is a clear improvement, and the fine-tuned model still answers most assumption questions incorrectly under the study's grading. It's also worth knowing that AfterQuery describes itself as curating data for foundation-model development [4]. The fine-tuning result supports the paper's case for better training data, but it comes from one run on one benchmark.
What the evidence can and can't tell you
FinanceQA is useful because it tests a skill that simpler question sets can skip: reasoning when information is incomplete. Within its setup, the weakness shows up across all four models.
Its limits are just as specific. Every tactical question comes from one company's filing, and the authors note the results may not carry over to industries with different financial structures. Spreadsheet modeling, where much financial work happens, wasn't tested [1]. Each model had a single attempt under strict grading. The assumption group is also small: the public release has 46 such questions, so one answer is worth about two percentage points. And the four models were among those available in early 2025; newer systems weren't part of the study.
The finding is narrower than a verdict on AI in finance, and more practical. When an answer depends on an assumption, check whether the model made one and said so.
LinkedIn adaptation
Speculative adaptation for Dorian Hartwood's own LinkedIn; not posted on the platform.
A filing can contain every number an analyst needs except one.
That gap is what FinanceQA, a January 2025 benchmark from AfterQuery, was built to test. Its questions were written by people with hedge fund, private equity and investment banking experience. Some are calculations from Costco's annual report. Some can't be answered without an assumption the filing doesn't supply. The rest are conceptual.
The four models tested, GPT-4o, o1, Claude 3.5 Sonnet and Llama 3.3 70B, did best on the conceptual questions. On questions that needed an assumption, each scored under 5%.
Two caveats matter. Every tactical question comes from a single company, and the four models were among those available in early 2025. Newer systems weren't tested.
The habit I'd take from it applies to any AI-assisted analysis: when a number depends on an assumption, look for the assumption. If it isn't stated, the work isn't finished.
Paper: arxiv.org/abs/2501.18062
X adaptation
Speculative adaptation for a neutral research account, written in the third person; not posted on the platform. It doesn't speak for AfterQuery or its founders.
1/5 When the answer isn't in the filing, an analyst makes an assumption and says so. FinanceQA, a January 2025 benchmark, tested whether language models do the same. Across four models, that's where they struggled most.
2/5 The benchmark: questions written by people with hedge fund, PE and banking experience. Basic calculations and assumption questions use Costco's 10-K; conceptual questions come with no document.
3/5 One attempt each, exact-match grading. Overall scores ran from 31.1% to 48.7%. Conceptual questions went best. On assumption questions, GPT-4o, o1, Claude 3.5 Sonnet and Llama 3.3 70B each scored under 5%.
4/5 Estimating Costco's variable lease assets means noticing the filing doesn't report them, choosing a defensible ratio and stating it. The question doesn't say which figure is missing. The authors' fine-tuned GPT-4o reached 15.2% on those questions, up from 2.2%.
5/5 Limits worth keeping: one company's filing, models from early 2025, no spreadsheet modeling. Paper: arxiv.org/abs/2501.18062
References
- Mateega, S., Georgescu, C., & Tang, D. (2025). FinanceQA: A Benchmark for Evaluating Financial Analysis Capabilities of Large Language Models. arXiv:2501.18062v1, submitted January 30, 2025. AfterQuery. ↩
- AfterQuery. FinanceQA dataset card, Hugging Face, Apache-2.0. Viewed October 4, 2026. ↩
- Hugging Face. FinanceQA dataset viewer, test split: 148 rows (38 basic, 46 assumption, 64 conceptual), with category counts from the viewer's statistics service. Viewed October 4, 2026 (Pacific time). ↩
- Y Combinator. AfterQuery company profile ("Applied research lab curating data solutions for foundation model development"). Viewed October 4, 2026. ↩