Why Should Large Language Models Be Good at Mathematics?
Published:
Why Should Large Language Models Be Good at Mathematics?
If we expect large language models to become general cognitive systems rather than extraordinarily powerful language imitators, then mathematics should not be treated as merely one more subject they ought to master.
Mathematics is one of the cleanest tests of whether a model possesses rationality in the real world.
A model may be remarkably knowledgeable, fluent, and persuasive while still being weak at mathematics. But in that case, it remains difficult to know whether the model has learned principles of reasoning that extend beyond statistical regularities in language.
The reason is simple: mathematics demands something fundamentally different from plausible prediction.
Natural language largely teaches a model what the world tends to look like. Mathematics asks what must follow from a set of assumptions.
In ordinary language modeling, we are often interested in something like \(P(y\mid x),\)
the probability that one statement, token, or event follows another.
Mathematics instead frequently takes the form
\[A \Rightarrow B.\]This does not mean that (B) is merely likely given (A). It means that, under the relevant axioms and rules, (B) cannot fail to follow.
This is a transition from plausibility to necessity.
And that transition is central to rationality.
Mathematics forces a model to discover structure
Suppose a model observes thousands of examples of addition: \(2+3=5,\qquad 7+8=15,\qquad 13+21=34.\) It could, in principle, memorize statistical regularities across those examples.
But mathematical understanding requires something else. The model must recover the operation behind them: \(+:\mathbb{N}\times\mathbb{N}\rightarrow\mathbb{N}.\) The individual numbers are incidental. The operation is invariant.
This distinction is profound because much of science works in exactly the same way. Scientific reasoning attempts to move from observations to invariants, and from invariants to laws: \(\text{observations} \rightarrow \text{invariants} \rightarrow \text{laws}.\) A mathematically capable model must therefore learn to distinguish surface variation from underlying structure.
In classical philosophical language, it must move from particulars to universals.
That is much closer to what we normally mean by abstract reasoning than simply predicting the next likely statement.
Mathematics is extreme compression
There is another way to see the same issue.
A system could memorize millions of arithmetic facts: \(1+1=2,\quad 1+2=3,\quad 1+3=4,\quad \ldots\) Or it could learn the rule that generates them.
The second representation is dramatically shorter.
In this sense, mathematics is one of the most extreme forms of intellectual compression humans have developed. A small number of equations can explain enormous classes of phenomena.
Newtonian mechanics compresses vast families of physical trajectories. Maxwell’s equations compress a huge range of electromagnetic behavior. Probability theory compresses regularities across uncertain systems.
From an information-theoretic perspective, intelligence can often be viewed as the ability to discover short generative descriptions of large bodies of observation: \(\text{intelligence} \approx \text{finding compact programs that explain complex data}.\) Mathematical ability therefore tests whether a model can do more than store knowledge.
It tests whether it can compress knowledge into principles.
Mathematics allows knowledge beyond observed data
This may be the most important property of all.
Ordinary empirical learning roughly follows \(D\rightarrow f,\) where observations (D) are used to infer some pattern or function (f).
Mathematics allows a different process: \(\text{axioms} \rightarrow \text{deduction} \rightarrow \text{new knowledge}.\) A theorem does not need to have appeared explicitly in the training data for it to be derivable from known premises.
That gives mathematics an unusual epistemic power: \(\boxed{ \text{new knowledge without new empirical observation} }\) For artificial intelligence, this marks an important conceptual boundary.
A system that mainly interpolates between known examples remains constrained by the accumulated record of human experience.
A system capable of constructing formal abstractions and deriving their consequences can, at least in principle, generate conclusions that have never previously been written down.
At that point, the system begins to resemble not merely a database or an assistant, but a researcher.
Mathematics is also a language of counterfactuals
Real intelligence is not only about answering:
What is the world like?
It must also answer:
What would happen if the world were different?
This transition, \(\text{What is} \rightarrow \text{What if},\) is fundamental to science, engineering, planning, economics, medicine, and causal reasoning.
Mathematics is naturally built around such questions.
Suppose \(x>0.\) What follows?
What changes if instead \(x<0?\) What happens if one assumption is removed? What happens when a boundary condition changes? Which conclusions remain invariant under a transformation?
Mathematical reasoning repeatedly trains the same cognitive pattern: \(\text{assumptions} \rightarrow \text{logical consequences}.\) That pattern is not confined to mathematics. It is one of the foundations of rational decision-making in the real world.
Mathematics makes it difficult to hide mistakes behind language
Large language models have a peculiar strength and weakness: they are exceptionally good at producing plausible text.
In natural language, an incorrect idea can still sound sophisticated.
A model can write an elegant paragraph about interacting social, economic, and historical factors while saying very little that is actually testable.
Mathematics is less forgiving.
If \(17\times 19=323,\) then the answer is 323.
If a proof contains an invalid inference, stylistic fluency cannot repair it.
Mathematics therefore imposes an unusually strict epistemic discipline: $$ [\text{validity}
\text{plausibility}, ]
and
[ \text{coherence}
\text{eloquence}.] $$ This is especially valuable for language models because their native objective rewards plausible continuation.
Mathematics asks them to learn something almost opposite:
Do not tell me what is most likely to come next. Tell me what must come next.
Mathematics may also be unusually close to the structure of reality
There is a final, more philosophical argument.
Modern science repeatedly discovers that deep physical regularities can be expressed through mathematical structures:
- symmetry,
- geometry,
- groups,
- topology,
- differential equations,
- probability,
- optimization,
- information.
Again and again, understanding a physical system means understanding not merely the objects in it, but the relations and transformations that remain stable.
This is close to the philosophical position known as structural realism: perhaps what science captures most reliably is not the intrinsic essence of things, but the structure of relations between them.
If something like this is true, then a sufficiently deep model of reality cannot consist merely of a catalogue of objects and descriptions.
It must understand \(\text{relations}, \quad \text{constraints}, \quad \text{symmetries}, \quad \text{transformations}.\) Mathematics is our most precise language for exactly these things.
Mathematical intelligence is not arithmetic performance
For this reason, being “good at mathematics” should not mean scoring highly on arithmetic benchmarks, solving routine integrals, or answering competition problems.
Those are useful measurements, but they are proxies.
What we ultimately care about is something deeper: $$ \boxed{ \text{Mathematical Intelligence}=
\text{Abstraction} + \text{Invariance} + \text{Constraint} + \text{Deduction} + \text{Counterfactual Reasoning} + \text{Compression} } $$ These capacities are not merely mathematical skills.
They are much of the architecture of general reasoning itself.
We therefore need mathematically capable large language models not because the real world is filled with mathematics exercises.
Quite the opposite.
We need them because the most important problems in the real world rarely come with an answer key.
When an answer cannot simply be retrieved from past human experience, an intelligent system needs something deeper than the ability to reproduce what people usually say.
It needs to reason over structure.
Language allows machines to enter human culture.
Mathematics may be what allows them to enter rationality itself.
