Thursday, July 23, 2026

News

Not Just RedNote: More AI Models Score Perfect Marks at Math Olympiad

ModelsPatryk Raba
Not Just RedNote: More AI Models Score Perfect Marks at Math Olympiad
Fot. Z3144228, Wikimedia Commons (CC BY-SA 4.0)

An independent test revealed that alongside China's RedNote, models from OpenAI, Anthropic, Kimi K3, and startup Axiom Math also scored 42 out of 42 points at the 2026 International Mathematical Olympiad. Huawei joined the group too, with its Celia system.

Contents
  1. Test Independent of the Developers
  2. Huawei Adds Its Own Result
  3. Three Years of Progress
  4. What It Means for the Industry

It turns out the result posted by China's RedNote, the first to announce a perfect score at the 2026 International Mathematical Olympiad (IMO), was not an isolated case. An independent test conducted by a Menlo Ventures investor showed that the exact same result, 42 out of 42 points, was also achieved by models from OpenAI, Anthropic, China's Moonshot AI, and startup Axiom Math. On top of that came Huawei's Celia system, which the company announced separately.

Test Independent of the Developers

The key difference from RedNote's earlier announcement lies in where the information about the other four models came from. Deedy Das, a partner at Menlo Ventures, personally gave this year's IMO problems to four leading AI models and checked their solutions. All four, the OpenAI and Anthropic models, startup Axiom Math, and Moonshot AI's Kimi K3, achieved the maximum score.

The AI capability frontier has officially moved well past IMO-level math - Deedy Das, partner, Menlo Ventures

This is an unofficial test, different from the process RedNote went through, which submitted its result directly as a participant in the AI systems round. Still, the result is consistent with what the company announced, and shows that achieving a perfect score was not a one-off feat by a single team but a phenomenon spanning several of the industry's top players at once.

Huawei Adds Its Own Result

Separately, Huawei also announced a perfect score, with its system called Celia said to have demonstrated well-rounded problem-solving abilities across various branches of mathematics. It is now the second Chinese tech giant besides RedNote to claim a maximum IMO score this year.

RedNote's dots-note-3.0 model, from the company known for the Xiaohongshu platform, took part in the IMO for the first time and immediately scored a perfect result. The company stressed that no large language model had previously achieved a perfect score under the IMO's official grading process, which counts not just a correct answer but a rigorous proof free of logical errors.

Three Years of Progress

A comparison of results from the last three editions shows the pace at which language models have been catching up to the best human mathematicians. In 2024, Google needed 2-3 days to solve 4 of 6 problems, a silver-medal-level result. A year later, Google and OpenAI models reached gold-medal level, 35 out of 42 points, but still trailed the five human contestants who scored a perfect result. In 2026, that line was crossed by several systems at once.

The IMO competition in Shanghai brought together 666 contestants from various countries, all under the age of 20. Only seven of them scored the maximum 42 points, showing how difficult the olympiad problems remain even for the most talented young mathematicians in the world.

What It Means for the Industry

For AI companies, the math olympiad has for years served as an informal frontier benchmark, showing models' ability to reason through multiple steps and construct proofs rather than just produce ready-made answers. The fact that in 2026 six different systems from labs in the US and China achieved a perfect score at the same time suggests that the math abilities of top models converged faster than last year's results indicated.

One distinction matters here: the RedNote and Huawei results are official statements from the companies themselves, while the OpenAI, Anthropic, Axiom Math, and Kimi K3 results come from an independent test run by an investor, not from an official submission by those labs to the IMO's AI round. That difference in methodology, however, doesn't change the overall picture: the line between human and machine performance at this level of mathematics has just blurred.

Share: