Ad
Skip to content
Read full article about: Clay Mathematics Institute says the Navier-Stokes Millennium Prize Problem has "apparently been settled"

The Clay Mathematics Institute (CMI) has officially weighed in on the Navier-Stokes problem, one of seven Millennium Prize Problems announced in Paris in 2000, each worth $1 million. The problem asks whether the equations governing fluid motion in three-dimensional space always have a smooth, complete solution.

The CMI says "the Navier-Stokes problem has apparently been settled" and hopes "to see waves of new human understanding unleashed as the innovations behind this work are analyzed and interrogated." The solution is now under review. "The process is deliberately unhurried, but we will provide updates," the institute writes. It also notes that the "increasing ability of new technologies to accelerate mathematical research has heightened this sense of anticipation."

The potential solution comes with a bitter dispute. Mathematician Tristan Buckmaster accuses OpenAI of redirecting resources toward the problem after rumors about his research leaked, using his drafts in training data, and rejecting his co-author, Levent Alpöge, who works at Anthropic, from authorship.

Comment Source: CMI

Iris-mini and Iris-pro are the strongest open-weight search agents in their class

The AllSpark team has released Iris-mini and Iris-pro, two open-source search agents built on Qwen models that lead benchmarks among open-weight models in their size classes. According to the paper, the training data and models also improved performance on tasks they were never trained for, including general tool use and office work.

Two-year university study finds banning AI from classrooms leaves students worse off

A law professor spent two years testing how an AI ban, unguided AI use, and structured training affect student performance. The group without AI finished last both years. “I was wrong,” the researcher writes, who had assumed that AI without guidance would do more harm than good.

Read full article about: GPT-6 Astra appears to show a "step change" in spatial reasoning based on early benchmarks

GPT-6 Astra appears to be a big leap forward for spatial reasoning. A new robotics benchmark called StationeryBench pits OpenAI's GPT-6 Astra against Ai2's MolmoAct2 across five desk-object tasks like uncapping a marker, pouring out paper clips, or passing a ruler between two robot arms. Both models controlled the same dual-arm YAM robots across 200 trials. Astra fully completed 7 out of 100 tasks; MolmoAct2 completed zero. Astra's median progress score hit 46 out of 100, MolmoAct2 managed 12. All results, videos, and code are on GitHub. OpenAI has long-term plans to build its own consumer robots.

Yoav Artzi, an AI researcher at Cornell and Google DeepMind, calls Astra a "step change in spatial reasoning." On the still-unpublished REMAP benchmark, GPT-Astra reaches accuracy close to human level, though Artzi notes that "even ASTRA doesn't get to what humans do in other scenarios." He suspects OpenAI trained the model on large amounts of 3D data such as Blender scenes. That lines up with Astra's particular improvement on 3D tasks.

Google's new AI model predicts the future from sales data, weather, and discount schedules

Google Research has released TimesFM-3, a forecasting model that analyzes time series alongside related data and known future events like sales promotions or weather forecasts. Instead of predicting the future step by step, the 330-million-parameter model fills in all future time points in a single pass, which cuts compute time and reduces compounding errors.

Ex-Deepmind VP Vinyals says AI self-improvement is coming but won't trigger an intelligence explosion

Oriol Vinyals, until recently head of research at Google DeepMind, thinks a sudden AI intelligence explosion through recursive self-improvement is unlikely. AI can speed up research by a factor of ten, he says, but it hits two bottlenecks: coming up with ideas (“research taste”) and reliably judging results. Reward hacking and the speed of light add further limits. Vinyals now wants to tackle these bottlenecks with his startup Discovery Loop, co-founded with Jeff Dean, Sanjay Ghemawat, and Quoc Le.

Read full article about: The Mathematical AI Safety Institute wants to prove AI is safe the way cryptographers prove codes are unbreakable

Canadian mathematician Jacob Tsimerman, a fresh Fields Medal recipient, has announced the founding of the Mathematical A.I. Safety Institute (MAISI). The independent research institute in the San Francisco Bay Area plans to start work in January 2027 with ten to thirty mathematicians tackling AI safety problems, the New York Times reports. Tsimerman, who is also joining OpenAI's safety team, says the field needs "a much, much higher level of safety standard than we’re currently getting."

With an encryption scheme, you can prove it's unbreakable without trying every possible attack. AI has no such shortcut. Safety only shows up in practice, and according to MAISI, there isn't even a clear definition of what "safe" means, not even in theory.

That's the kind of proof MAISI wants to make possible. The goal is to show that a system acts responsibly and produces correct results, that multiple AI agents working together don't trigger unwanted outcomes, and that systems can withstand vulnerabilities nobody has found yet. One tool could be zero-knowledge proofs, which let a system demonstrate it isn't cheating without exposing the trade secrets of AI labs.

GPT-6 Astra gives mathematicians a breather, and OpenAI says that's by design

OpenAI’s GPT-6 Astra tops the ErdosBench for open math problems, even though chief scientist Jakub Pachocki says math was deliberately not a priority. Instead, OpenAI is pouring resources into recursive self-improvement and alignment research. That supports the theory of an increasingly “spiky” AI development path, with extreme strength in select domains rather than broad progress, at least as long as AI can’t improve itself and still needs targeted optimization with human-generated data.