September 17, 2026
A 1965 paper argued that the first genuinely ultra-intelligent machine would be the last thing humans ever needed to invent. The logic is seductive: once a system is good enough to improve itself, every improvement makes it better at improving, and the curve runs away from you. The idea has a name, recursive self-improvement, and it has been the field's favourite daydream ever since.
Two papers landed in the past week that both claim to move it along. One, from 33 researchers across several Chinese labs, lays out a five-stage roadmap ending with an AI that rewrites the process it used to improve itself. The other, from Google DeepMind and the University of Maryland, takes a much narrower swing, and we think a more interesting one.
It improves the loop without ever touching the model.
Every AI mathematics result this year has run the same shape of loop, popularised by AlphaEvolve: hand a coding agent a problem and a scoring function, then let it propose, evaluate, read the feedback and try again a few thousand times. Reported progress on the Jacobian conjecture, on Navier-Stokes and on the Riemann hypothesis all came out of harnesses like this.
The part nobody talks about is the exploration policy: the rule deciding what the agent tries next. If an attempt scores slightly better, do you build on it or start over with something that might be better still? If it crashes before producing anything at all, was the idea wrong or just the implementation?
Until now, that policy was hardcoded by whoever set up the loop.
The insight is almost embarrassingly practical. If you log everything from every attempt (the code it wrote, the score it got, whether it crashed), you can replay those logs against a brand new policy without running a single fresh experiment. The candidate policy reads the cached history and says where it would have gone instead.
Replay costs essentially nothing, so you can test thousands of candidate policies against the same run, keep whichever would have reached the best result in the fewest attempts, deploy that one on the next real run, log it too, and repeat. The paper calls this dreaming.
Pointed at eight problems spanning algorithm design and mathematics.
Benchmarked against the identical setup running a fixed policy.
The standout: a LASSO solver beating Python's standard machine learning library in roughly 300 attempts, against 550 for the static policy and a previously reported 51,000 for the prior record.
Our favourite detail is the prompt, which reportedly pleads with the agent to read every past attempt before writing any code, to stop making tiny tweaks to the same idea over and over, and to please not kill any processes. Frontier research, same as the rest of us.
Roughly 300 attempts where the previous record needed about 51,000. The model never changed. The search did.
By the 1965 definition, no. The whole point there is that the thing doing the improving gets smarter each round. Here, the model writing each new exploration policy is the same model throughout, so it can never find a solution it was not already capable of writing. It just finds it faster, with far less waste.
That caveat applies to every AI maths result this year, though. Static weights, wrapped in a custom harness, with sub-agents, swarms and orchestration doing the heavy lifting. The weights only improve when a human goes back and trains the next model on whatever the swarm turned up.
Call it a very good search algorithm with caching. That is not nothing.
This is the part we keep coming back to, because it matches what we see in delivery work: the gap between a mediocre agent and a good one is rarely the model. It is the harness around it.
Log everything, not just the wins. Failed attempts, crash traces and mediocre scores are the training data for a better search. Most teams throw them away.
Treat the exploration policy as a component. If it is buried in a prompt nobody versions, you cannot improve it. Make it a thing you can swap, measure and roll back.
Replay beats retrying. Evaluating a strategy against cached history is close to free. Running it live is not. That ratio should shape how you test agent behaviour.
Keep a person at the top of the loop. Nothing here removes the human who decides which problem is worth 300 attempts in the first place.
None of this requires a frontier lab budget. Structured logs, a replayable evaluation harness and the discipline to version your search strategy are ordinary engineering, and they are where most of the gains in agentic work actually come from.