0:00
/
Generate transcript
A transcript unlocks clips, previews, and editing.

Will AI build itself in just 2 years?

A recording from Kobe Yank-Jacobs' live video

Both OpenAI and Anthropic have suggested that, in two years or less, they may no longer depend on their employees to improve ChatGPT and Claude. Instead, the work will be done by AI agents themselves.

“You can imagine this as being equivalent to Anthropic [today], except it has 100 or 1,000 times more researchers than the company actually has,” Epoch AI senior researcher JS Denain told me, working through the hypothetical scenario in a Substack Live with The Argument Wednesday. “And also, all those researchers are working at 10 times, 100 times the speed.”

That scenario, which AI researchers call recursive self-improvement, would not necessarily mean that humans lose control of AI. That said, most of the scary, sci-fi-adjacent scenarios do at least start from this premise.

So, just how close are we to self-improving AI?

One way to try to answer the question is to look at the recent overall arc of AI progress: How long and complex are the tasks that AI completes? How many resources are dedicated to this project? In both of those cases, progress has grown exponentially, lending weight to bold predictions about what could happen if such trends continue.

Another thing you could do is look at what’s happening inside the labs: OpenAI recently had its latest model optimize a smaller model. Anthropic reported that it tested Claude’s decision-making at certain critical junctures in the AI research process. Successive models made better and better decisions until Claude Mythos Preview beat human researchers 64% of the time.

Successfully outsourcing AI research decisions to Mythos would be, to some, an obvious sign that agents could lead research soon. With this evidence in mind, aggressive predictions that we’ll have automated AI researchers in just two years do not sound quite so far-fetched.

(Source: Anthropic)

But making key decisions isn’t the same as running a whole scientific research process. That involves novel thinking, critical judgment, and resource management.

To figure out whether AI is up to these tasks, Denain has proposed treating AI research the way economists treat any other automatable job in the economy: Itemize all the tasks involved in the job, then figure out how well AI can do each of them specifically.

In other words, what does it actually look like to be a scientist who is building AI? As I put it in our conversation: “What is the equivalent of an AI researcher dropping a drop of solution like a chemist does to run their own experiment?”

Denain and I tried to separate things AI currently does well in the research process from the things it can’t yet do well.

It’s common to say that AI succeeds at defined tasks but struggles at more open-ended ones. This discussion shows there are still layers of complexity below that. You have to ask what particular aspects of open-ended tasks AI struggles with: Can it set new research directions? Can it incorporate critiques? Can it manage a budget?

Recent research suggests that, for all three, the answer, so far, is “no.” This complicates the picture of steady progress coming out of the AI labs. Watch the interview to learn more about why.

Discussion about this video

User's avatar

Ready for more?