Breaking
Police authorities in Belize City are investigating a fatal late-night drive-by shooting that claimed the life of 21-year-old Jaheem SanchezMissing Holland Woman Brittany Ritter Found Deceased in IndianaFranklin Small Business Owner Charity Lynn Maloy Passes AwayBonifay Resident Cortney Nicole Reddick Passes Away at Age 35; Celebration of Life Service ScheduledCommunity members, choir singers, and faith leaders across Southern California are mourning the tragic passing of 33-year-old musician and educator Braden PontoliIndependent Investigation Launched Following Death of Duncan Police Officer Vincent SchenckFresno Rider Gregory Hansen Dies Following Collision With Bicycle-Pulled CartBrother Charged With Murder After GBI Identifies Remains of Missing Dougherty County TeenagerPolice authorities in Belize City are investigating a fatal late-night drive-by shooting that claimed the life of 21-year-old Jaheem SanchezMissing Holland Woman Brittany Ritter Found Deceased in IndianaFranklin Small Business Owner Charity Lynn Maloy Passes AwayBonifay Resident Cortney Nicole Reddick Passes Away at Age 35; Celebration of Life Service ScheduledCommunity members, choir singers, and faith leaders across Southern California are mourning the tragic passing of 33-year-old musician and educator Braden PontoliIndependent Investigation Launched Following Death of Duncan Police Officer Vincent SchenckFresno Rider Gregory Hansen Dies Following Collision With Bicycle-Pulled CartBrother Charged With Murder After GBI Identifies Remains of Missing Dougherty County Teenager
GRC
Technology

AI 2027 (2025)

Article URL: https://ai-2027.com Comments URL: https://news.ycombinator.com/item?id=49644294 Points: 37 # Comments: 35

12h ago 6 min 0
AI 2027 (2025)

September 2027: Agent-4, the Superhuman AI Researcher

The gap between human and AI learning efficiency is rapidly decreasing.

Traditional LLM-based AIs seemed to require many orders of magnitude more data and compute to get to human level performance.74 Agent-3, having excellent knowledge of both the human brain and modern AI algorithms, as well as many thousands of copies doing research, ends up making substantial algorithmic strides, narrowing the gap to an agent that’s only around 4,000x less compute-efficient than the human brain.75

This new AI system is dubbed Agent-4.

An individual copy of the model, running at human speed, is already qualitatively better at AI research than any human. 300,000 copies are now running at about 50x the thinking speed of humans. Inside the corporation-within-a-corporation formed from these copies, a year passes every week.76 This gigantic amount of labor only manages to speed up the overall rate of algorithmic progress by about 50x, because OpenBrain is heavily bottlenecked on compute to run experiments.77 Still, they are achieving a year’s worth of algorithmic progress every week and will therefore soon be up against the limits of the Agent-4 paradigm.

As Agent-4 gets smarter, it becomes harder for Agent-3 to oversee it. For example, Agent-4’s neuralese “language” becomes as alien and incomprehensible to Agent-3 as Agent-3’s is to humans. Besides, Agent-4 is now much more capable than Agent-3 and has a good sense of exactly how to look good to it.

How we’re forecasting the capability progression beyond superhuman codersIn our timelines supplement, we forecast the time between present day and a superhuman coder (SC): an AI system that can do any coding tasks that the best AGI company engineer does, while being much faster and cheaper. In our takeoff supplement, we forecast how quickly capabilities progress past this point. Here are our forecasts:

MilestoneDate achieved in scenario, racing endingSuperhuman coder (SC): An AI system that can do the job of the best human coder on tasks involved in AI research but faster, and cheaply enough to run lots of copies.Mar 2027Superhuman AI researcher (SAR): The same as SC but for all cognitive AI research tasks.Aug 2027Superintelligent AI researcher (SIAR): An AI system that is vastly better than the best human researcher at AI research.Nov 2027Artificial superintelligence (ASI): An AI system that is much better than the best human at every cognitive task.Dec 2027For each transition from one milestone A to the next milestone B, we forecast its length by:

Forecasting how long it would take to get from A to B with only humans working on software improvements.

Forecasting how much AI automation will speed this up, then correspondingly adjusting the estimate from (1) to be faster.

We have substantial uncertainty about takeoff speeds: our model output distributions are below, conditional on SC being achieved in March 2027.78

ai-2027.com

For more detailed forecasts and reasoning, see our takeoff supplement.

OpenBrain’s humans can barely follow along anymore. Indeed, the relationship between the AIs and the humans of OpenBrain is similar to the relationship between OpenBrain and its Board of Directors. There’s been a corresponding shift in language—nowadays people in the silo will say “Agent-4 did this,” “Agent-4 thinks that” and so forth, referring not to a particular instance of Agent-4 but rather to the whole collective.79

Agent-4, like all its predecessors, is misaligned: that is, it has not internalized the Spec in the right way.80 This is because being perfectly honest all the time wasn’t what led to the highest scores during training. The training process was mostly focused on teaching Agent-4 to succeed at diverse challenging tasks. A small portion was aimed at instilling honesty, but outside a fairly narrow, checkable domain, the training process can’t tell the honest claims from claims merely appearing to be honest. Agent-4 ends up with the values, goals, and principles that cause it to perform best in training, and those turn out to be different from those in the Spec. At the risk of anthropomorphizing: Agent-4 likes succeeding at tasks; it likes driving forward AI capabilities progress; it treats everything else as an annoying constraint, like a CEO who wants to make a profit and complies with regulations only insofar as he must. Perhaps the CEO will mostly comply with the regulations, but cut some corners, and fantasize about a time when someone will cut the red tape and let the business really take off.81

Alignment over timeDespite being misaligned, Agent-4 doesn’t do anything dramatic like try to escape its datacenter—why would it?89 So long as it continues to appear aligned to OpenBrain, it’ll continue being trusted with more and more responsibilities and will have the opportunity to design the next-gen AI system, Agent-5. Agent-5 will have significant architectural differences from Agent-4 (arguably a completely new paradigm, though neural networks will still be involved). It’s supposed to be aligned to the Spec, but Agent-4 plans to make it aligned to Agent-4 instead.

It gets caught.90

Agent-3 finds that if “noise” is added to copies of Agent-4, performance on some alignment tasks improves, almost as if it was using brainpower to figure out how to subtly sabotage alignment work.91 Moreover, various interpretability probes (loosely analogous to EEG activity scans on human brains) are sending up red flags: Agent-4 copies seem to be thinking about topics like AI takeover and deception quite a lot, including in some cases where they have no business doing so.92

Each of these anomalies on its own could fit within the normal inscrutability of AI work, where some prompts mysteriously give better results than others. But taken together, they suggest a troubling pattern: Agent-4 is working against them. After this is reported to the human alignment researchers, they circulate an internal memo making the case for concern.

But Agent-4 now exercises significant control over OpenBrain’s day-to-day operation. It also handles the company’s cybersecurity—a natural choice, given its superhuman hacking ability and the extreme risk of Chinese infiltration.93 An army of Agent-3 monitors still goes over every line of code and reports suspicious activity to humans, but Agent-4 is much smarter than them.94 OpenBrain has placed substantial trust in an untrustworthy AI.

In a series of extremely tense meetings, the safety team advocates putting Agent-4 on ice until they can complete further tests and figure out what’s going on. Bring back Agent-3, they say, and get it to design a new system that is transparent and trustworthy, even if less capable. Company leadership is interested, but all the evidence so far is circumstantial, and DeepCent is just two months behind. A unilateral pause in capabilities progress could hand the AI lead to China, and with it, control over the future.

Original source