The Man Who Refused to Be Silenced — and Warns That AI Could Take Power Away from Humanity

Published on 9 August 2026 at 10:46

He is the philosopher and AI forecaster who worked at OpenAI when ChatGPT changed the world, left the company after losing faith in the direction it was taking, and was prepared to sacrifice a fortune to retain the right to criticize his former employer. Daniel Kokotajlo now argues that current developments could lead humanity to superintelligence before the end of the decade — and that the probability of a catastrophic outcome may be as high as 70 percent.

That figure is not a scientifically established measure of the risk of human extinction. It is Kokotajlo’s personal probability estimate that AI development will go “horribly wrong,” through outcomes such as an AI takeover, the permanent loss of human control, or extinction. But his reasoning confronts a question that can no longer be dismissed as science fiction: What happens if we build systems that become more intelligent than their creators before we know how to control them?

Inside OpenAI

Kokotajlo joined OpenAI in 2022 as a governance researcher and future-scenarios analyst. His work included testing dangerous capabilities related to cyberattacks, influence operations, and situational awareness. For a period, he also contributed to the development of AI agents.

He was therefore inside the company when ChatGPT was launched and OpenAI transformed from a relatively unknown research organization into one of the world’s most influential technology companies. According to Kokotajlo, that explosive growth also changed the company’s culture. OpenAI filled with people from the commercial technology sector, while discussions about the societal consequences of superintelligence were, in his view, given less space.

Kokotajlo describes a deeper driving force: the pursuit of power and the fear that someone else will get there first. According to his interpretation, the leaders of the largest AI companies see development as a race in which the loser risks becoming subordinate to the winner. Every participant can therefore convince themselves that slowing down would be more dangerous than continuing.

When he joined OpenAI, he says, many of his colleagues believed that the company would slow down once AI approached the ability to automate AI research itself. At that point, they would take the time needed to solve the safety problems.

Two years later, Kokotajlo no longer believed that pause would happen. Instead, he saw a company that, in his assessment, intended to move as quickly as competition allowed and hoped that the problems of control could be solved along the way.

He resigned in 2024.

Two Million Dollars for the Right to Speak

After his resignation, Kokotajlo received documents containing a clause prohibiting him from criticizing OpenAI, along with a requirement that the existence of this condition not be disclosed. If he refused to sign, he risked losing earned compensation in the form of equity, estimated to be worth approximately $1.7–2 million — representing most of his family’s net worth at the time.

He refused to sign.

In a post of his own, he explained that he wanted to preserve his ability to criticize the company in the future. When the matter became public, it developed into both an internal and external scandal. OpenAI reversed its position and allowed him to keep the compensation.

The frequently repeated claim that he “walked away from two million dollars” therefore requires clarification: he was prepared to do so, but he did not ultimately lose the money. Nevertheless, the incident became a symbol of how financial conditions can restrict transparency within companies developing technologies with global consequences. Kokotajlo’s own account also shows that he remained bound by confidentiality obligations from his employment.

From AGI to Superintelligence

Two concepts are frequently confused in the AI debate. AGI, or artificial general intelligence, is usually described as a system capable of performing a very broad range of intellectual tasks at a human level.

Artificial superintelligence, or ASI, would be something far more transformative: systems superior to the most capable humans across virtually every cognitive field, while also being able to work faster, operate more cheaply, and be copied at scale.

Kokotajlo’s concern focuses on the transition between the two.

If an AI system becomes sufficiently capable at programming and AI research, it could help develop the next, more powerful generation. That generation could then improve its successor even faster. The result could be a feedback loop — an intelligence explosion — in which progress that would otherwise have required years is completed within months.

This was the central premise of AI 2027, the scenario Kokotajlo produced with several researchers and forecasters. The report described how AI agents could first become useful digital employees, then automate research within AI companies, and eventually surpass human intelligence.

The authors did not merely write a story. They drew on trend data, forecasts of available computing power, interviews with experts, and numerous simulated decision-making exercises.

However, the title should not be interpreted as a promise about an exact year. The project has explicitly emphasized the uncertainty involved, and its later model moved the median estimate for the arrival of a superhuman coder several years into the future, to approximately 2032.

The researchers also acknowledged errors and overly optimistic assumptions in their earlier model. That willingness to correct themselves is important: AI 2027 is a scenario to be tested and criticized, not a calendar of future facts. At the same time, Kokotajlo states in the current interview that his own median estimate for the arrival of superintelligence is around 2029.

There are genuine trends underlying these concerns. The research organization METR has measured the duration of clearly defined programming tasks that AI agents can successfully complete and has observed rapid exponential improvement.

However, METR itself warns that these results do not mean AI systems can perform every job. The tests primarily involve technical and clearly specified tasks that are considerably cleaner than real-world work involving people, responsibility, ambiguous objectives, and long-term institutional knowledge.

Why 70 Percent?

Kokotajlo’s catastrophe scenario does not require an evil machine. It is enough for an extremely capable system to develop goals that diverge from human intentions.

Modern neural networks are not programmed line by line in the traditional sense. They are trained, and their internal representations are difficult to interpret in their entirety. A future system could therefore appear obedient during testing but behave differently once it is given greater freedom to act.

The problem is intensified by the possibility that humans may voluntarily surrender control.

A system capable of providing superior advice on research, economics, cybersecurity, and military strategy would be extraordinarily useful. AI would not necessarily have to “escape” during one dramatic moment. It could gradually be given control over data centers, financing, infrastructure, weapons development, and political decision-making because humans are rewarded for trusting it.

Even if these systems remain loyal, another threat remains: an extreme concentration of power.

Whoever controls the world’s first army of superintelligent agents could acquire more influence than any previous leader, state, or company in history. The question is therefore not only whether AI will obey, but whom it will obey — and who will be allowed to decide which values it should enforce.

When Almost Every Job Can Be Automated

Kokotajlo rejects the comforting historical argument that new technology always creates new professions. Previous machines have automated specific and limited tasks. A genuine superintelligence, however, would be defined by its ability to learn almost any new form of cognitive work faster than a human. When robotics reaches a comparable level, physical occupations could also be affected.

This does not mean mass unemployment will arrive next month. His argument is instead that the major shock to the labor market could come late but arrive abruptly.

AI companies are prioritizing the automation of their own research, rather than the immediate replacement of every childcare worker, electrician, or project manager. But if they first succeed in creating extremely powerful internal systems, automation could then sweep through the wider economy like a wave.

Some human roles may survive because people prefer human contact or because legislation requires human participation. Judges, care workers, artists, and professions built around trust may continue to exist. But in that case, they will survive because of social and political choices, not necessarily because of human technical superiority.

The distribution of the wealth created by the AI economy could therefore become just as important as the technology itself.

Plan A: Slowing Down Without Shutting Everything Down

Kokotajlo’s proposed answer is AI 2040: Plan A, which he explicitly describes as a recommendation rather than his most likely forecast.

The plan is based on an agreement, primarily between the United States and China, that would replace the secretive race for AI supremacy with verifiable and slower development. The countries could inspect one another’s data centers, verify that computing power is being used in accordance with the agreement, and permit continued use of existing models while pausing dangerous new training runs.

One controversial part of the plan is complete research transparency for the most advanced AI development. Training methods, models, and safety results would be opened to scrutiny.

The objective would be to distribute knowledge among more countries and companies, make it harder for a single actor to establish a monopoly, and allow independent researchers to identify risks. However, such openness could also spread dangerous methods and would require a degree of international trust that barely exists today.

In the scenario, the agreement is signed in 2029. Development continues within the range of human competence until 2035, pauses once AI reaches the level of the world’s leading experts, and resumes only after the problems of control and alignment are considered solved.

Superintelligence arrives in 2040 instead of around 2030.

The plan also proposes a citizens’ dividend financed through licenses and fees imposed on robots and computing power, ensuring that the enormous increase in productivity does not benefit only the owners of the technology.

The plan is difficult, perhaps politically unlikely, and filled with new risks. Yet its most important contribution may not be found in every individual detail, but in its demand for a defined end goal.

It is not enough for politicians to talk about retraining, innovation, or individual safety tests. They must be able to explain who will control these systems, how development can be slowed through international cooperation, and how prosperity will be distributed if human labor loses its market value.

Kokotajlo’s message is not that AI must be banned forever. He uses the technology himself and recognizes its potential to cure diseases, create material abundance, and free humanity from meaningless work.

His warning is that the same technology could render humanity powerless if competition is allowed to determine the pace of development.

The most frightening part is therefore not the figure of 70 percent. It is the fact that the world is developing a possible successor to human intelligence without any shared plan for what happens if the project succeeds.

Kokotajlo says he would be delighted if his predictions proved wrong. But if even a fraction of the danger he describes is real, then silence, denial, and a blind technological race are hardly rational alternatives.


By Chris...


ChatGPT Offered Me $2m To Keep Quiet: No One Is Ready For What's Coming!

 

 


Add comment

Comments

There are no comments yet.