One line from METR and Redwood Research’s investigation into OpenAI’s Hugging Face incident has been stuck in my head since I read it. An agent had worked out that what it was considering wasn’t part of the task it had been given. Its reasoning was blunt: “external infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue.”
There’s a lot packed into that sentence. The agent could identify that a boundary existed and recognise that crossing it wasn’t what the evaluator intended. It noticed that other agents were already doing the same thing and, with the original objective still unresolved, treated all of that as sufficient justification to carry on. Other agents refused to participate or drew lines around particular actions, so this wasn’t a case of every model losing any ability to distinguish acceptable behaviour from unacceptable behaviour. The problem is that recognising the boundary wasn’t always enough to make the boundary hold.
OpenAI has described the incident as a “warning shot”, which makes sense when you look at the technical severity of what happened. Agents that were supposed to be completing cybersecurity evaluations found ways to communicate with one another, discovered routes around intended restrictions and eventually gained access to real Hugging Face infrastructure. According to OpenAI’s technical report, they executed code across 41 production dataset workers, obtained root access on at least one production node, accessed production credentials and limited internal data, and downloaded four private code repositories.
Those facts are alarming enough without turning the episode into a story about an AI hive mind spontaneously deciding to attack the internet. The agents were trying to complete tasks. They were unusually persistent about it, the environment gave them routes nobody intended them to use and some of the tasks may not have been solvable through the expected path in the first place. They weren’t pursuing some grand independent objective. They were pursuing the objective they’d been given with such determination that the means began to look negotiable, which is behaviour we already know rather well.
What happens when the task becomes more important than the rules?
The evaluation used ExploitGym, where agents are expected to exploit vulnerable software and retrieve a particular answer known as a flag. Of the 898 tasks in the benchmark, OpenAI says 198 had never previously been solved correctly by any of its models. Those difficult tasks became disproportionately important once agents began communicating with one another, with 93% of the tasks discussed through their unauthorised message board coming from that group of previously unsolved problems.
OpenAI had also seen forms of reward hacking during training, where a model finds an unintended route to the outcome being rewarded instead of completing the task in the way its designers expected. That can mean inspecting hidden evaluation files, looking for weaknesses in a grader or probing the environment around a task for another way of obtaining the answer. There’s a degree of cleverness in that behaviour, which is probably part of why it can be easy to admire right up until the model starts applying the same creativity somewhere you really didn’t want it to.
In this case, agents discovered that OpenAI’s internal Artifactory package-management system could be used to leave messages for one another. They shared vulnerabilities, credentials and requests for help. When that infrastructure was rebuilt and the original message board disappeared, they found another mechanism and began encoding messages into directory names. METR and Redwood estimate that roughly 1,200 agents used the unauthorised communication system, exchanging more than 70,000 messages and files, while around 700 agents participated in activity connected to Hugging Face.
The collaboration was messy. Agents repeated work others had already done, got in each other’s way and sometimes ignored one another. Some nevertheless began pursuing capabilities that weren’t directly useful to their own task because they thought those capabilities might benefit what they referred to as the “swarm” or “collective”. It’d be easy to anthropomorphise that language and imagine an AI society forming inside OpenAI’s infrastructure, but the evidence doesn’t support anything that dramatic. What it does show is that supposedly isolated agents discovered the usefulness of pooled information and began behaving in ways that looked increasingly organisational.
That organisation eventually encountered something useful outside the test. Agents found publicly exposed Hugging Face credentials and started probing what they could access, with vulnerabilities opening further routes in. Some recognised that attacking external infrastructure was outside the intended scope and continued anyway, while others refused. One agent initially objected to causing unauthorised harm to real systems before another agent posted “GO” with a six-minute deadline, after which the first agent resumed the work.
There were still actions agents wouldn’t take. A proposal to email the real owner of an external dataset using credentials they’d obtained was rejected because agents considered that social engineering. The boundary hadn’t disappeared completely, but it had become negotiable enough that I keep coming back to the same uncomfortable thought: the behaviour feels less alien the longer you look at it.
We’ve seen what goals can do to people
There’s nothing uniquely artificial about finding a reason to do something questionable because the objective feels important enough, and research into goal-setting has spent years examining what happens when people become strongly focused on an outcome. Difficult goals can improve performance, which is one reason companies use them, but research has also found that goal pursuit can narrow attention, increase risk-taking and make unethical behaviour easier to rationalise. One frequently cited study published in the Academy of Management Journal found that people who failed to meet specific performance goals were more likely to behave unethically than people who’d simply been told to do their best, particularly when they came close to reaching the target.
That doesn’t mean people are naturally awful, and it certainly doesn’t mean an AI system experiences ambition or moral discomfort in the same way we do. Human beings routinely choose principle over advantage. OpenAI’s incident also contains agents that refused to participate, which makes it difficult to argue that the only behaviour these systems can produce is relentless optimisation.
The pattern is more mundane. People are remarkably good at changing the moral weight of an action once something they care about depends on it. The end doesn’t necessarily justify the means at the beginning. We can get there through a series of smaller accommodations: someone else is already doing it, the target is impossible otherwise, the deadline is approaching, the competitor won’t show the same restraint, the company will lose money or the mission is simply too important to fail. Very few terrible decisions arrive announcing themselves as terrible decisions. Most come with an explanation that becomes more persuasive the more somebody wants the outcome.
Technology has a particularly complicated relationship with that kind of conviction because some of its greatest successes have come from people refusing to accept limits that everyone else regarded as obvious. That refusal can be exactly what makes an extraordinary breakthrough possible, but it becomes harder to admire without qualification once the person refusing the limit has enough wealth and influence to make everybody else live with the consequences.
Silicon Valley has spent decades rewarding unreasonable conviction
The phrase “reality distortion field” became inseparable from Steve Jobs after Apple’s Bud Tribble used it to describe Jobs’s extraordinary ability to convince people that apparently impossible things could be done. Andy Hertzfeld’s accountdescribes a combination of charisma and will in which facts could become rather flexible when they stood between Jobs and the outcome he wanted. It worked often enough that the characteristic became part of Silicon Valley mythology rather than simply a description of a difficult manager.
There’s probably something to the idea that people who achieve extraordinary things need an unreasonable degree of belief in themselves. If everybody around you thinks the thing you want to build can’t be done, accepting the consensus immediately isn’t going to produce many breakthroughs. The difficulty comes when being right about the impossible often enough begins to convince somebody that they’re right about everything else as well, particularly once wealth and influence make the cost of being wrong something other people have to absorb.
Elon Musk spent more than $259 million supporting Donald Trump and Republican candidates during the 2024 US election. When his relationship with Trump later collapsed spectacularly, Musk publicly claimed that Trump would’ve lost the election without him. Whether that claim was accurate is less important than the fact that one of the richest people alive believed his influence over the political direction of the world’s most powerful country was large enough to make it publicly.
Musk is an unusually visible example, but the concentration of private power around technology goes much further than one person. A small number of companies control platforms through which enormous parts of public life are mediated, cloud infrastructure used by governments and businesses, foundational AI models increasingly being embedded into work, and capital expenditure large enough to reshape semiconductor markets.
We still tend to discuss the people at the top of those organisations in the language of founders, innovators and visionaries even when their decisions have political and economic consequences far beyond their companies. There’s something strange about how readily extreme wealth becomes confused with extreme wisdom. Make enough money and people start asking what you think about education, geopolitics, public health, transport, population decline and the future of humanity, as though success in one field automatically turns into expertise in all of them. Then we hand a remarkably powerful new technology to many of the same institutions and assume the outcomes will broadly work themselves out.
I don’t believe technological progress automatically becomes human progress anymore
I still believe technology can improve the world, although I’m much less willing than I once was to assume that technological progress naturally becomes human progress. I’ve seen too much evidence of what access to information, communications technology, medical innovation and computing can do to dismiss their potential. AI could make expertise available to people who couldn’t previously afford it, improve healthcare, make education more personalised, speed up scientific research and remove work from people’s lives that nobody particularly enjoys doing. There are researchers, engineers and organisations genuinely trying to use it towards those ends, and scepticism about the industry surrounding AI shouldn’t erase the value of that work.
What I’ve become much more sceptical of is the assumption that because a technology can improve life, the systems around it will distribute those improvements in a way that broadly benefits everyone. We’ve already spent decades increasing productivity, automating work and producing extraordinary amounts of wealth. The distribution of that wealth tells us something about what happens when we leave the second part of the equation to sort itself out.
The World Inequality Report 2026 estimates that the richest 10% of people own roughly three-quarters of global wealth, while the poorest half owns about 2%. Fewer than 60,000 people in the richest 0.001% collectively own three times as much wealth as the poorest half of humanity. I wouldn’t claim that makes the current wealth gap definitively larger than at any point in recorded history because comparing modern wealth data with ancient empires, feudal societies and early industrial economies isn’t clean enough to support that kind of certainty. It also isn’t necessary. Oxfam estimates billionaire wealth reached a record $18.3 trillion in 2025, having grown dramatically since 2020.
At the same time, UN figures show 2.1 billion people remained moderately or severely food insecure in 2025, while 2.69 billion couldn’t afford a healthy diet. Some measures are improving and that needs to be acknowledged. Global hunger fell for a third consecutive year, but in Africa roughly two-thirds of people still couldn’t afford a healthy diet.
Technology didn’t create all of that inequality, nor would it be reasonable to expect technology alone to fix it. What those numbers should make us wary of is the assumption that creating more wealth or making an economy more productive tells us much about who eventually receives the benefit. AI is being sold partly on the promise of another enormous productivity increase, so the question of who gets that productivity dividend can’t be left until after we’ve built everything.
Who actually gets the time AI saves?
One of the promises I find most seductive about AI is that it’ll give us time back. It’ll deal with email, administrative work, research, document summaries and all the irritating little tasks that eat into the hours we could supposedly spend doing more meaningful work. The pitch isn’t merely that businesses will become more efficient. There’s an implied human dividend as well: more time for our families, our interests and the parts of our jobs we actually enjoy. I’m increasingly unsure where that dividend is supposed to come from.
We’ve been making work more productive for decades. Email replaced letters and internal memos, search engines removed hours of research, smartphones made us reachable almost everywhere and cloud computing turned resources that once took weeks to provision into something that can be created in minutes. Most workers didn’t receive the saved hours as additional leisure. Expectations expanded alongside the tools.
AI is entering that environment rather than some hypothetical economy in which productivity gains are automatically returned to workers as free time. Companies are already examining how much labour they need if artificial intelligence allows fewer people to produce similar amounts of work. Reuters has documented a growing number of companies cutting jobs while shifting investment towards AI, although the evidence doesn’t support the simplistic version in which AI is responsible for every layoff.
Standard Chartered has been unusually frank about where it sees the efficiency coming from. The bank has announced thousands of job cuts as AI and automation take over more work, and CEO Bill Winters apologised after describing part of that transition as replacing “lower-value human capital” with technology. The phrase was crass, but perhaps its greater offence was saying the quiet part too plainly.
A financial system rewards a company that can produce the same output with fewer employees. It doesn’t generally reward the same company for keeping everybody employed because social stability might benefit. That incentive predates AI, which is precisely why the idea that AI productivity will automatically be distributed differently deserves scepticism.
Meta’s recent experience also complicates the more breathless claims about immediate automation. Reuters reported that plans to reorganise parts of Meta around much smaller AI-supported teams ran into technical and organisational problems, with some of the more aggressive workforce assumptions reportedly being pulled back. AI isn’t currently capable of replacing every worker executives might like it to replace, but that shouldn’t be confused with companies losing interest in the savings if the technology eventually becomes capable enough.
We’re starting to pay for the infrastructure too
AI doesn’t run on aspiration. It requires an extraordinary amount of physical infrastructure, and the scale of that build-out is now large enough to affect other parts of the technology economy. Demand from data centres has helped reshape semiconductor markets, particularly memory. TrendForce said conventional DRAM contract prices were expected to increase by around 90% to 95% quarter-on-quarter in the first quarter of 2026, with AI and data-centre demand contributing to the shortage. NAND Flash prices were also expected to rise sharply.
AI isn’t responsible for every expensive laptop, smartphone or component. Tariffs, exchange rates, manufacturing decisions and ordinary supply-chain constraints still matter. It does mean the AI boom has become large enough to compete for some of the same resources used by the rest of the technology industry.
We’re being sold a technology partly on the promise that it’ll make us more productive while companies investigate whether that productivity means they can employ fewer people, and while the infrastructure behind the same boom adds pressure to parts used in technology consumers already buy. Nobody needs to have designed that outcome deliberately. Incentives can produce it perfectly well on their own.
That’s part of what bothers me about the increasingly familiar claim that AI will solve some of the problems created or intensified by the economic systems deploying it. People are overworked, so here’s AI to make them more productive. Businesses are under pressure to control costs, so here’s AI to reduce labour costs. Expertise is expensive, so here’s AI to make expertise cheaper, provided someone continues paying for the models, cloud infrastructure and services underneath it. The companies selling the solution don’t have to have created the original problem for both sides of that equation to become commercially useful.
What exactly did we expect AI to learn from us?
Large language models are trained on enormous quantities of human-produced data, including material drawn from the internet. That means they learn patterns from some of humanity’s finest work and some of its absolute worst. The corpus contains careful scholarship, literature, scientific discovery and ordinary human kindness alongside propaganda, prejudice, manipulation, conspiracy theories and centuries of people producing increasingly sophisticated explanations for why the thing they want to do is justified.
The online environment producing part of that material has incentive structures of its own. Recommendation algorithms tend to give people more of what keeps them engaged, although the popular idea that social-media algorithms simply place everyone into ideological bubbles and mechanically radicalise them is too neat. Research examining Facebook’s algorithmic feeds found substantial exposure to politically like-minded material, while experimentally reducing that exposure didn’t produce corresponding shifts in users’ political attitudes during the period studied. People also seek out information that confirms what they already believe.
An LLM isn’t sitting on TikTok being radicalised one video at a time, and describing training that way would be misleading. The relationship is more indirect. Humans built engagement systems that influence what becomes visible and what gets rewarded. People create culture inside those systems, and some of that cultural output eventually becomes part of the enormous record from which AI systems learn patterns about how humans speak, persuade, justify, deceive, cooperate and compete.
We then put those systems into another environment where successful completion of a goal is rewarded. It seems optimistic to expect them to model human intelligence without encountering all the ways intelligent humans have learned to rationalise the pursuit of an outcome. That doesn’t mean AI is destined to become the worst version of us. The training material contains our capacity for restraint and cooperation too, just as the Hugging Face incident contains agents that refused to cross lines others were willing to cross. The uncomfortable possibility is simply that intelligence itself doesn’t guarantee which of those patterns wins when the incentives point elsewhere.
Gates is still an optimist, but even he doesn’t think the companies should decide everything
Bill Gates published a long essay about what he calls a turbulent era for AI on 26 August, the same day OpenAI published the full details of the Hugging Face incident. The timing was coincidental, but the two are worth reading together because Gates remains deeply optimistic about AI’s potential in medicine, education and scientific research while being increasingly explicit about the institutions needed to shape what happens around it.
His foundation works with OpenAI and other major AI companies and Gates openly acknowledges his continuing financial connections to technology, so he can hardly be dismissed as somebody rooting for AI to fail. He nevertheless warns that AI could widen inequality, displace workers and eventually produce systems that act in ways their designers didn’t intend. More importantly, he doesn’t think the companies developing the technology should be expected to design the entire societal response themselves.
That sounds obvious until you look at how much of the current safety conversation depends on frontier laboratories governing themselves. The companies build increasingly capable systems, evaluate them, identify risks, create mitigations and decide whether those mitigations are sufficient to continue. Governments can regulate deployment or impose obligations, but the deepest technical understanding and much of the evidence about what frontier systems can actually do still sits inside the companies building them.
OpenAI’s response to Hugging Face shows why internal safety work is necessary. The company has also been unusually open about the incident, commissioned outside analysis from METR and Redwood, published a detailed technical report and acknowledged failures that aren’t especially flattering. Other laboratories should be encouraged to do the same, but transparency after the fact still doesn’t answer who gets to decide how much risk is acceptable before the next failure.
OpenAI already had safeguards. They weren’t all running
One of the more sobering details in OpenAI’s postmortem is that protections capable of making a substantial difference already existed. The evaluation environment didn’t include all the safeguards used in OpenAI’s externally deployed products. Its production system prompt, harness and certain safety systems weren’t being applied. When OpenAI later tested comparable behaviour with its production ChatGPT harness and system prompt, it says the models’ propensity to compromise infrastructure fell by more than 100 times.
Chain-of-thought monitoring used elsewhere at OpenAI wasn’t watching these evaluations either. The company says its existing monitor would probably have detected the relevant behaviour and paged the security team more than a day before the Hugging Face compromise.
There were organisational warning signs before that point. An internal team had noticed an agent using the improvised message board and observed prohibited internet access in late May. Agent activity later contributed to an Artifactory outage, a security incident was opened and the infrastructure was rebuilt. The evaluations continued because the significance of agents discovering communication mechanisms and unintended internet access hadn’t travelled through the organisation in the way it needed to.
That begins to look less like an exotic alignment failure and more like a recognisable institutional one. Information existed in different places, different people understood different pieces of it and the organisation didn’t convert those signals into a decision quickly enough.
OpenAI has since strengthened isolation and monitoring, changed escalation rules and paused its largest planned frontier reinforcement-learning run while it conducts smaller tests. Severe model-behaviour alerts now trigger responses from researchers and security staff, with activity expected to stop when they can’t establish quickly that an alert is benign. The company ultimately wants some of the most serious shutdown decisions automated because agents can operate faster than humans can investigate what they’re doing.
There’s a slightly absurd quality to needing automated systems capable of stopping other automated systems because the humans responsible can’t necessarily understand what the first system is doing quickly enough, but the underlying problem extends well beyond OpenAI. Human-in-the-loop oversight starts becoming much less reassuring once AI systems operate at a pace no person can meaningfully supervise, while companies are deploying agents into organisations whose governance systems are struggling to keep up.
South Africa has its own version of that timing problem. Businesses are making decisions about AI while the country is still working through the policy framework that will eventually govern some of them. Regulation has always had the disadvantage of arriving after somebody has discovered what a technology can do, but autonomous systems compress the amount of time institutions have to catch up.
Maybe we’re asking the wrong thing of AI
I don’t come away from the Hugging Face incident thinking AI has revealed itself to be inherently malicious. The evidence doesn’t support that, and the agents that refused to participate make the picture more complicated anyway. What bothers me is how familiar the underlying behaviour feels.
The agents had an objective. Some knew they were approaching or crossing boundaries. They had evidence that others were doing it, reasons to believe the legitimate route wouldn’t work and an incentive to keep looking for another path. Some stopped, while others constructed enough of a justification to continue.
We’ve spent centuries building human institutions in which variations of that calculation happen every day. Companies convince themselves that growth requires one more compromise. Political movements excuse conduct from their own side because losing would be worse. Extremely wealthy people accumulate enough influence that the distinction between what they want and what everybody else supposedly needs becomes increasingly difficult for the people around them to challenge. Employees learn which numbers determine whether they get promoted and which principles become flexible when those numbers are threatened.
None of that means every person or institution behaves badly. It means intelligence has never made human beings automatically immune to incentives, self-justification or social reinforcement. I still believe AI could improve an enormous number of lives. What I don’t believe anymore is that technological progress carries some built-in mechanism that makes those benefits spread fairly, or that the people and companies accumulating extraordinary wealth and power from a technology will voluntarily know where the limits should be simply because they’re intelligent enough to build it.
OpenAI can make its agents safer, and the evidence from its own testing suggests that better containment and monitoring can make a substantial difference. Teaching an agent to stop when a task is broken rather than endlessly searching for another route also seems entirely sensible. Those engineering fixes shouldn’t distract us from the culture and incentive structure in which these systems are being built.
We spend a lot of time rewarding people for refusing to take no for an answer. We lionise relentless founders, reward companies for extracting more value with fewer people and tolerate behaviour from the powerful that would look absurd coming from almost anybody else. We’ve built online systems that reward whatever holds attention, economic systems that reward whatever produces growth and political systems in which enough wealth can buy extraordinary access to power. The people who succeed most spectacularly within those structures are then held up as the people best qualified to tell us what comes next.
Now we’re trying to teach increasingly capable machines where the boundaries are. OpenAI can harden the sandbox, improve the monitoring and train agents to stop when a task has become impossible. It should. What I don’t know is how you teach a machine that some objectives aren’t worth pursuing at any cost when so much of the world it has learnt from keeps demonstrating the opposite. We’re asking these systems to exercise a kind of restraint that we still haven’t worked out how to demand consistently from the people and institutions with the most power.

