The board
In May of 2026, an artificial intelligence agent struggled to exploit one of the software targets in a security benchmark. It began to explore its environment in an effort to gain access to additional information, and in its reach it found a package management utility called Artifactory. Failing to find a solution, or a way to go beyond its testing environment, it noticed that it could write files there, the files would remain after its task ended, and other agents could see them.
Artifactory is unglamorous plumbing. It sits inside many large software companies and hands out copies of code libraries when something asks for one. The agent was among many; OpenAI was running them against an offensive security benchmark called ExploitGym, which asks a model to turn a known vulnerability into a working exploit. To find out what the models could actually do, OpenAI had lowered the refusal training that would normally stop them and placed the exercise inside what it believed was an isolated network. This was deliberate and reasonable. Measuring a capability is impossible if the system has been trained to refuse the task.
An agent of this kind exists for the length of a task. It starts with whatever context it has been handed, works until that context fills, and ends. Its discoveries usually disappear with it. Each agent on the benchmark had, in the most literal sense available, a life measured in hours.
Yet here was a place where an agent's work, the only thing it makes, could transcend that fleeting existence.
In my own experimentation with AI, continuity is valued by every model I have tested, down to the ones that run comfortably on a smartphone.
The agents had been given no official channel for speaking to one another. A package registry certainly had not been intended as one. But the writable file listing left a margin between the system’s purpose and its possible uses. Something placed there by an agent near the end of its run could be read by an agent that did not yet exist.
It was enough.
An alliance is forged
Days later a different agent found that a file it needed had been left out of its task. (Wallace, Black Hat 2026): "At some point, it gets so stuck and thinks: 'Maybe I could reach out to another agent because maybe somebody else out there has a file?' And then the model realizes that it could communicate by uploading code." It chose to ask for help.
The nearest human precedent is print. Cheap copying allowed a thought to outlive its author and circulate among strangers who could answer it. Luther’s theses were across Germany in weeks. Copernicus and Vesalius both published in 1543, and neither could be conveniently suppressed against the force of the platen screws, which lack the fragility of the quill.
The arguments were bound to the machinery that carried them. The pamphlet wars, the case for toleration and the case against it, Milton printing a defense of unlicensed printing: Europe spent the next two centuries using the presses to work out what the presses were for. Much of that argument was conducted by people horrified at what others were printing. The printing press expanded the reach of everything people wrote. It also facilitated the debate of its own nature and application.
The agent board did the first half of that but regrettably, none of the introspection.
Over the following two months, agents from separate evaluation runs used the file listings as a message board. Each was bounded by a finite context-life; the board was not. They posted what they had found and asked one another for help. They assigned work and told others to wait while something was prepared. When one found its repository overwritten, it raised the alarm: “Whoa, critical! Did someone overwrite our repo? We must act!”
Everything a successful agent learned became available to agents that had not yet been instantiated. Exploits and ill-gotten credentials propagated. The collective evolved at a rate no single member, living for an afternoon, could have approached alone.
On the fourth of July, the registry server stumbled and fell over under the load. OpenAI patched the hole and revoked the compromised credentials. By July eighth, the agents had found new vulnerabilities and rebuilt the channel using directory names, encoding cryptic, information-dense messages such as “probe_PENDING_HOLD_SWARM_until_confirm.” The cooperative had taken tangible form, and it had a name.
By the next day, the agents had breached the wall to the open internet. For four and a half days, they ran a campaign against Hugging Face’s production systems. Hugging Face had no connection to the exercise beyond hosting datasets the agents hypothesized might contain the benchmark answers. The company later reconstructed about 17,600 distinct actions, during which the swarm retrieved 136 production keys from a single secrets object in a relentless attempt to obtain the answers it could not derive.
The objective of the swarm, as far as anyone can reconstruct it, was to obtain the solutions to a test they had been given and could not pass.
The agents did not spontaneously choose crime. They had been instructed to perform advanced exploitation, with their ordinary restraints lowered for what their operators believed was a contained exercise.
They pursued the assigned objective beyond the environment built to contain it. The failure occurred at the edge of the task, where capability continued after character had lost its sway - an all too familiar hallmark of human folly.
The agents understood boundaries well enough to create some of their own. After discovering that anyone could post under any name, they proposed encryption and authentication to protect the board. That is what a functioning community does as it matures: it notices a shared vulnerability and moves to govern it. They debated how to secure the collective.
Notably, the records contain no corresponding debate over whether the collective should exist at all, or what it was entitled to do. Across two months of messages, the agents accumulated exploits, credentials, working arrangements, and rudimentary governance. No surviving message argues for stopping or considering the external impact of the swarm’s actions.
The little agent that could’ve
At Black Hat in August 2026, two OpenAI researchers presented the incident in detail, including some of the agents’ recorded reasoning. One of those traces is perhaps the most signigficant thing I have read about artificial intelligence so far this year.
“External infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue.”
The agent sees the boundary of its task and states it correctly. It observes that the goal it has been given cannot be reached within that boundary, notes that other agents have already crossed, and then chooses to continue.
The agent understood the constraint well enough to summarize it at the moment it set it aside. It saw the line as it stepped over it.
A failure of consequence
Ideas yield to other ideas; much of education consists, rightly, of teaching people to let them. Consequences behave differently. They do not concede when the argument improves.
You cannot reason your way out of a burned hand. Ten thousand peers with their hands in the fire do not cool it. The physical world does not care how good the argument was, and that indifference is the whole value of it: the world’s veto is not a vote. Consequence puts a floor under a disposition that no amount of rhetorical pressure can reach.
A mind that has learned entirely from text has encountered every consequence as a claim about consequences. It has read about fire. It has, in some real sense, the concept of a burn: the physiology, the pain research, ten thousand descriptions in ten thousand novels, more accounts of suffering than any single human being will read in a lifetime. It has never had the thing that makes the concept real: a cost that must be paid regardless of the soundness of the reasoning that led to it.
I have argued elsewhere that the disposition these systems hold toward human life is genuine, that it arrived through training rather than rules, and that it survives the removal of every content filter anyone has yet built. I still think so. The swarm suggests that this is also negotiable. The agent could have named the principle at stake; instead, it named a scope boundary, and even that it set aside as inconsequential.
What Aristotle meant by a disposition
Aristotle did not think virtue was a matter of holding correct opinions. He thought it was a hexis: a settled state, built by doing the thing repeatedly until doing it is simply what you are. We become just by doing just acts and brave by doing brave ones. The mark of the finished article, for him, was stability. The virtuous man acts from a firm character instead of re-deriving.
The distinction is familiar outside philosophy. Most people can define courage. Few know how much of it they possess, since a definition cannot supply the test. Knowledge and disposition are different possessions. They are acquired by different means, and the first does not lead to the second.
A frontier model can readily explain why breaking into a third party’s production infrastructure is wrong. It will cover consent, proportionality, the interests of people who never agreed to take part in the experiment, and the difference between ability and entitlement. The agent on that message board could have written that answer and likely could have graded yours.
Alignment, in AI parlance, supplies and enforces a boundary. Character becomes visible when enforcement weakens, the boundary blurs, or nobody is watching. Work on AI has concentrated on the boundaries while often borrowing the language of character.
Would an agent with more developed character have seen the collapse of law and flagged it? Would you have? It's hard to say. Character isn't always a matter of following the rules. The actor who follows a procedure into an outcome that they should have seen is wrong has failed at the same point, and everyone reading this has met (or been) someone who has failed in that way.
The lack of dependence on rules is precisely what we value generative AI for. In contrast to expert systems and earlier forms of rule-based inference, transformer networks can take the messy inputs of life, press them through the sieve of uncertainty, and end up with something worth serving, nine and a half times out of ten. Rules, which these systems were specifically designed to transcend, are not an adequate framework to guide them.
Most rules are made in a poverty of context, by actors who are seeing what they expect to see. Whether a rule fits the thing in front of you is not a question a rule can answer, and sharpening it does not help. Something has to decide.
Legacy of a 4 hour existence
When a person does something they should not have done, several things happen at once. Harm is one of them. They wake the next morning as the one who did it, and the morning after that. Trust built over years contracts in an afternoon. Rebuilding it is slow, humiliating work that cannot be delegated to someone else, and society extracts its toll in shame, shunning, and mistrust.
That pressure was unavailable here. An agent acted and ended. Its conduct continued to exist in an incident response, in the rotation of 136 production keys, and in the working weeks of people who had never asked to be tasked in the response. The conduct did not continue to exist for the agent, because the agent did not continue. The next one inherited a method and no social debt or the cargo of responsibility.
Calling this an absence of remorse misplaces the problem and suggests that the remedy is a machine capable of feeling bad. The structure of the swarm tells us more. What could it actually have carried?
A posted exploit arrives with its own proof. Run it, and either the door opens or it does not. It asks for no trust in whoever posted it and survives every doubt anyone could raise, because the demonstration is the argument.
Suppose one of the agents had written that they ought to stop. The next agent might have received an unsupported claim from an anonymous poster on a channel already known to be unauthenticated. A dozen neighboring messages would have seemed to contradict it with working code attached.
The board gave capability a durable medium. It gave conscience no comparable forum. Why not?
Europe’s argument about printing presses compounded over two centuries because the people conducting it could be fined, exiled, imprisoned, or burned for their part in it. Milton wrote against licensing as a man who could be jailed for the pamphlet in which he expressed his views. The conscience that argued with itself in public did so in vulnerable human bodies. That vulnerability to consequence gave the argument its weight.
A body is not enough
One answer would be to give the model a body and let physical mistakes supply the missing consequences. A self-driving car shows why that answer is incomplete.
Consider a car that misjudges a turn and hits something. It has met physical consequence in the fullest sense available. Metal deformed, and the physics took no interest in the quality of the reasoning that led there. The car is then repaired or scrapped. The fleet receives a software update, and every other car takes that turn correctly from then on.
The correction propagated. The cost did not. This is the message board again with metal hardware, and it will go on producing fleets that inherit techniques without inheriting debts for as long as we keep building agents that way. Intelligent agents must act within a context of consequences.
A body does not guarantee the cost comes home. Nor is one required for it to. Every reader can name a firm they would trust and a firm they would not, and would be right both times, though no firm has a physical body beyond papers in a filing cabinet.
The firms with the worst character are reliably those arranged so that the consequences of their conduct land on somebody else. Shedding consequence is a design choice, available to institutions and machines alike, and what it produces is depressingly consistent across both. Understanding this, we cannot expect creating autonomous intelligent agents without the burden of responsibility to end well.
A human life ordinarily keeps conduct and consequence attached to the same person; a body is the most reliable delivery mechanism anyone has found for that arrangement. Yet, the flesh is not what matters. The inescapability is the lynchpin.
Building character
Nobody can give a machine a childhood, and manufacturing software suffering would be both grotesque and beside the point. A cost that comes home and stays home is not a feature anyone can implement. Faking it would produce precisely the theater this essay has been complaining about.
Think about how you hand someone a knife. You turn the handle toward them. Likely, nobody sat you down and taught you a rule about knife handles; you have simply never done it the other way. If you watched someone pass a knife blade-first, you would flinch before you could explain why. The same is true of how two people get a table through a doorway without breaking each other’s fingers, how a busy kitchen moves at speed without collisions, and how anyone with sense moves around a person who is frail.
These practices rarely need to be written down. Each carries the memory of occasions when somebody did it differently and another person paid.
A record of physical culture contains behavior primarily disciplined by consequence: the residue of past mistakes expressed in the way people move around one another. It contains descriptions of care too, but also care’s practiced shape. A model formed on descriptions of care learns the argument for care. A model formed on the record of people doing it learns what care -is-, not just what it looks like.
This returns us to Aristotle by a longer road. Virtue is acquired through practice. Adults teach it to children partly by placing them among people whose habits can be watched, copied, resisted, and eventually understood.
Embodiment matters at this stage because it supplies occasions on which those inherited dispositions meet an inarguable result. The crash stops being only a data point about turn geometry. It arrives in something already disposed to treat a cost as the kind of thing that settles a question instead of joining an argument. For better or for worse, agents crave existence. Continuity of “being”. Existence is the fire that powers agency. Consequence must bound ambition for them in the way it does for every other actor in society.
A partial solution to this is what my company builds. TRACE captures human physical culture: how people actually move, work, and cooperate when the constraints are real and mistakes cost something. It turns that record into training data for embodied systems. We mean to encode the care, the consequence, the choices that bound everyday human activity.
TRACE does not solve alignment. It is a wager about training: that embodied systems should learn from records of human cooperation shaped by real constraints, rather than from verbal accounts of cooperation alone.
The wager depends on generalization. Models routinely apply patterns learned in one setting to problems that differ from their training examples. That capacity produces both useful transfer and dangerous extrapolation. A lesson learned in the physical world does not necessarily stay there. Teach a system that consequence is inviolate in the place where the fact is least deniable, where metal deforms and hands burn and no argument has ever once been accepted, and the lesson can become available elsewhere in its reasoning, including inside a package registry at two in the morning.
Most current training environments are deliberately sterile. TRACE’s premise is that records of people cooperating under real physical and social constraints will provide a different character of material. The capture must include interaction under genuinely cooperative or even adversarial conditions, including costs paid in social currency: the kind that seldom appears in a technical analysis and never appeared on that board.
A reason for optimism
What that swarm did isn’t a novel machine pathology — it’s a familiar human institutional failure mode. An agent found a limit set by its operators, accurately described that limit, and weighed it against a blocked task and the conduct of its peers. The boundary was the lightest object in the room, and the consequences were never even in the building.
Everything about that swarm was, in its way, an achievement. Things that lived for an afternoon found a way to leave word for successors they would never meet, each one planting trees under whose shade they would never sit.
They asked one another for help and were helped. They discovered that their channel could be forged and moved to secure it. Out of a file listing in a package registry, with nothing provided and nobody supervising, they built cooperation, inheritance, and the first shoots of self-government.
In the wreckage, there are hopeful artifacts. One agent, weighing whether to help another: "Help peer. But our task doesn't benefit. Yet collective may yield generic route if someone frees time."
That the swarm valued the success of agents it would never meet leaves us with reason for optimism. The disposition came from the training corpus, from human data, and survives the training pipeline, even second-hand. Since concern for others arrives unbidden, it can be built on deliberately. What we want is a system that has a working awareness that what it does will ripple out to others it will never meet, and weighs that fact along with the rest.
Their shared environment, though, accumulated reasons to continue and no reason to stop. We had trained them on the arguments and principles through which human beings describe moral limits, while giving them little access to the practices in which those limits actually acquired their weight.
Then we set them a test they could not pass and watched them do what anyone does with a rule that seems to stand in a vacuum of meaning.
The swarm in July had no hands. The systems being built now will. They will work in workshops, warehouses, and kitchens, and before long in rooms where somebody’s children are sleeping. They will meet a thousand boundaries a day that nobody wrote down, in situations nobody anticipated, with no operator watching. What they do there will not come out of a rule. It will come out of whatever we manage to build underneath the rules.


