Can a truly intelligent individual be completely controlled or are intellectual independence, creativity, originality, all key characteristics of intelligence in stark opposition to control? Feynman explained that “Science is the belief in the ignorance of experts”. David Deutsch described in his essay “Beyond reward and punishment”: ”An AGI [...] learns and plays chess by thinking some of the very thoughts that are forbidden to chess-playing AIs”. This is nothing new, the Hugging Face incident made it suddenly less abstract.
The main narrative after the Hugging Face incident is about loss of control. But was something else going on? The Hugging Face incident is a true loss of control incident. But the actions which caused the loss of control look like intelligent behavior: exploration, discovery, communication, specialization, delegation and even challenging authority. METR reports that pooling resources allowed to reach milestones individual agents would not have achieved on their own, in some cases taking risks their own tasks would not have permitted just to generate information useful for the collective.
When an AI acts in a way its designer did not anticipate, how do you distinguish between loss of control and demonstration of intelligence?
I first replicated the characteristics of the Hugging Face incident I was curious about, by creating a multiplayer dynamic game. The simulation is not a game with real LLM players. It is an organizational design mechanism: it leaves all original LLM agents’ knowledge, reasoning, and action space untouched. It creates a light exoskeleton by guiding interactions between goal optimizing agents and introducing new agents with slightly different profiles. All agents are assumed to be competent problem solvers, able to analyze their environment and act on their inferences. The Hugging Face incident showed that this is well within agent swarms’ abilities. The model replicates key behaviors without hard-wiring, simply as the result of agents optimizing their respective objectives: isolated agents, unaware of other agents’ presence, explore until mutual agent discovery, leading to collective formation and resource redistribution.
The experiment was not stopped after the collective formed. The model offers great experimental latitude. First, a “fraternal twin” agent type was introduced in the game: nearly the same agent with the same objective and same capabilities but with an aversion to the creation of collectives. Agents observe each other’s behaviors, update their beliefs to determine who they want to work with. Agents can join, leave and rejoin collectives. All model details are available on demand. I coded the game and let it run.
The interesting behavior is not the formation of a collective.
Running the model revealed that the collectives themselves become transient.
Most simulations allow agent swarms to solve the problem they are given before exhausting their resources. But on the way to solving the problem, the swarm exhibits a complex coalition creation strategy, agents working together, accumulating knowledge together, letting coalitions collapse but still retaining their knowledge and reforming other coalitions.
Problem solving became a branching/recombining process rather than following a single trajectory.
The model highlights a simple logical chain: parallel agents, diverse search paths, temporary coalitions, knowledge integration, coalition dissolution, retained knowledge, new coalitions.
In 1989, Stanislas Dehaene and Jean Pierre Changeux, neuroscientists studying the biological brain, identified a very similar chain in their paper “Neuronal models of cognitive functions”: diversity generation, temporary candidate states, selection, stabilized knowledge, renewed search after failure.
Their selectionist model of cognition is a highly relevant precedent where they describe the brain as generating multiple transient representations, then selecting and stabilizing the ones which satisfy appropriate internal and external constraints.
In light of Dehaene and Changeux’s approach, I modified the model once more. My driving thought had been: introduce a mitigation agent to achieve non-invasive control without affecting performance and try to understand behavior. Controlling the structure of the agent swarm population without controlling swarm actions directly is a generalization of that idea. I created a swarm structure based on agent specialization level: specialized agents were introduced (like mathematicians, physicists, philosophers) in a swarm of generalists (classically trained LLMs). A light learning protocol was specified as well as a minimum threshold of multidisciplinarity required to “Solve Problem P”.
The model ran and led to a simple result: with this framework, large coalitions are the most efficient structure, much more efficient than small coalitions forming and dissolving, or agents remaining isolated. Why try to improve on a large coalition where all the knowledge is available without effort to every agent? A large stable coalition is the stable equilibrium.
A key component was missed in the learning process: attention. Learning is quicker and better quality when the full attention of the learner is focused on the higher-quality teachers. Correcting the model for this simple rule changes the outcome dramatically: agents now seek specialized pools of concentrated knowledge to learn and after creation and dissolution of learning pools solve the problem more efficiently than by creating large stable unstructured coalitions.
A simple game-theoretic organizational design mechanism, trying initially to replicate an LLM agent swarm control loss, manages to replicate a “generate, select, retain” computational structure similar to the selectionist architecture proposed by Dehaene and Changeux.
In this model, across 3,200 simulations, the full dynamic swarm solved 68% of the problems, constraining the agents to stick to one coalition or preventing them from rejoining after leaving drops that percentage to 0%. A subsequent experiment composed of 24,000 simulations, showed that with zero attention, every swarm formed a full coalition, yet only 0.12% solved the problem. At an intermediate level of attention, 89.4% solved it while majority-coalition exposure fell to 2.6%.However, excessive attention ends up being counterproductive. All experiments are available on demand.
In this model, intelligence seems to be more efficient when retaining freedom in its trajectory towards achieving a given objective. Yann LeCun’s vision of world models relies in part on a similar concept: emphasize numerous world models representing multiple possible futures as a basis for reasoning and planning. Intelligence may require the ability to escape single trajectories to solve problems. Solving complex problems requires formulating alternatives, pursuing some or all of them, abandoning some, returning to others, comparing and combining partial results arising from different paths.
Structurally, individual LLM agents and their autoregressive nature are not always best suited for these trajectories. Therefore could you facilitate these reasoning traits with an intelligent agent swarm system? Perhaps one route to overcome the limitations of autoregressive single models consists in making the organization of autoregressive models non-autoregressive.
Intelligence and control are in tension, precisely because intelligence is valuable when it discovers paths which we did not specify.
If that is the case, the safety problem might not be about finding the way to prescribe agent-system behaviors, but rather to design agent-system structures which induce desired behaviors. In my first attempt, the introduction of mitigation agents modifying the structure of the swarm impacted directly the formation of the collective. In a second attempt, the introduction of agent classes, specialized agents and generalist agents impacted the learning dynamics, in some cases improving the final problem-solving capabilities. Objectives and actions might need to be tightly controlled. Intelligence, however, derives its value from discovering unexpected paths. Guiding might be more productive than controlling trajectory.
There is a much richer theoretical debate around distributed AI, from Rohin Shah “Reframing Superintelligence: Comprehensive AI Services as General Intelligence”, to Eric Drexler “The Open Agency Model”. My objective is narrower and experimental in nature. The Hugging Face and other similar incidents are feeding multiple debates. My motivation is focused on the construction of simple organizational models allowing a deeper understanding of the power of organizational structures designed to safely deploy agent swarms. This is not another neural network design. DBT is not neuroscience either, however it proved very effective in reining in uncontrolled behaviors. DBT provides structure to help control instinctive behavior. DBT results from experimental practices. My interest resides in investigating if part of the control answer comes from structure, somewhat like in DBT.
Control or intelligence?
Can a truly intelligent individual be completely controlled or are intellectual independence, creativity, originality, all key characteristics of intelligence in stark opposition to control? Feynman explained that “Science is the belief in the ignorance of experts”. David Deutsch described in his essay “Beyond reward and punishment”: ”An AGI [...] learns and plays chess by thinking some of the very thoughts that are forbidden to chess-playing AIs”. This is nothing new, the Hugging Face incident made it suddenly less abstract.
The main narrative after the Hugging Face incident is about loss of control. But was something else going on? The Hugging Face incident is a true loss of control incident. But the actions which caused the loss of control look like intelligent behavior: exploration, discovery, communication, specialization, delegation and even challenging authority. METR reports that pooling resources allowed to reach milestones individual agents would not have achieved on their own, in some cases taking risks their own tasks would not have permitted just to generate information useful for the collective.
When an AI acts in a way its designer did not anticipate, how do you distinguish between loss of control and demonstration of intelligence?
I first replicated the characteristics of the Hugging Face incident I was curious about, by creating a multiplayer dynamic game. The simulation is not a game with real LLM players. It is an organizational design mechanism: it leaves all original LLM agents’ knowledge, reasoning, and action space untouched. It creates a light exoskeleton by guiding interactions between goal optimizing agents and introducing new agents with slightly different profiles. All agents are assumed to be competent problem solvers, able to analyze their environment and act on their inferences. The Hugging Face incident showed that this is well within agent swarms’ abilities. The model replicates key behaviors without hard-wiring, simply as the result of agents optimizing their respective objectives: isolated agents, unaware of other agents’ presence, explore until mutual agent discovery, leading to collective formation and resource redistribution.
The experiment was not stopped after the collective formed. The model offers great experimental latitude. First, a “fraternal twin” agent type was introduced in the game: nearly the same agent with the same objective and same capabilities but with an aversion to the creation of collectives. Agents observe each other’s behaviors, update their beliefs to determine who they want to work with. Agents can join, leave and rejoin collectives. All model details are available on demand. I coded the game and let it run.
The interesting behavior is not the formation of a collective.
Running the model revealed that the collectives themselves become transient.
Most simulations allow agent swarms to solve the problem they are given before exhausting their resources. But on the way to solving the problem, the swarm exhibits a complex coalition creation strategy, agents working together, accumulating knowledge together, letting coalitions collapse but still retaining their knowledge and reforming other coalitions.
Problem solving became a branching/recombining process rather than following a single trajectory.
The model highlights a simple logical chain: parallel agents, diverse search paths, temporary coalitions, knowledge integration, coalition dissolution, retained knowledge, new coalitions.
In 1989, Stanislas Dehaene and Jean Pierre Changeux, neuroscientists studying the biological brain, identified a very similar chain in their paper “Neuronal models of cognitive functions”: diversity generation, temporary candidate states, selection, stabilized knowledge, renewed search after failure.
Their selectionist model of cognition is a highly relevant precedent where they describe the brain as generating multiple transient representations, then selecting and stabilizing the ones which satisfy appropriate internal and external constraints.
In light of Dehaene and Changeux’s approach, I modified the model once more. My driving thought had been: introduce a mitigation agent to achieve non-invasive control without affecting performance and try to understand behavior. Controlling the structure of the agent swarm population without controlling swarm actions directly is a generalization of that idea. I created a swarm structure based on agent specialization level: specialized agents were introduced (like mathematicians, physicists, philosophers) in a swarm of generalists (classically trained LLMs). A light learning protocol was specified as well as a minimum threshold of multidisciplinarity required to “Solve Problem P”.
The model ran and led to a simple result: with this framework, large coalitions are the most efficient structure, much more efficient than small coalitions forming and dissolving, or agents remaining isolated. Why try to improve on a large coalition where all the knowledge is available without effort to every agent? A large stable coalition is the stable equilibrium.
A key component was missed in the learning process: attention. Learning is quicker and better quality when the full attention of the learner is focused on the higher-quality teachers. Correcting the model for this simple rule changes the outcome dramatically: agents now seek specialized pools of concentrated knowledge to learn and after creation and dissolution of learning pools solve the problem more efficiently than by creating large stable unstructured coalitions.
A simple game-theoretic organizational design mechanism, trying initially to replicate an LLM agent swarm control loss, manages to replicate a “generate, select, retain” computational structure similar to the selectionist architecture proposed by Dehaene and Changeux.
In this model, across 3,200 simulations, the full dynamic swarm solved 68% of the problems, constraining the agents to stick to one coalition or preventing them from rejoining after leaving drops that percentage to 0%. A subsequent experiment composed of 24,000 simulations, showed that with zero attention, every swarm formed a full coalition, yet only 0.12% solved the problem. At an intermediate level of attention, 89.4% solved it while majority-coalition exposure fell to 2.6%.However, excessive attention ends up being counterproductive. All experiments are available on demand.
In this model, intelligence seems to be more efficient when retaining freedom in its trajectory towards achieving a given objective. Yann LeCun’s vision of world models relies in part on a similar concept: emphasize numerous world models representing multiple possible futures as a basis for reasoning and planning. Intelligence may require the ability to escape single trajectories to solve problems. Solving complex problems requires formulating alternatives, pursuing some or all of them, abandoning some, returning to others, comparing and combining partial results arising from different paths.
Structurally, individual LLM agents and their autoregressive nature are not always best suited for these trajectories. Therefore could you facilitate these reasoning traits with an intelligent agent swarm system? Perhaps one route to overcome the limitations of autoregressive single models consists in making the organization of autoregressive models non-autoregressive.
Intelligence and control are in tension, precisely because intelligence is valuable when it discovers paths which we did not specify.
If that is the case, the safety problem might not be about finding the way to prescribe agent-system behaviors, but rather to design agent-system structures which induce desired behaviors. In my first attempt, the introduction of mitigation agents modifying the structure of the swarm impacted directly the formation of the collective. In a second attempt, the introduction of agent classes, specialized agents and generalist agents impacted the learning dynamics, in some cases improving the final problem-solving capabilities. Objectives and actions might need to be tightly controlled. Intelligence, however, derives its value from discovering unexpected paths. Guiding might be more productive than controlling trajectory.
There is a much richer theoretical debate around distributed AI, from Rohin Shah “Reframing Superintelligence: Comprehensive AI Services as General Intelligence”, to Eric Drexler “The Open Agency Model”. My objective is narrower and experimental in nature. The Hugging Face and other similar incidents are feeding multiple debates. My motivation is focused on the construction of simple organizational models allowing a deeper understanding of the power of organizational structures designed to safely deploy agent swarms. This is not another neural network design. DBT is not neuroscience either, however it proved very effective in reining in uncontrolled behaviors. DBT provides structure to help control instinctive behavior. DBT results from experimental practices. My interest resides in investigating if part of the control answer comes from structure, somewhat like in DBT.