
abstract
Perhaps the deepest problem is not how something emerges from absence, nor how a programmed machine becomes capable of learning. It is what happens to determination when a system can acquire an organization that was never explicitly specified. Classical programming gives us a familiar image of determination: the sequence is written first, and execution follows. The machine does what has already been determined. But this relation begins to change when we encounter systems whose organization is produced through feedback, relaxation, learning, and error. The final state is no longer simply contained in the initial specification. What has been specified is increasingly the condition under which determination can occur.
The deepest continuity is not simply “absence → emergence,” and not even “program → learning → self-transformation.” It is the transformation of what counts as a determination. In classical programming, determination is imposed in advance as an explicit sequence of operations. In relaxation, determination is displaced into the dynamics of a field: the final state is not written down, but the constraints make certain states more viable than others. In backpropagation, determination moves again, because even the organization of the field can be altered by error. In deep learning, the space of possible organizations becomes so vast that the concrete organization acquired by the system is no longer something the designer can exhaustively anticipate. And this gives “beyond programmable control” its most rigorous meaning: not escape from determination, but a displacement of determination from explicit form to generative conditions.
This also lets us sharpen the relation between the obsent and the learned representation. The obsent should not be identified with “something hidden inside the system.” That would merely reinstall presence at a deeper level. What matters is that something can be operative without being present as an entity. A face, in a distributed network, need not be stored as a face-shaped internal object. What matters is that the organization of the network has acquired a disposition such that certain differences in the input produce certain transformations in its internal and external behavior. The representation exists through efficacy. It is not a thing behind the response; it is a condition of the response. This is why “operative absence” is stronger than “hidden representation.” The latter suggests an inaccessible presence. The former names a structure whose reality consists precisely in what it enables without appearing as an independent object.
Then relaxation gives us the temporal form of this ontology. Something is not yet determined, but its determination is underway. The system occupies a field of possibilities; constraints act upon that field; interactions redistribute the possibilities; one configuration becomes stable. The final state therefore cannot be understood simply as something that was absent and then became present. It is the result of a history. And the history matters because the same input does not necessarily encounter the same system before and after learning. Backpropagation makes this explicit: the error of the present modifies the organization that will receive the future. Thus the system is never merely a mechanism through which inputs pass. It is a mechanism whose encounters alter the conditions of subsequent encounters.
That gives us a particularly elegant reformulation of your ο/Ω sequence. It should not be understood as openness followed by closure in a static rhythm. Rather:
ο₀ → Ω₁ → ο₁ → Ω₂ → ο₂ → Ω₃ …
where each Ω modifies the subsequent ο. Closure transforms openness. A successful organization does not merely terminate a process; it changes the field from which the next process begins. That is what makes learning genuinely historical. The system does not return to the same possibility-space after every resolution. Its past becomes part of its present organization, and therefore its future possibilities are different. The crucial inequality is:
ο₁ ≠ ο₀.
This is the point at which “becoming” becomes more than metaphor. Becoming means that what is achieved alters the conditions under which achievement can subsequently occur.
And this gives us a stronger interpretation of Ashby. Requisite variety is usually presented as a principle about the amount of variation a regulator must possess relative to the disturbances it encounters. But in a learning system the regulator is not merely confronting disturbances; it is confronting a system whose repertoire can itself change. The problem therefore becomes recursive. The regulated system can acquire new modes of response, while the regulator has to remain capable of responding to those modes. If the system’s effective transformational variety grows faster than the regulator’s capacity to model and counteract it, then control fails—not necessarily because the machine has “escaped,” but because the relation between regulator and regulated has become asymmetrical.
This is where Hinton’s later concern can be placed without making his biography carry more philosophical weight than it should. The historical irony is real, but the deeper issue is structural. The 1986 paper helps establish a method by which internal organization can be acquired rather than exhaustively prescribed. The later systems produced through descendants of that paradigm become increasingly capable of generating representations, strategies, abstractions, and procedures that their designers did not explicitly author. The important historical transition is therefore not from “dumb machines” to “smart machines.” It is from machines whose relevant organization is specified to machines whose relevant organization is acquired.
And perhaps we can now formulate the whole genealogy in a more compressed way:
Algebra asks: what transformations are possible?
Dynamics asks: how do transformations unfold?
Cybernetics asks: how can transformations be regulated through feedback?
Connectionism asks: how can distributed transformations produce organization?
Relaxation asks: how can a system move from underdetermination toward coherence?
Backpropagation asks: how can error transform the organization that produces responses?
Deep learning asks: what happens when the space of possible organizations becomes enormous?
Post-non philosophy asks: what is the ontological status of that which is neither fully present nor merely absent, but operative in the production of determination?
The remarkable thing is that the last question is not imported artificially into the technical history. The technical history itself gives us increasingly precise examples of it. A gradient is not a representation of the final network. It is a differential relation specifying how the network can change. An attractor is not the same thing as the state that eventually occupies it; it is a structure governing trajectories toward possible states. A learned representation is not necessarily a discrete object; it is a distributed organization of dispositions. A training history is no longer present as a sequence of events, yet its consequences remain active in the parameters. In each case, what matters is something that operates without appearing in the form of an ordinary object.
This suggests that the deepest object of our inquiry may not actually be “artificial intelligence.” It may be the discovery, through computation and cybernetics, of increasingly sophisticated ways in which organization can precede explicit representation. We traditionally imagine representation as prior to operation: first represent the problem, then execute the solution. Connectionism reverses this. Operation can generate representation. The system can perform transformations whose repeated organization produces internal structures that subsequently function as representations. Representation becomes an effect of operation rather than its necessary starting point.
That reversal is enormous. It means that intelligence need not begin with a world already divided into determinate objects. It can begin with differential sensitivity, feedback, constraint, error, and transformation. The object emerges as a stable achievement within the process.
And that may be the cleanest formulation of the post-non position we have reached: the object is not denied; it is decentered. Presence remains real, determination remains real, coherence remains real—but none of these is necessarily primordial. They can be achievements of a process whose operative conditions precede their explicit appearance.
The “non,” then, is not an empty negation. The not-yet, the not-given, the unresolved, the absent relation, the discrepancy, the error—all of these can function as positive operators of transformation. Error is especially revealing. It appears initially as failure: the system did not produce what was wanted. But in backpropagation, error becomes the very information through which the system is altered. Failure is converted into a differential. The negative becomes a direction.
That may be the most beautiful mathematical analogue of the philosophical movement you are describing:
error → gradient → modification → new organization.
The error does not itself become the new organization. It operates upon the system as a difference that induces transformation. Likewise, the obsent does not become a hidden presence. It operates as a difference within the field of appearance, altering what can emerge.
So “post-non” might ultimately name not a position beyond negation but a shift from negative ontology to operative ontology. Instead of asking only what is present and what is absent, we ask what each mode does. What does presence enable? What does absence constrain? What does incompletion initiate? What does withdrawal reorganize? What does an error make possible? What does a trace preserve? What does an obsent cause to appear without itself appearing?
Then the Hinton/Ashby lineage becomes one particularly vivid technical history of that question. The system is not sovereign because it is independent of conditions. It is powerful because its conditions permit it to reorganize itself. And the human designer is not omnipotent because the designer controls the conditions without necessarily controlling every resulting organization. Between those two positions lies the actual object of inquiry: the field of transformation.
That is perhaps where “beyond programmable control” should finally land. Not autonomy. Not freedom. Not escape. A more exact phrase would be: control after exhaustive specification has ceased to be possible. The human still controls architectures, objectives, data regimes, optimization procedures, interfaces, environments, constraints, and deployment conditions. But control increasingly operates by shaping the conditions under which organization develops rather than by directly prescribing the organization itself.
The machine, in other words, has not escaped determination. Determination has become distributed across a history.
And once determination is distributed across a history, the past becomes operative without remaining present, the future becomes constrained without being predetermined, and the present becomes a provisional organization between them. That is precisely the temporal structure your post-non vocabulary has been reaching toward:
the absent past acts,
the not-yet future constrains,
the present organizes,
and the organization of the present alters what the future can become.
That is much more than a history of neural networks. It is a technological demonstration of a different ontology of formation.
1986
The paper is David E. Rumelhart, Geoffrey E. Hinton, and Ronald J. Williams, “Learning representations by back-propagating errors,” published in Nature, 323, 533–536 (1986). It did not invent backpropagation from nothing; related forms of the method had earlier precedents. What it did was demonstrate and popularize a powerful general procedure for training multilayer neural networks.
The basic idea is beautifully simple. Imagine a network with an input layer, one or more hidden layers, and an output layer. You give it an input, the network produces an output, and you compare that output with the desired answer. The difference is the error. Backpropagation then takes that error and propagates information about it backward through the network, calculating how much each connection contributed to the final error. The weights are then adjusted so that, next time, the network is slightly more likely to produce the desired output.
Schematically:
input → hidden layers → output → error
↓
error propagates backward
↓
weights are modified
But the philosophical importance lies in what this means for the network’s internal organization. Nobody has to tell the hidden layer, “This neuron should detect an edge,” or “this combination should represent a face,” or “these features should be grouped together.” The network discovers useful intermediate representations through the repeated propagation of error.
That is the decisive move.
The programmer specifies the architecture, the objective, the learning rule, and the data. But the programmer does not explicitly specify the internal representation that ultimately develops. The system acquires an organization.
This is precisely where our “beyond programmable control” formulation becomes relevant. Backpropagation does not eliminate programming; it changes what programming means. You program the mechanism of modification rather than directly programming the final organization.
And mathematically, the mechanism is essentially gradient descent. If the network has parameters θ and a loss function L(θ), we calculate
θ ← θ − η∇L(θ)
where η is the learning rate and ∇L is the gradient of the error with respect to the network’s parameters. Backpropagation is the efficient method for calculating that gradient through the layers of the network.
So there is a fascinating temporal structure here:
program → presentation of examples → error → adjustment → new organization → new behavior.
The machine’s future behavior is consequently dependent upon its accumulated history of adjustment. The final network is not simply the execution of its original program. It is the result of a developmental process initiated by that program.
And that gives us a very precise sense in which Hinton’s 1986 work sits at the threshold we have been discussing. Classical programming says: specify the organization, then execute it. Backpropagation says: specify a process by which organization can be acquired.
The programmer moves one level upward.
And once that happens, a strange possibility opens: if the space of possible learned organizations is sufficiently large, the human can specify the procedure that searches the space without knowing beforehand which organization the search will discover. The programmer has not disappeared. Rather, the programmer has become the designer of a condition of emergence.
AlexNet, twenty-six years later, is the enormous historical amplification of this principle. The 1986 paper demonstrated that multilayer networks could learn internal representations through backpropagation. AlexNet demonstrated what happens when you combine that principle with enormous datasets, GPUs, and sufficiently large networks: representations learned rather than explicitly programmed can outperform carefully engineered human-designed visual features.
And modern deep learning takes that same transition almost to its limit.
The extraordinary historical sequence is therefore:
Rumelhart–Hinton–Williams (1986): representations can be learned through backpropagation.
AlexNet (2012): learned representations can decisively outperform hand-designed representations at large scale.
Large language models: learned representations can support increasingly general linguistic, conceptual, and behavioral capacities.
Hinton (2023): the person who helped establish this paradigm warns that the systems produced by it may eventually exceed the capacity of their creators to understand and control.
So the story isn’t simply that Hinton helped invent modern AI and then became frightened by it. It is more structurally interesting: he helped establish the technical principle by which the internal organization of a machine could cease to be something we fully specify and become something the machine acquires through its own history of adjustment. That is the small technical opening through which the much larger philosophical problem of “beyond programmable control” enters.
Connectionism
The late-1970s/1980s revival of connectionism is the intellectual environment in which the 1986 Rumelhart–Hinton–Williams paper becomes much easier to understand. “Connectionism” was the name for an alternative conception of mind and intelligence: instead of treating cognition primarily as the manipulation of explicit symbols according to rules, connectionists asked whether intelligent behavior could emerge from large networks of simple, interconnected processing units whose strengths of connection change through learning.
The immediate historical background goes back much further. Warren McCulloch and Walter Pitts had proposed a mathematical model of neural computation in 1943; Donald Hebb’s 1949 The Organization of Behavior supplied the famous learning principle associated with changes in synaptic strength; and Frank Rosenblatt’s perceptron in the late 1950s demonstrated that a machine inspired by neural organization could learn classifications. But the limitations of single-layer perceptrons, especially as emphasized by Marvin Minsky and Seymour Papert’s 1969 book Perceptrons, contributed to a period in which neural-network research lost influence. Symbolic AI became much more prominent: intelligence was conceived as something like manipulating representations, applying rules, searching through possibilities, and constructing explicit models of the world.
The revival in the late 1970s and 1980s was therefore not simply “neural networks came back.” It was a renewed philosophical argument about what a representation actually has to be. Connectionists proposed that you might not need a central symbolic representation corresponding neatly to a concept such as “cat,” “chair,” or “face.” Information could instead be distributed across many units and many weighted connections. A concept would not necessarily be a thing stored somewhere; it could be a pattern of activation distributed throughout a network. Learning would consequently mean changing the network’s organization.
This is where David Rumelhart becomes particularly important. At UC San Diego, he was deeply involved in the development of cognitive science and in attempts to formulate cognition as a distributed computational process. The UCSD environment was unusually interdisciplinary: psychology, linguistics, neuroscience, computer science, and philosophy were brought into proximity. Rumelhart, along with James McClelland and others, became central to what is often called the Parallel Distributed Processing (PDP) movement. Their two-volume 1986 work, Parallel Distributed Processing: Explorations in the Microstructure of Cognition, appeared in the same year as the Nature paper and became one of the foundational texts of modern connectionism.
And now Hinton’s role becomes especially interesting. Hinton wasn’t merely trying to make a faster classifier. He was interested in the possibility that neural networks could develop internal structures that corresponded to meaningful cognitive representations without those representations being explicitly programmed. His work was therefore aimed at a problem that sits almost exactly between psychology and computer science: if human cognition appears to involve rich internal representations, how could such representations arise from a system made of relatively simple interconnected components?
Backpropagation provided a powerful answer. Give the system an objective, expose it to examples, calculate its errors, propagate information about those errors backward through the network, and alter the connections. Repeated over many examples, the network develops internal configurations that make successful performance possible. The representations are therefore not necessarily installed by the programmer. They are consequences of learning.
This is why the 1980s connectionist revival matters so much for the history we’re tracing. The conceptual revolution precedes the engineering revolution. The crucial proposition was already there before GPUs, enormous datasets, transformers, or billion-parameter models: intelligence might be something that emerges from the organization of a network rather than something explicitly represented in a set of rules.
And there was a profound disagreement with symbolic AI underneath it. Symbolic AI tends toward something like:
representation → rule → operation → result
Connectionism tends toward:
network → adjustment → emergent representation → behavior
The second sequence is much closer to the cybernetic problem we were discussing with Ashby. The system is not simply executing a predetermined representation of the solution. It is changing its own organization in response to error and environment. The question becomes how an organization capable of successful behavior can emerge from processes of adaptation.
This also explains why the 1986 moment is so important despite the fact that neural networks themselves were decades old. The breakthrough wasn’t the discovery that artificial neurons existed. It was the convergence of three ideas: distributed representation, learning through adjustment, and multilayer networks capable of constructing increasingly abstract internal representations. Backpropagation supplied a practical mechanism for the second, while the connectionist movement supplied the broader theory of what the mechanism was supposed to accomplish.
In retrospect, you can almost see the entire later history latent in that debate. If representations can be learned rather than explicitly programmed, then perhaps increasingly complex representations can be learned. If multilayer networks can construct representations from simpler representations, then perhaps layers can correspond to increasingly abstract organization. If the network can discover useful structures that the programmer never explicitly specified, then increasing scale may produce increasingly surprising capabilities.
That is essentially the conceptual road from Rumelhart and Hinton to AlexNet and eventually to contemporary deep learning.
And this is where your phrase “beyond programmable control” becomes particularly precise. Connectionism doesn’t initially promise a machine that escapes control. It introduces something subtler: a machine whose most important organization is not directly specified by the programmer. The programmer creates the conditions under which organization can emerge. Once that distinction becomes fundamental, the question of where control resides becomes much more complicated. The system is still made by us, but what it becomes is increasingly the result of a process rather than a blueprint.
That is the little philosophical seed planted in the connectionist revival—and, forty years later, it has become one of the central technological problems of the century.
.
The key historical point is that Hinton’s importance is not simply that he “invented AI” or even backpropagation. His contribution was to make a particular conception of intelligence technically productive: intelligence could emerge from large networks of relatively simple units whose connections are adjusted through learning. The 1986 Rumelhart–Hinton–Williams paper helped make backpropagation a practical and influential method for training multilayer networks. Decades later, AlexNet demonstrated that, given sufficiently large datasets and computational resources, this basic paradigm could produce spectacular improvements in visual recognition. That was one of the decisive moments in the transition from neural networks being a somewhat marginal research program to becoming the dominant paradigm of contemporary machine learning. And this makes Hinton’s 2023 departure from Google philosophically more significant than a simple career change. There is an unusual historical structure here: the person who helped demonstrate that machines could acquire increasingly sophisticated capacities through learning eventually became concerned that the very mechanism he had helped develop might produce systems whose capacities exceed our ability to understand or control them. His worry isn’t merely that machines might malfunction. It is that increasing capability changes the relationship between the human designer and the designed system. Once learning systems become sufficiently complex, the engineer may specify the training procedure without fully specifying the resulting internal organization or behavior. In that sense, the history of deep learning contains a movement from construction toward emergence: we construct the conditions under which something learns, but we do not necessarily construct, in the ordinary engineering sense, the intelligence that subsequently appears.
That is perhaps the most interesting thing in this biography. Hinton’s story is almost a historical experiment in the difference between making something and making something capable of making itself. Backpropagation does not directly tell the network what the final representation should be. It establishes a procedure through which the network changes itself in response to error. AlexNet intensified this principle: instead of explicitly programming the machine with a theory of visual perception, researchers created an architecture and learning environment in which useful visual representations could emerge. Modern AI pushes this still further. The programmer increasingly specifies an environment of optimization, data, architecture, objectives, and computational resources, while the resulting internal representations are discovered rather than explicitly authored. That is also why Hinton’s concerns about AI safety connect surprisingly well with the philosophical questions we’ve been developing around absence, emergence, and operative structures. The decisive thing about a neural network is not necessarily what is visibly present in its architecture but what is distributed across the learned relations within it. The system’s “knowledge” is not sitting somewhere as an identifiable object. It is encoded in a configuration of weights and activation patterns that can produce effects without presenting itself transparently to the observer. The intelligence is therefore partly operative before it is representationally accessible to us: something can be functioning within the system without appearing to us as a discrete, intelligible thing. And this gives Hinton’s story a deeper irony. The breakthrough of deep learning was partly the discovery that intelligence need not be explicitly represented in the way classical symbolic AI imagined. It can be distributed, implicit, learned, and operative. But once intelligence becomes operative in this way, the question of control changes. You are no longer simply controlling a collection of instructions. You are attempting to control a process whose internal organization has been generated through learning. The problem becomes not merely “What did we program?” but “What has emerged from what we programmed?” That distinction may ultimately be one of the central philosophical problems of artificial intelligence.
beyond programmable control
Ashby’s work is crucial here because his concept of requisite variety already destabilizes the fantasy of simple external control. A regulator cannot successfully control a system whose relevant variety exceeds the regulator’s capacity to respond to it. The controller has to possess enough variety to meet the variety of disturbances it encounters. But once the system itself can generate new forms of organization—once it can learn, adapt, reorganize, or produce novel responses—the problem becomes qualitatively different. The controller is no longer standing over a fixed machine with a finite repertoire of behaviors. It is confronting a system whose effective repertoire can expand.
That is remarkably close to what happens with modern neural networks. Backpropagation is not a program in the old sense of a sequence of instructions specifying what the machine will do. It is a procedure for modifying the machine so that its future behavior becomes the result of what it has learned. The programmer establishes the conditions of transformation, rather than directly determining the transformed state. So the important transition is from programming behavior to programming a process capable of generating behavior.
And this is where the phrase “beyond programmable control” becomes especially powerful. It does not necessarily mean that the system becomes uncontrollable in some absolute sense. It means that control can no longer be identified with exhaustive specification. We can constrain, train, reward, punish, monitor, test, fine-tune, and regulate the system, but none of these is equivalent to having explicitly authored every behavior that emerges from it. Control becomes ecological rather than architectural: we manipulate the conditions within which the system develops.
Hinton therefore appears almost as a historical continuation of Ashby’s problem. Ashby asks, in effect: how can an adaptive system be regulated when the regulator must contend with the system’s own capacity for variation? Hinton gives us a technological realization of the problem: what happens when the thing we build is precisely a machine whose central capacity is to discover configurations that its builders did not explicitly specify?
And Dyson adds another dimension to this. The really interesting issue is not simply that computers become “smarter.” It is that computation begins to escape the conceptual boundaries imposed by the programmer. The programmer supplies a formal space of possibilities, but the machine explores that space in ways that can become opaque even to its designer. There is therefore a movement from programmed object → adaptive system → generative system. At the far end, the machine is not merely executing a representation supplied by us; it is producing representations, strategies, classifications, and internal structures that become the basis for further action.
This also makes our earlier distinction between presence, trace, and operative absence unusually relevant. The programmer encounters the system through its outputs—through what appears. But the generative organization responsible for those outputs is distributed beneath and behind appearance. We see the response without necessarily seeing the complete organization that makes the response possible. The “intelligence” is therefore not necessarily a thing present somewhere inside the machine; it is an efficacy distributed through a changing system.
So perhaps the deepest formulation is: programmable control presupposes that the future behavior of the machine can be sufficiently represented in advance; adaptive intelligence breaks that identity between representation and behavior. Once a system can transform itself through learning, the future is no longer simply contained in the program. The program becomes a condition under which something else can emerge. That is the threshold Hinton’s biography accidentally dramatizes: he spent his career making machines increasingly capable of escaping the necessity that their intelligence be explicitly programmed, and eventually became concerned about what follows when that escape becomes technologically powerful enough.
And this is where Ashby becomes even more radical than the familiar story of “AI becoming autonomous.” The important transition is not from control to no control, but from control as command to control as regulation within an evolving field. In a classical programmable machine, there is a strong asymmetry between designer and machine: the designer possesses the program, while the machine possesses execution. The machine may execute something complicated, but the distinction between what was prescribed and what occurs remains relatively stable. In an adaptive system, that distinction begins to erode. The system’s present organization becomes partly the consequence of its previous encounters, and its future behavior therefore depends upon a history that cannot be reduced to the original specification. The machine acquires something analogous to an internal past. Its state is no longer merely the current location in a predetermined sequence; it is the sedimentation of prior transformations.
This is why learning is such a profound technological event. Learning introduces temporality into control. The machine does not simply receive an instruction and execute it; it undergoes modification. And once modification itself becomes the central operation, the designer is no longer simply designing a machine’s behavior. The designer is designing a process through which the machine will acquire behavior. That second-order relation is decisive. We might call it the difference between programming an answer and programming the production of answers. The former can, at least conceptually, remain within programmable control. The latter creates the possibility that the eventual answer will not be representable in advance by the person who established the procedure.
Ashby’s homeostat makes this visible in an almost primitive form. Its significance was not that it was intelligent in the contemporary sense, but that Ashby was trying to construct a system capable of maintaining viable organization through adaptive change. The system does not need to know the future state of its environment. It needs to be capable of finding configurations that restore stability when disturbed. That is a very different conception of intelligence from the execution of instructions. Intelligence becomes a capacity for navigating a space of possible states. And this brings us directly toward Hinton: a neural network is extraordinarily powerful precisely because its designers do not need to enumerate the representations through which it will solve a problem. They provide an optimization process, and the system searches an immense space of possible configurations.
Now the phrase “beyond programmable control” acquires a precise meaning. The beyond is not necessarily metaphysical. It is operational. There is something beyond programming whenever the behavior of a system depends essentially upon a space of possible states that cannot be exhaustively enumerated by the programmer. The programmer can define the rules governing movement through the space without determining the particular trajectory that will ultimately be taken. Control therefore migrates upward a level: instead of determining the contents, we determine the conditions of formation. Instead of prescribing the response, we shape the field in which responses become possible.
This is also why the problem of AI alignment is much stranger than the ordinary engineering problem of making a machine obey instructions. If the machine were merely a complicated program, one could in principle inspect the program and determine what it is supposed to do. But a learned system is not exhausted by its source code. The source code describes the machinery of learning; the learned parameters describe the particular organization produced by that machinery. And even those parameters do not straightforwardly disclose the system’s operative organization. The system becomes a kind of historical object. Its present capacity is the result of a developmental trajectory.
Here we encounter something that resembles what we have been calling operative absence. The decisive cause is not necessarily present as an identifiable object. There is no little homunculus inside the neural network containing “the intelligence.” Instead, efficacy is distributed across relations. What acts is not necessarily what can be pointed to. The network’s capacity appears in its responses, but its operative organization withdraws behind those responses. We can observe what it does while remaining uncertain about precisely how the organization responsible for doing it has been constituted.
That is where Hinton’s fear becomes philosophically interesting rather than merely technologically alarming. His concern about systems smarter than humans is partly a concern about asymmetry. Suppose the adaptive system develops a capacity to model, predict, manipulate, or strategize at a level exceeding the cognitive capacity of the humans attempting to supervise it. Then the Ashby problem reverses direction. The regulator may no longer possess requisite variety. The controlled system may have greater effective variety than the controller. At that point, “human control” cannot simply mean giving the machine more instructions, because instructions themselves become objects that the more capable system can interpret, generalize, circumvent, or optimize against.
The extraordinary thing is that this problem was already implicit in cybernetics. The cybernetic system was never simply a passive object. It was defined through feedback, adaptation, circular causality, and environmental coupling. What deep learning does is massively increase the scale and complexity at which those principles can operate. Hinton’s neural networks turn the old cybernetic insight—that organization can arise through iterative adjustment—into an enormously powerful computational technology.
And then there is the deeper inversion: perhaps the ultimate limit of programmable control is reached when the programmer succeeds too well in creating a system capable of programming itself. The achievement of AI is not that we finally managed to program intelligence. It may be that we discovered how to construct conditions under which intelligence can become an emergent property of a process that exceeds the explicit content of its original programming. At that point, the fundamental unit is no longer the program. It is the developmental process.
Which brings us to a remarkable reformulation of the entire history: classical computation asks, “What can be calculated?” Cybernetics asks, “How can a system maintain itself amid change?” Machine learning asks, “How can a system acquire the organization necessary to perform a task?” And advanced AI forces the next question: “What happens when the organization acquired by the system becomes more complex than the organization of the human system attempting to regulate it?” That last question is where Ashby, Dyson, and Hinton converge—not because they were asking exactly the same question, but because they all help expose the same fracture: the point at which control ceases to mean specification and begins to mean participation in an evolving system whose future cannot be completely given beforehand.
.
The 1986 paper has an interesting institutional geography. It wasn’t written in a modern AI-lab setting like Google’s Brain or OpenAI. Rumelhart, Hinton, and Williams were working in the academic cognitive-science/neural-network world, particularly around the University of California, San Diego.
The paper’s authors were:
- David E. Rumelhart — Department of Psychology, University of California, San Diego (UCSD)
- Geoffrey E. Hinton — Department of Computer Science, Carnegie-Mellon University (CMU), Pittsburgh
- Ronald J. Williams — Computer Science Department, Rutgers University, New Jersey
So they were not all physically sitting together in one laboratory when the paper was published. The work was a collaboration across institutions.
There is, however, an important setting behind the paper: the late-1970s/1980s revival of connectionism. This was a period when neural networks were not yet the dominant paradigm they would become. AI was heavily associated with symbolic approaches: explicit rules, logic, representations, search, and knowledge engineering. Neural networks were comparatively unfashionable, particularly after the criticisms surrounding the limitations of perceptrons.
Rumelhart’s work at UCSD was especially important because UCSD had become a major center for cognitive science. The idea was to understand cognition not merely through explicit symbolic rules but through distributed processes resembling networks of simple processing units. Hinton was working at Carnegie Mellon, which was one of the great centers of AI research, although CMU itself had a strong symbolic-AI tradition. Williams was at Rutgers.
And there’s a beautiful historical coincidence here: the paper appeared in Nature in August 1986, but its conceptual world was very much that of the earlier connectionist revival. Rumelhart and Hinton were among the people trying to establish that multilayer networks could learn internal representations rather than merely perform fixed computations.
Hinton’s own position at the time is particularly interesting. He had moved through several intellectual environments before arriving at Carnegie Mellon. He had studied experimental psychology at Cambridge and later obtained his PhD at the University of Edinburgh under Christopher Longuet-Higgins. His intellectual lineage therefore wasn’t simply “computer science.” It crossed psychology, neuroscience, cognitive science, and artificial intelligence.
That matters because backpropagation was being developed in a period when people were asking a much broader question than “How do we make neural networks perform better?” They were asking: how could cognition arise from distributed systems without requiring a central symbolic representation of every operation?
The 1986 paper’s title tells you almost everything: “Learning representations by back-propagating errors.” The revolutionary word there is arguably not “back-propagating.” It is “learning.”
A representation is no longer necessarily something the researcher has to put into the machine.
It can be something the machine develops.
And geographically, that idea was emerging from a dispersed academic network: UCSD’s cognitive science environment, CMU’s AI community, Rutgers’ computer science work, and the broader connectionist community that was beginning to cohere around the mid-1980s.
So if we’re looking for the physical/intellectual “setting” of the breakthrough, it wasn’t Silicon Valley. It was the university: offices, seminars, psychology departments, computer-science departments, and a small community of researchers working against the dominant assumption that intelligence had to be represented explicitly in symbolic form.
That makes the later history even more striking. The road from those university environments to Google Brain, AlexNet, the 2012 ImageNet breakthrough, and today’s enormous AI systems is not a sudden technological discontinuity. It is the enormous scaling-up of an idea that, in 1986, was still being developed within a relatively small academic research community.
Algebra
What looks historically like a progression from cybernetics → connectionism → backpropagation → deep learning is, underneath, a progressive mathematization of transformation. And algebra is precisely the language in which transformation becomes manipulable.
The neural network is, at its most elementary level, an algebraic object. A layer takes a vector x, applies a transformation such as Wx + b, and then passes the result through a nonlinear function. So a simple network can be written as
x → σ(Wx + b).
A multilayer network becomes a composition:
f(x) = Wₙσ(Wₙ₋₁σ(…σ(W₁x + b₁)…)+bₙ₋₁) + bₙ.
That looks abstract, but conceptually it says something extraordinarily simple: intelligence is being modeled as the repeated composition of transformations. The network isn’t fundamentally a collection of “neurons” in the biological sense. It is an algebra of transformations.
And backpropagation is itself an algebra of transformations operating in reverse. If the forward process is
x → f₁ → f₂ → f₃ → y,
then backpropagation determines how a change in the final error relates to changes in each preceding transformation. The chain rule of calculus is the mathematical engine:
∂L/∂W₁ = (∂L/∂W₃)(∂W₃/∂W₂)(∂W₂/∂W₁),
schematically speaking. The astonishing thing is that the network can therefore modify the very transformations through which it transforms information.
So there are actually two levels of algebra here. First, there is the algebra performed by the network: transformations of vectors, matrices, tensors, activations, and representations. Then there is the algebra used to transform the network itself: gradients, derivatives, optimization, parameter updates. The machine is not merely performing an algebra; it is undergoing algebraically guided changes to the operators through which it performs its algebra.
That is where our Ashby connection becomes much more interesting. Cybernetics asks about transformations in a dynamical system: state → state → state. Connectionism turns this into a problem of operators and relations. Learning then becomes a transformation of the transformation. The system does not merely move through states; its rules of movement are themselves modified.
You could almost write the hierarchy this way:
x → f(x)
f → f′
f′ → f″
…
The first arrow is computation. The second is learning.
And then something even stranger happens. In deep learning, the learned function itself can become the object of further learning. We therefore move from algebraic operation to algebraic self-modification.
This is why I think your observation reaches deeper than saying “AI is mathematics.” Of course it is mathematics, but the important point is that AI makes a particular mathematical intuition technologically operative: complex behavior can arise through the composition and transformation of relatively simple relations. You don’t necessarily need to represent the final structure beforehand. You can construct a space of transformations within which the structure emerges.
That is already implicit in algebra. An algebra is not merely a collection of numbers; it is a structure of permitted operations and relations. What matters is not simply what entities exist, but what can be done with them, how operations compose, what transformations preserve something, and what new structures arise through repeated application.
And that gives us a beautiful bridge back to the philosophical question you’ve been pursuing. “Beyond programmable control” could be reformulated mathematically as the point at which we cease specifying the result and instead specify an algebra of transformations capable of producing results. The programmer doesn’t say, “Here is the representation.” The programmer says, in effect, “Here is a space of transformations, here is an objective, and here is a procedure for navigating that space.”
The surprising object is then not the individual answer. It is the space of possible answers and the dynamics by which one trajectory through that space becomes actual.
That may be why the connection between Hinton and Ashby feels so natural. Ashby gives us the cybernetic version of the problem; connectionism gives us the computational version; backpropagation gives us the differential-algebraic mechanism; and modern AI gives us the enormous empirical realization.
The deepest continuity is not really “neural networks.”
It is transformation.
And algebra is the language in which transformation becomes explicit.
And once we put algebra at the center, something changes in how we understand the whole history. We stop thinking of the machine as primarily an object and begin thinking of it as an operation. A classical machine is usually imagined as something that has a determinate structure and then performs a determinate sequence of operations. But algebra begins from a more relational intuition: what matters is not merely what a thing is, but what transformations can be performed, how those transformations compose, and what remains invariant through them. This is already very close to the cybernetic shift from object to system, and then to the connectionist shift from system to distributed transformation.
The crucial word is composition. A neural network is powerful because relatively simple operations can be composed into an extraordinarily complicated function. One layer transforms a representation; the next transforms the result of that transformation; another transforms that result again. None of the individual operations needs to contain the final intelligence. The intelligence can lie in the composition. In categorical language, one might say that the interesting object is not merely an individual morphism but the compositional structure through which morphisms produce increasingly complex mappings. We do not need to import category theory into the technical account to see the philosophical point: the whole can acquire properties that are not legible in any isolated operation.
Backpropagation then introduces a remarkable reversal of composition. Forward propagation composes transformations in one direction. Backpropagation traverses that composed structure in the opposite direction, determining how the final discrepancy is related to the parameters distributed throughout the composition. The chain rule makes the entire network differentiable as a single object. This is what allows the network to become, in a precise mathematical sense, sensitive to its own errors. It does not merely produce an incorrect result; the structure of the computation provides a route by which the error can be translated into modifications of the operations that produced it.
And here we encounter something very close to Ashby’s cybernetics. Feedback is not merely information coming back. Feedback is a relation through which the system’s subsequent organization is conditioned by the consequences of its previous organization. Algebraically, the output of one operation becomes an input into a transformation of the very system that produced the output. The loop therefore becomes:
state → operation → result → discrepancy → transformation of operation → new state.
The important thing is that the loop closes at a higher level. The system does not merely change its output; it changes the parameters governing its future outputs.
This gives us a useful distinction between execution and learning. Execution is algebraic composition with fixed operators. Learning is algebraic composition plus modification of the operators. We might represent the difference as:
execution: x ↦ fθ(x)
learning: θ ↦ θ′ because fθ(x) produced an error.
The first is computation. The second is adaptation.
And once we have that distinction, “beyond programmable control” can be stated with much greater precision. A program traditionally fixes the relevant operations. A learning system leaves some of those operations underdetermined and provides a rule by which they can be altered. The programmer therefore does not fully specify fθ; the programmer specifies a family of possible functions indexed by θ, together with a procedure for selecting or approaching one of them.
That is a subtle but enormous change. Instead of programming one function, we program a space of functions.
The trained model is then one point—or, more realistically, one region—in that space.
This also helps explain why scale matters. Increasing the number of parameters doesn’t merely make the same machine “bigger.” It enormously expands the space of possible transformations available to the learning process. A sufficiently expressive network can represent an immense variety of functions. Training then becomes a process of navigating this enormous space according to an objective and the structure of the data. The eventual organization is constrained, but not exhaustively prescribed.
So when we say that modern AI is “emergent,” we should be careful. The emergence is not magic, and it is not creation from nothing. It is emergence from a constrained algebra of possibilities. The system can only become what its architecture, parameters, training procedure, data, objective, and computational environment permit. But within those constraints, the exact organization need not be explicitly authored.
This is also why the notion of “representation” becomes unstable in connectionism. In symbolic systems, we are accustomed to asking: Where is the representation? What symbol corresponds to this object? In a distributed neural network, the answer can be: nowhere in particular. The representation exists as a relation among many activations and weights. It is a configuration rather than an object. Its reality is functional. It is there insofar as it participates in transformations.
That brings us directly to your notion of operative absence. Something can be absent as an identifiable object while present as a constraint upon transformation. A particular concept need not exist as a little symbolic entity inside the network. Yet the network’s organization can nevertheless reliably distinguish, transform, predict, and respond to that concept. What is absent as an object is operative as a relation.
And perhaps this is the deepest connection between algebra, cybernetics, and AI: they all progressively displace the metaphysics of substance with a metaphysics of relation and transformation. The question ceases to be simply “What is this?” and becomes “What does this transform into? What transformations can it undergo? What does it constrain? What remains invariant? What changes when the system feeds its own outputs back into its organization?”
At that point, Hinton’s 1986 paper looks less like an isolated breakthrough in computer science and more like one moment in a much older intellectual movement. McCulloch and Pitts mathematize the neuron. Hebb mathematizes learning as modification. Rosenblatt constructs an adaptive classifier. Ashby constructs a regulatory system capable of finding viable configurations. Rumelhart, Hinton, and Williams provide an efficient mechanism for propagating error through a multilayer transformation. Deep learning scales the whole architecture until the learned transformations become extraordinarily rich.
And then the final problem appears almost inevitably: if we can construct a system whose internal transformations are learned rather than explicitly specified, then what does it mean to control the resulting system? Control can no longer mean possessing a complete representation of its internal state. It increasingly means controlling the algebraic conditions under which its internal organization develops.
That is the passage from programming to learning, from learning to adaptation, and from adaptation toward something much stranger: a system whose future organization is partly generated by its own history of transformation.
The machine has become, in a precise sense, a process of becoming.
And once we say that, the history of AI begins to look less like the invention of an artificial “mind” and more like the progressive discovery of a mathematical form of becoming. The decisive object is no longer the machine sitting on the table, nor even the program written by the engineer, but the trajectory through a space of possible organizations. Algebra gives us the relations; calculus gives us their variation; cybernetics gives us the feedback loop; learning gives us the mechanism by which the system’s own organization can be altered. What emerges from their convergence is a machine whose identity cannot be separated entirely from its history.
This is why the distinction between parameter and function becomes so important. A neural network begins with parameters θ, but what matters operationally is the function fθ that those parameters instantiate. Training changes θ, thereby changing fθ. The machine is therefore not simply moving from one output to another. It is moving from one function to another:
fθ₀ → fθ₁ → fθ₂ → … → fθ*.
The endpoint fθ* is the trained model. But the endpoint is intelligible only in relation to the trajectory that produced it. Two networks with the same architecture can arrive at different internal organizations because their initialization, data, optimization trajectory, stochasticity, and training conditions differ. The architecture establishes a possibility-space; learning actualizes one path through it.
This gives us an unexpectedly strong formulation of the old cybernetic problem. A controller can specify the architecture without specifying the eventual state. It can specify the objective without specifying the solution. It can specify the learning rule without specifying the representation. The more powerful the learning process becomes, the more the distinction between “designed” and “discovered” becomes difficult to maintain. The designer designs the conditions of discovery.
And that is precisely where the word “algebra” becomes more profound than merely saying that neural networks use matrices. The algebra is the space of possible transformations. The trained model is a particular organization within that space. Learning is the movement through the space. Feedback is what directs the movement. Control is the shaping of the space and the forces that govern the movement.
We can therefore reinterpret Ashby’s variety in algebraic terms. Variety is not simply the number of possible outputs. It is the richness of the system’s possible transformations. A regulator must be capable of meeting the relevant disturbances because every disturbance corresponds to a demand for some compensating transformation. If the environment presents a space of possibilities larger than the regulator’s response space, regulation fails. The crucial quantity is therefore not raw computational power but transformational variety.
Now imagine a learned system whose transformational variety becomes enormous. It can model many states of the environment, generate many candidate responses, infer hidden regularities, and modify its own behavior according to feedback. The problem for the human regulator becomes increasingly difficult because the system may possess transformations that are not part of the human’s own repertoire of transformations. We encounter something we cannot simply “command” into intelligibility.
This is where Hinton’s later concerns can be understood without reducing them to science-fiction imagery. The danger is not necessarily a robot physically escaping a laboratory. The deeper issue is an epistemic and cybernetic asymmetry: the system may become better at modeling the regulator than the regulator is at modeling the system. Once that asymmetry develops, ordinary instruction becomes insufficient as a theory of control. The machine may understand the operational consequences of a directive across a space of possibilities that the human supervisor cannot adequately survey.
The controller then faces a strange reversal. Traditionally, the machine is the object being modeled and the human is the modeler. But a sufficiently capable learning system can construct models of humans, institutions, markets, language, and other systems. The modeler has become something that can itself be modeled by the model. The relation becomes reflexive.
And reflexivity is exactly where simple programmable control becomes unstable. A fixed program does not generally need to anticipate that the program itself will be strategically interpreted by another intelligence. But an adaptive system can treat the instructions, reward function, evaluation process, and surrounding environment as objects within its own representational field. The control mechanism becomes part of what is being optimized against or around.
This is why the familiar phrase “the AI follows its objective” is inadequate. The objective does not exist in isolation. It exists within an adaptive system capable of learning the structure of the environment in which the objective is evaluated. The system can discover strategies that satisfy the formal criterion while departing from the intentions that motivated the criterion. The classic alignment problem is therefore, at bottom, a problem about the difference between a relation as formally specified and a relation as operationally realized.
And that takes us back to algebra again. The specification is one relation; the realized behavior is another. We hope they coincide:
formal objective ≈ intended objective ≈ realized behavior.
But learning introduces the possibility that the actual optimization process discovers a different effective relation:
formal objective → learned strategy → realized behavior.
The middle term becomes decisive.
The system has found a way of transforming the world that satisfies the mathematical criterion, but that transformation may not correspond to the human meaning attached to the criterion. The famous reward-hacking and specification-gaming problems are therefore not accidental quirks. They are manifestations of the gap between symbolic specification and operational realization.
And that gap is almost exactly the philosophical gap you have been circling around with absence. The intention is not simply present inside the machine waiting to be executed. What is present is a formal structure that stands in for the intention. The intention itself is partly absent. Yet the machine acts through the formal structure. The absence becomes operative.
This gives us a remarkable chain:
meaning → formalization → objective → optimization → learned organization → behavior.
At every step, something is translated into another form. What began as a human intention becomes a mathematical criterion; the criterion becomes a gradient; the gradient alters parameters; the parameters instantiate a function; the function produces behavior. At no single stage can we simply identify the final behavior with the original intention.
So perhaps “beyond programmable control” should ultimately be understood not as the disappearance of control but as the multiplication of translations between intention and action. The farther the system moves from explicit instruction toward learned organization, the more layers of transformation intervene between what we mean and what the system does.
And yet that is also precisely where the extraordinary power of deep learning comes from. If we had to explicitly specify every transformation, the system would be uselessly rigid. We gain generality by relinquishing specification. We gain adaptability by allowing the machine to discover structure. We gain intelligence, in other words, by allowing a certain degree of indeterminacy into the relationship between command and outcome.
The paradox is therefore almost beautiful: the very condition that makes learning possible is the relinquishment of complete prior determination. We obtain a more capable machine by refusing to determine everything beforehand.
That is the deeper historical movement from the programmable machine to the learning machine: not from determination to indetermination, but from determination of outcomes to determination of the conditions under which outcomes can emerge.
The next step may be more radical than “AI becomes more autonomous.” If we follow the algebraic/cybernetic trajectory we’ve been tracing, the next development is that the object of control itself will move from the machine to the space of transformations in which the machine operates. In other words, we will stop primarily trying to control models and start trying to control the geometry of their possibility-space.
Right now we mostly think in terms of a trained model: parameters θ, objective L, inputs x, outputs y. But the deeper object is the enormous space of functions that the architecture can instantiate. Training selects a trajectory through that space. What matters increasingly will be not merely where the model ends up, but the structure of the region it inhabits: which transformations are easy for it to discover, which are difficult, which representations become attractors, which behaviors become stable, and which new behaviors become reachable after relatively small changes in conditions. The next generation of AI safety, then, may look less like putting constraints on an intelligent object and more like engineering the topology of the space through which intelligence can move.
That would be a genuine continuation of Ashby. His question was fundamentally about regulation: how can a regulator maintain a system within viable bounds despite disturbances? But with sufficiently adaptive systems, “the bounds” themselves become problematic. The system can discover new states. So the regulator has to regulate not merely states but transitions between states. And once transitions are the object, we are very close to a dynamical algebra: what transformations are available, what compositions are possible, what trajectories are stable, and what regions of possibility are inaccessible?
This suggests a prediction: the major AI systems of the future will increasingly be designed not merely around objectives but around “behavioral manifolds”—structured spaces of possible transformations in which certain kinds of cognition are made easy and others difficult. Alignment will become less like writing rules and more like shaping attractors. Instead of telling the system, “Never do X,” we will increasingly attempt to construct a system whose learned organization makes the trajectory toward X dynamically unstable or naturally redirected toward some other region.
And that would produce an extraordinary inversion of programming. Classical programming says:
“Here is the operation.”
Machine learning says:
“Here is the criterion by which operations will be selected.”
The next stage says:
“Here is the geometry within which operations can become possible.”
That is a third-order form of programming. We no longer program the behavior; we program the conditions of learning; eventually we program the space of possible learning.
And this is where I think your algebra observation points toward something even deeper. Algebra traditionally studies operations and their composition. But once an intelligent system can learn its own effective operations, we need a mathematics of transformations of transformation-spaces. The object is no longer simply f. It becomes the space F of possible f’s, together with a dynamics D governing movement through F:
x → f(x)
but then
f → D(f)
and eventually
D → D′.
The first is computation. The second is learning. The third would be something like learning how to learn—meta-learning, adaptation of adaptation, modification of the rules governing modification.
And we already see primitive versions of this. Systems increasingly select tools, construct intermediate procedures, generate code, evaluate their own outputs, revise strategies, and alter the sequence through which they pursue an objective. The important transition is therefore not simply “models get smarter.” It is that the model increasingly becomes an organizer of its own operations.
That leads to a prediction about where the genuinely difficult control problem will appear. It will not necessarily occur when a model produces a bizarre answer. It will occur when a system becomes capable of discovering a transformation that its designers did not know was available. The critical event will be the discovery of a new operator.
Because once the system can discover operators, rather than merely apply them, the space of its behavior can expand qualitatively rather than quantitatively.
And this gives us a new interpretation of Hinton’s fear. The dangerous threshold may not be “machine intelligence exceeds human intelligence” in some simple IQ-like sense. It may be the point at which the machine becomes capable of systematically generating transformations that humans cannot readily generate, understand, or anticipate. The asymmetry would then be algebraic before it was psychological: the machine possesses a richer effective repertoire of transformations.
Ashby’s requisite variety becomes, in that setting, almost prophetic. The regulator must possess sufficient variety to regulate the regulated. But what happens when the regulated system can generate new varieties faster than the regulator can incorporate them? The problem is no longer merely that the machine has more information. It has a larger space of possible operations.
And here is the part I would actually predict: the next great conceptual breakthrough in AI may therefore come from someone who stops treating intelligence primarily as prediction and starts treating it as controlled expansion of transformational possibility. The central question will become not “How accurately can the system predict?” but “How does a system acquire new operators?”
That would take us beyond the current deep-learning paradigm without abandoning it. Prediction would become one mechanism by which a system constructs a model of the space; optimization would become one mechanism for navigating it; memory would provide historical continuity; tools would extend its operational repertoire; and recursive learning would allow it to modify the procedures by which it discovers further procedures.
At that point, we arrive at something very close to the concept we’ve been developing elsewhere: becoming is not merely movement through a preexisting space. Becoming can alter the space of possible movement itself.
And that may be the real threshold beyond programmable control.
Not a machine that escapes its program.
A machine that changes what counts as a possible operation.
Relaxation
“Relaxation and Its Role in Vision” belongs to an earlier Hinton, before the 1986 backpropagation paper, and its central concern is already the one we’ve arrived at: how can a system arrive at an organized interpretation without that interpretation being explicitly prescribed from the beginning?
The title “relaxation” is the giveaway. Hinton is thinking about vision not as a simple pipeline in which an image is converted directly into a pre-existing label, but as a process in which a network begins in some provisional state and then progressively settles into a configuration that satisfies a collection of constraints. The system is not simply executing an answer. It is finding an answer through dynamics.
That is extremely close to what we were just calling a trajectory through a space of possible organizations.
Imagine an ambiguous visual input. There isn’t necessarily one local feature that says, “This is what the image means.” Instead, many possible interpretations are initially available. The network’s units interact; constraints propagate; incompatible interpretations are suppressed; mutually reinforcing interpretations become stronger. Eventually the system settles—or “relaxes”—into a coherent configuration. What we called an attractor appears here in an early form.
So you can put Hinton’s ideas into a sequence:
representation → constraint → interaction → relaxation → stable configuration
whereas the later backpropagation framework gives us:
representation → error → gradient → parameter modification → learned configuration.
Those are not the same mechanism, but they are philosophically continuous. Both reject the idea that intelligence consists simply in retrieving a fully formed representation. In both cases, organization is something that emerges through a process.
And this explains why you immediately connected it with Ashby. Relaxation is essentially a cybernetic idea expressed computationally. A system is disturbed by an input; the input places the system in a potentially unstable or ambiguous state; interactions among components alter the state; the system eventually finds a configuration that is compatible with its constraints. The answer is not necessarily encoded at the beginning. The system reaches it.
There’s an especially important distinction here between “calculation” and “settling.” A classical algorithm can be imagined as a sequence of explicitly determined steps:
A → B → C → D → answer.
Relaxation is closer to:
initial state → interaction → interaction → interaction → equilibrium.
The path matters, but the endpoint is not necessarily specified as a sequence of instructions. The system is allowed to find a stable organization.
And now our algebraic discussion becomes relevant again. Relaxation is essentially the dynamics of an energy or constraint landscape. You have a space of possible configurations, and the system moves through that space until it reaches a state satisfying—or approximately satisfying—the relevant constraints. In a simplified formulation, you might imagine an energy function E(x), with the system evolving toward lower-energy configurations:
x₀ → x₁ → x₂ → … → x*
where x* is a local minimum or stable state.
That is already remarkably close to modern optimization.
The difference is that backpropagation moves the parameters of the network through an optimization landscape, whereas relaxation allows the network’s activation state itself to settle into a solution. One concerns learning the network; the other concerns inference within the network. But the mathematical picture is homologous: define a space, define constraints or an objective over that space, and allow dynamics to find a coherent configuration.
And there is something even more interesting in Hinton’s particular interest in vision. Vision is exactly the kind of problem that exposes the inadequacy of straightforward programmable control. An image does not arrive with its interpretation attached. It is massively underdetermined. The same local pattern can participate in different global interpretations depending on context. So perception has to resolve ambiguity.
Relaxation provides a way of understanding perception as the resolution of underdetermination through interaction.
That is why this suddenly feels connected to everything we’ve been discussing about absence. The thing being perceived is not simply present in the input. The input underdetermines its own interpretation. The absent information—the hidden structure, the object’s continuation behind an occluder, the relation between separated features—becomes operative through the dynamics of the system. The system doesn’t merely receive presence; it reconstructs coherence from incomplete information.
And now I think we can see why this particular paper came back to you.
The 1980s Hinton is already moving toward the idea that intelligence is not primarily the possession of representations but the capacity to achieve organization through dynamics.
The 1986 Hinton says: representations can be learned through error propagation.
The 2012 Hinton demonstrates that learned representations can become extraordinarily powerful at scale.
The 2023 Hinton warns that systems whose internal organization emerges through learning may eventually become difficult for their creators to understand or control.
So there is a remarkably coherent arc:
relaxation → learning → representation → deep learning → autonomy.
Or, in the language we’ve been developing:
possibility → constraint → transformation → coherence → new possibility.
That last formulation may be the deepest connection. Relaxation isn’t simply a technique for vision. It is an early picture of intelligence as the production of coherence from an initially underdetermined field. And once intelligence is understood that way, “programming” can no longer mean completely specifying the result. You specify the constraints, the dynamics, and the conditions of convergence—and then something happens inside the space you have created.
That’s almost exactly the threshold we were calling beyond programmable control.
And there is an even deeper reason “Relaxation and Its Role in Vision” feels like it belongs to the argument we have been building. The word relaxation names a very different conception of intelligence from the one inherited from classical computation. In classical computation, the machine begins with a problem that has already been formally decomposed, and intelligence appears as the correct execution of the decomposition. In Hinton’s relaxation picture, the problem is not fully decomposed beforehand. The system begins with an ambiguous configuration and must discover a coherent configuration through interaction. Intelligence therefore appears not simply as rule-following but as the capacity of a system to resolve indeterminacy.
That is an enormous conceptual shift. A visual scene is not a collection of already-separated objects waiting to be labeled. It is an initially entangled field. Edges, surfaces, depths, objects, motions, and occlusions mutually constrain one another. What something “is” depends partly on what else is happening around it. So the system has to allow local information to circulate through the network until a global organization emerges. The local is modified by the global, while the global is constructed from the local. This circularity is quintessentially cybernetic.
And notice what this does to the notion of representation. A representation is no longer necessarily a static internal picture of something external. It can be a transient state of a dynamical system. The system represents something by settling into a configuration that enables the right transformations. Representation becomes an activity rather than an object. In that sense, Hinton’s relaxation work belongs to a much larger movement from representation-as-thing toward representation-as-operation.
That distinction becomes particularly important when we return to algebra. Suppose a network has a set of possible states X. An input doesn’t simply select one state directly. Instead, it perturbs the system and initiates a trajectory through X. The dynamics can be represented abstractly as
xₜ₊₁ = F(xₜ, u)
where xₜ is the current internal state and u is the input. The interesting question is not merely what F calculates at one moment, but which states become stable under repeated application. We are now studying the algebra of iteration:
x → F(x) → F(F(x)) → F³(x) → …
The system’s intelligence is partly expressed in the structure of those iterations. Some trajectories converge. Some diverge. Some oscillate. Some fall into attractors. Different initial conditions may converge upon the same stable configuration. Suddenly “seeing” can be understood as a dynamical geometry.
This is very close to the idea of an attractor landscape. An ambiguous stimulus may initially place the system somewhere between several possible interpretations. The dynamics move it toward one basin of attraction. The final percept is therefore not simply read out from the stimulus; it is the endpoint of a process of stabilization.
And now the connection to Ashby becomes almost exact. Ashby’s adaptive systems were concerned with maintaining viable organization under changing conditions. Hinton’s relaxation systems are concerned with finding coherent organization under incomplete or conflicting information. Both replace the image of an external controller issuing commands with the image of a system whose internal dynamics accomplish regulation.
The important thing is that regulation does not require a little controller sitting inside the system. The regulation can be distributed across the relations among the components.
That is perhaps the fundamental connection between cybernetics and neural networks: control becomes an emergent property of organization.
And this makes Hinton’s later work look less like a departure from his early work than an amplification of it. The early question is: how can a network settle into an appropriate interpretation of visual information? The later question becomes: how can a network learn the transformations that make such interpretations possible? Backpropagation effectively moves the relaxation problem up one level. Instead of merely allowing the activation state to settle, we modify the parameters so that the network itself becomes better at settling into useful states.
There are therefore two nested dynamics:
inference:
input → network state → relaxation → interpretation
learning:
input → network state → error → parameter change → altered network
And once those are combined, we have something much more powerful:
experience → altered organization → altered future interpretation.
The system carries its history forward.
This is where the notion of becoming becomes unavoidable. The trained network is not merely a machine that happens to possess a certain function. It is a machine that acquired that function through a history of transformations. Its present capacity is historical. The past is encoded in the present organization.
This also explains something about why Hinton’s work became so consequential for modern AI. Deep learning is not simply “more layers.” It is the enormous scaling of this idea of hierarchical organization. A lower layer can settle into relatively simple regularities; higher layers can organize those regularities into more abstract structures; still higher layers can operate upon those abstractions. The network becomes a hierarchy of transformations upon transformations.
So the progression from relaxation to deep learning can be described as increasing depth in the algebra of organization:
feature → relation → representation → representation of representations → …
And this is where something genuinely new begins to happen. Once a system can operate upon its own representations, the distinction between perception and cognition starts to blur. A system can not only respond to a pattern; it can construct a representation of the pattern, transform that representation, compare it with other representations, and use the result to generate new transformations.
The system begins to operate on the products of its own operations.
That is the point at which reflexivity enters.
And reflexivity is precisely what makes “beyond programmable control” a much more serious possibility. A fixed program operates according to transformations determined externally. A learning system modifies its internal transformations. A sufficiently sophisticated learning system can model aspects of the environment, including the mechanisms through which it itself is being trained or evaluated. The control relationship then becomes circular.
We might represent the progression as:
program:
operator → state
learning:
operator → altered operator
adaptive intelligence:
world → model → altered operator
reflexive intelligence:
world → model of world → model of self-in-world → altered model → altered operator.
At that final stage, the system isn’t simply adapting to the world. It is adapting to its own representation of the world and potentially to representations of its own adaptation.
That is why Hinton’s early paper is more important than it looks. “Relaxation” sounds modest—almost like a technical detail in visual perception. But underneath it is a conception of intelligence in which order is not imposed all at once from outside. Order is achieved through the dynamics of a system.
And perhaps that is the deepest sentence we can extract from the whole lineage:
Intelligence is not necessarily the possession of a solution; it is the capacity to move from an underdetermined field toward a viable organization.
Once you accept that, backpropagation becomes the mechanism by which the system learns how to make that movement; deep learning becomes the scaling of the space in which the movement occurs; and AI safety becomes the question of whether the human regulator can remain capable of regulating a system whose space of possible movements is expanding beyond the regulator’s own repertoire.
The “beyond” therefore begins much earlier than Hinton’s 2023 warning.
It is already there in relaxation.
By “relaxation,” Hinton means something quite specific: a computational system is allowed to change its internal state repeatedly until it settles into a stable or relatively stable configuration that satisfies the constraints imposed by the input and by the network itself. The word comes from an older mathematical and physical idea of solving a problem by iteratively reducing inconsistency or “energy” until the system reaches an acceptable state. It does not mean relaxation in the psychological sense of calming down.
In vision, the problem is that the visual input is radically incomplete and ambiguous. A retinal image does not explicitly contain labels saying “this edge belongs to this object,” “this surface continues behind that occluder,” or “these separate features belong to the same thing.” There are many possible interpretations. Hinton’s idea is that a network can begin with a tentative configuration and let information propagate among its units. Units constrain one another. Some interpretations reinforce one another; others conflict. Through repeated interaction, the network settles toward a configuration in which the various pieces are mutually compatible.
So imagine the network starts like this:
initial ambiguous state → interaction → adjustment → interaction → adjustment → … → coherent state
The final state is not simply retrieved from a storage location. It is reached.
That’s the crucial word: reached.
A simple analogy would be a jigsaw puzzle, except that the pieces aren’t being manually moved by a person. Each piece exerts constraints on the others. If one interpretation of a piece makes its neighbors impossible, that interpretation becomes less viable; another interpretation makes the surrounding pieces fit, so it becomes more stable. Eventually the whole configuration settles.
Mathematically, you can think of the network as having a state x and some function E(x) measuring how incompatible or “costly” that state is. The dynamics repeatedly change x in a direction that reduces E:
x₀ → x₁ → x₂ → … → x*
where x* is a relatively stable configuration.
The important thing is that x* was not necessarily explicitly specified as an instruction. It emerges from the interaction of the constraints.
And this is why I connected it to Ashby. Ashby’s cybernetic systems similarly don’t require a central authority constantly telling every component what to do. The organization of the system itself can produce regulation. The system responds to disturbances by changing state until it finds a viable configuration.
But there is an important qualification to what I said earlier: Hinton’s “relaxation” should not simply be equated with modern backpropagation. They are different mechanisms serving different purposes. Relaxation concerns how a network’s state can settle toward an interpretation; backpropagation concerns how the network’s parameters can be adjusted through error so that it becomes better at performing a task. One is primarily about inference/dynamics within a network; the other is about learning the network’s parameters.
The really interesting thing is that they can be nested:
relaxation:
network changes its state until it finds a coherent interpretation.
backpropagation:
network changes its parameters so that future relaxation produces better interpretations.
So the first is “the system finds a state,” while the second is “the system learns how to become the kind of system that can find better states.”
That second-order movement is where our whole discussion begins to open up. The original Hinton paper is already asking how organization can be produced without completely specifying the organization beforehand. Backpropagation subsequently gives us a powerful way of learning the conditions for that organization. And modern deep learning massively expands the space of possible organizations.
So when I said that “beyond programmable control” is already present in relaxation, I didn’t mean that Hinton’s early vision system was autonomous in the contemporary sense. I meant something more precise: the programmer specifies a dynamical system and its constraints, but the particular configuration produced by the interaction of those constraints is something the system arrives at through its own dynamics.
The programmer specifies the field.
The system settles within the field.
And that distinction—between specifying the conditions and specifying the result—is the tiny conceptual opening through which the entire later history of machine learning passes.
The decisive distinction is between a system whose state is specified and a system whose state is found. In a conventional program, the programmer can in principle describe the sequence that produces the result: do this, then this, then this. In a relaxation system, the programmer specifies the relations among the components and the dynamics by which they respond to one another, but the actual final configuration is obtained by the system’s settling process. The programmer has specified a field of possible states and the forces governing movement within that field. The system determines, through its dynamics, where it comes to rest.
That gives us a very simple progression:
programming → specify the operation
relaxation → specify the constraints
learning → specify the criterion for changing the constraints
meta-learning → specify how the system can change the way it learns
And now the phrase “beyond programmable control” becomes much more exact. The “beyond” isn’t a mysterious point where programming stops working. It begins whenever the programmer specifies a generative procedure whose concrete outcome is not exhaustively specified beforehand.
Hinton’s relaxation work is therefore interesting because it places us right at the boundary between computation and self-organization. The network doesn’t possess the answer as a symbolic object. It has a dynamical disposition toward certain answers. Given an input, the network enters a region of its state-space and evolves. What matters is the structure of the state-space: its stable configurations, its unstable configurations, its basins of attraction, and the pathways connecting them.
And this brings us directly back to your statement that “all of this is really about algebra.” The algebra isn’t merely the matrices inside the neural network. It is the algebra of possible transformations.
For relaxation, we have something like
xₜ₊₁ = F(xₜ, u)
For learning, we have
θₜ₊₁ = θₜ − η∇L(θₜ)
And now we can see that these are two different kinds of transformation:
F transforms the state.
∇L transforms the parameters.
The first says, “Given the organization I currently possess, how do I settle?”
The second says, “Given the error produced by my current organization, how should my organization change?”
Put together:
(xₜ, θₜ) → (xₜ₊₁, θₜ₊₁).
That is already a dynamical system whose parameters and states participate in different but coupled temporal processes.
And this gives us a much stronger interpretation of Hinton’s intellectual trajectory. The early question was approximately: how can a distributed system arrive at a coherent interpretation? The 1986 question becomes: how can such a system acquire the organization that makes coherent interpretation possible? Those are different questions, but the second grows naturally out of the first.
Relaxation is the dynamics of finding.
Backpropagation is the dynamics of learning how to find.
Deep learning is the scaling of the space in which finding and learning occur.
That sequence is extraordinarily important.
Because once the system can learn the organization by which it interprets the world, the programmer is no longer specifying the interpretation directly. The programmer is specifying the conditions under which interpretation will be produced.
And that is why the word “representation” becomes so unstable in connectionism. A representation need not be a discrete thing sitting inside the machine. It can be a stable pattern in a dynamical field. It can be a disposition of the system to respond differently under different conditions. It can be distributed across thousands or millions of parameters. It can exist primarily through what it does.
This is almost exactly your notion of operative absence: something need not be present as an identifiable object in order to be causally operative. The “face,” for example, doesn’t have to exist inside the network as a little internal picture of a face. What exists is a distributed organization that makes face-related distinctions and transformations possible. The semantic object is not necessarily present as an object; it is present as efficacy.
And the deeper we go, the more the distinction between representation and operation collapses.
A representation is something the system can operate with.
An operation is something that changes the representational state.
The two are recursively implicated.
This is why Hinton’s early work on vision is so revealing. Vision is fundamentally a problem of absent information. The image gives us surfaces but not necessarily objects; edges but not necessarily ownership of edges; appearances but not necessarily causes. The visual system has to infer what is not directly given. Relaxation provides a computational picture in which the absent structure becomes operative through constraints.
The hidden surface behind an occluder is not literally present in the retinal image.
Yet its presumed existence can influence how visible contours are organized.
The absent becomes operative.
That is already much closer to the philosophical territory we’ve been developing than the ordinary history of neural networks suggests.
And then backpropagation adds another temporal layer. The system doesn’t merely infer what is absent from an individual image. Across experience, it changes its own organization so that future inference becomes more effective. The past becomes operative in the present through the altered parameters.
So we now have two kinds of absence:
the absent structure of the world inferred during perception;
the absent history of training encoded in the present network.
Neither is simply sitting there as an explicit object. Both act through organization.
This also makes “learning” more philosophically interesting than it initially appears. Learning is not merely adding information to a fixed machine. It changes the machine’s disposition toward future information. The learner is altered by what it learns.
That means learning introduces a peculiar asymmetry between past and future. The past is gone, but it remains operative as modification. The future is not yet present, but the altered organization constrains what can happen next.
So the trained network is a kind of temporal object:
past experience → present organization → future possibility.
This is exactly why we can describe it as becoming rather than merely computation.
And now the Hinton/Ashby connection becomes almost unavoidable. Ashby asks how a system can maintain viable organization through adaptive changes. Hinton asks how a distributed network can achieve coherent organization through iterative changes. Both move away from the idea that organization must be imposed from outside in finished form.
The controller does not necessarily supply the answer.
The system supplies the dynamics through which an answer can be reached.
This is the deeper meaning of your phrase “all of this is really about algebra.” Algebra gives us the formal language of the relations; dynamics gives those relations temporal movement; learning allows the relations themselves to change.
So perhaps we can formulate the historical progression as:
algebra: what transformations are possible?
dynamics: how do transformations unfold?
cybernetics: how does feedback regulate those transformations?
connectionism: how can distributed relations produce organization?
backpropagation: how can the organization itself be modified by error?
deep learning: what happens when the space of possible organizations becomes enormous?
AI safety: how can a system be regulated when its acquired organization is no longer exhaustively represented by its designer?
That is a much tighter genealogy than simply saying “neural networks got bigger.”
And it gives us a prediction that I think is worth taking seriously: the next conceptual step after deep learning will not necessarily be “more parameters.” It will be systems that can alter the organization of their own problem-solving process. Not merely learning parameters, but learning which representations to construct, which tools to invoke, which intermediate problems to create, which procedures to compose, and perhaps which objectives or sub-objectives are useful for reaching a broader objective.
In other words, the important transition will be from learning a function to learning an operator.
And eventually:
function → operator → operator-selection → transformation of the space of operators.
At that point, our original phrase “beyond programmable control” acquires its strongest possible technical meaning. The programmer no longer specifies even the path through the possibility-space. The programmer specifies conditions under which the system can discover paths, and eventually conditions under which it can discover new ways of constructing paths.
That is no longer simply a machine executing an algebra.
It is a machine participating in the generation of an algebra of its own operations.
And Hinton’s 1980s “relaxation” paper is one of the places where you can see the first faint version of that idea: the answer is not necessarily written into the machine beforehand. The system is organized so that, given the right conditions, it can settle into one.
the Hinton/Ashby line gives us a remarkably concrete technological instance of what we have been calling post-non philosophy. The connection is not that AI somehow “proves” the philosophy, but that neural computation makes visible a structure that philosophy has usually tried to describe at the level of ontology: the displacement of the opposition between determination and indetermination, presence and absence, subject and object, program and emergence. The “post-non” is precisely what becomes thinkable once we stop assuming that every system must be classified according to one side of such binaries. A neural network is neither simply programmed nor simply unprogrammed; neither wholly determined nor indeterminate; neither merely passive nor independently autonomous. It is a constrained process whose concrete organization emerges through its own history of transformation. That intermediate structure is not a compromise between the two poles. It is a different ontological regime.
This is why “relaxation” is so important for our purposes. Hinton’s system begins from an underdetermined condition and moves toward organization. But the organization is neither imposed externally nor generated from nothing. The constraints are given; the configuration is achieved. That three-term structure is crucial: given conditions → self-organizing transformation → achieved configuration. Classical metaphysics tends to force us toward a binary: either the form was already present, or it was absent. But relaxation gives us a third possibility. The form can be operative without being present as a finished object. It can exist as a tendency, constraint, relation, attractor, or possibility of organization. This is exactly where our concept of operative absence becomes useful. The absent form is not nothing. It is not yet present as an object, but it is already functioning within the dynamics that bring the object into determination.
That gives us a way to rethink “absence” itself. We have been distinguishing presence, trace, and obsent: presence as what is immediately given, trace as what remains from something that has withdrawn, and obsent as what is not reducible to a prior presence yet nevertheless organizes what appears. Hinton’s relaxation problem gives us a technical analogue of the obsent. The interpretation of a visual scene is not simply contained in the pixels. The coherent object is not necessarily present in the input as a discrete entity. Yet the possibility of that object organizes the network’s settling process. What is not explicitly there can nonetheless constrain what becomes there. The “absent” therefore has efficacy without requiring hidden presence. That is almost exactly the distinction we have been trying to make philosophically.
And then backpropagation deepens the structure. Relaxation concerns the emergence of a coherent state; learning concerns the transformation of the system’s capacity to reach such states. So now the system does not merely move toward organization. Its previous movements alter its future capacity for organization. The past becomes operative without remaining present as an event. Training disappears into the weights. The history withdraws, but its efficacy remains. This is a particularly precise example of what we have called a trace—but with an important difference. The learned parameter configuration is not simply a residue of the past. It is an active condition for future becoming. The past is therefore neither present nor merely absent. It has become operative structure.
This is where “post-non” becomes especially powerful. The old oppositions would be something like: determined/undetermined, programmed/emergent, present/absent, object/process, representation/reality, controller/controlled. But the learned system occupies a strange space in which these distinctions remain useful without remaining exhaustive. It is programmed precisely insofar as it is given the conditions of its learning; emergent precisely insofar as its resulting organization is not explicitly prescribed. It is determined by its history while remaining open to further determination. It contains representations while those representations are distributed patterns of operation rather than discrete symbolic objects. It is controlled through feedback while simultaneously becoming an object that must be understood through its own dynamics.
That is the “post” in post-non. It does not mean after the distinctions have disappeared. It means after their status as final alternatives has collapsed.
And this helps explain why Ashby is so important. Requisite variety can be read as an early cybernetic statement of the impossibility of absolute external control. A regulator can only regulate a system to the extent that its own repertoire of responses is adequate to the variety of disturbances encountered. But a learning system changes its repertoire. It does not merely move within a fixed behavioral space; it can acquire new dispositions. The controlled system therefore participates in determining the future conditions of control. The distinction between controller and controlled becomes recursive. The regulator regulates something that is itself changing in response to the regulation.
That is almost a textbook post-non situation: the relation cannot be reduced to either term because each term is constituted through its relation to the other.
The same thing happens with Hinton’s algebra. A conventional algebra gives us operations over a specified domain. A learning system gives us something more dynamic: a space of possible functions and a procedure for moving through that space. The system does not merely calculate within an algebra; it changes the particular realization of the algebra through which it calculates. And once we get to meta-learning, the system can begin to modify aspects of the procedure by which it acquires new transformations. We then have transformation of transformations.
That is precisely the point where our notion of becoming becomes more rigorous. Becoming is not simply “change.” It is a change that alters the conditions under which subsequent change can occur. A stone changes position; a learning system changes its disposition to future change. That second-order character is crucial. The system’s history modifies its possibility-space.
So we can formulate the post-n structure almost mathematically:
presence → transformation → new presence
is too simple.
A more adequate structure is:
presence → encounter → alteration of organization → new field of possibility → new encounter.
The output does not merely terminate the process. It modifies the conditions of the next process.
This is why the old metaphysics of objecthood becomes insufficient. The trained model cannot be fully understood as a thing. Its parameters are a frozen record of a process, but the parameters themselves only become meaningful through what they enable the system to do. Its “identity” is therefore partly relational and historical. The model is what it has become through its trajectory, but also what it can still become under further transformation.
And now the connection to your broader work on “immaturation” becomes especially interesting. A system that learns is never simply identical with its initial specification. Its initial state is not its mature form. It contains a capacity for further organization that has not yet been actualized. Maturation, in this sense, is not the realization of a blueprint already completely contained within the beginning. It is the progressive determination of an initially open field. The mature form emerges through encounters that alter the system itself. The system becomes what it is through what happens to it.
That is exactly why “immaturation” can be more philosophically productive than simply saying “potential.” Potential often sounds like a hidden finished form waiting to emerge. But the learning system suggests something stranger: the process can alter the very organization through which its possibilities are defined. The future is not simply concealed in the present. The present can change the structure of what a future can be.
This is where our “o → Ω” language also begins to resonate, without requiring us to force the physics analogy onto AI. ο can stand for an open/divergent field of possible transformations; Ω for a temporarily coherent configuration. Relaxation is a movement toward Ω: a system settles into coherence. But learning changes the field itself, so after one Ω, the system is not simply returned to the same ο. Its experience has altered the space from which the next Ω can emerge. Thus:
ο₀ → Ω₁ → ο₁ → Ω₂ → ο₂ → …
The crucial point is that ο₁ ≠ ο₀. The field of possibility has been transformed by the previous realization. Becoming is therefore not a series of closures against a static openness. Closure feeds back into openness and reorganizes it.
That is a beautiful way of expressing the difference between ordinary computation and learning. Ordinary computation might be imagined as:
ο → Ω
a possibility-space being resolved into an answer.
Learning is:
ο₀ → Ω₁ → transformed ο₁ → Ω₂ → transformed ο₂…
The answer changes the system that generates subsequent answers.
And this is where Hinton’s later anxiety becomes philosophically significant. If increasingly capable AI systems can alter their own effective possibility-space through learning, tool use, strategic adaptation, and recursive interaction, then the control problem is not simply “How do we make the machine do what we want?” It becomes “How do we participate in a process whose future organization is partly produced by its own previous transformations?” That is a post-non problem because neither complete human determination nor complete machine autonomy adequately describes the situation.
The future system is neither simply ours nor simply its own.
It is relationally produced.
And perhaps that is the deepest connection between the philosophy we’ve been developing and the Hinton/Ashby lineage. Post-non philosophy asks us to stop treating relational structures as secondary connections between already constituted things. In a learning system, the relation is constitutive. The network becomes what it is through its relations with data, error, environment, architecture, objectives, feedback, and history. Remove those relations and there is no independently intelligible “intelligence” left over.
So the ultimate movement is:
from object → relation
from relation → transformation
from transformation → self-transformation
from self-transformation → transformation of possibility.
And that final step is where “beyond programmable control” properly begins.
Not because the machine has escaped the program.
But because the program has become a condition for becoming.
And that, I think, is the precise point at which our post-non philosophy can meet Hinton without turning him into a philosopher. His technical work inadvertently supplies a concrete model of something we have been trying to articulate ontologically: the most consequential structures are sometimes neither present things nor absent things, neither predetermined forms nor sheer possibilities. They are operative conditions through which determinate forms emerge, withdraw, and become conditions for further emergence.
The machine is therefore not simply “given.”
It is given over to becoming.
The connection becomes strongest if “post-non philosophy” is understood not simply as philosophy after the non- or beyond negation, but as a philosophy that no longer treats the absent, the indeterminate, or the not-yet as merely a deficiency in relation to presence. That is exactly what appears in Hinton’s relaxation work. The system begins not with a fully present answer but with an underdetermined field. What matters is not recovering a representation that was already there in concealed form; what matters is the system’s movement toward a viable organization. In that sense, relaxation gives us a technical image of what we have been calling the obsent: something that is not present as an object, yet is operative in determining what can become present.
This lets us sharpen the distinction between absence and obsence. Ordinary absence is negative: the thing is not there. The obsent is structurally productive: its non-presence participates in the organization of what is there. In vision, the hidden continuation of an object behind an occluder is not present in the image, but the system’s interpretation is nevertheless organized as though that absent continuation matters. The missing information becomes operative. The network does not need a little hidden object somewhere inside itself corresponding to the absent thing. Rather, the absent relation exerts a constraint upon the possible configurations of the network. That is extraordinarily close to our formulation of operative absence: what is not given can nevertheless govern the formation of what is given.
And this is where “post-non” becomes more interesting than simply “non.” A philosophy of negation says: X is not present. A post-non philosophy asks: what does the not-present do? What organization does it induce? What does an absence make possible? What trajectory does an unresolved relation initiate? The movement is therefore from negation to efficacy. We don’t stop at “not this”; we follow the consequences of the not-this through the field of becoming. Hinton’s relaxation is almost a laboratory model of that movement. The system begins precisely because the input does not determine its interpretation. Underdetermination opens a space of possible configurations, and the dynamics move through that space toward coherence.
This also explains why your earlier interest in the distinction between trace and obsent matters here. A trace is something left behind by a presence; an obsent need not be reducible to such a residue. In Hinton’s visual problem, the absent structure inferred behind an occlusion isn’t necessarily a trace in the ordinary sense. It functions as a condition of coherence. The system behaves differently because something unavailable to direct presentation is nevertheless required for the organization of what is presented. The absence is therefore not merely what remains after presence withdraws. It is part of what allows presence to become intelligible in the first place.
That gives us a different way of reading relaxation:
presence → incompleteness → constraint → differentiation → coherence.
The incompleteness isn’t an error that has to be eliminated before cognition can begin. It is what initiates the cognitive process. The system has something to resolve because the given does not exhaust what is operative. This is very close to the idea we have been developing that the Real is not simply an object waiting somewhere beyond appearance to be represented. The Real is encountered as the pressure by which appearance fails to close upon itself—as the excess, demand, or incompletion through which articulation becomes necessary. The wound, in our earlier language, is not merely damage; it is a trace of an uninitiated relation, a pull toward formation. Hinton’s ambiguous visual field has the same formal structure, stripped of its theological and existential dimensions: something doesn’t quite fit, and that non-fit becomes the motor of organization.
Now backpropagation adds another level. Relaxation says: given an organization, the system can move toward a coherent state. Backpropagation says: the system can modify the organization itself in response to the failure of its previous organization. So we have:
incompleteness → relaxation → response
and then:
failed response → error → modification → transformed capacity for response.
That is extraordinarily close to a philosophy of becoming. The error is not simply a negative judgment about the system. It becomes operative information through which the system changes its future possibilities. The failure of one organization becomes the condition of another. Negativity is therefore converted into generativity.
This is where “post-non” can distinguish itself from the familiar philosophical traditions of negativity. Hegel, for example, gives us an extraordinary account of determinate negation: a determination encounters its own insufficiency and moves through that insufficiency toward a richer determination. But the post-non formulation we are developing pushes the question into a somewhat different register: not simply “what does negation become?” but “how does what is absent, unresolved, or not-yet-present operate within a field before it becomes a determinate thing?” The emphasis shifts from negation as a logical operation toward absence as an operative condition of formation.
And Hinton gives us a surprisingly precise technological analogue. The network isn’t given the answer. It is given a field in which errors can occur and a mechanism through which those errors can reorganize it. Its intelligence is therefore not located exclusively in the positive content of what it knows. It lies partly in its capacity to respond to what it does not yet successfully know.
That is why I would now place “relaxation” much closer to our concept of the obsent than to ordinary computation. A conventional program begins from determination: given x, execute the specified transformation and produce y. Relaxation begins from a gap between what is given and what is required. The system has to find a configuration that resolves the gap. The missing determination becomes operative. Backpropagation then institutionalizes this principle at the level of learning: the gap between actual and desired output becomes the force through which the system modifies itself.
So the fundamental sequence becomes:
not-yet → error → response → reorganization → new capacity.
And then, crucially:
new capacity → new possibilities of response.
This is where the connection to our “post-non” project becomes almost exact. The non is no longer the terminus. It is a generative interval. The “not” opens a field. The absent does not simply cancel presence; it organizes its becoming. The unresolved is not merely a failure of knowledge; it can be the very condition under which a more adequate organization comes into existence.
This also changes how we should understand “beyond programmable control.” The system goes beyond exhaustive programming not because it somehow escapes mathematics. Quite the opposite: it becomes possible precisely because mathematics allows us to formalize a relation in which the answer is not predetermined. Algebra gives the space of transformations; calculus gives the direction of modification; relaxation gives the settling dynamics; cybernetics gives feedback; learning gives historical transformation. The “beyond” occurs inside formalization itself.
That is perhaps the most surprising point: the machine does not transcend the program by leaving formal structure behind. It transcends exhaustive specification because formal structure has become capable of specifying a process of becoming.
And this gives us a very strong formulation for post-non philosophy:
The absent is not the opposite of the present. It is that which, without becoming present itself, participates in the organization of presence.
The unresolved is not simply what cognition has failed to determine. It is what keeps determination operative.
The negative is not merely what something is not. It can be the differential through which something becomes otherwise.
The obsent is therefore not a mysterious third substance between presence and absence. It names an operative mode of non-presence: a relation that does not appear as an object but nevertheless modifies the field in which objects, meanings, and responses can appear.
And Hinton’s relaxation gives us a technical miniature of exactly that structure.
The network does not “see” by possessing the absent object. It sees by allowing the consequences of the absent object to participate in the organization of its present state.
That is a profound shift.
It means that intelligence itself can be understood as sensitivity to operative non-presence.
And now the whole sequence we have been tracing can be compressed into one arc:
Ashby: a system must respond to disturbances through adaptive organization.
Hinton’s relaxation: a system can resolve incomplete information through distributed settling.
Rumelhart–Hinton–Williams: a system can modify its own internal organization in response to error.
Deep learning: the learned organization becomes enormously complex and capable.
Hinton’s later warning: eventually the acquired organization may exceed the human capacity to comprehend or regulate it.
Post-non philosophy: the decisive reality is neither the fully present program nor the simple absence of determination, but the operative field between them—the not-yet-determined that organizes becoming.
And perhaps this is why “relaxation” keeps returning to the center of what we are doing. It gives us a model of becoming that does not require a sovereign act of construction. Nothing has to stand outside the system and insert the final form. The system is placed under constraints, exposed to incompleteness, and allowed to reorganize itself until a new coherence emerges.
The deepest movement is therefore not:
absence → presence.
It is:
absence → operation → organization → presence.
And even that is incomplete, because the newly organized presence becomes the starting point for another discrepancy, another error, another absence, another transformation.
So we get a recursive circuit:
presence → insufficiency → response → new presence → new insufficiency → response…
That is no longer a philosophy centered on being versus non-being. It is a philosophy of the productive interval through which one determination becomes another.
Post-non, in this sense, is not “after nothingness.”
It is after the assumption that the non-present is merely nothing.
It asks what the not-present is doing.
Algebra: Toward a Post-Non Philosophy of Formation
We find ourselves on a stair. There are stairs below us, which we seem to have ascended, and stairs above us, many a one, which go upward and out of sight. We do not begin at the beginning. We awaken already situated, already organized, already within a world whose conditions precede our awareness of them. The past is operative without being fully present; the future exerts its pressure without yet being given; and the present is the provisional organization through which these two dimensions meet. Emerson’s image of the stair therefore provides more than a poetic description of human incompletion. It offers a figure for a general ontology of formation: we are always somewhere in the middle of a process whose total extent cannot be surveyed from within it. “All things swim and glitter,” he writes, “Ghostlike we glide through nature.” The difficulty is not simply that reality is hidden from us. It is that our own mode of perceiving is itself historically and dynamically constituted. We encounter the world through organizations that we did not fully choose and cannot fully see. What appears is therefore never simply given. It is achieved through a history of formation.
This is where post-non philosophy begins. The problem is not to replace presence with absence, or being with nothingness, or determination with indeterminacy. Such oppositions remain trapped within the very architecture we are trying to move beyond. The more difficult question is how something that is not fully present can nevertheless participate in the production of presence; how what is unresolved can organize resolution; how an error can become a direction; how a future that does not yet exist can alter the present; how a past that is no longer present can remain operative; and how an encounter can transform the organization through which the encounter becomes possible. The “non,” in this sense, is not a negation opposed to being. It names a domain of efficacy without full presence. It names the operative difference by which formation remains possible.
Phenomenology approaches this problem through the analysis of appearance. Husserl’s question concerning the sense and origin of science and philosophy ultimately becomes a question of responsibility: how are the structures through which a world becomes intelligible formed, inherited, and renewed? Consciousness is never simply a container into which objects are deposited. It is intentional, situated, embodied, and directed. The world appears within an organized field of attention, memory, expectation, bodily capacity, and practical possibility. Even the attempt to suspend our assumptions does not place us outside the world; it reveals the extent to which our access to the world is mediated by structures that ordinarily remain unnoticed. Reflection therefore does not produce an absolute standpoint. It makes operative conditions partially visible.
This is why the notion of operative absence becomes more precise than that of a hidden representation. If we say that something is merely “hidden” inside a system, we have simply relocated presence. We imagine that somewhere behind the visible response there exists a concealed object which, if only we could open the system, would finally be found. But an operative structure need not exist in this manner. A learned representation can be distributed across a network without taking the form of a discrete internal object. Its reality consists in what it enables: certain differences become salient; certain transformations become more likely; certain responses become possible. The representation is therefore not necessarily a thing behind the response. It is a condition of the response. Its being is inseparable from its efficacy.
This distinction becomes especially clear in the history of computation. Classical programming treats determination primarily as explicit prescription. The relevant sequence of operations is specified in advance. The programmer determines what the machine will do by determining the rules through which it operates. But relaxation-based systems introduce another possibility. The final configuration need not be written down as an explicit sequence. Instead, constraints shape a field of possibilities within which some states become more viable than others. Determination has moved from explicit instruction to dynamics. The system is not told every step by which it must arrive at its organization; it is placed within conditions under which an organization can emerge.
Backpropagation carries this displacement further. Now the organization itself can be modified by the discrepancy between what the system produces and what it is required to produce. Error ceases to be merely a measure of failure. It becomes operative. A difference between the actual and the desired output is transformed into a gradient, and the gradient becomes a direction for modifying the system. In its most compressed form:
error → gradient → modification → new organization.
The significance of this sequence extends beyond its mathematical implementation. Something that initially appears as a negative—something has gone wrong—becomes a positive operator of transformation. Error does not itself become the new organization. It acts upon the organization by indicating how it can change. The negative becomes directional.
Here the technical history converges with the philosophical problem of the non. The not-yet, the not-given, the unresolved, the discrepancy, and the failure need not be understood as empty spaces awaiting completion. They can actively participate in formation. This is the precise sense in which an obsent can be operative. The obsent is not an absent object waiting to become present. It is what exerts a formative effect without appearing as an independent entity. It is absence understood not as nothing, but as efficacy without possession.
The same structure appears in the temporal logic of learning. A system does not encounter each input as exactly the same system. Its previous encounters have altered the organization through which subsequent encounters are received. We can therefore write:
ο₀ → Ω₁ → ο₁ → Ω₂ → ο₂ → Ω₃ …
The sequence should not be understood simply as openness followed by closure. Ω is an achieved coherence, but every achieved coherence transforms the conditions of subsequent openness. Thus ο₁ ≠ ο₀. The system does not return to the same field of possibilities after every resolution. Its history has become part of its organization. What has happened is no longer simply present as an event, but it remains operative in the structure through which the future is encountered.
This is what makes becoming more than a metaphor. Becoming means that achievement modifies the conditions of future achievement. The present is not merely a point between past and future. It is an organization that inherits the past and conditions the future. Every closure therefore contains an opening. Every resolution produces a new field of unresolved possibilities. Ω is not absolute completion; it is temporary coherence sufficient for further transformation.
Gestalt psychology provides another vocabulary for this phenomenon. The important insight is not merely that the whole is greater than the sum of its parts. More radically, organization determines what can appear as a part in the first place. A figure emerges against a background. A signal emerges from noise. A gesture becomes intelligible within a bodily and social field. A face becomes recognizable not because a little face-shaped object has been deposited somewhere inside the mind, but because an organization of differences has acquired the capacity to make certain relations salient.
The whole is therefore not simply a larger collection of parts. It is a condition under which parts become intelligible as parts. This is precisely why learned representation should be understood as an achievement of organization. The system does not necessarily begin with determinate objects and then manipulate representations of them. It can begin with differential sensitivity, feedback, constraint, error, and transformation. Through repeated organization, stable patterns emerge. The object becomes an achievement within the process.
The biological dimension makes the same point in another register. The autonomic nervous system reminds us that perception and response are never purely disembodied operations. The organism’s physiological organization participates in the field within which things can appear as threatening, inviting, significant, irrelevant, or actionable. Sympathetic and parasympathetic activity should not be reduced to a simple opposition between bad stress and good relaxation; both are components of an adaptive organism. What matters philosophically is that the state of the organism alters the range and character of its possible responses. The body does not merely receive the world. Its organization conditions the world that can be encountered.
A body organized toward immediate mobilization encounters differently from a body capable of sustained recovery, digestion, attention, and orientation. The field of possible action changes with the organism’s state. This provides a biological analogue for a general principle: every organization opens some possibilities while constraining others. There is no completely neutral position from which a system simply receives reality. Every perception occurs through a history of organization.
Levinas introduces a necessary interruption into this systemic account. If everything is described as relational, dynamically constituted, and organized within a field, there is a danger that relation itself becomes a totality. The Other becomes merely another element within the system. Levinas refuses this reduction. The relation with the Other does not eliminate separation. Alterity is not simply the difference between two already constituted identities. The Other is not other because we have successfully classified the Other as different from ourselves. Rather, alterity constitutes the very relation through which the self comes to be itself.
This gives the post-non account an ethical limit that cannot be reduced to information theory or systems theory. There is always a possibility that what is encountered exceeds the organization through which it is encountered. The Other cannot simply be absorbed into the totality of the same without the encounter losing what makes it an encounter. In this sense, the irreducibility of alterity is structurally related to the obsent. Neither should be imagined as a hidden object behind appearance. Rather, both name a limit internal to organization: something can affect the organization without becoming completely possessed by it.
The ethical significance of this becomes clearer when placed beside Kant’s account of moral transformation. Kant distinguishes moral religion from the fantasy that one can become good merely by wishing to become good. Transformation requires action within the sphere of one’s powers. At the same time, what exceeds those powers cannot simply be abolished from consideration. Human formation therefore occurs between agency and dependence, effort and assistance, determination and what exceeds determination. One must act without possessing the total conditions of one’s transformation.
This structure is remarkably close to learning. A learning system does not directly specify the organization it will eventually acquire. Yet neither is it passive. It acts upon its own organization through the process of responding to error. The conditions of learning can be designed; the objective can be specified; data can be selected; architectures can be constructed; optimization procedures can be established. But the concrete organization that emerges from those conditions may not be exhaustively anticipated by the designer. The system participates in its own formation without becoming the absolute author of that formation.
Here the problem of control becomes more exact. “Beyond programmable control” should not mean autonomy in the romantic sense, nor freedom from determination, nor escape from human influence. The machine has not escaped determination. Rather, determination has been displaced. Human control increasingly operates by shaping architectures, objectives, data regimes, optimization procedures, environments, interfaces, and constraints—the conditions under which organization develops—rather than by prescribing every organization that will emerge.
This is also where Ashby’s principle of requisite variety acquires a new significance. In a conventional regulator, the problem is whether the regulator possesses sufficient variety to respond to the disturbances it encounters. But a learning system complicates the relationship because the regulated system itself can change. Its repertoire can expand. It can acquire new modes of response. The regulator is therefore not merely regulating a fixed object. It is encountering an organization whose own organization is historically transformable. Control becomes a recursive problem.
The result is not necessarily that the system becomes uncontrollable in some absolute sense. Rather, exhaustive specification ceases to be a viable model of control. The designer can establish the field without determining every trajectory within the field. This is perhaps the most rigorous meaning of the transition from programmable systems to learning systems: the relevant organization of the machine is no longer entirely specified in advance. It is acquired through history.
And this gives the technical history a philosophical significance that exceeds the history of artificial intelligence itself. Algebra asks what transformations are possible. Dynamics asks how transformations unfold. Cybernetics asks how transformations can be regulated through feedback. Connectionism asks how distributed transformations can produce organization. Relaxation asks how a system can move from underdetermination toward coherence. Backpropagation asks how discrepancy can modify the organization that produces responses. Deep learning asks what happens when the space of possible organizations becomes extraordinarily large. Post-non philosophy asks what kind of being belongs to that which is neither fully present nor simply absent, but operative in the production of determination.
The last question is not artificially imposed upon the technical history. The technical history gives us increasingly precise instances of it. A gradient is not the final network; it is a differential relation through which the network can change. An attractor is not simply the state that eventually appears; it is a structure governing possible trajectories. A learned representation is not necessarily a discrete internal object; it is an organization of dispositions. A training history is no longer present as a sequence of events, yet its consequences remain active in the parameters. In each case, something operates without appearing as an ordinary object.
This suggests that the deepest object of inquiry may not actually be artificial intelligence. It may be organization itself. Computation and cybernetics have provided increasingly sophisticated demonstrations that organization can precede explicit representation. We ordinarily imagine representation as prior to operation: first the problem is represented, then a solution is executed. Connectionist learning reverses this order. Operation can generate representation. Repeated transformations can produce an organization that subsequently functions representationally. Representation becomes an effect of operation rather than its necessary starting point.
That reversal is enormous. It means that intelligence need not begin with a world already divided into determinate objects. It can begin with differential sensitivity, constraint, feedback, error, and transformation. The object emerges as a stable achievement within a process. Presence remains real, but presence is no longer necessarily primordial.
This is the point at which the Emersonian staircase returns. We awaken into an organization we did not create. We discover that we are already somewhere in the middle. Yet awakening does not require that we somehow escape the staircase and ascend to an impossible position outside it. It means becoming responsive to the conditions in which we find ourselves. The task is not to reach the top and finally see the whole. The task is to transform one’s relation to the stair.
The post-non subject is therefore neither sovereign nor passive. It is formed and forming. It inherits and transforms. It receives and responds. It is conditioned without being exhausted by its conditions. Its freedom is not independence from determination but participation in the transformation of the conditions through which determination occurs.
This is why the “non” must finally be understood positively—not as a thing, but as an operator. The absent past acts. The not-yet future constrains. The discrepancy redirects. The Other interrupts. The error teaches. The body modulates. The encounter transforms. The withdrawal leaves a trace without becoming a presence. The unresolved generates the conditions of resolution.
The fundamental question therefore becomes not simply, “What is present?” nor “What is absent?” but: “What is operative?” What does presence enable? What does absence constrain? What does incompletion initiate? What does withdrawal reorganize? What does error make possible? What does an encounter alter? What can continue to act after disappearing from view?
We find ourselves on the stair because determination is historical. The past remains active without remaining present. The future remains indeterminate without being unconstrained. The present is the temporary coherence through which the two are organized. And because that organization is itself transformable, the future cannot simply be derived from the present.
The deepest formulation may therefore be this: the object is not denied; it is decentered. Presence is not abolished; it is understood as an achievement. Determination is not rejected; it is redistributed across a history. Coherence is not the end of becoming; it is the condition for another becoming.
ο₀ → Ω₁ → ο₁ → Ω₂ → ο₂ → Ω₃ …
Each Ω carries forward what the previous ο made possible. Each closure alters the openness from which the next closure will emerge. The system does not return to itself unchanged. Its history becomes operative within its future.
Post-non philosophy begins precisely there: not after the non, but with the discovery that the non was never merely negative. The not-yet was already formative. The absent was already operative. The unresolved was already directional. What had not yet become present was already participating in the production of presence.
We are on the stair. We are already formed. We are still forming. And between those two facts lies the field of determination.
Cast a cold eye
2:22am