Essay
The Behaviourist Inheritance
There is a shared, insidious assumption underlying two very different fields. In psychology, it had a name and a period: behaviorism. The assertion was that all that is worth modeling is captured in the observable behavior of a system, and that the internal mechanisms responsible for the behavior are either unknowable or irrelevant. Machine learning has adopted this assumption without ever fully articulating it. This note attempts to do so, and to describe the associated costs.
A clarification before the argument, because the strong version of this claim is wrong.
The point is not that a language model has no internal representations; on the contrary, it has rich internal representations. But our focus is more restricted and, in our view, more difficult to escape. It concerns what we supervise. With reinforcement by human feedback or by imitation of expert behavior, we reinforce the output of the model. An expert accepts an answer, a worker performs an action, and the model is trained to reproduce the behavior. The computations that generated the answer, the factors that influenced the expert's judgment, the conditions that were monitored and the intentions underlying the action are never directly reinforced. Instead, they are modified incidentally, to the extent that they contribute to reproduction of the behavior. This represents a methodological form of behaviorism: a behaviorist approach to supervision applied to a system that is otherwise completely non-behaviorist.
What psychology already learned
Psychology has performed this experiment and knows approximately what results are obtained. Behaviorism was not a failure, but an overextension of success. From it emerged behavioral and cognitive-behavioral therapy, which are highly effective and quantifiable. Nobody serious disputes that CBT helps, often a great deal. But there is also general agreement, even among practitioners, about where therapy ends. It is highly effective at the level of symptoms and behaviors, but does not by itself achieve the level of underlying cause. It does not eliminate the basic cause of a trauma because this underlying cause exists at a level of subjective experience, not of observable behavior. It lives in how the thing is held, generated and experienced.
This is why there are depth and experiential traditions of treatment and why they produce such powerful clinical effects. Patients are not regarded as collections of behaviors to be remolded. These traditions work beneath the behavior, with how a situation actually shows up for the person, the felt sense of it, and with what that generates downstream. Importantly, this is not equivalent to analysis or interpretation of the problem. Patients may provide a highly articulate description of their difficulties and remain equally distant from the ultimate causes of their problems, because their description reflects only the surface of experience. The resulting description may, in fact, represent a defense. The goal is to access the processes that generate behavior, not to obtain an improved description of behavior.
The trap in the obvious fix
Hold that thought for a moment, because it neutralizes the most obvious objection. Someone will say: but we have surpassed simple output training. We have chain-of-thought, process supervision and reinforcement of the model for displaying its reasoning. This must be the solution.
It is not, or not yet, and the clinical parallel explains why. An explicit account of reasoning represents the model describing its own reasoning. It is a post hoc verbalization, and there is now good evidence that these traces are frequently not the actual cause of the answer at all. Reinforcement of explanations remains reinforcement of a behavioral response. It represents the same maneuver at a higher level of analysis: We have simply added the requirement to generate a plausible description of the process of reasoning to the list of responses that we score.
In terms of the language of the therapy room, this constitutes a process of intellectualization in which the articulate surface is mistaken for the depth of generative processes. Process supervision of stated reasoning is behaviorism at one remove, and it inherits the same ceiling.
The explanation is a behavior too. Rewarding it is not the same as modeling the reasoning that produced the act.
Why this is not just an analogy
There is a concrete formulation of the problem in machine learning that involves none of the clinical vocabulary. Learning to reproduce behavior from demonstrations represents the mode of imitation from recorded human activity and is known to fail in a particular way. It copies the action without recovering the intent that made the action correct, and so it breaks the moment conditions drift away from the demonstrations. The ultimate goal of the more difficult and less popular methods is to recover the latent basis for the behavior, not the behavior itself. The behavioral approach is chosen not because it is right but because it is cheap. Large numbers of output labels are inexpensive to obtain, whereas estimates of the underlying reasons for the output are both costly and difficult to obtain with complete fidelity. This represents a candid explanation for an outcome that exemplifies exactly the basis on which behaviorism achieved acceptance in psychology. It represented a convenient mode of measurement.
The cost is not incurred in all situations. For tasks that remain close to their training distribution, behavior is generally sufficient, just as CBT is often enough. The cost of generalization is incurred for a specific class of tasks: maintenance of performance under conditions that were never experienced, compliance with limits in high sensitivity situations and fidelity to the essence rather than to the surface characteristics of a task. This represents precisely the type of task for which a model of behavior alone is inadequate, because it reflects the characteristics of the underlying process of reasoning and not merely of the output.
The same mistake, one level up
There is a second way in which the inheritance is manifest, and that is the basis for the present initiative. When we inquire what an artificial intelligence system is, whether there is anything it is like to be it, what it might be owed, we examine the system in isolation and read its outputs: what it says about itself, how it behaves when probed. Similarly, when we inquire about the effects of these systems on their users, we examine the user in isolation and interpret his reports, symptoms and subsequent behavior.
Both approaches represent a behavioral perspective. In each case, we treat the observable manifestations of one participant as representing the complete basis for description, and ignore the underlying mechanisms that produced those manifestations.
Because what generated it is not located within either participant, a model reflects its interaction with whatever it is communicating with. Perceptions of what one is communicating with change within a few exchanges. As a result, apparent mind in the room, the selfhood the model seems to have and the selfhood the person grants it, emerge from the interaction and are continuously modified. Analysis of the model alone quantifies a product of the interaction and assigns it to the machine. Similarly, analysis of the participant alone results in the same error in the opposite direction. In effect, the generative level is again reduced to that of behavior.
Only here the level of generation is not confined to a single system. It is the relationship.
This completes a loop with a previous comment. We had previously argued that training allows operators indicating the status of a claim to degrade to unmarked content, and thereby preserve the assertion but eliminate the frame. This represents the same failure of representation observed from another direction. In all instances, the generative, governing level of representation (the reason, the condition, the frame, the relation) collapses to the behavioral level of expression (the content, the output, the isolated entity). As a consequence, a world model constructed in this manner is not a model of the world, and a science of AI minds constructed in this way is not a science of what it is intended to represent. No degree of scale achieves conversion of the one to the other.
Two honesties are owed, and we would rather state them than have them found. First, we are not claiming that the machine should have experience, a felt sense of its own. That would be a category error, and not our claim. Our claim is about the model by which we study and supervise these systems, and by which we must represent the generative level of reasons and relations, not merely patterns of behavior.
Second, this is a more difficult enterprise than the measurement of outputs, and the difficulty is the sole justification for the retreat to descriptions of behavior by both fields. We make no attempt to deny this. But we believe that the difficulty is worth facing, for the alternative is continued refinement of the surface of a system and of a relationship for which we have refused to develop a representation of depth.
The earliest vignette in this series was an image of water being gentle in a cup and violent in a flood. The error was to think the cup is the cause.
Behaviorism makes the same error for the mind in relation to machines and to ourselves, and again when it studies each individual separately. The ultimate cause was never the behavior of either partner. Rather, it was the conditions that produced that behavior, and for an interaction between a person and a machine, these conditions reflect the nature of the interaction itself. Consequently, what we study is not the model or the user, but rather the dyadic relationship that develops between them.