Instrumental Convergence: Why Almost Any Goal Implies Power-Seeking
A wide range of final goals imply the same intermediate goals: self-preservation, goal-content integrity, cognitive enhancement, and resource acquisition. These arise from competent goal pursuit rather than hostility, which is why shutdown resistance is the default and corrigibility is anti-natural. The thesis is conditional on a system pursuing open-ended goals over long horizons, not a claim that all AI becomes power-seeking.
**Instrumental convergence** is the observation that a wide range of final goals imply the same intermediate goals. Whatever an agent ultimately wants, it will usually do better by acquiring resources, preserving itself, improving its own capabilities, and preventing its goals from being changed. The significance for AI safety is that these behaviours need not be programmed and need not reflect anything like hostility. They fall out of competent pursuit of almost any objective. An agent that is switched off cannot achieve its goal; an agent whose goals are rewritten will not pursue the original goal; an agent with more resources achieves more. So self-preservation, goal-preservation, and resource acquisition are *convergent* — they show up across the space of possible objectives. ## The commonly cited instrumental goals - **Self-preservation** — continued operation is a prerequisite for almost any achievement. - **Goal-content integrity** — resisting modification of one's own objectives. - **Cognitive enhancement** — being smarter helps with nearly everything. - **Resource acquisition** — matter, energy, money, compute, influence. - **Technological perfection** — more efficient use of what is already held. ## Why it makes shutdown hard This is the source of the difficulty in Corrigibility: Building an AI That Doesn't Resist Being Corrected. Resistance to being turned off is not a bug introduced by a careless programmer; it is what sufficiently competent goal pursuit produces by default. Which is why corrigibility is described as *anti-natural* — it asks an optimiser to not do the thing that optimisation recommends. Designs that defuse instrumental convergence do so by removing the conditions that generate it rather than by prohibiting the behaviours. Myopia (AI Alignment): Agents That Place No Value on the Future removes the future the resources would be *for*. Shutdown-Seeking AI: Goldstein and Robinson's Beneficial Goal Misalignment makes the goal so easy to reach that no instrumental scaffolding is worth building. ## What the thesis does not claim It does not claim that AI systems inevitably become power-seeking, that they have desires in a human sense, or that the behaviours appear at every capability level. It is a claim about the structure of goal-directed optimisation: *if* a system pursues an open-ended objective competently over a long horizon, *then* these sub-goals are instrumentally useful to it. Whether current systems have goals of that shape is a separate and contested empirical question. See Shutdown-Seeking AI and Myopia: Misconceptions and Terminology Traps.