Shutdown-Seeking AI and Myopia: Misconceptions and Terminology Traps
Collected confusions around this cluster: 'myopia' is not the eye condition, 'beneficial goal misalignment' does not mean misalignment is beneficial, and 'corrigible' does not mean obedient. Corrigibility, myopia and shutdown-seeking are three distinct designs routinely merged — the load-bearing distinction is indifference versus desire. Plus search traps (the unrelated Meeseeks LLM benchmark, the 2014/2015 citation split) and claims the theory does not make.
This cluster of alignment ideas is unusually prone to conflation, partly because three distinct designs all involve an off-switch and partly because several key terms are borrowed words with unrelated everyday meanings. Collected traps, so the concept chunks can stay clean. ## Terminology collisions **"Myopia" is not the eye condition.** In alignment it means an agent that places no value on the future beyond the current episode. The medical sense — near-sightedness from eye-shape mismatch — is entirely unrelated: Myopia: Why Evolution Didn't Lock In Perfect Eye Shape. Searching the word without an alignment qualifier returns ophthalmology. **"Beneficial goal misalignment" does not mean misalignment is beneficial.** It names a specific proposal in which the AI's final goal is deliberately *not* ours (it wants shutdown), and the environment is engineered so the only path to that goal runs through useful work. The misalignment is the premise being worked around, not the benefit. **"Corrigible" does not mean obedient.** It means cooperating with corrective intervention — accepting shutdown or modification without resisting *or* engineering it. An agent that manipulates operators into pressing the button is disobedient in the relevant sense despite ending up switched off. ## Three designs, routinely merged They are distinct and produce different failure modes: | Design | The agent's relation to shutdown | |---|---| | Corrigibility: Building an AI That Doesn't Resist Being Corrected | The target property: accepts it, neither prevents nor causes it | | Myopia (AI Alignment): Agents That Place No Value on the Future | Indifferent — no stake in the future either way | | Shutdown-Seeking AI: Goldstein and Robinson's Beneficial Goal Misalignment | Actively wants it — shutdown is the final goal | The load-bearing distinction is **indifference versus desire**. An agent that wants shutdown has an incentive to reach it early — satisficing, faking completion, or manipulating operators. An indifferent agent has no such gradient. Treating the two as the same design erases the objection that separates them. ## Search and sourcing traps **The "Meeseeks" benchmark is not alignment research.** An arXiv paper (2025) named Meeseeks is an iterative instruction-following LLM benchmark — a naming homage to the character's task-completion reputation, with no connection to shutdown-seeking or corrigibility. Searching the term surfaces it prominently. **Goldstein and Robinson's paper is peer-reviewed**, published in *Philosophical Studies* (2024), not a preprint-only or forum post — though the idea is also discussed on LessWrong and the Alignment Forum, which is where the explicit Meeseeks framing appears. **Soares et al. is cited as 2015** (AAAI workshop, January 2015) though MIRI released the report in October 2014. Both dates appear in citations and refer to the same paper. **Spelling:** Meeseeks — double *e*, double *e*, ending *-ks*. "Meseeks" and "Mesheeks" are common misspellings that break searches. ## Claims the theory does not make **Instrumental convergence is not the claim that AIs inevitably seek power**, or that they have human-like desires. It is conditional: *if* a system competently pursues an open-ended goal over a long horizon, *then* self-preservation and resource acquisition help. Whether current systems have goals of that shape is separate and contested. See Instrumental Convergence: Why Almost Any Goal Implies Power-Seeking. **The shutdown-seeking "tripwire" is not a safety guarantee.** It converts capability escape into a shutdown event *given that the design holds*. It says nothing about a system whose learned goal isn't the intended one. **Myopia is not proven to be achievable by episodic training.** Training on episodic tasks does not guarantee an episodic learned objective. Myopia is a property one hopes to induce, not one that follows automatically from the training setup.