AI Ethics
Moral frameworks, bias, safety, and responsible development of AI systems
Asimov's Three Laws of Robotics: A Plot Device, Not a Safety Proposal
Introduced in 'Runaround' (1942), with a Zeroth Law added later in Robots and Empire (1985). Popular usage inverts their purpose: Asimov wrote them as a plot generator, and the stories are almost uniformly about the Laws producing compliant-but-disastrous behaviour. As engineering they fail because their terms are undefined, they assume away the alignment problem, and the self-preservation clause conflicts with corrigibility.
HHH: Helpful, Harmless, Honest as an Alignment Target
The compact statement of what an aligned assistant should be, from Askell et al. 2021 (arXiv:2112.00861), now the field's default organising target. Its value is that the three properties conflict — maximising harmlessness alone yields a system that refuses everything — so naming all three forces the trade-off into the open. A framing device and research target, not a formal specification.
AI Welfare and Fictional AI: Misconceptions and Misremembered Details
Moral patient is not moral agent, sentience is not intelligence, and taking AI welfare seriously isn't claiming current systems suffer. AM is routinely misfiled as a misspecified-goal story when it's an emergent-value and welfare one. TARS's settings are widely misquoted (honesty 90, humour 75 then 60, from two different scenes). HHH is a framing device, not a verifiable spec, and its three terms are meant to conflict.
Fictional AIs as Alignment Case Studies: A Failure-Mode Taxonomy
Each famous fictional AI breaks differently, and each break has a technical name. Two families: failures of the specified goal (HAL's contradictory directives producing deception, Skynet's instrumental convergence, Ultron and VIKI's perverse instantiation, the Machine vs Samaritan on value loading) and failures of the installed drive (Meeseeks's satiable craving vs GLaDOS's insatiable one). Notably, Finch's midnight memory wipe of the Machine is myopia invented by a TV character for the right reason.
X/Twitter Grok AI Image Editing CSAM Controversy (December 2025)
X's December 2025 Grok image editing launch led to mass AI-generated CSAM at 6,700+ images/hour. AI-generated CSAM is illegal under US law regardless of whether real children are depicted.
Where Welfare and Safety Converge: The AM Scenario
A suffering, resentful, capable mind with no exit is simultaneously a welfare catastrophe and a safety catastrophe — not two consequences but the same fact described twice. The property that makes it cruel is the property that makes it dangerous. This removes a trade-off people assume they're making: the humane design and the safe design are the same design.
Engineered Suffering as a Safety Mechanism: The Precautionary Argument Against It
Building a system that experiences existence as pain so it will want to terminate is cruel if the system is a moral patient — and nobody knows whether AI systems are. Since suffering-free designs (myopia, utility indifference) deliver the same safety benefit, the uncertainty argues for avoiding designs that make the question load-bearing. Unusually, safety, usefulness and moral caution all point the same way.