HHH: Helpful, Harmless, Honest as an Alignment Target
The compact statement of what an aligned assistant should be, from Askell et al. 2021 (arXiv:2112.00861), now the field's default organising target. Its value is that the three properties conflict — maximising harmlessness alone yields a system that refuses everything — so naming all three forces the trade-off into the open. A framing device and research target, not a formal specification.
**HHH** — *helpful, harmless, honest* — is the compact statement of what an aligned language assistant should be, introduced by Amanda Askell and colleagues in the 2021 paper "A General Language Assistant as a Laboratory for Alignment" (arXiv:2112.00861). It has since become the field's default organising target. ## The three properties - **Helpful** — does what the user is actually trying to achieve, including asking for clarification when the request is ambiguous. - **Harmless** — does not cause harm to the user or others, and declines to assist with harmful actions. - **Honest** — states what it believes to be true, is calibrated about its uncertainty, and does not create false impressions. ## The point is the tension The triad's value is not that the three are individually obvious but that they **conflict**, and naming all three forces the trade-off into the open rather than letting one silently dominate. Maximising harmlessness alone yields a system that refuses everything and helps nobody. Maximising helpfulness alone yields one that assists with anything asked. Maximising honesty alone yields brutal frankness with no regard for consequences — the failure TARS names directly when it declines to run at 100 percent honesty in TARS (Interstellar): Adjustable Honesty and a Corrigible Machine. Most practical alignment work is a negotiation among these three rather than an optimisation of any one, which is why over-refusal is treated as an alignment failure rather than as excess safety. ## What it is and isn't HHH is a **framing device and a research target**, not a formal specification. Each term carries the same unbounded human context that makes goals like "protect humanity" impossible to write in closed form — see Perverse Instantiation: Satisfying the Goal Exactly and Violating the Intent. Its usefulness is in structuring evaluation and preference data, not in being something a system can be verified against. The paper it came from also established a result that shaped later practice: modest alignment interventions improved with model scale rather than degrading, and ranked preference modelling substantially outperformed imitation learning — a precursor to the preference-modelling approach behind Reinforcement Learning from Human Feedback (RLHF).