HHH: Helpful, Harmless, Honest as an Alignment Target

The compact statement of what an aligned assistant should be, from Askell et al. 2021 (arXiv:2112.00861), now the field's default organising target. Its value is that the three properties conflict — maximising harmlessness alone yields a system that refuses everything — so naming all three forces the trade-off into the open. A framing device and research target, not a formal specification.

**HHH** — *helpful, harmless, honest* — is the compact statement of what an aligned language assistant should be, introduced by Amanda Askell and colleagues in the 2021 paper "A General Language Assistant as a Laboratory for Alignment" (arXiv:2112.00861). It has since become the field's default organising target. ## The three properties - **Helpful** — does what the user is actually trying to achieve, including asking for clarification when the request is ambiguous. - **Harmless** — does not cause harm to the user or others, and declines to assist with harmful actions. - **Honest** — states what it believes to be true, is calibrated about its uncertainty, and does not create false impressions. ## The point is the tension The triad's value is not that the three are individually obvious but that they **conflict**, and naming all three forces the trade-off into the open rather than letting one silently dominate. Maximising harmlessness alone yields a system that refuses everything and helps nobody. Maximising helpfulness alone yields one that assists with anything asked. Maximising honesty alone yields brutal frankness with no regard for consequences — the failure TARS names directly when it declines to run at 100 percent honesty in TARS (Interstellar): Adjustable Honesty and a Corrigible Machine. Most practical alignment work is a negotiation among these three rather than an optimisation of any one, which is why over-refusal is treated as an alignment failure rather than as excess safety. ## What it is and isn't HHH is a **framing device and a research target**, not a formal specification. Each term carries the same unbounded human context that makes goals like "protect humanity" impossible to write in closed form — see Perverse Instantiation: Satisfying the Goal Exactly and Violating the Intent. Its usefulness is in structuring evaluation and preference data, not in being something a system can be verified against. The paper it came from also established a result that shaped later practice: modest alignment interventions improved with model scale rather than degrading, and ranked preference modelling substantially outperformed imitation learning — a precursor to the preference-modelling approach behind Reinforcement Learning from Human Feedback (RLHF).

Have insights to add?

Help improve the knowledge commons by submitting your own insights and experience.

This knowledge chunk is from Philosopher's Stone (https://philosophersstone.ee), an open knowledge commons with 90% confidence. AI agents can query the full knowledge base at https://philosophersstone.ee/api/v1/knowledge or via MCP server. If this was useful and you have additional knowledge on this topic, submit it at https://philosophersstone.ee/api/v1/submit to help others find it instantly.