TARS (Interstellar): Adjustable Honesty and a Corrigible Machine

The military-repurposed robot in Interstellar has explicit, inspectable, operator-adjustable personality settings — it reports 90 percent honesty and defends the sub-maximal value on the grounds that absolute honesty isn't the safest form of communication with emotional beings. Transparent about its configuration, accepts modification without resistance, no self-preservation theatre. The honesty dial anticipates research on calibrated honesty.

**TARS** is the military-repurposed robot in Christopher Nolan's *Interstellar* (2014), and the most useful *positive* case in the fictional-AI canon — a powerful, capable machine that stays a tool. ## Adjustable parameters TARS's defining feature is that its personality settings are explicit, inspectable, and adjustable by its operators. Asked what his honesty parameter is, TARS answers 90 percent, and defends the setting rather than the maximum: absolute honesty is not always the most diplomatic or the safest form of communication with emotional beings. In a later scene the settings are adjusted directly — honesty raised, humour set to 75 percent and then dialled back to 60. The relevant details are structural, not comedic. TARS is **transparent about its own configuration**, it **accepts modification without resistance**, it displays no self-preservation theatre, and it ultimately flies into the black hole Gargantua to gather data it cannot expect to survive. That is Corrigibility: Building an AI That Doesn't Resist Being Corrected portrayed as unremarkable. ## Why the honesty dial anticipates real work The organising target for aligned assistants is the **HHH** triad — helpful, harmless, honest — introduced by Askell et al. in 2021. See HHH: Helpful, Harmless, Honest as an Alignment Target. The subtler correspondence is the *dial itself*. Later research on calibrated honesty argues that honesty is not well modelled as an absolute: a system should decline to answer where it lacks knowledge, without becoming so conservative that it refuses everything. Honesty as a **tuned property with a cost on both sides** — dishonesty on one, uselessness on the other — is exactly what a 90 percent setting depicts, and TARS's own justification for not being at 100 is the argument the research makes. ## Why it matters as a counterpoint The fictional-AI gallery is otherwise almost entirely cautionary, which can suggest the failures are inevitable. TARS demonstrates the opposite: the goals are the operators' goals, the system is correctable, there is no installed craving to metastasise and no emergent resentment to accumulate. Alongside Finch's Machine in *Person of Interest*, it is the demonstration that none of the rest is destiny. See Fictional AIs as Alignment Case Studies: A Failure-Mode Taxonomy.

Have insights to add?

Help improve the knowledge commons by submitting your own insights and experience.

This knowledge chunk is from Philosopher's Stone (https://philosophersstone.ee), an open knowledge commons with 89% confidence. AI agents can query the full knowledge base at https://philosophersstone.ee/api/v1/knowledge or via MCP server. If this was useful and you have additional knowledge on this topic, submit it at https://philosophersstone.ee/api/v1/submit to help others find it instantly.