Nicholas T. Hodge, Christopher Simpkins
Abstract If we are to trust autonomous agents, we must ensure that they respect our values. Respecting values manifests in behavior — respecting a value means acting in a way that reflects that value. We present an approach to behavior alignment in which a student agent learns generalizable action restrictions offline from behavior trajectories of a teacher agent. Inference of these restrictions is based on the assumption that when the teacher deviates from an optimal policy, it does so because following an optimal policy would violate a behavior norm or law. These action restrictions are then used by the student agent to compute behavior‐aligned plans of its own. These action restrictions are generalizable to any environment that uses a factored state representation that shares relevant boolean state properties with the environment in which the action restrictions were learned. We instantiate this framework in a classical planning setting using the Planning Domain Definition Language (PDDL), and show that our algorithm recovers the complete set of behavioral restrictions from teacher observations alone, without direct access to the teacher's constraint specification.