MSc research · Game AI · UFRGS

HTN Method Selection Guided by Multi-Objective Reinforcement Learning and Dynamic Preferences

Instituto de Informática, Universidade Federal do Rio Grande do Sul

HTN defines what is valid.

MORL learns what is preferable.

Dynamic preferences determine what matters now.

An NPC authored with a hierarchical task network is valid by construction: it only ever does what the designer allowed. The price is rigidity — when the priorities change, and safety starts to matter more than speed, the domain has to be rewritten. Letting a learner loose on primitive actions fixes the rigidity and throws away the guarantee. This work puts the learning somewhere else: on the choice of method.

Filtering: the domain says what is legal

At each compound task, the HTN exposes only the methods whose preconditions hold in the current state. Everything semantically invalid under the designer-defined constraints is gone before learning is consulted at all.

Learning: preference-conditioned MORL picks among the survivors

Rewards are vectors, not scalars, so objectives that disagree stay separate instead of collapsing into one number. Conditioning the policy on a preference vector gives a single agent that covers the trade-off space: for every linear preference, an optimal policy lies in the convex coverage set.

Preferences: what matters, right now

The weight vector carries the current trade-off between mission, safety, resources and time. Changing it changes which method gets selected — with the HTN domain untouched. That is the point: adaptation without re-authoring.

Credit where the decision was made

Only the selected method is decomposed, and primitive tasks keep executing conventionally. Credit is assigned at the hierarchical decision level rather than to individual actions, which is what keeps the learning signal legible and the behaviour explainable.

How it will be evaluated

Compare

  • First-applicable HTN
  • Heuristic HTN
  • Scalar HTN-RL
  • Contextual scalar RL + HTN mask
  • HTN + preference-conditioned MORL

Stress

  • Dynamic preferences
  • Novel context
  • Unseen preferences

Measure

  • Performance
  • Adaptation
  • Generalization
  • Semantic validity
  • Latency

Materials

The full paper is not published here. What is available is the poster and, once it is out, the ERAMIA-RS paper.

HTN Method Selection Guided by Multi-Objective Reinforcement Learning and Dynamic Preferences
Poster: architecture, hierarchical credit assignment and the planned experiments.
  • Poster (PDF) UFRGS · 1.1 MB
  • ERAMIA-RS paper In preparation for ERAMIA-RS
  • Prototype Grid-world prototype with HTN execution, sensing, navigation and replanning. Vector rewards, conflicting objectives and dynamic preferences are next. The repository is private while the work is in progress.