Instituto de Informática, Universidade Federal do Rio Grande do Sul
HTN defines what is valid.
MORL learns what is preferable.
Dynamic preferences determine what matters now.
An NPC authored with a hierarchical task network is valid by construction: it only ever does what the designer allowed. The price is rigidity — when the priorities change, and safety starts to matter more than speed, the domain has to be rewritten. Letting a learner loose on primitive actions fixes the rigidity and throws away the guarantee. This work puts the learning somewhere else: on the choice of method.
Filtering: the domain says what is legal
At each compound task, the HTN exposes only the methods whose preconditions hold in the current state. Everything semantically invalid under the designer-defined constraints is gone before learning is consulted at all.
Learning: preference-conditioned MORL picks among the survivors
Rewards are vectors, not scalars, so objectives that disagree stay separate instead of collapsing into one number. Conditioning the policy on a preference vector gives a single agent that covers the trade-off space: for every linear preference, an optimal policy lies in the convex coverage set.
Preferences: what matters, right now
The weight vector carries the current trade-off between mission, safety, resources and time. Changing it changes which method gets selected — with the HTN domain untouched. That is the point: adaptation without re-authoring.
Credit where the decision was made
Only the selected method is decomposed, and primitive tasks keep executing conventionally. Credit is assigned at the hierarchical decision level rather than to individual actions, which is what keeps the learning signal legible and the behaviour explainable.
How it will be evaluated
Compare
First-applicable HTN
Heuristic HTN
Scalar HTN-RL
Contextual scalar RL + HTN mask
HTN + preference-conditioned MORL
Stress
Dynamic preferences
Novel context
Unseen preferences
Measure
Performance
Adaptation
Generalization
Semantic validity
Latency
Materials
The full paper is not published here. What is available is the poster and, once it is out, the ERAMIA-RS paper.
Poster: architecture, hierarchical credit assignment and the planned experiments.
PrototypeGrid-world prototype with HTN execution, sensing, navigation and replanning. Vector rewards, conflicting objectives and dynamic preferences are next. The repository is private while the work is in progress.