arXiv:2606. 26397v1 Announce Type: new Abstract: Real-world decision-making often requires balancing multiple conflicting objectives, a challenge that standard Reinforcement Learning (RL) frequently addresses by aggregating rewards into a single scalar signal.
Paper
Deterministic Pareto-Optimal Policy Synthesis for Multi-Objective Reinforcement Learning
Unreadunread