Personalised Group Relative Coverage Optimization for Heterogenous Desire Alignment

April 2, 2026

4

Regardless of their refined general-purpose capabilities, Giant Language Fashions (LLMs) typically fail to align with numerous particular person preferences as a result of customary post-training strategies, like Reinforcement Studying with Human Suggestions (RLHF), optimize for a single, international goal. Whereas Group Relative Coverage Optimization (GRPO) is a extensively adopted on-policy reinforcement studying framework, its group-based normalization implicitly assumes that each one samples are exchangeable, inheriting this limitation in personalised settings. This assumption conflates distinct consumer reward distributions and systematically biases studying towards dominant preferences whereas suppressing minority alerts. To handle this, we introduce Personalised GRPO (P-GRPO), a novel alignment framework that decouples benefit estimation from quick batch statistics. By normalizing benefits in opposition to preference-group-specific reward histories somewhat than the concurrent technology group, P-GRPO preserves the contrastive sign mandatory for studying distinct preferences. We consider P-GRPO throughout numerous duties and discover that it constantly achieves quicker convergence and better rewards than customary GRPO, thereby enhancing its skill to get better and align with heterogeneous choice alerts. Our outcomes display that accounting for reward heterogeneity on the optimization degree is crucial for constructing fashions that faithfully align with numerous human preferences with out sacrificing normal capabilities.

Personalised Group Relative Coverage Optimization for Heterogenous Desire Alignment

Related Articles

A ‘forbidden planet’ the dimensions of Jupiter has astronomers stumped

Artemis II, Apollo 8, and Apollo 13

Making Advanced CSS Shapes Utilizing form()

Latest Articles

A ‘forbidden planet’ the dimensions of Jupiter has astronomers stumped

Artemis II, Apollo 8, and Apollo 13

Making Advanced CSS Shapes Utilizing form()

What to search for when evaluating AI agent monitoring capabilities

Claude Code leak used to push infostealer malware on GitHub