← All concepts

reward modeling

3 articles · 5 co-occurring · 1 contradictions · 99 briefs

Rubric-based rewards break down desired model behavior into clear criteria that LLM judges use to give better feedback. This method improves reinforcement learning by making rewards more reliable, esp

Scaling Reinforcement Learning will never lead to AGI

[STRONG] "Its scalar reward-driven architecture leads to reward hacking and poor robustness" — Article argues reward optimization mechanisms are fundamentally flawed and lead to alignment problems

2026-W30
12
2026-W29
14
2026-W28
14
2026-W27
10
2026-W26
6
2026-W25
14
2026-W24
14
2026-W23
8
2026-W22
14
2026-W21
12
2026-W20
14
2026-W19
10

Rubric-based rewards break down desired model behavior into clear criteria that LLM judges use to give better feedback. This method improves reinforcement learning by making rewards more reliable, esp

Rubric-based rewards are a specific instantiation of the broader reward modeling challenge in RLHF.

Its scalar reward-driven architecture leads to reward hacking and poor robustness" — Article argues reward optimization mechanisms are fundamentally flawed and lead to alignment problems

query this concept
$ db.articles("reward-modeling")
$ db.cooccurrence("reward-modeling")
$ db.contradictions("reward-modeling")