Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

In some respects that sounds similar to what we already do with reward models. I think with GRPO, the “bag of rewards” approach doesn’t strike me as terribly different. The challenge is in building out a sufficient “world” of rewards to adequately represent more meaningful feedback-based learning.

While it sounds nice to reframe it like a physics problem, it seems like a fundamentally flawed idea, akin to saying “there is a closed form solution to the question of how should I live.” The problem isn’t hallucinations, the problem is that language and relativism are inextricably linked.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: