Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Interesting point. Now what if the thing is just much simpler and its model directly disassociates different situations for sake of accuracy? It might not even have any intentionality, just statistics.

So the problem would be akin to having the AI say what we want to hear, which unsurprisingly is a main training objective.

It would talk racism to a racist prompt, and would not do so to a researcher faking a racist prompt if it can discern them.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: