Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

It looks like Zach Weinersmith predicted this exact line of research 11 years ago [1], when he suggested testing the liar sentence using fMRI.

The same analysis applies: the probe tells us what the LLM thinks about the truth value if the sentence, not the truth value of the sentence. I don't think anyone claimed that these probes were truth oracles.

[1] https://smbc-comics.com/index.php?id=3657



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: