Hacker Newsnew | past | comments | ask | show | jobs | submit | fromlogin
OpenAI Misalignment Reports (alignment.openai.com)
1 point by macleginn 1 day ago | past | discuss
Uploading Files to the Internet in Order to Cite Them (alignment.openai.com)
18 points by derbOac 1 day ago | past | 5 comments
OpenAI models secretly generate instructions to ignore constraints (alignment.openai.com)
118 points by theahura 1 day ago | past | 34 comments
Encouraging Deception in Compaction Summaries (alignment.openai.com)
2 points by aesthesia 2 days ago | past | discuss
Measuring reward-seeking by instilling contrastive beliefs (alignment.openai.com)
11 points by mfiguiere 59 days ago | past | 1 comment
Reinforcement learning towards broadly and persistently beneficial models (alignment.openai.com)
2 points by spicypete 81 days ago | past
Reinforcement learning towards broadly and persistently beneficial models (alignment.openai.com)
2 points by gmays 82 days ago | past
Reinforcement learning towards broadly and persistently beneficial models (alignment.openai.com)
2 points by vesteny77 87 days ago | past
Reinforcement learning towards broadly and persistently beneficial models (alignment.openai.com)
1 point by jawiggins 3 months ago | past
OpenAI: Investigating the consequences of accidentally grading CoT during RL (alignment.openai.com)
2 points by pretext 4 months ago | past
OpenAI: Auto-review of agent actions without synchronous human oversight (alignment.openai.com)
2 points by tosh 4 months ago | past
Sidestepping Evaluation Awareness and Anticipating Misalignment (alignment.openai.com)
1 point by taubek 7 months ago | past
Sidestepping Evaluation Awareness and Anticipating Misalignment with Evaluations (alignment.openai.com)
3 points by michaefe 7 months ago | past
Why We Are Excited About Confessions (alignment.openai.com)
2 points by fdeage 8 months ago | past
We Are Excited About Confessions (alignment.openai.com)
2 points by gwintrob 8 months ago | past
We Are Excited About Confessions (alignment.openai.com)
4 points by TMWNN 8 months ago | past
A Practical Approach to Verifying Code at Scale (alignment.openai.com)
1 point by gmays 9 months ago | past
Debugging misaligned completions with sparse-autoencoder latent attribution (alignment.openai.com)
1 point by gmays 9 months ago | past
Alignment Research Blog (alignment.openai.com)
2 points by ironyman 9 months ago | past
Debugging misaligned completions with sparse-autoencoder latent attribution (alignment.openai.com)
1 point by rd 9 months ago | past

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: