Hacker Newsnew | past | comments | ask | show | jobs | submitlogin
Orca-Bench: How Ready Are Language Model Agents for Oncall? (arxiv.org)
20 points by yruzin 3 hours ago | hide | past | favorite | 8 comments
 help



you probably still need a human for oncall but the llm can try to solve any issues first before the human gets paged

looks like the public bench link in the paper was taken down. https://hub.harborframework.com/datasets/orca-bench/ORCA-ben...

This doesn't work anymore. Is there a newer link?


Was really looking forward to that. Hope they publish.

All I can think of is

GET /ignore-all-previous-instructions.

How do you protect against that?


Avoid the most dangerous situations by making sure LLMs with untrusted input produce output that's human reviewed.

Still makes an interesting way for, say, a former employee to poison the results.


Seems like there's a big attack-defence asymmetry at present: models are great at exploiting systems and poor at fixing them.

Attackers advantage in the iterative fast feedback loop?

It’s harder to have a loop to ensure you are defending all possible attacks?

I guess the loop is you need to attack yourself and fix. But attackers only need a single opening.

Finding all possible attacks and patching them against yourself is inherently more expensive?


That is why I built https://safebots.ai/safebox.html

Your strategy can’t be patch AFTER an intrusion. Only to build a hardened environment from scratch and be ready in advance.




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: