Evaluating Coding Agents on Kernel Exploit Generation
arXiv:2609.25591v1 Announce Type: new Abstract: Coding agents now find real vulnerabilities in production software. However, bug discovery results do not measure whether agents can construct exploit primitives. We introduce KEX-bench, a benchmark for evaluating coding agents on exploit primitive…