arXiv:2606. 20668v1 Announce Type: new Abstract: LLM supervision systems, namely input/output moderation filters and jailbreak detectors, are the primary safeguard against misuse in deployed AI applications, yet existing benchmarks are often vendor-biased, omit cost and latency, and rarely compare specialized guardrails against repurposed generalist LLMs.
Paper
BELLS-O: Evaluating the Operational Trade-offs of LLM Supervision Systems
Unreadunread