Security-Fidelity Tradeoffs: The Hidden Cost of Prompt Injection Defense
SecFid reveals that prompt-injection defenses can gain security by suppressing untrusted text, trading away fidelity on tasks that must preserve it.
3 entries
Papers and preprints on language model behavior, evaluation, and earlier work in scientific computing. Google Scholar ↗
SecFid reveals that prompt-injection defenses can gain security by suppressing untrusted text, trading away fidelity on tasks that must preserve it.
A study of how language models handle conflicts between learned programming knowledge and prompt-provided updates, with probes and activation steering for code generation.
A reproducible method for comparing crystal packings, with scalable implementation work and analysis for crystal structure similarity.