JA EN
Glossary › attacks

GLOSSARY

attacks

appears in 3 paper titles

Definition

Inputs crafted to make a model fail or to bypass its constraints — adversarial perturbations in vision, prompt injection and jailbreaks in language. The recurring lesson is that defences are only ever evaluated against known attacks, so robustness claims expire. Most of these are not implementation bugs but consequences of what the model learned, which is why patching them one at a time rarely generalises.

Explainers using this term