Breaking certified defenses: Semantic adversarial examples with spoofed robustness certificates.

Amin Ghiasi,Ali Shafahi,Tom Goldstein

Breaking certified defenses: Semantic adversarial examples with spoofed robustness certificates.

2020

Amin Ghiasi
Ali Shafahi
Tom Goldstein

To deflect adversarial attacks, a range of "certified" classifiers have been proposed. In addition to labeling an image, certified classifiers produce (when possible) a certificate guaranteeing that the input image is not an $\ell_p$-bounded adversarial example. We present a new attack that exploits not only the labelling function of a classifier, but also the certificate generator. The proposed method applies large perturbations that place images far from a class boundary while maintaining the imperceptibility property of adversarial examples. The proposed "Shadow Attack" causes certifiably robust networks to mislabel an image and simultaneously produce a "spoofed" certificate of robustness.

Keywords:

Robustness (computer science)
Exploit
Artificial intelligence
Shadow
Certification
Spoofing attack
Mathematics
Machine learning
Adversarial system
Classifier (linguistics)
Certificate

Correction
Source
Cite
Save
Machine Reading By IdeaReader

References

Citations