Towards Better Understanding of Training Certifiably Robust Models against Adversarial Examples

Citations

WEB OF SCIENCE

2
Citations

SCOPUS

20

초록

We study the problem of training certifiably robust models against adversarial examples. Certifiable training minimizes an upper bound on the worst-case loss over the allowed perturbation, and thus the tightness of the upper bound is an important factor in building certifiably robust models. However, many studies have shown that Interval Bound Propagation (IBP) training uses much looser bounds but outperforms other models that use tighter bounds. We identify another key factor that influences the performance of certifiable training: smoothness of the loss landscape. We find significant differences in the loss landscapes across many linear relaxation-based methods, and that the current state-of-the-arts method often has a landscape with favorable optimization properties. Moreover, to test the claim, we design a new certifiable training method with the desired properties. With the tightness and the smoothness, the proposed method achieves a decent performance under a wide range of perturbations, while others with only one of the two factors can perform well only for a specific range of perturbations. Our code is available at https://github.com/sungyoon-lee/LossLandscapeMatters. © 2021 Neural information processing systems foundation. All rights reserved.

제목
Towards Better Understanding of Training Certifiably Robust Models against Adversarial Examples
저자
Lee, SungyoonLee, WoojinPark, JinseongLee, Jaewook
발행일
2021
유형
Conference Paper
저널명
Advances in Neural Information Processing Systems
2
페이지
953 ~ 964