LYU Fan, LI Linyan, Victor S. Sheng, et al., “Multi-label Image Classification via Coarse-to-Fine Attention,” Chinese Journal of Electronics, vol. 28, no. 6, pp. 1118-1126, 2019, doi: 10.1049/cje.2019.07.015
Multi-label Image Classification via Coarse-to-Fine Attention

doi: 10.1049/cje.2019.07.015
Funds:  This work is supported by the National Natural Science Foundation of China (No.61876121, No.61472267, No.61728205, No.61502329, No.61672371), Primary Research & Developement Plan of Jiangsu Province (No.BE2017663), Natural Science Foundation of the Higher Education Institutions of Jiangsu Province (No.19KJB520054), and Foundation of Key Laboratory in Science and Technology Development Project of Suzhou(No.SZS201609, No.SZS201813).
  • Corresponding author: HU Fuyuan (corresponding author) was a postdoctoral researcher at Vrije Universiteit Brussel,Belgium,a Ph.D.student at Northwestern Polytechnical University,and a visiting Ph.D.student at the City University of Hong Kong.He is a professor at Suzhou University of Science and Technology.His research interests include graphical models,structured learning,and tracking.(
  • Received Date: 2018-09-05
  • Rev Recd Date: 2019-07-23
  • Publish Date: 2019-11-10
  • Great efforts have been made by using deep neural networks to recognize multi-label images. Since multi-label image classification is very complicated, many studies seek to use the attention mechanism as a kind of guidance. Conventional attention-based methods always analyzed images directly and aggressively, which is difficult to well understand complicated scenes. We propose a global/local attention method that can recognize a multi-label image from coarse to fine by mimicking how human-beings observe images. Our global/local attention method first concentrates on the whole image, and then focuses on its local specific objects. We also propose a joint max-margin objective function, which enforces that the minimum score of positive labels should be larger than the maximum score of negative labels horizontally and vertically. This function further improve our multi-label image classification method. We evaluate the effectiveness of our method on two popular multi-label image datasets (i.e., Pascal VOC and MS-COCO). Our experimental results show that our method outperforms state-of-the-art methods.
