Home / Publications / E-library page
Only AES members and Institutional Journal Subscribers can download
In this paper, we propose a pixel-based attention (PBA) module for acoustic scene classification (ASC). By performing feature compression on the input spectrogram along the spatial dimension, PBA can obtain the global information of the spectrogram. Besides, PBA applies attention weights to each pixel of each channel through two convolutional layers combined with global information. In addition, the spectrogram applied after the attention weights is multiplied by the gamma coefficient and superimposed with the original spectrogram to obtain more effective spectrogram features for training the network model. Furthermore, this paper implements a convolutional neural network (CNN) based on PBA (PB-CNN) and compares its classification performance on task 1 of Detection and Classification of Acoustic Scenes and Events (DCASE) 2016 Challenge with CNN based on time attention (TB-CNN), CNN based on frequency attention (FB-CNN), and pure CNN. The experimental results show that the proposed PB-CNN achieves the highest accuracy of 89.2% among the four CNNs, 1.9% higher than that of TB-CNN (87.3%), 2.2% higher than that of FB-CNN (86.6%), and 3% higher than that of pure CNN (86.2%). Compared with DCASE 2016’s baseline system, the PB-CNN improved by 12%, and its 89.2% accuracy was the highest among all submitted single models.
Author (s): Wang, Xingmei; Xu, Yichao; Shi, Jiahao; Teng, Xuyang
Affiliation:
College of Computer Science and Technology, Harbin Engineering University, Harbin, 150001, People’s Republic of China; College of Communication Engineering, Hangzhou Dianzi University, Hangzhou, 310018, People’s Republic of China
(See document for exact affiliation information.)
Publication Date:
2020-11-06
Import into BibTeX
Permalink: https://aes2.org/publications/elibrary-page/?id=20998
(588KB)
Click to purchase paper as a non-member or login as an AES member. If your company or school subscribes to the E-Library then switch to the institutional version. If you are not an AES member Join the AES. If you need to check your member status, login to the Member Portal.
Wang, Xingmei; Xu, Yichao; Shi, Jiahao; Teng, Xuyang; 2020; Acoustic Scene Classification Using Pixel-Based Attention [PDF]; College of Computer Science and Technology, Harbin Engineering University, Harbin, 150001, People’s Republic of China; College of Communication Engineering, Hangzhou Dianzi University, Hangzhou, 310018, People’s Republic of China; Paper ; Available from: https://aes2.org/publications/elibrary-page/?id=20998
Wang, Xingmei; Xu, Yichao; Shi, Jiahao; Teng, Xuyang; Acoustic Scene Classification Using Pixel-Based Attention [PDF]; College of Computer Science and Technology, Harbin Engineering University, Harbin, 150001, People’s Republic of China; College of Communication Engineering, Hangzhou Dianzi University, Hangzhou, 310018, People’s Republic of China; Paper ; 2020 Available: https://aes2.org/publications/elibrary-page/?id=20998
@article{wang2020acoustic,
author={wang xingmei and xu yichao and shi jiahao and teng xuyang},
journal={journal of the audio engineering society},
title={acoustic scene classification using pixel-based attention},
year={2020},
volume={68},
issue={11},
pages={843-855},
month={november},}