You are currently logged in as an
Institutional Subscriber.
If you would like to logout,
please click on the button below.
Home / Publications / E-library page
Only AES members and Institutional Journal Subscribers can download
This paper investigates the use of Convolutional Neural Networks for spatial audio classification. In contrast to traditional methods that use hand-engineered features and algorithms, we show that a Convolutional Network in combination with generic preprocessing can give good results and allows for specialization to challenging conditions. The method can adapt to e.g. different source distances and microphone arrays, as well as estimate both spatial location and audio content type jointly. For example, with typical single-source material in a simulated reverberant room, we can achieve cross-validation accuracy of 94.3% for 40-ms frames across 16 classes (eight spatial directions, content type speech vs. music).
Author (s): Hirvonen, Toni
Affiliation:
Dolby Laboratories, Stockholm, Sweden
(See document for exact affiliation information.)
AES Convention: 138
Paper Number:9294
Publication Date:
2015-05-06
Import into BibTeX
Session subject:
Sound Localization and Separation
Permalink: https://aes2.org/publications/elibrary-page/?id=17718
(851KB)
Click to purchase paper as a non-member or login as an AES member. If your company or school subscribes to the E-Library then switch to the institutional version. If you are not an AES member Join the AES. If you need to check your member status, login to the Member Portal.
Hirvonen, Toni; 2015; Classification of Spatial Audio Location and Content Using Convolutional Neural Networks [PDF]; Dolby Laboratories, Stockholm, Sweden; Paper 9294; Available from: https://aes2.org/publications/elibrary-page/?id=17718
Hirvonen, Toni; Classification of Spatial Audio Location and Content Using Convolutional Neural Networks [PDF]; Dolby Laboratories, Stockholm, Sweden; Paper 9294; 2015 Available: https://aes2.org/publications/elibrary-page/?id=17718