You are currently logged in as an
Institutional Subscriber.
If you would like to logout,
please click on the button below.
Home / Publications / E-library page
Only AES members and Institutional Journal Subscribers can download
In the context of music information retrieval, similarity-based approaches are useful for a variety of tasks that benefit from a query-by-example approach. Music however, naturally decomposes into a set of semantically meaningful factors of variation. Current representation learning strategies pursue the disentanglement of such factors from deep representations, and result in highly interpretable models. This allows to model the perception of music similarity, which is highly subjective and multi-dimensional. While the focus of prior work is on metadata driven similarity, we suggest to directly model the human notion of multi-dimensional music similarity. To achieve this, we propose a multi-input deep neural network architecture, which simultaneously processes mel-spectrogram, CENSchromagram and tempogram representations in order to extract informative features for different disentangled musical dimensions: genre, mood, instrument, era, tempo, and key. We evaluated the proposed music similarity approach using a triplet prediction task and found that the proposed multi-input architecture outperforms a state of the art method. Furthermore, we present a novel multi-dimensional analysis to evaluate the influence of each disentangled dimension on the perception of music similarity.
Author (s): Ribecky, Sebastian; Abeßer, Jakob; Lukashevich, Hanna
Affiliation:
Semantic Music Technologies Group, Fraunhofer IDMT, Ilmenau, Germany
(See document for exact affiliation information.)
AES Convention: 152
Paper Number:10568
Publication Date:
2022-05-06
Import into BibTeX
Session subject:
Sound Classification
Permalink: https://aes2.org/publications/elibrary-page/?id=21681
(1679KB)
Click to purchase paper as a non-member or login as an AES member. If your company or school subscribes to the E-Library then switch to the institutional version. If you are not an AES member Join the AES. If you need to check your member status, login to the Member Portal.
Ribecky, Sebastian; Abeßer, Jakob; Lukashevich, Hanna; 2022; Multi-Input Architecture and Disentangled Representation Learning for Multi-Dimensional Modeling of Music Similarity [PDF]; Semantic Music Technologies Group, Fraunhofer IDMT, Ilmenau, Germany; Paper 10568; Available from: https://aes2.org/publications/elibrary-page/?id=21681
Ribecky, Sebastian; Abeßer, Jakob; Lukashevich, Hanna; Multi-Input Architecture and Disentangled Representation Learning for Multi-Dimensional Modeling of Music Similarity [PDF]; Semantic Music Technologies Group, Fraunhofer IDMT, Ilmenau, Germany; Paper 10568; 2022 Available: https://aes2.org/publications/elibrary-page/?id=21681
@article{ribecky2022multi-input,
author={ribecky sebastian and abeßer jakob and lukashevich hanna},
journal={journal of the audio engineering society},
title={multi-input architecture and disentangled representation learning for multi-dimensional modeling of music similarity},
year={2022},
number={10568},
month={may},}