A hybrid CNN-autoencoder-SVM/XGBoost model for polyphonic orchestral instrument classification
Abstract
Polyphonic orchestral recordings pose significant challenges in music information retrieval (MIR) due to their overlapping frequency ranges and timbral similarities among instrument families, which complicate multi-label instrument classification. Prior studies have explored the integration of convolutional neural networks (CNN)-based feature extraction with classical machine learning (ML) classifiers, often on monophonic or simpler datasets like IRMAS. But the integration of deep learning (DL) feature extraction, Autoencoder (AE)-based dimensionality reduction, and ML classifiers for polyphonic orchestral instrument recognition remains underexplored. This study proposes a hybrid framework utilizing a pre-trained Inception V3 CNN for feature extraction from mel-spectrograms, followed by an optional 50% dimensionality reduction via AE, and finally, classification with support vector machines (SVM) or extreme gradient boosting (XGBoost). Experiments were run on two polyphonic datasets, OpenMIC-2018 and Orchset. The results demonstrate that non-AE configurations generally outperform AE variants. These results extend prior studies such using polyphonic datasets. The results highlight the practical value of hybrid CNN-ML pipelines and the trade-offs of feature compression in MIR.
Keywords
Autoencoder; CNN information retrieval; Music; Orchestral instrument classification polyphonic audio; SVM; XGBoost
Full Text:
PDFDOI: http://doi.org/10.11591/ijeecs.v43.i1.pp114-126
Refbacks
- There are currently no refbacks.

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
Indonesian Journal of Electrical Engineering and Computer Science (IJEECS)
p-ISSN: 2502-4752, e-ISSN: 2502-4760
This journal is published by the Institute of Advanced Engineering and Science (IAES).