PSLA: Improving Audio Tagging with Pretraining, Sampling, Labeling, and Aggregation

doi:doi:10.57702/62ys4vwk

PSLA: Improving Audio Tagging with Pretraining, Sampling, Labeling, and Aggregation

Audio tagging is an active research area and has a wide range of applications. Since the release of AudioSet, great progress has been made in advancing model performance, which mostly comes from the development of novel model architectures and attention modules. However, we find that appropriate training techniques are equally important for building audio tagging models with AudioSet, but have not received the attention they deserve.

BibTex:

@dataset{Yuan_Gong_and_Yu-An_Chung_and_James_Glass_2024,
    abstract = {Audio tagging is an active research area and has a wide range of applications. Since the release of AudioSet, great progress has been made in advancing model performance, which mostly comes from the development of novel model architectures and attention modules. However, we find that appropriate training techniques are equally important for building audio tagging models with AudioSet, but have not received the attention they deserve.},
    author = {Yuan Gong and Yu-An Chung and James Glass},
    doi = {10.57702/62ys4vwk},
    institution = {No Organization},
    keyword = {'Aggregation', 'Audio Tagging', 'AudioSet', 'Labeling', 'Pretraining', 'Sampling'},
    month = {dec},
    publisher = {TIB},
    title = {PSLA: Improving Audio Tagging with Pretraining, Sampling, Labeling, and Aggregation},
    url = {https://service.tib.eu/ldmservice/dataset/psla--improving-audio-tagging-with-pretraining--sampling--labeling--and-aggregation},
    year = {2024}
}