Easy to use Audio Tagging in PyTorch

Last update: Dec 22, 2022

Overview

Audio Classification, Tagging & Sound Event Detection in PyTorch

Progress:

Model Zoo

AudioSet Pretrained Models

Model	Task	mAP ^(%)	Sample Rate ^(kHz)	Window Length	Num Mels	Fmax	Weights
CNN14	Tagging	43.1	32	1024	64	14k	download
CNN14_16k	Tagging	43.8	16	512	64	8k	download

CNN14_DecisionLevelMax	SED	38.5	32	1024	64	14k	download

Note: These models will be used as a pretrained model in the fine-tuning tasks below. Check out audioset-tagging-cnn, if you want to train on AudioSet dataset.

Fine-tuned Classification Models

Model	Dataset	Accuracy ^(%)	Sample Rate ^(kHz)	Weights
CNN14	ESC50 (Fold-5)	95.75	32	download
CNN14	FSDKaggle2018 (test)	93.56	32	download
CNN14	SpeechCommandsv1 (val/test)	96.60/96.77	32	download

Fine-tuned Tagging Models

Model	Dataset	mAP(%)	AUC	d-prime	Sample Rate ^(kHz)	Config	Weights
CNN14	FSDKaggle2019	-	-	-	32	-	-

Fine-tuned SED Models

Model	Dataset	F1	Sample Rate ^(kHz)	Config	Weights
CNN14_DecisionLevelMax	DESED	-	32	-	-

Supported Datasets

Dataset	Task	Classes	Train	Val	Test	Audio Length	Audio Spec	Size
ESC-50	Classification	50	2,000	5 folds	-	5s	44.1kHz, mono	600MB
UrbanSound8k	Classification	10	8,732	10 folds	-	<=4s	Vary	5.6GB
FSDKaggle2018	Classification	41	9,473	-	1,600	300ms~30s	44.1kHz, mono	4.6GB
SpeechCommandsv1	Classification	30	51,088	6,798	6,835	<=1s	16kHz, mono	1.4GB
SpeechCommandsv2	Classification	35	84,843	9,981	11,005	<=1s	16kHz, mono	2.3GB

FSDKaggle2019*	Tagging	80	4,970+19,815	-	4,481	300ms~30s	44.1kHz, mono	24GB
MTT*	Tagging	50	19,000	-	-	-	-	3GB

DESED*	SED	10	-	-	-	10	-	-

Notes: * datasets are not available yet. Classification dataset are treated as multi-class/single-label classification and tagging and sed datasets are treated as multi-label classification.

Dataset Structure (click to expand)

Download the dataset and prepare it into the following structure.

datasets
|__ ESC50
    |__ audio

|__ Urbansound8k
    |__ audio

|__ FSDKaggle2018
    |__ audio_train
    |__ audio_test
    |__ FSDKaggle2018.meta
        |__ train_post_competition.csv
        |__ test_post_competition_scoring_clips.csv

|__ SpeechCommandsv1/v2
    |__ bed
    |__ bird
    |__ ...
    |__ testing_list.txt
    |__ validation_list.txt

Augmentations (click to expand)

Currently, the following augmentations are supported. More will be added in the future. You can test the effects of augmentations with this notebook

WaveForm Augmentations:

Spectrogram Augmentations:

Time Masking
Frequency Masking
Filter Augmentation

Usage

Requirements (click to expand)

python >= 3.6
pytorch >= 1.8.1
torchaudio >= 0.8.1

Other requirements can be installed with pip install -r requirements.txt.

Configuration (click to expand)

Create a configuration file in configs. Sample configuration for ESC50 dataset can be found here.
Copy the contents of this and then edit the fields you think if it is needed.
This configuration file is needed for all of training, evaluation and prediction scripts.

Training (click to expand)

To train with a single GPU:

$ python tools/train.py --cfg configs/CONFIG_FILE_NAME.yaml

To train with multiple gpus, set DDP field in config file to true and run as follows:

$ python -m torch.distributed.launch --nproc_per_node=2 --use_env tools/train.py --cfg configs/CONFIG_FILE_NAME.yaml

Evaluation (click to expand)

Make sure to set MODEL_PATH of the configuration file to your trained model directory.

$ python tools/val.py --cfg configs/CONFIG_FILE.yaml

Audio Classification/Tagging Inference

Set MODEL_PATH of the configuration file to your model's trained weights.
Change the dataset name in DATASET >> NAME as your trained model's dataset.
Set the testing audio file path in TEST >> FILE.
Run the following command.

$ python tools/infer.py --cfg configs/CONFIG_FILE.yaml

## for example
$ python tools/infer.py --cfg configs/audioset.yaml

You will get an output similar to this:

Class                     Confidence
----------------------  ------------
Speech                     0.897762
Telephone bell ringing     0.752206
Telephone                  0.219329
Inside, small room         0.20761
Music                      0.0770325

Sound Event Detection Inference

Set MODEL_PATH of the configuration file to your model's trained weights.
Change the dataset name in DATASET >> NAME as your trained model's dataset.
Set the testing audio file path in TEST >> FILE.
Run the following command.

$ python tools/sed_infer.py --cfg configs/CONFIG_FILE.yaml

## for example
$ python tools/sed_infer.py --cfg configs/audioset_sed.yaml

You will get an output similar to this:

Class                     Start    End
----------------------  -------  -----
Speech                      2.2    7
Telephone bell ringing      0      2.5

The following plot will also be shown, if you set PLOT to true:

References (click to expand)

Citations (click to expand)

@misc{kong2020panns,
      title={PANNs: Large-Scale Pretrained Audio Neural Networks for Audio Pattern Recognition}, 
      author={Qiuqiang Kong and Yin Cao and Turab Iqbal and Yuxuan Wang and Wenwu Wang and Mark D. Plumbley},
      year={2020},
      eprint={1912.10211},
      archivePrefix={arXiv},
      primaryClass={cs.SD}
}

@misc{gong2021ast,
      title={AST: Audio Spectrogram Transformer}, 
      author={Yuan Gong and Yu-An Chung and James Glass},
      year={2021},
      eprint={2104.01778},
      archivePrefix={arXiv},
      primaryClass={cs.SD}
}

@misc{nam2021heavily,
      title={Heavily Augmented Sound Event Detection utilizing Weak Predictions}, 
      author={Hyeonuk Nam and Byeong-Yun Ko and Gyeong-Tae Lee and Seong-Hu Kim and Won-Ho Jung and Sang-Min Choi and Yong-Hwa Park},
      year={2021},
      eprint={2107.03649},
      archivePrefix={arXiv},
      primaryClass={eess.AS}
}

You might also like...

TorchMetrics is a collection of 25+ PyTorch metrics implementations and an easy-to-use API to create custom metrics.

Machine learning metrics for distributed, scalable PyTorch applications.

1.2k Jan 6, 2023

TorchFlare is a simple, beginner-friendly, and easy-to-use PyTorch Framework train your models effortlessly.

TorchFlare TorchFlare is a simple, beginner-friendly and an easy-to-use PyTorch Framework train your models without much effort. It provides an almost

85 Dec 26, 2022

A more easy-to-use implementation of KPConv based on PyTorch.

A more easy-to-use implementation of KPConv This repo contains a more easy-to-use implementation of KPConv based on PyTorch. Introduction KPConv is a

36 Dec 29, 2022

Use MATLAB to simulate the signal and extract features. Use PyTorch to build and train deep network to do spectrum sensing.

Deep-Learning-based-Spectrum-Sensing Use MATLAB to simulate the signal and extract features. Use PyTorch to build and train deep network to do spectru

10 Dec 14, 2022

Fast image augmentation library and easy to use wrapper around other libraries. Documentation: https://albumentations.ai/docs/ Paper about library: https://www.mdpi.com/2078-2489/11/2/125

Albumentations Albumentations is a Python library for image augmentation. Image augmentation is used in deep learning and computer vision tasks to inc

11.4k Jan 9, 2023

Fast, flexible and easy to use probabilistic modelling in Python.

Please consider citing the JMLR-MLOSS Manuscript if you've used pomegranate in your academic work! pomegranate is a package for building probabilistic

3k Dec 29, 2022

High performance, easy-to-use, and scalable machine learning (ML) package, including linear model (LR), factorization machines (FM), and field-aware factorization machines (FFM) for Python and CLI interface.

What is xLearn? xLearn is a high performance, easy-to-use, and scalable machine learning package that contains linear model (LR), factorization machin

3k Jan 3, 2023

High performance, easy-to-use, and scalable machine learning (ML) package, including linear model (LR), factorization machines (FM), and field-aware factorization machines (FFM) for Python and CLI interface.

What is xLearn? xLearn is a high performance, easy-to-use, and scalable machine learning package that contains linear model (LR), factorization machin

2.8k Feb 12, 2021

A fast and easy to use, moddable, Python based Minecraft server!

PyMine PyMine - The fastest, easiest to use, Python-based Minecraft Server! Features Note: This list is not always up to date, and doesn't contain all

144 Dec 30, 2022

Releases(v0.2.0)

v0.2.0(Aug 17, 2021)
This release includes the following:

Fine-tuned on ESC50, FSDKaggle2018, SpeechCommandsv1

Add waveform augmentations

Add spectrogram augmentations

Add augmentation testing notebook

Add tagging metrics

Source code(tar.gz)
Source code(zip)
v0.1.0(Aug 13, 2021)
Add the following datasets:

ESC50

UrbanSound8k

FSDKaggle2018

SpeechCommandsv1/v2

Release fine-tuned model on ESC50.
Source code(tar.gz)
Source code(zip)

Owner

sithu3

AI Developer

GitHub Repository

TensorFlow implementation of the paper "Hierarchical Attention Networks for Document Classification"

Hierarchical Attention Networks for Document Classification This is an implementation of the paper Hierarchical Attention Networks for Document Classi

83 Dec 05, 2022

Learning Features with Parameter-Free Layers (ICLR 2022)

Learning Features with Parameter-Free Layers (ICLR 2022) Dongyoon Han, YoungJoon Yoo, Beomyoung Kim, Byeongho Heo | Paper NAVER AI Lab, NAVER CLOVA Up

65 Dec 07, 2022

Robust Video Matting in PyTorch, TensorFlow, TensorFlow.js, ONNX, CoreML!

6.5k Jan 04, 2023

Code for the paper "Combining Textual Features for the Detection of Hateful and Offensive Language"

The repository provides the source code for the paper "Combining Textual Features for the Detection of Hateful and Offensive Language" submitted to HA

3 Aug 04, 2022

Scaling Vision with Sparse Mixture of Experts

Scaling Vision with Sparse Mixture of Experts This repository contains the code for training and fine-tuning Sparse MoE models for vision (V-MoE) on I

290 Dec 25, 2022

NuPIC Studio is an all-in-one tool that allows users create a HTM neural network from scratch

NuPIC Studio is an all-in-one tool that allows users create a HTM neural network from scratch, train it, collect statistics, and share it among the members of the community. It is not just a visual

93 Sep 30, 2022

TDN: Temporal Difference Networks for Efficient Action Recognition

TDN: Temporal Difference Networks for Efficient Action Recognition Overview We release the PyTorch code of the TDN(Temporal Difference Networks).

326 Dec 13, 2022

September-Assistant - Open-source Windows Voice Assistant

September - Windows Assistant September is an open-source Windows personal assis

9 Nov 22, 2022

[CVPR 2021] Counterfactual VQA: A Cause-Effect Look at Language Bias

Counterfactual VQA (CF-VQA) This repository is the Pytorch implementation of our paper "Counterfactual VQA: A Cause-Effect Look at Language Bias" in C

94 Dec 03, 2022

Hardware accelerated, batchable and differentiable optimizers in JAX.

JAXopt Installation | Examples | References Hardware accelerated (GPU/TPU), batchable and differentiable optimizers in JAX. Installation JAXopt can be

621 Jan 08, 2023

Classical OCR DCNN reproduction based on PaddlePaddle framework.

Paddle-SVHN Classical OCR DCNN reproduction based on PaddlePaddle framework. This project reproduces Multi-digit Number Recognition from Street View I

1 Nov 12, 2021

Python lib to talk to pylontech lithium batteries (US2000, US3000, ...) using RS485

python-pylontech Python lib to talk to pylontech lithium batteries (US2000, US3000, ...) using RS485 What is this lib ? This lib is meant to talk to P

26 Dec 28, 2022

Jax/Flax implementation of Variational-DiffWave.

jax-variational-diffwave Jax/Flax implementation of Variational-DiffWave. (Zhifeng Kong et al., 2020, Diederik P. Kingma et al., 2021.) DiffWave with

37 Dec 16, 2022

Data stream analytics: Implement online learning methods to address concept drift in data streams using the River library. Code for the paper entitled "PWPAE: An Ensemble Framework for Concept Drift Adaptation in IoT Data Streams" accepted in IEEE GlobeCom 2021.

PWPAE-Concept-Drift-Detection-and-Adaptation This is the code for the paper entitled "PWPAE: An Ensemble Framework for Concept Drift Adaptation in IoT

162 Dec 16, 2022

Easy to use Audio Tagging in PyTorch

Related tags

Overview

Audio Classification, Tagging & Sound Event Detection in PyTorch

Model Zoo

Supported Datasets

Usage

You might also like...

TorchMetrics is a collection of 25+ PyTorch metrics implementations and an easy-to-use API to create custom metrics.

TorchFlare is a simple, beginner-friendly, and easy-to-use PyTorch Framework train your models effortlessly.

A more easy-to-use implementation of KPConv based on PyTorch.

Use MATLAB to simulate the signal and extract features. Use PyTorch to build and train deep network to do spectrum sensing.

Fast image augmentation library and easy to use wrapper around other libraries. Documentation: https://albumentations.ai/docs/ Paper about library: https://www.mdpi.com/2078-2489/11/2/125

Fast, flexible and easy to use probabilistic modelling in Python.

High performance, easy-to-use, and scalable machine learning (ML) package, including linear model (LR), factorization machines (FM), and field-aware factorization machines (FFM) for Python and CLI interface.

High performance, easy-to-use, and scalable machine learning (ML) package, including linear model (LR), factorization machines (FM), and field-aware factorization machines (FFM) for Python and CLI interface.

A fast and easy to use, moddable, Python based Minecraft server!

Releases(v0.2.0)

v0.2.0(Aug 17, 2021)

v0.1.0(Aug 13, 2021)

Owner

sithu3

TensorFlow implementation of the paper "Hierarchical Attention Networks for Document Classification"

Learning Features with Parameter-Free Layers (ICLR 2022)

Robust Video Matting in PyTorch, TensorFlow, TensorFlow.js, ONNX, CoreML!

Code for the paper "Combining Textual Features for the Detection of Hateful and Offensive Language"

Scaling Vision with Sparse Mixture of Experts

NuPIC Studio is an all­-in-­one tool that allows users create a HTM neural network from scratch

TDN: Temporal Difference Networks for Efficient Action Recognition

September-Assistant - Open-source Windows Voice Assistant

[CVPR 2021] Counterfactual VQA: A Cause-Effect Look at Language Bias

Hardware accelerated, batchable and differentiable optimizers in JAX.

Classical OCR DCNN reproduction based on PaddlePaddle framework.

Python lib to talk to pylontech lithium batteries (US2000, US3000, ...) using RS485

Jax/Flax implementation of Variational-DiffWave.

Data stream analytics: Implement online learning methods to address concept drift in data streams using the River library. Code for the paper entitled "PWPAE: An Ensemble Framework for Concept Drift Adaptation in IoT Data Streams" accepted in IEEE GlobeCom 2021.

git《Investigating Loss Functions for Extreme Super-Resolution》(CVPR 2020) GitHub:

Not Suitable for Work (NSFW) classification using deep neural network Caffe models.

This repository contains all the code and materials distributed in the 2021 Q-Programming Summer of Qode.

This is a simple plugin for Vim that allows you to use OpenAI Codex.

“Data Augmentation for Cross-Domain Named Entity Recognition” (EMNLP 2021)

Compositional Sketch Search

NuPIC Studio is an all-in-one tool that allows users create a HTM neural network from scratch