Pytorch implementation of "MOSNet: Deep Learning based Objective Assessment for Voice Conversion"

Last update: Nov 18, 2022

Related tags

Overview

MOSNet

pytorch implementation of "MOSNet: Deep Learning based Objective Assessment for Voice Conversion" https://arxiv.org/abs/1904.08352

Dependency

Linux Ubuntu 20.04

GPU: GeForce RTX 2080 Ti
CUDA version: 10.0

Python 3.7

pytorch==1.4.0
numpy==1.19.5
tqdm
scipy==1.6.2
pandas==1.2.4
matplotlib
librosa==0.6.0

Usage

Reproducing results in the paper

cd ./data and run bash download.sh to download the VCC2018 evaluation results and submitted speech. (downsample the submitted speech might take some times)
Run python mos_results_preprocess.py to prepare the evaluation results. (Run python bootsrap_estimation.py to do the bootstrap experiment for intrinsic MOS calculation)
Run python utils.py to extract .wav to .h5
Run python train.py -c config.json to train a CNN-BLSTM version of MOSNet.
Run python test.py -c config.json --epoch BEST_EPOCH --is_fp16 to test a CNN-BLSTM version of MOSNet.

Note

Thanks to the authors of the paper MOSNet and the code is based on their tensorflow implementation https://github.com/lochenchou/MOSNet. However, my workstation will show OOM errors even with BATCH_SIZE=4 under tensorflow2.0 and RTX 2080 Ti. Therefore I implement the code with pytorch. Currently only 7700MiB memory is used when BATCH_SIZE=64. If you find any problem with my code, you can write a issue.

Citation

If you find this work useful in your research, please consider citing:

@inproceedings{mosnet,
  author={Lo, Chen-Chou and Fu, Szu-Wei and Huang, Wen-Chin and Wang, Xin and Yamagishi, Junichi and Tsao, Yu and Wang, Hsin-Min},
  title={MOSNet: Deep Learning based Objective Assessment for Voice Conversion},
  year=2019,
  booktitle={Proc. Interspeech 2019},
}

License

This work is released under MIT License (see LICENSE file for details).

VCC2018 Database & Results

The model is trained on the large listening evaluation results released by the Voice Conversion Challenge 2018.
The listening test results can be downloaded from here
The databases and results (submitted speech) can be downloaded from here

Pytorch implementation of "MOSNet: Deep Learning based Objective Assessment for Voice Conversion"

Related tags

Overview

MOSNet

Dependency

Usage

Reproducing results in the paper

Note

Citation

License

VCC2018 Database & Results

Owner

Neighborhood Contrastive Learning for Novel Class Discovery

Send text to girlfriend in the morning

A collection of papers about Transformer in the field of medical image analysis.

Real-Time-Student-Attendence-System - Real Time Student Attendence System

Implement face detection, and age and gender classification, and emotion classification.

Implementation for paper MLP-Mixer: An all-MLP Architecture for Vision

Immortal tracker

Bootstrapped Unsupervised Sentence Representation Learning (ACL 2021)

Train an imgs.ai model on your own dataset

Implementation detail for paper "Multi-level colonoscopy malignant tissue detection with adversarial CAC-UNet"

Code for "LoRA: Low-Rank Adaptation of Large Language Models"

PlaidML is a framework for making deep learning work everywhere.

Data and extra materials for the food safety publications classifier

Language Models for the legal domain in Spanish done @ BSC-TEMU within the "Plan de las Tecnologías del Lenguaje" (Plan-TL).

3D Avatar Lip Syncronization from speech (JALI based face-rigging)

Revisiting Video Saliency: A Large-scale Benchmark and a New Model (CVPR18, PAMI19)

Publication describing 3 ML examples at NSLS-II and interfacing into Bluesky

The codes I made while I practiced various TensorFlow examples

Open-AI's DALL-E for large scale training in mesh-tensorflow.

RaceBERT -- A transformer based model to predict race and ethnicty from names