Codebase for ECCV18 "The Sound of Pixels"

Last update: Dec 20, 2022

Overview

Sound-of-Pixels

Codebase for ECCV18 "The Sound of Pixels".

*This repository is under construction, but the core parts are already there.

Environment

The code is developed under the following configurations.

Hardware: 1-4 GPUs (change [--num_gpus NUM_GPUS] accordingly)
Software: Ubuntu 16.04.3 LTS, CUDA>=8.0, Python>=3.5, PyTorch>=0.4.0

Training

Prepare video dataset.

a. Download MUSIC dataset from: https://github.com/roudimit/MUSIC_dataset

b. Download videos.

Preprocess videos. You can do it in your own way as long as the index files are similar.

a. Extract frames at 8fps and waveforms at 11025Hz from videos. We have following directory structure:

data
├── audio
|   ├── acoustic_guitar
│   |   ├── M3dekVSwNjY.mp3
│   |   ├── ...
│   ├── trumpet
│   |   ├── STKXyBGSGyE.mp3
│   |   ├── ...
│   ├── ...
|
└── frames
|   ├── acoustic_guitar
│   |   ├── M3dekVSwNjY.mp4
│   |   |   ├── 000001.jpg
│   |   |   ├── ...
│   |   ├── ...
│   ├── trumpet
│   |   ├── STKXyBGSGyE.mp4
│   |   |   ├── 000001.jpg
│   |   |   ├── ...
│   |   ├── ...
│   ├── ...

b. Make training/validation index files by running:

python scripts/create_index_files.py

It will create index files train.csv/val.csv with the following format:

./data/audio/acoustic_guitar/M3dekVSwNjY.mp3,./data/frames/acoustic_guitar/M3dekVSwNjY.mp4,1580
./data/audio/trumpet/STKXyBGSGyE.mp3,./data/frames/trumpet/STKXyBGSGyE.mp4,493

For each row, it stores the information: AUDIO_PATH,FRAMES_PATH,NUMBER_FRAMES

Train the default model.

./scripts/train_MUSIC.sh

During training, visualizations are saved in HTML format under ckpt/MODEL_ID/visualization/.

Evaluation

(Optional) Download our trained model weights for evaluation.

./scripts/download_trained_model.sh

Evaluate the trained model performance.

./scripts/eval_MUSIC.sh

Reference

If you use the code or dataset from the project, please cite:

    @InProceedings{Zhao_2018_ECCV,
        author = {Zhao, Hang and Gan, Chuang and Rouditchenko, Andrew and Vondrick, Carl and McDermott, Josh and Torralba, Antonio},
        title = {The Sound of Pixels},
        booktitle = {The European Conference on Computer Vision (ECCV)},
        month = {September},
        year = {2018}
    }

Codebase for ECCV18 "The Sound of Pixels"

Related tags

Overview

Sound-of-Pixels

Environment

Training

Evaluation

Reference

Owner

Hang Zhao

ScaleNet: A Shallow Architecture for Scale Estimation

Lane follower: Lane-detector (OpenCV) + Object-detector (YOLO5) + CAN-bus

Tensorflow Tutorials using Jupyter Notebook

Official code for UnICORNN (ICML 2021)

Model Zoo for AI Model Efficiency Toolkit

HistoKT: Cross Knowledge Transfer in Computational Pathology

Codes of paper "Unseen Object Amodal Instance Segmentation via Hierarchical Occlusion Modeling"

Code for Recurrent Mask Refinement for Few-Shot Medical Image Segmentation (ICCV 2021).

Zsseg.baseline - Zero-Shot Semantic Segmentation

PyTorch implementation of the paper Dynamic Data Augmentation with Gating Networks

[PyTorch] Official implementation of CVPR2021 paper "PointDSC: Robust Point Cloud Registration using Deep Spatial Consistency". https://arxiv.org/abs/2103.05465

Practical and Real-world applications of ML based on the homework of Hung-yi Lee Machine Learning Course 2021

Gesture Volume Control Using OpenCV and MediaPipe

Py4fi2nd - Jupyter Notebooks and code for Python for Finance (2nd ed., O'Reilly) by Yves Hilpisch.

FindFunc is an IDA PRO plugin to find code functions that contain a certain assembly or byte pattern, reference a certain name or string, or conform to various other constraints.

Official Pytorch implementation for video neural representation (NeRV)

(CVPR2021) Kaleido-BERT: Vision-Language Pre-training on Fashion Domain

Code for the paper titled "Generalized Depthwise-Separable Convolutions for Adversarially Robust and Efficient Neural Networks" (NeurIPS 2021 Spotlight).

Mask2Former: Masked-attention Mask Transformer for Universal Image Segmentation in TensorFlow 2

HyDiff: Hybrid Differential Software Analysis