A lossless neural compression framework built on top of JAX.

Last update: Mar 14, 2022

Related tags

Overview

Kompressor

Branch	CI	Coverage
`main` (active)
`main`
`development`

A neural compression framework built on top of JAX.

Install

setup.py assumes a compatible version of JAX and JAXLib are already installed. Automated build is tested for a cuda:11.1-cudnn8-runtime-ubuntu20.04 environment with jaxlib==0.1.76+cuda11.cudnn82.

git clone https://github.com/rosalindfranklininstitute/kompressor.git
cd kompressor
pip install -e .

# Run tests
python -m pytest --cov=src/kompressor tests/

Install & Run through Docker environment

Docker image for the Kompressor dependencies are provided in the quay.io/rosalindfranklininstitute/kompressor:main Quay.io image.

# Run the container for the Kompressor environment
docker run --rm quay.io/rosalindfranklininstitute/kompressor:main \
    python -m pytest --cov=/usr/local/kompressor/src/kompressor /usr/local/kompressor/tests

Install & Run through Singularity environment

Singularity image for the Kompressor dependencies are provided in the rosalindfranklininstitute/kompressor/kompressor:main cloud.sylabs.io image.

singularity pull library://rosalindfranklininstitute/kompressor/kompressor:main
singularity run kompressor_main.sif \
    python -m pytest --cov=/usr/local/kompressor/src/kompressor /usr/local/kompressor/tests

Comments

Refactor map tuples to dicts

Closes #14. Functions which currently return an ordered tuple of maps (lrmap, udmap, cmap, ...) now return keyed dictionaries { 'lrmap': lrmap, 'udmap': udmap, 'cmap': cmap, ... } so that order/usage is explicitly enforced.

List comprehensions over the tuples now use jax.tree_map and jax.tree_multimap to ensure key safety.

@GMW99, this will break the current implementation of the Metrics Callback class which iterates over a zip of the hardcoded map names and the maps tuple. This iteration can be replaced by iterating over maps.items() since it is now a dict already.
enhancement

opened by JossWhittle 1
Ensure jax.jit static_argnums is refactored to static_argnames

Functions that currently mark static_argnums=(0, 1, 2) should be updated to use the safer static_argnames=('tom', 'dick', 'harry') that is now available.
enhancement high priority

opened by JossWhittle 1
Update development examples
Splits docker image into JAX base image and Kompressor dependency and install image

JAX image installs JAX from source to ensure correct CUDA / CUDNN versions

Adjust setup.py to install dependencies from requirement.txt

Refactors a how submodules are imported (within the kom.image submodule. Need to check volumes matches)

Add kom.image.data submodule for dealing with tensorflow data pipelines

Fixed pooling in the total variation losses (used as metrics in the example notebooks)

Move all the encoding/decoding functions for the maps into a kom.mapping submodule

Add within-k and run-length metrics to kom.image.metrics for example notebooks

Added example notebooks for interacting with the maps and training a basic Haiku compression model

feature
opened by JossWhittle 0
Add mapping encode/decode functions for float32 data

Will need a bit of thinking to get right. We probably need to consider similar tricks that we used for applying Radix Sort on float32 data to make the compression numerically stable and portable between machines.
enhancement low priority

opened by JossWhittle 0
Add mapping encode/decode functions for uint32 data

Some of our data is uint32 volumes.

Will need to trace through the full compression implementation and make sure intermediate value dtypes are large enough to avoid uint32 overflow when needed.
enhancement low priority

opened by JossWhittle 0
Modify core encode decode functions to pass a dict to the prediction function
Currently the lowres inputs are passed directly to the prediction_fn as the only input.

Modify to accept a dict that has at least one key for the lowres input.

Provide boolean flag to also pass a positional encoding tensor along with the lowres which the model can use if needed.

Chunked encode decode will need to generate the correct chunks of the positional encoding for the current chunk.

Model can choose how to use positional encodings.

Image case would receive (B, H, W, 2) tensor containing the Y and X coordinates of each pixel in the trailing axis.

Volume case would receive (B, D, H, W, 3) tensor containing the Z, Y, and X coordinates of each voxel in the trailing axis.

enhancement high priority
opened by JossWhittle 0
Look at decompressing sliced chunks
Decompress sliced chunk of image or volume without needing to decompress the entire data element.

May require applying secondary compression in blocks to avoid needing to decompress the full level maps, only to apply the predictor to the target slice.

Instead unpack just the blocks needed for the slice then trim.

A kompressor (or stack of) trained to secondary compress the maps from the primary kompressor (or stack of) would be able to naturally handle slice chunked decoding.

Could such a secondary compressor be shared between levels? Between multiple kompressors in the primary stack?

experiment low priority
opened by JossWhittle 0
Look at compressing timeseries data
Experiment with implementing the 1D case for compressing signals.

Video as sequence of 2D frames using the 3D volume code directly.

Look at compressing within timestep using information from neighbouring timesteps without actually compressing (dropping frames) the temporal axis.

experiment low priority
opened by JossWhittle 0

Releases(v0.0.0)

v0.0.0(Feb 14, 2022)

Base Kompressor release to test CI pipeline.
Source code(tar.gz)
Source code(zip)

Owner

Rosalind Franklin Institute

The Rosalind Franklin Institute is dedicated to transforming life science through interdisciplinary research and technology development

GitHub Repository

MADT: Offline Pre-trained Multi-Agent Decision Transformer

MADT: Offline Pre-trained Multi-Agent Decision Transformer A link to our paper can be found on Arxiv. Overview Official codebase for Offline Pre-train

51 Dec 21, 2022

Code and data of the Fine-Grained R2R Dataset proposed in paper Sub-Instruction Aware Vision-and-Language Navigation

Fine-Grained R2R Code and data of the Fine-Grained R2R Dataset proposed in the EMNLP2020 paper Sub-Instruction Aware Vision-and-Language Navigation. C

34 Nov 15, 2022

A PyTorch implementation of "Predict then Propagate: Graph Neural Networks meet Personalized PageRank" (ICLR 2019).

APPNP ⠀ A PyTorch implementation of Predict then Propagate: Graph Neural Networks meet Personalized PageRank (ICLR 2019). Abstract Neural message pass

329 Dec 30, 2022

Official repository for: Continuous Control With Ensemble DeepDeterministic Policy Gradients

Continuous Control With Ensemble Deep Deterministic Policy Gradients This repository is the official implementation of Continuous Control With Ensembl

4 Dec 06, 2021

A video scene detection algorithm is designed to detect a variety of different scenes within a video

Scene-Change-Detection - A video scene detection algorithm is designed to detect a variety of different scenes within a video. There is a very simple definition for a scene: It is a series of logical

1 Jan 04, 2022

Reference code for the paper CAMS: Color-Aware Multi-Style Transfer.

CAMS: Color-Aware Multi-Style Transfer Mahmoud Afifi1, Abdullah Abuolaim*1, Mostafa Hussien*2, Marcus A. Brubaker1, Michael S. Brown1 1York University

36 Dec 04, 2022

Matching python environment code for Lux AI 2021 Kaggle competition, and a gym interface for RL models.

Lux AI 2021 python game engine and gym This is a replica of the Lux AI 2021 game ported directly over to python. It also sets up a classic Reinforceme

74 Nov 03, 2022

Easy and Efficient Object Detector

EOD Easy and Efficient Object Detector EOD (Easy and Efficient Object Detection) is a general object detection model production framework. It aim on p

381 Jan 01, 2023

Band-Adaptive Spectral-Spatial Feature Learning Neural Network for Hyperspectral Image Classification

258 Dec 29, 2022

DeepHyper: Scalable Asynchronous Neural Architecture and Hyperparameter Search for Deep Neural Networks

What is DeepHyper? DeepHyper is a software package that uses learning, optimization, and parallel computing to automate the design and development of

214 Jan 08, 2023

Codebase for the self-supervised goal reaching benchmark introduced in the LEXA paper

LEXA Benchmark Codebase for the self-supervised goal reaching benchmark introduced in the LEXA paper (Discovering and Achieving Goals via World Models

36 Dec 22, 2022

Codes for [NeurIPS'21] You are caught stealing my winning lottery ticket! Making a lottery ticket claim its ownership.

You are caught stealing my winning lottery ticket! Making a lottery ticket claim its ownership Codes for [NeurIPS'21] You are caught stealing my winni

8 Nov 01, 2022

Introducing neural networks to predict stock prices

IntroNeuralNetworks in Python: A Template Project IntroNeuralNetworks is a project that introduces neural networks and illustrates an example of how o

637 Jan 04, 2023

Pi-NAS: Improving Neural Architecture Search by Reducing Supernet Training Consistency Shift (ICCV 2021)

Π-NAS This repository provides the evaluation code of our submitted paper: Pi-NAS: Improving Neural Architecture Search by Reducing Supernet Training

18 Aug 18, 2022

The LaTeX and Python code for generating the paper, experiments' results and visualizations reported in each paper is available (whenever possible) in the paper's directory

This repository contains the software implementation of most algorithms used or developed in my research. The LaTeX and Python code for generating the

3 Jan 03, 2023

A lossless neural compression framework built on top of JAX.

Related tags

Overview

Kompressor

Install

Install & Run through Docker environment

Install & Run through Singularity environment

Comments

Refactor map tuples to dicts

Ensure jax.jit static_argnums is refactored to static_argnames

Update development examples

Add mapping encode/decode functions for float32 data

Add mapping encode/decode functions for uint32 data

Modify core encode decode functions to pass a dict to the prediction function

Look at decompressing sliced chunks

Look at compressing timeseries data