Flexible HDF5 saving/loading and other data science tools from the University of Chicago

Last update: Dec 10, 2022

Overview

https://travis-ci.org/uchicago-cs/deepdish.svg?branch=master

https://img.shields.io/badge/license-BSD%203--Clause-blue.svg?style=flat

deepdish

Flexible HDF5 saving/loading and other data science tools from the University of Chicago. This repository also host a Deep Learning blog:

http://deepdish.io

Installation

pip install deepdish

Alternatively (if you have conda with the conda-forge channel):

conda install -c conda-forge deepdish

Main feature

The primary feature of deepdish is its ability to save and load all kinds of data as HDF5. It can save any Python data structure, offering the same ease of use as pickling or numpy.save. However, it improves by also offering:

Interoperability between languages (HDF5 is a popular standard)
Easy to inspect the content from the command line (using h5ls or our specialized tool ddls)
Highly compressed storage (thanks to a PyTables backend)
Native support for scipy sparse matrices and pandas DataFrame, Series and Panel
Ability to partially read files, even slices of arrays

An example:

import deepdish as dd

d = {
    'foo': np.ones((10, 20)),
    'sub': {
        'bar': 'a string',
        'baz': 1.23,
    },
}
dd.io.save('test.h5', d)

This can be reconstructed using dd.io.load('test.h5'), or inspected through the command line using either a standard tool:

$ h5ls test.h5
foo                      Dataset {10, 20}
sub                      Group

Or, better yet, our custom tool ddls (or python -m deepdish.io.ls):

$ ddls test.h5
/foo                       array (10, 20) [float64]
/sub                       dict
/sub/bar                   'a string' (8) [unicode]
/sub/baz                   1.23 [float64]

Documentation

http://deepdish.readthedocs.io/

Flexible HDF5 saving/loading and other data science tools from the University of Chicago

Related tags

Overview

deepdish

Installation

Main feature

Documentation

Owner

UChicago - Department of Computer Science

A Pythonic introduction to methods for scaling your data science and machine learning work to larger datasets and larger models, using the tools and APIs you know and love from the PyData stack (such as numpy, pandas, and scikit-learn).

INFO-H515 - Big Data Scalable Analytics

MIR Cheatsheet - Survival Guidebook for MIR Researchers in the Lab

Amundsen is a metadata driven application for improving the productivity of data analysts, data scientists and engineers when interacting with data.

DefAP is a program developed to facilitate the exploration of a material's defect chemistry

Demonstrate a Dataflow pipeline that saves data from an API into BigQuery table

Pipetools enables function composition similar to using Unix pipes.

Exploratory Data Analysis of the 2019 Indian General Elections using a dataset from Kaggle.

Educational project on how to build an ETL (Extract, Transform, Load) data pipeline, orchestrated with Airflow.

A project consists in a set of assignements corresponding to a BI process: data integration, construction of an OLAP cube, qurying of a OPLAP cube and reporting.

A fast, flexible, and performant feature selection package for python.

Calculate multilateral price indices in Python (with Pandas and PySpark).

Accurately separate the TLD from the registered domain and subdomains of a URL, using the Public Suffix List.

Hangar is version control for tensor data. Commit, branch, merge, revert, and collaborate in the data-defined software era.

Developed for analyzing the covariance for OrcVIO

A Python package for the mathematical modeling of infectious diseases via compartmental models

SparseLasso: Sparse Solutions for the Lasso

Parses data out of your Google Takeout (History, Activity, Youtube, Locations, etc...)

A crude Hy handle on Pandas library

Elasticsearch tool for easily collecting and batch inserting Python data and pandas DataFrames