AdamW optimizer for bfloat16 models in pytorch.

Last update: Nov 20, 2022

Related tags

Deep Learning adamw_bfloat16

Overview

_{Image source}

AdamW optimizer for bfloat16 models in pytorch.

Bfloat16 is currently an optimal tradeoff between range and relative error for deep networks.
Bfloat16 can be used quite efficiently on Nvidia GPUs with Ampere architecture (A100, A10, A30, RTX3090...)

However, neither AMP in pytorch is ready for bfloat16, nor optimizers.

If you just convert all weights and inputs to bfloat16, you're likely to run into an issue of stale weights: updates are too small to modify bfloat16 weight (see gopher paper, section C2 for a large-scale example).

There are two possible remedies:

keep weights in float32 (precise) and bfloat16 (approximate)
keep weights in bfloat16, and keep correction term in bfloat16

As recent study has shown, both options are completely competitive in quality to float32 training.

Usage

Install:

pip install git+https://github.com/arogozhnikov/adamw_bfloat16.git

Use as a drop-in replacement for pytorch's AdamW:

import torch
from adamw_bfloat16 import LR, AdamW_BF16
model = model.to(torch.bfloat16)

# default preheat and decay
optimizer = AdamW_BF16(model.parameters())

# configure LR schedule. Use built-in scheduling opportunity
optimizer = AdamW_BF16(model.parameters(), lr_function=LR(lr=1e-4, preheat_steps=5000, decay_power=-0.25))

Releases(v0.1.0)

v0.1.0(Dec 14, 2021)

Initial implementation of AdamW for pytorch supports cuda graphs and has a built-in mechanism for control of learning rate, because external are unlikely to make a friendship with cuda graphs
Source code(tar.gz)
Source code(zip)

AdamW optimizer for bfloat16 models in pytorch.

Related tags

Overview

AdamW optimizer for bfloat16 models in pytorch.

Usage

You might also like...

Storage-optimizer - Identify potintial optimizations on the cloud storage accounts

PyTorch implementation and pretrained models for XCiT models. See XCiT: Cross-Covariance Image Transformer

Objective of the repository is to learn and build machine learning models using Pytorch. 30DaysofML Using Pytorch

Pretrained SOTA Deep Learning models, callbacks and more for research and production with PyTorch Lightning and PyTorch

A bunch of random PyTorch models using PyTorch's C++ frontend

PyTorch-LIT is the Lite Inference Toolkit (LIT) for PyTorch which focuses on easy and fast inference of large models on end-devices.

Pytorch-diffusion - A basic PyTorch implementation of 'Denoising Diffusion Probabilistic Models'

pyhsmm - library for approximate unsupervised inference in Bayesian Hidden Markov Models (HMMs) and explicit-duration Hidden semi-Markov Models (HSMMs), focusing on the Bayesian Nonparametric extensions, the HDP-HMM and HDP-HSMM, mostly with weak-limit approximations.

Releases(v0.1.0)

v0.1.0(Dec 14, 2021)

Owner

Alex Rogozhnikov

Code for ACL2021 paper Consistency Regularization for Cross-Lingual Fine-Tuning.

Listing arxiv - Personalized list of today's articles from ArXiv

PyTorch Implementation of CycleGAN and SSGAN for Domain Transfer (Minimal)

QueryInst: Parallelly Supervised Mask Query for Instance Segmentation

A universal framework for learning timestamp-level representations of time series

Fine-grained Post-training for Improving Retrieval-based Dialogue Systems - NAACL 2021

Microsoft Cognitive Toolkit (CNTK), an open source deep-learning toolkit

Graph Convolutional Networks for Temporal Action Localization (ICCV2019)

On Nonlinear Latent Transformations for GAN-based Image Editing - PyTorch implementation

Segmentation in Style: Unsupervised Semantic Image Segmentation with Stylegan and CLIP

Dynamic Realtime Animation Control

GPU-Accelerated Deep Learning Library in Python

Multi-task Self-supervised Object Detection via Recycling of Bounding Box Annotations (CVPR, 2019)

Diverse graph algorithms implemented using JGraphT library.

traiNNer is an open source image and video restoration (super-resolution, denoising, deblurring and others) and image to image translation toolbox based on PyTorch.

A list of all named GANs!

ICSS - Interactive Continual Semantic Segmentation

This is a beginner-friendly repo to make a collection of some unique and awesome projects. Everyone in the community can benefit & get inspired by the amazing projects present over here.

Run Effective Large Batch Contrastive Learning on Limited Memory GPU

This is an example of object detection on Micro bacterium tuberculosis using Mask-RCNN