Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework

Last update: Jan 08, 2023

Overview

OFA

OFA is a unified multimodal pretrained model that unifies modalities (i.e., cross-modality, vision, language) and tasks (e.g., image generation, visual grounding, image captioning, image classification, text generation, etc.) to a simple sequence-to-sequence learning framework. For more information, please refer to our paper: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework.

News

2022.2.13: Released the demo of image captioning. Have fun!
2022.2.11: Released the Colab notebook for image captioning . Enjoy!
2022.2.11: Released the pretrained checkpoint of OFA-Large and the complete (2-staged) finetuning code for image captioning.
2022.2.10: Released the inference code & finetuned checkpoint for image captioning, which can reproduce the results on COCO Karparthy test split (149.6 CIDEr)

TODO

To release finetuning and inference codes for multimodal downstream tasks soon, including image captioning, VQA, text-to-image generation, SNLI-VE, Referring expression, comprehension, etc.
To release codes for pretraining soon.

Approach

Requirements

python 3.7.4
pytorch 1.8.1
torchvision 0.9.1
JAVA 1.8 (for COCO evaluation)

Installation

git clone https://github.com/OFA-Sys/OFA
pip install -r requirements.txt

Datasets and Checkpoints

See datasets.md and checkpoints.md.

Pretraining

To release soon:)

Finetuning & Inference

Below we provide methods for fintuning and inference on different downstream tasks.

Caption

Download data and files and put them in the correct directory
Train

cd run_scripts/caption
nohup sh train_caption_stage1.sh &  # stage1, train with cross-entropy loss
nohup sh train_caption_stage2.sh &  # stage2, load the best ckpt of stage1 and train with CIDEr optimization

Inference

cd run_scripts/caption ; sh evaluate_caption.sh  # inference & evaluate

Gallery

Below we provide examples of OFA in text-to-image generation and open-ended VQA. Also, we demonstrate its performance in unseen task (Grounded QA) as well as unseen domain (Visual Grounding on images from unseen domains).

Text-to-Image Generation (normal query)

Text-to-Image Generation (counterfactual query)

Open-Ended VQA

Grounded QA (unseen task)

Viusal Grounding (unseen domain)

Citation

Please cite our paper if you find it helpful :)

@article{wang2022OFA,
  title={Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework},
  author={Wang, Peng and Yang, An and Men, Rui and Lin, Junyang and Bai, Shuai and Li, Zhikang and Ma, Jianxin and Zhou, Chang and Zhou, Jingren and Yang, Hongxia},
  journal={arXiv e-prints},
  pages={arXiv--2202},
  year={2022}
}

Related Codebase

Fairseq

License

Apache-2.0

Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework

Related tags

Overview

OFA

News

TODO

Approach

Requirements

Installation

Datasets and Checkpoints

Pretraining

Finetuning & Inference

Caption

Gallery

Text-to-Image Generation (normal query)

Text-to-Image Generation (counterfactual query)

Open-Ended VQA

Grounded QA (unseen task)

Viusal Grounding (unseen domain)

Citation

Related Codebase

License

Owner

OFA Sys

Fast, Attemptable Route Planner for Navigation in Known and Unknown Environments

This is a collection of all challenges in HKCERT CTF 2021

Sync2Gen Code for ICCV 2021 paper: Scene Synthesis via Uncertainty-Driven Attribute Synchronization

Repo for paper "Dynamic Placement of Rapidly Deployable Mobile Sensor Robots Using Machine Learning and Expected Value of Information"

DGCNN - Dynamic Graph CNN for Learning on Point Clouds

Orbivator AI - To Determine which features of data (measurements) are most important for diagnosing breast cancer and find out if breast cancer occurs or not.

Dcf-game-infrastructure-public - Contains all the components necessary to run a DC finals (attack-defense CTF) game from OOO

This code implements constituency parse tree aggregation

Implementation of a Transformer, but completely in Triton

Experiments for Fake News explainability project

Generic image compressor for machine learning. Pytorch code for our paper "Lossy compression for lossless prediction".

Groceries ARL: Association Rules (Birliktelik Kuralı)

Source code for paper "ATP: AMRize Than Parse! Enhancing AMR Parsing with PseudoAMRs" @NAACL-2022

PyTorch implementations of the paper: "Learning Independent Instance Maps for Crowd Localization"

Sign-to-Speech for Sign Language Understanding: A case study of Nigerian Sign Language

A pytorch-based real-time segmentation model for autonomous driving

[제 13회 투빅스 컨퍼런스] OK Mugle! - 장르부터 멜로디까지, Content-based Music Recommendation

Geometric Deep Learning Extension Library for PyTorch

Event queue (Equeue) dialect is an MLIR Dialect that models concurrent devices in terms of control and structure.

Implementation of Diverse Semantic Image Synthesis via Probability Distribution Modeling