Use PaddlePaddle to reproduce the paper：mT5: A Massively Multilingual Pre-trained Text-to-Text Transformer

Last update: Oct 17, 2021

Related tags

Text Data & NLP MT5_paddle

Overview

MT5_paddle

Use PaddlePaddle to reproduce the paper：mT5: A Massively Multilingual Pre-trained Text-to-Text Transformer

English | 简体中文

mT5: A Massively Multilingual Pre-trained Text-to-Text Transformer

Abstract： The recent “Text-to-Text Transfer Transformer” (T5) leveraged a unified text-to-text format and scale to attain state-of-the-art results on a wide variety of English-language NLP tasks. In this paper, we introduce mT5, a multilingual variant of T5 that was pre-trained on a new Common Crawl-based dataset covering 101 languages. We detail the design and modified training of mT5 and demonstrate its state-of-the-art performance on many multilingual benchmarks. We also describe a simple technique to prevent “accidental translation” in the zero-shot setting, where a generative model chooses to (partially) translate its prediction into the wrong language. All of the code and model checkpoints used in this work are publicly available.

This project is an open source implementation of MT5 on Paddle 2.x.

Environment Installation

label	value
python	>=3.6
GPU	V100
Frame	PaddlePaddle2.1.2
Cuda	10.1
Cudnn	7.6

Cloud platform used in this recurrence：https://aistudio.baidu.com/

# Clone the repository
git clone https://github.com/27182812/MT5_paddle
# Enter the root directory
cd MT5_paddle
# Install the necessary python libraries locally
pip install -r requirements.txt

"test.ipynb" has run results display.

Quick Start

（一）Tokenizer Accuracy Alignment

### 对齐tokenizer
text = "Welcome to use paddle and paddlenlp!"
torch_tokenizer = PTT5Tokenizer.from_pretrained("./mt5-large")
paddle_tokenizer = PDT5Tokenizer.from_pretrained("./mt5-large")
torch_inputs = torch_tokenizer(text)
paddle_inputs = paddle_tokenizer(text)
print(torch_inputs)
print(paddle_inputs)

（二）Model Accuracy Alignment

run python compare.py，Comparing the accuracy between huggingface and paddle.

python compare.py
# MT5-large-pytorch vs paddle MT5-large-paddle
mean difference: tensor(2.0390e-06)
max difference: tensor(0.0004)

(三）Weights Transform

run python convert.py，transform weights of huggingface model to weights of paddle model. The weight path needs to be replaced

(四）Downstream task fine-tuning

run python train.py. "args.py" is for parameter.

Reference

大佬的T5代码：https://github.com/JunnYu/paddle_t5

@unknown{unknown,
author = {Xue, Linting and Constant, Noah and Roberts, Adam and Kale, Mihir and Al-Rfou, Rami and Siddhant, Aditya and Barua, Aditya and Raffel, Colin},
year = {2020},
month = {10},
pages = {},
title = {mT5: A massively multilingual pre-trained text-to-text transformer}
}

Use PaddlePaddle to reproduce the paper：mT5: A Massively Multilingual Pre-trained Text-to-Text Transformer

Related tags

Overview

MT5_paddle

Environment Installation

Quick Start

（一）Tokenizer Accuracy Alignment

（二）Model Accuracy Alignment

(三）Weights Transform

(四）Downstream task fine-tuning

Reference

Owner

本项目是作者们根据个人面试和经验总结出的自然语言处理(NLP)面试准备的学习笔记与资料，该资料目前包含自然语言处理各领域的面试题积累。

Accurately generate all possible forms of an English word e.g "election" --> "elect", "electoral", "electorate" etc.

Paradigm Shift in NLP - "Paradigm Shift in Natural Language Processing".

SentAugment is a data augmentation technique for semi-supervised learning in NLP.

Generate vector graphics from a textual caption

Mycroft Core, the Mycroft Artificial Intelligence platform.

Faster, modernized fork of the language identification tool langid.py

GPT-3: Language Models are Few-Shot Learners

A2T: Towards Improving Adversarial Training of NLP Models (EMNLP 2021 Findings)

A python project made to generate code using either OpenAI's codex or GPT-J (Although not as good as codex)

Production First and Production Ready End-to-End Keyword Spotting Toolkit

NL. The natural language programming language.

Semantic search through a vectorized Wikipedia (SentenceBERT) with the Weaviate vector search engine

Official implementation of Meta-StyleSpeech and StyleSpeech

Chinese Pre-Trained Language Models (CPM-LM) Version-I

TalkNet: Audio-visual active speaker detection Model

Journalism AI – Quotes extraction for modular journalism

Multilingual Emotion classification using BERT (fine-tuning). Published at the WASSA workshop (ACL2022).

Biterm Topic Model (BTM): modeling topics in short texts

This repository contains the official release of the model "BanglaBERT" and associated downstream finetuning code and datasets introduced in the paper titled "BanglaBERT: Combating Embedding Barrier in Multilingual Models for Low-Resource Language Understanding".

Use PaddlePaddle to reproduce the paper：mT5: A Massively Multilingual Pre-trained Text-to-Text Transformer

Related tags

Overview

MT5_paddle

Environment Installation

Quick Start

（一）Tokenizer Accuracy Alignment

（二）Model Accuracy Alignment

(三）Weights Transform

(四）Downstream task fine-tuning

Reference

Owner

本项目是作者们根据个人面试和经验总结出的自然语言处理(NLP)面试准备的学习笔记与资料，该资料目前包含 自然语言处理各领域的 面试题积累。

Accurately generate all possible forms of an English word e.g "election" --> "elect", "electoral", "electorate" etc.

Paradigm Shift in NLP - "Paradigm Shift in Natural Language Processing".

SentAugment is a data augmentation technique for semi-supervised learning in NLP.

Generate vector graphics from a textual caption

Mycroft Core, the Mycroft Artificial Intelligence platform.

Faster, modernized fork of the language identification tool langid.py

GPT-3: Language Models are Few-Shot Learners

A2T: Towards Improving Adversarial Training of NLP Models (EMNLP 2021 Findings)

A python project made to generate code using either OpenAI's codex or GPT-J (Although not as good as codex)

Production First and Production Ready End-to-End Keyword Spotting Toolkit

NL. The natural language programming language.

Semantic search through a vectorized Wikipedia (SentenceBERT) with the Weaviate vector search engine

Official implementation of Meta-StyleSpeech and StyleSpeech

Chinese Pre-Trained Language Models (CPM-LM) Version-I

TalkNet: Audio-visual active speaker detection Model

Journalism AI – Quotes extraction for modular journalism

Multilingual Emotion classification using BERT (fine-tuning). Published at the WASSA workshop (ACL2022).

Biterm Topic Model (BTM): modeling topics in short texts

This repository contains the official release of the model "BanglaBERT" and associated downstream finetuning code and datasets introduced in the paper titled "BanglaBERT: Combating Embedding Barrier in Multilingual Models for Low-Resource Language Understanding".

本项目是作者们根据个人面试和经验总结出的自然语言处理(NLP)面试准备的学习笔记与资料，该资料目前包含自然语言处理各领域的面试题积累。