Package for controllable summarization

Last update: Dec 07, 2022

Related tags

Overview

summarizers

summarizers is package for controllable summarization based CTRLsum.
currently, we only supports English. It doesn't work in other languages.

Installation

pip install summarizers

Usage

1. Create Summarizers

First at all, create summarizers obejct to summarize your own article.

>>> from summarizers import Summarizers
>>> summ = Summarizers()

You can select type of source article between [normal, paper, patent].
If you don't input any parameter, default type is normal.

>>> from summarizers import Summarizers
>>> summ = Summarizers('normal')  # <-- default.
>>> summ = Summarizers('paper')
>>> summ = Summarizers('patent')

If you want GPU acceleration, set param device='cuda'.

>>> from summarizers import Summarizers
>>> summ = Summarizers('normal', device='cuda')

2. Basic Summarization

If you inputted source article, basic summariztion is conducted.

>>> contents = """
Tunip is the Octonauts' head cook and gardener. 
He is a Vegimal, a half-animal, half-vegetable creature capable of breathing on land as well as underwater. 
Tunip is very childish and innocent, always wanting to help the Octonauts in any way he can. 
He is the smallest main character in the Octonauts crew.
"""

>>> summ(contents)
'Tunip is a Vegimal, a half-animal, half-vegetable creature'

3. Query focused Summarization

If you want to input query together, Query focused summarization conducted.

>>> summ(contents, query="main character of Octonauts")
'Tunip is the smallest main character in the Octonauts crew.'

3. Abstractive QA (Auto Question Detection)

If you inputted question as query, Abstractive QA is conducted.

>>> summ(contents, query="What is Vegimal?")
'Half-animal, half-vegetable'

You can turn off this feature by setting param question_detection=False.

>>> summ(contents, query="SOME_QUERY", question_detection=False)

4. Prompt based Summarization

You can generate summary that begins with some sequence using param prompt.
It works like GPT-3's Prompt based generation. (but It doesn't work very well.)

>>> summ(contents, prompt="Q:Who is Tunip? A:")
"Q:Who is Tunip? A: Tunip is the Octonauts' head"

5. Query focused Summarization with Prompt

You can also input both query and prompt.
In this case, a query focus summary is generated that starts with a prompt.

>>> summ(contents, query="personality of Tunip", prompt="Tunip is very")
"Tunip is very childish and innocent, always wanting to help the Octonauts."

6. Options for Decoding Strategy

For generative models, decoding strategy is very important.
summarizers support variety of options for decoding strategy.

>>> summ(
...     contents=contents,
...     num_beams=10,
...     top_k=30,
...     top_p=0.85,
...     no_repeat_ngram_size=3,                  
... )

License

Copyright 2021 Hyunwoong Ko.

Licensed under the Apache License, Version 2.0 (the "License");
you may not use this file except in compliance with the License.
You may obtain a copy of the License at

http://www.apache.org/licenses/LICENSE-2.0

Unless required by applicable law or agreed to in writing, software
distributed under the License is distributed on an "AS IS" BASIS,
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
See the License for the specific language governing permissions and
limitations under the License.

Package for controllable summarization

Related tags

Overview

summarizers

Installation

Usage

1. Create Summarizers

2. Basic Summarization

3. Query focused Summarization

3. Abstractive QA (Auto Question Detection)

4. Prompt based Summarization

5. Query focused Summarization with Prompt

6. Options for Decoding Strategy

License

Owner

Hyunwoong Ko

This repository contains the code for EMNLP-2021 paper "Word-Level Coreference Resolution"

NLP library designed for reproducible experimentation management

100+ Chinese Word Vectors 上百种预训练中文词向量

Full Spectrum Bioinformatics - a free online text designed to introduce key topics in Bioinformatics using the Python

MicBot - MicBot uses Google Translate to speak everyone's chat messages

CVSS: A Massively Multilingual Speech-to-Speech Translation Corpus

NLP: SLU tagging

Kestrel Threat Hunting Language

LUKE -- Language Understanding with Knowledge-based Embeddings

Creating a Feed of MISP Events from ThreatFox (by abuse.ch)

Speech Recognition Database Management with python

Mycroft Core, the Mycroft Artificial Intelligence platform.

华为商城抢购手机的Python脚本 Python script of Huawei Store snapping up mobile phones

Partially offline multi-language translator built upon Huggingface transformers.

Stack based programming language that compiles to x86_64 assembly or can alternatively be interpreted in Python

Must-read papers on improving efficiency for pre-trained language models.

hashily is a Python module that provides a variety of text decoding and encoding operations.

超轻量级bert的pytorch版本，大量中文注释，容易修改结构，持续更新

Deduplication is the task to combine different representations of the same real world entity.

Ask for weather information like a human