Learning Spatio-Temporal Transformer for Visual Tracking

Last update: Jan 04, 2023

Related tags

Text Data & NLP Stark

Overview

STARK

The official implementation of the paper Learning Spatio-Temporal Transformer for Visual Tracking

Highlights

The strongest performances

Tracker	LaSOT (AUC)	GOT-10K (AO)	TrackingNet (AUC)
STARK	67.1	68.8	82.0
TransT	64.9	67.1	81.4
TrDiMP	63.7	67.1	78.4
Siam R-CNN	64.8	64.9	81.2

Real-Time Speed

STARK-ST50 and STARK-ST101 run at 40FPS and 30FPS respectively on a Tesla V100 GPU.

End-to-End, Post-processing Free

STARK is an end-to-end tracking approach, which directly predicts one accurate bounding box as the tracking result.
Besides, STARK does not use any hyperparameters-sensitive post-processing, leading to stable performances.

Purely PyTorch-based Code

STARK is implemented purely based on the PyTorch.

Install the environment

Option1: Use the Anaconda

conda create -n stark python=3.6
conda activate stark
bash install.sh

Option2: Use the docker file

We provide the complete docker at here

Data Preparation

Put the tracking datasets in ./data. It should look like:

${STARK_ROOT}
 -- data
     -- lasot
         |-- airplane
         |-- basketball
         |-- bear
         ...
     -- got10k
         |-- test
         |-- train
         |-- val
     -- coco
         |-- annotations
         |-- images
     -- trackingnet
         |-- TRAIN_0
         |-- TRAIN_1
         ...
         |-- TRAIN_11
         |-- TEST

Run the following command to set paths for this project

python tracking/create_default_local_file.py --workspace_dir . --data_dir ./data --save_dir .

After running this command, you can also modify paths by editing these two files

lib/train/admin/local.py  # paths about training
lib/test/evaluation/local.py  # paths about testing

Train STARK

Training with multiple GPUs using DDP

# STARK-S50
python tracking/train.py --script stark_s --config baseline --save_dir . --mode multiple --nproc_per_node 8  # STARK-S50
# STARK-ST50
python tracking/train.py --script stark_st1 --config baseline --save_dir . --mode multiple --nproc_per_node 8  # STARK-ST50 Stage1
python tracking/train.py --script stark_st2 --config baseline --save_dir . --mode multiple --nproc_per_node 8 --script_prv stark_st1 --config_prv baseline  # STARK-ST50 Stage2
# STARK-ST101
python tracking/train.py --script stark_st1 --config baseline_R101 --save_dir . --mode multiple --nproc_per_node 8  # STARK-ST101 Stage1
python tracking/train.py --script stark_st2 --config baseline_R101 --save_dir . --mode multiple --nproc_per_node 8 --script_prv stark_st1 --config_prv baseline_R101  # STARK-ST101 Stage2

(Optionally) Debugging training with a single GPU

python tracking/train.py --script stark_s --config baseline --save_dir . --mode single

Test and evaluate STARK on benchmarks

LaSOT

python tracking/test.py stark_st baseline --dataset lasot --threads 32
python tracking/analysis_results.py # need to modify tracker configs and names

GOT10K-test

python tracking/test.py stark_st baseline_got10k_only --dataset got10k_test --threads 32
python lib/test/utils/transform_got10k.py --tracker_name stark_st --cfg_name baseline_got10k_only

TrackingNet

python tracking/test.py stark_st baseline --dataset trackingnet --threads 32
python lib/test/utils/transform_trackingnet.py --tracker_name stark_st --cfg_name baseline

VOT2020
Before evaluating "STARK+AR" on VOT2020, please install some extra packages following external/AR/README.md

cd external/vot20/<workspace_dir>
export PYTHONPATH=<path to the stark project>:$PYTHONPATH
bash exp.sh

VOT2020-LT

cd external/vot20_lt/<workspace_dir>
export PYTHONPATH=<path to the stark project>:$PYTHONPATH
bash exp.sh

Test FLOPs, Params, and Speed

# Profiling STARK-S50 model
python tracking/profile_model.py --script stark_s --config baseline
# Profiling STARK-ST50 model
python tracking/profile_model.py --script stark_st2 --config baseline
# Profiling STARK-ST101 model
python tracking/profile_model.py --script stark_st2 --config baseline_R101

Model Zoo

The trained models, the training logs, and the raw tracking results are provided in the model zoo

Acknowledgments

Thanks for the great PyTracking Library, which helps us to quickly implement our ideas.
We use the implementation of the DETR from the official repo https://github.com/facebookresearch/detr.

Learning Spatio-Temporal Transformer for Visual Tracking

Related tags

Overview

STARK

Highlights

The strongest performances

Real-Time Speed

End-to-End, Post-processing Free

Purely PyTorch-based Code

Install the environment

Data Preparation

Train STARK

Test and evaluate STARK on benchmarks

Test FLOPs, Params, and Speed

Model Zoo

Acknowledgments

Owner

Multimedia Research

APEACH: Attacking Pejorative Expressions with Analysis on Crowd-generated Hate Speech Evaluation Datasets

A simple command line tool for text to image generation, using OpenAI's CLIP and a BigGAN

FireFlyer Record file format, writer and reader for DL training samples.

Neural network sequence labeling model

Speech Recognition Database Management with python

TaCL: Improve BERT Pre-training with Token-aware Contrastive Learning

Use fastai-v2 with HuggingFace's pretrained transformers

HAIS_2GNN: 3D Visual Grounding with Graph and Attention

Différents programmes créant une interface graphique a l'aide de Tkinter pour simplifier la vie des étudiants.

Global Rhythm Style Transfer Without Text Transcriptions

This is a really simple text-to-speech app made with python and tkinter.

🏆 • 5050 most frequent words in 109 languages

Use the state-of-the-art m2m100 to translate large data on CPU/GPU/TPU. Super Easy!

Simple python code to fix your combo list by removing any text after a separator or removing duplicate combos

:P Some basic stuff I'm gonna use for my upcoming Agile Software Development and Devops

GSoC'2021 | TensorFlow implementation of Wav2Vec2

easySpeech is an open-source Python wrapper for google speech to text API that doesn't require PyAudio(So you especially windows user don't have to deal with the errors while installing PyAudio) and also works with hugging face transformers

Code for "Semantic Role Labeling as Dependency Parsing: Exploring Latent Tree Structures Inside Arguments".

Chinese NER(Named Entity Recognition) using BERT(Softmax, CRF, Span)

[ICCV 2021] Instance-level Image Retrieval using Reranking Transformers