Learning RGB-D Feature Embeddings for Unseen Object Instance Segmentation

Last update: Dec 13, 2022

Related tags

Overview

Unseen Object Clustering: Learning RGB-D Feature Embeddings for Unseen Object Instance Segmentation

Introduction

In this work, we propose a new method for unseen object instance segmentation by learning RGB-D feature embeddings from synthetic data. A metric learning loss functionis utilized to learn to produce pixel-wise feature embeddings such that pixels from the same object are close to each other and pixels from different objects are separated in the embedding space. With the learned feature embeddings, a mean shift clustering algorithm can be applied to discover and segment unseen objects. We further improve the segmentation accuracy with a new two-stage clustering algorithm. Our method demonstrates that non-photorealistic synthetic RGB and depth images can be used to learn feature embeddings that transfer well to real-world images for unseen object instance segmentation. arXiv, Talk video

License

Unseen Object Clustering is released under the NVIDIA Source Code License (refer to the LICENSE file for details).

Citation

If you find Unseen Object Clustering useful in your research, please consider citing:

@inproceedings{xiang2020learning,
    Author = {Yu Xiang and Christopher Xie and Arsalan Mousavian and Dieter Fox},
    Title = {Learning RGB-D Feature Embeddings for Unseen Object Instance Segmentation},
    booktitle = {Conference on Robot Learning (CoRL)},
    Year = {2020}
}

Required environment

Ubuntu 16.04 or above
PyTorch 0.4.1 or above
CUDA 9.1 or above

Installation

Install PyTorch.
Install python packages
```
pip install -r requirement.txt
```

Download

Download our trained checkpoints from here, save to $ROOT/data.

Running the demo

Download our trained checkpoints first.
Run the following script for testing on images under $ROOT/data/demo.
```
./experiments/scripts/demo_rgbd_add.sh
```

Training and testing on the Tabletop Object Dataset (TOD)

Download the Tabletop Object Dataset (TOD) from here (34G).
Create a symlink for the TOD dataset
```
cd $ROOT/data
ln -s $TOD_DATA tabletop
```

Training and testing on the TOD dataset

cd $ROOT

# multi-gpu training, we used 4 GPUs
./experiments/scripts/seg_resnet34_8s_embedding_cosine_rgbd_add_train_tabletop.sh

# testing, $GPU_ID can be 0, 1, etc.
./experiments/scripts/seg_resnet34_8s_embedding_cosine_rgbd_add_test_tabletop.sh $GPU_ID $EPOCH

Testing on the OCID dataset and the OSD dataset

Download the OCID dataset from here, and create a symbol link:
```
cd $ROOT/data
ln -s $OCID_dataset OCID
```
Download the OSD dataset from here, and create a symbol link:
```
cd $ROOT/data
ln -s $OSD_dataset OSD
```

Check scripts in experiments/scripts with name test_ocid or test_ocd. Make sure the path of the trained checkpoints exist.

experiments/scripts/seg_resnet34_8s_embedding_cosine_rgbd_add_test_ocid.sh
experiments/scripts/seg_resnet34_8s_embedding_cosine_rgbd_add_test_osd.sh

Running with ROS on a Realsense camera for real-world unseen object instance segmentation

Python2 is needed for ROS.

Make sure our pretrained checkpoints are downloaded.

# start realsense
roslaunch realsense2_camera rs_aligned_depth.launch tf_prefix:=measured/camera

# start rviz
rosrun rviz rviz -d ./ros/segmentation.rviz

# run segmentation, $GPU_ID can be 0, 1, etc.
./experiments/scripts/ros_seg_rgbd_add_test_segmentation_realsense.sh $GPU_ID

Our example:

Learning RGB-D Feature Embeddings for Unseen Object Instance Segmentation

Related tags

Overview

Unseen Object Clustering: Learning RGB-D Feature Embeddings for Unseen Object Instance Segmentation

Introduction

License

Citation

Required environment

Installation

Download

Running the demo

Training and testing on the Tabletop Object Dataset (TOD)

Testing on the OCID dataset and the OSD dataset

Running with ROS on a Realsense camera for real-world unseen object instance segmentation

Owner

NVIDIA Research Projects

Fast and customizable reconnaissance workflow tool based on simple YAML based DSL.

UnivNet: A Neural Vocoder with Multi-Resolution Spectrogram Discriminators for High-Fidelity Waveform Generation

Vignette is a face tracking software for characters using osu!framework.

Code for ICLR 2020 paper "VL-BERT: Pre-training of Generic Visual-Linguistic Representations".

[CVPR 2022] Official PyTorch Implementation for "Reference-based Video Super-Resolution Using Multi-Camera Video Triplets"

Rethinking the U-Net architecture for multimodal biomedical image segmentation

This program automatically runs Python code copied in clipboard

Code for: Imagine by Reasoning: A Reasoning-Based Implicit Semantic Data Augmentation for Long-Tailed Classification

[arXiv22] Disentangled Representation Learning for Text-Video Retrieval

Megaverse is a new 3D simulation platform for reinforcement learning and embodied AI research

Tensorflow implementation of "Learning Deep Features for Discriminative Localization"

Repo for CVPR2021 paper "QPIC: Query-Based Pairwise Human-Object Interaction Detection with Image-Wide Contextual Information"

TorchGeo is a PyTorch domain library, similar to torchvision, that provides datasets, transforms, samplers, and pre-trained models specific to geospatial data.

YOLOv5 detection interface - PyQt5 implementation

This repository is the official implementation of Using Time-Series Privileged Information for Provably Efficient Learning of Prediction Models

Intrusion Detection System using ensemble learning (machine learning)

This project aims to explore the deployment of Swin-Transformer based on TensorRT, including the test results of FP16 and INT8.

Retinal Vessel Segmentation with Pixel-wise Adaptive Filters (ISBI 2022)

Implementation of Rotary Embeddings, from the Roformer paper, in Pytorch

Research on controller area network Intrusion Detection Systems