Home · Open data · Perception · Human video

DexYCB

Roughly 5.4 hours (582K RGB-D frames) of 10 people grasping 20 YCB objects, 8 synchronised RealSense cameras with 3D hand and 6D object pose labels.

Published by NVIDIA Research, 2021-04

A person lifting a can with the hand pose overlaid
Open data

Frame from the DexYCB dataset, NVIDIA Research, CC BY-NC 4.0.

Non-commercial only

CC BY-NC 4.0

PriceFree

The publisher's licence restricts use to research and other non-commercial purposes.

Get from publisher Opens dex-ycb.github.io in a new tab

Public Google Drive links on the project page: one 119 GB archive (dex-ycb-20210415.tar.gz) or ten per-subject archives of about 12 GB each, plus bop.tar.gz (1.2 GB), calibration.tar.gz (16 KB) and models.tar.gz (1.4 GB). Large Drive files need gdown or a browser confirm step. Licence is CC BY-NC 4.0 (non-commercial).

Specification, as published

Hours (derived)
5.4 h
Episodes (stated)
1,000 episodes
Scale
582K RGB-D frames over 1,000 sequences of 10 subjects grasping 20 different objects from 8 views
Task
Pick & place, Other
Environment
Indoor – studio
Modality
RGB, Depth, Hand pose, Object pose
Capture
CCTV Scene
In frame
Person hands
Captured in
North America
Embodiment
human
Frame rate
30 fps
Resolution
640x480
Formats
Images (JPEG/PNG), NPZ/NumPy, Other
Size
119 GB (single archive)
Hosted on
Google Drive

Figures are the publisher's own. Kinetic Blocks has not measured this dataset. Derived from the paper: 582K frames at 30 fps, summed across eight cameras. Unique recording time is about 50 minutes (1,000 sequences of 3 seconds).

About this dataset

DexYCB is NVIDIA's multi-view benchmark for hand-object pose estimation during grasping. Ten subjects pick up 20 YCB objects from a table in a calibrated rig of eight RealSense D415 cameras, producing 1,000 sequences of colour and depth frames. Every frame carries MANO hand parameters, 21 3D joints, 2D keypoints, 6D object poses and segmentation, packaged as JPEG, PNG and NPZ files in per-subject archives.

Publisher's own task names: hand grasping of tabletop objects, human-to-robot object handover grasp generation, 2D object and keypoint detection, 6D object pose estimation, 3D hand pose estimation.

What Kinetic Blocks did, and did not do

Indexed and linked from the publisher. Not hosted, verified or graded by Kinetic Blocks. The publisher's terms govern. The licence shown is the one stated on the publisher's page, linked above. Nothing on this page is a claim by Kinetic Blocks.

Want the verified ones too?

Request access to the marketplace, where every paid listing is verified against its files and graded before it goes on sale.