Home · Open data · Perception · Human video

Assembly101

513 hours of twelve-view recordings of 53 people assembling and taking apart 101 toy vehicles, with fine and coarse action labels, mistakes and 3D hand poses.

Published by Meta Reality Labs and NUS, 2022-03

Twelve synchronised views of a person assembling a toy vehicle
Open data

Frame from the Assembly101 dataset, Meta Reality Labs and NUS, CC BY-NC 4.0.

Non-commercial only

CC BY-NC 4.0

PriceFree

The publisher's licence restricts use to research and other non-commercial purposes.

Get from publisher Opens huggingface.co in a new tab

The host asks you to sign in to Hugging Face and accept the publisher's terms before it serves files.

From a terminal
huggingface-cli download --repo-type dataset cvml-nus/assembly101

Hugging Face repo is gated 'auto': log in, agree to share contact information, then use a token. 3.89 TB total; views and feature packs can be pulled individually with --include. Older Google Drive route needs a per-user access request and download.py with --views fixed|egocentric. CC BY-NC 4.0, no commercial use.

Specification, as published

Hours (stated)
513 h
Episodes (stated)
362 episodes
Scale
4321 videos of people assembling and disassembling 101 take-apart toy vehicles
Task
Assembly
Environment
Indoor – lab
Modality
RGB, Hand pose, Language
Capture
Egocentric
In frame
Person hands
Captured in
Multiple regions
Embodiment
human
Resolution
1920x1080 (static), 640x480 (egocentric monochrome)
Formats
MP4, Other
Size
3.89 TB
Hosted on
Hugging Face

Figures are the publisher's own. Kinetic Blocks has not measured this dataset.

About this dataset

Assembly101 records people building and dismantling 101 take-apart toy vehicles at a bench inside a capture rig with 8 fixed RGB cameras and 4 monochrome egocentric cameras on a custom headset. The 4,321 videos (513 hours) carry over 100K coarse and about 1M fine-grained action segments, mistake labels and 18M 3D hand poses. Benchmarks cover action recognition, anticipation, temporal segmentation and mistake detection. Licensed CC BY-NC 4.0 and hosted on Hugging Face behind a contact-information gate.

Publisher's own task names: assembly, disassembly, action recognition, action anticipation, temporal action segmentation, mistake detection, 3D hand pose estimation.

What Kinetic Blocks did, and did not do

Indexed and linked from the publisher. Not hosted, verified or graded by Kinetic Blocks. The publisher's terms govern. The licence shown is the one stated on the publisher's page, linked above. Nothing on this page is a claim by Kinetic Blocks.

Want the verified ones too?

Request access to the marketplace, where every paid listing is verified against its files and graded before it goes on sale.