Midcentury · datasets
Dataset documentation · Midcentury

Stereo Dataset Sample

A curated 1,000-clip egocentric manipulation dataset spanning 24 environments and three difficulty tiers, built for balanced coverage and near-zero idle time.

0%
Idle clips
Every clip stays active for its full duration
4.9%
Largest environment
4.9% share — no single environment dominates
22.8
Effective environments
22.8 of 24 — a near-flat spread

·Sample clips

Representative frames drawn from 20 clips in the delivered set — selected across the environment and difficulty spread detailed in Coverage below.

L
R
Food Processing

Shaping dough into flat discs and placing them on a baking tray

L
R
Healthcare

Sorting and packaging medicines according to a list

L
R
Repair Services

Repairing a motorcycle's front suspension

L
R
Creative Workshops

Painting a deity on paper

Example tasks in this set

Drawn directly from the delivered clip manifest — real category and instruction text, not illustrative examples.

CategoryDifficultyTask
AutomotiveHardRepairing car air conditioner components
HealthcareHardPreparing an injection from a vial
FashionMediumPreparing fabric pieces by aligning and folding them for stitching
CleaningEasyCleaning a hotel room
Creative WorkshopsHardPainting a deity on paper
Food and BeverageHardPreparing ingredients and cooking food in a commercial kitchen
AgricultureEasyWatering plants in a garden with a hose
Industrial ManufacturingHardSetting up and preparing an edge banding machine for operation

01Overview

1,000
Clips
64.9h
Total duration
180s
Median clip length
1.58 TiB
Total size
MCAP
Format — H.264 + IMU

Stereo egocentric capture — synchronized left/right cameras with calibration and a fixed stereo baseline, plus head-mounted IMU. Each clip ships with an .annotations.json sidecar carrying category, skill group, task, difficulty, operator, and subtask timings — see Files & annotations for the full schema.

02Coverage — environment × difficulty

EnvironmentEasyMediumHardTotalDistribution
Automotive274049
Beauty and Personal Care12112649
Construction454049
Creative Workshops154349
Factory11102849
Fashion3182849
Food and Beverage1426949
Other non-residential19141649
Repair Services2113649
Agriculture2424048
Healthcare1722948
Industrial Manufacturing12211548
Cleaning2124045
Administrative Work1726043
Food Processing1823041
Home1822040
Hospitality1920039
Metalworking6231039
Retail1424038
Household Tasks1816034
Sports and Recreation1519034
Entertainment1911030
Printing and Design116017
Laboratory Work312015
Total3004003001,000

All 24 environments in the corpus are represented. No environment exceeds a 5% share — the red tick marks where that ceiling sits.

Screening methodology

Category and difficulty balance are enforced by a set of hard caps applied during selection, not left to whatever the source corpus happened to contain:

  • No single environment exceeds 20% of the set; the top 5 environments combined stay under 60%, and the top 10 under 85%.
  • At least 15 distinct environments must be represented.
  • Easy-difficulty clips are capped at 30% of the set; Hard-difficulty clips are held to a minimum of 5%.
  • Individual tasks are capped per difficulty tier (6% of Hard, 4% of Medium, 1.5% of Easy) so no single repeated task dominates a tier.
  • Factory-setting footage specifically is capped at 20% of the set, with any overflow re-classified into Industrial Manufacturing, Creative Workshops, Metalworking, or Printing and Design rather than simply discarded.
  • Per-(operator, task, environment) hour caps limit how much any single contributor or repeated setup can contribute.

03Diversity

Reported as effective count (inverse Simpson, 1 / Σpᵢ²). It answers "how many groups does this dataset effectively contain," which a raw distinct count overstates whenever the distribution is lopsided.

Environments
22.8 / 24 distinct
Tasks
103.7 / 413 distinct
Workers
412.9 / 630 distinct
Venues
32.2 / 121 distinct
Sessions
132.0 / 334 distinct

Effective environments of 22.8 against 24 distinct means the distribution is close to flat — a deliberate design outcome. Tasks and venues are intentionally long-tailed by comparison — see 3b.

3a. Skill groups

291 distinct skill groups, effective count 40.7. The largest group is Tool Use at 8.9% of clips.

Skill groupClipsShareDistribution
Tool Use898.9%
Assembly585.8%
Cooking494.9%
Cleaning484.8%
Textile Production474.7%
Packaging393.9%
Metalworking313.1%
Repair242.4%
Food Preparation222.2%
Sorting and Packaging202.0%
Garment Care131.3%
Material Handling121.2%

3b. Tasks

413
Distinct tasks
103.7
Effective count
55
Largest task (5.5%, metalworking)
282
Tasks with a single clip

The distribution is deliberately long-tailed: the top 10 tasks account for 22% of the dataset and the top 50 for 52%, with 282 tasks appearing exactly once. Median clips per task: 1.

TaskEasyMediumHardTotalDistribution
metalworking0163955
assembling_electronic_components3111731
weaving_fabric_on_handloom002424
packaging_products194023
haircut_service101819
playing_board_game151016
preparing_flatbread212014
sorting_and_weighing_food_items85013
assembling_product_components112013
sorting_and_packaging_electronic_components57012
ironing_clothes111012
packaging_items_for_shipment111012
preparing_food_items29112
cooking_food010212
repairing_electronic_device001111
sorting_and_packaging_items29011
cleaning_hotel_room19010
processing_customer_order1809
managing_sales_and_inventory1809
automotive_bodywork0189
cleaning_kitchen8109
sewing_garments0729
packaging_food_products4509
using_smartphone8008
folding_clothes3508

The red tick marks a 5% / 50-clip ceiling. Only metalworking sits above it.

Cumulative concentration

22.0%
Top 10 tasks
220 clips
37.0%
Top 25 tasks
370 clips
51.8%
Top 50 tasks
518 clips
65.6%
Top 100 tasks
656 clips

04Quality

Every clip in the delivered set passes privacy, usability, orientation, and engagement screening, plus structural validation of its stereo rig data, before inclusion.

1
Privacy screeningSampled frames are run through face detection. A clip is rejected if faces dominate 50% or more of sampled frames, or if identifiable-person risk reaches 25% of frames.
2
CV usabilitySampled frames are checked for brightness, sharpness, and motion — catches unusable footage before it reaches annotation.
3
Orientation checkIMU data confirms the camera rig was upright and correctly oriented for the full clip — not upside-down or sideways.
4
Engagement screeningA model rubric flags standing idle, watching, waiting, or hands-not-engaged footage as unusable — scene motion alone doesn't count as active work.
5
Stereo rig validationEach clip must carry valid left and right camera calibration, an intact stereo baseline transform, and a valid subtask annotation before it's structurally admitted.

Known limitation: subtask timestamps are model-estimated and carry roughly ±5–10s of error on a 150s clip — treat subtask boundaries as approximate, not frame-exact.

05Files & annotations

Each clip is delivered as an MCAP file alongside a JSON annotation sidecar:

// per-clip MCAP topics
/top-left-camera/image-raw  // H.264 video, left camera
/top-right-camera/image-raw  // H.264 video, right camera
/top-left-camera/camera-info  // left intrinsics/calibration
/top-right-camera/camera-info  // right intrinsics/calibration
/tf-static  // fixed stereo baseline transform
/top-camera/imu  // head-mounted IMU
/subtask-annotation  // subtask timeline channel

The .annotations.json sidecar carries:

FieldDescription
category / skill_group / task_idWhere this clip sits in the taxonomy
difficulty_levelEasy / Medium / Hard
operator_idDe-identified operator reference, used for the Workers diversity metric
clip_usable / pii_riskResult of the privacy and engagement screens in Quality
camera_alignment / two_hands_in_frameFraming metadata for the manipulation task
subtasks[]Ordered {time_s, description} array — the subtask timeline (see Quality for its accuracy caveat)
clip_start_ns / clip_end_nsSource-recording timestamps this clip was cut from

06Access

This page is a live documentation view of the delivered sample — the frames in Sample clips and the tables above are drawn directly from the real 1,000-clip set.