A curated 1,000-clip egocentric manipulation dataset spanning 24 environments and three difficulty tiers, built for balanced coverage and near-zero idle time.
Representative frames drawn from 20 clips in the delivered set — selected across the environment and difficulty spread detailed in Coverage below.
Shaping dough into flat discs and placing them on a baking tray
Sorting and packaging medicines according to a list
Repairing a motorcycle's front suspension
Painting a deity on paper
Drawn directly from the delivered clip manifest — real category and instruction text, not illustrative examples.
| Category | Difficulty | Task |
|---|---|---|
| Automotive | Hard | Repairing car air conditioner components |
| Healthcare | Hard | Preparing an injection from a vial |
| Fashion | Medium | Preparing fabric pieces by aligning and folding them for stitching |
| Cleaning | Easy | Cleaning a hotel room |
| Creative Workshops | Hard | Painting a deity on paper |
| Food and Beverage | Hard | Preparing ingredients and cooking food in a commercial kitchen |
| Agriculture | Easy | Watering plants in a garden with a hose |
| Industrial Manufacturing | Hard | Setting up and preparing an edge banding machine for operation |
Stereo egocentric capture — synchronized left/right cameras with calibration and a fixed stereo baseline, plus head-mounted IMU. Each clip ships with an .annotations.json sidecar carrying category, skill group, task, difficulty, operator, and subtask timings — see Files & annotations for the full schema.
| Environment | Easy | Medium | Hard | Total | Distribution |
|---|---|---|---|---|---|
| Automotive | 2 | 7 | 40 | 49 | |
| Beauty and Personal Care | 12 | 11 | 26 | 49 | |
| Construction | 4 | 5 | 40 | 49 | |
| Creative Workshops | 1 | 5 | 43 | 49 | |
| Factory | 11 | 10 | 28 | 49 | |
| Fashion | 3 | 18 | 28 | 49 | |
| Food and Beverage | 14 | 26 | 9 | 49 | |
| Other non-residential | 19 | 14 | 16 | 49 | |
| Repair Services | 2 | 11 | 36 | 49 | |
| Agriculture | 24 | 24 | 0 | 48 | |
| Healthcare | 17 | 22 | 9 | 48 | |
| Industrial Manufacturing | 12 | 21 | 15 | 48 | |
| Cleaning | 21 | 24 | 0 | 45 | |
| Administrative Work | 17 | 26 | 0 | 43 | |
| Food Processing | 18 | 23 | 0 | 41 | |
| Home | 18 | 22 | 0 | 40 | |
| Hospitality | 19 | 20 | 0 | 39 | |
| Metalworking | 6 | 23 | 10 | 39 | |
| Retail | 14 | 24 | 0 | 38 | |
| Household Tasks | 18 | 16 | 0 | 34 | |
| Sports and Recreation | 15 | 19 | 0 | 34 | |
| Entertainment | 19 | 11 | 0 | 30 | |
| Printing and Design | 11 | 6 | 0 | 17 | |
| Laboratory Work | 3 | 12 | 0 | 15 | |
| Total | 300 | 400 | 300 | 1,000 |
All 24 environments in the corpus are represented. No environment exceeds a 5% share — the red tick marks where that ceiling sits.
Category and difficulty balance are enforced by a set of hard caps applied during selection, not left to whatever the source corpus happened to contain:
Reported as effective count (inverse Simpson, 1 / Σpᵢ²). It answers "how many groups does this dataset effectively contain," which a raw distinct count overstates whenever the distribution is lopsided.
Effective environments of 22.8 against 24 distinct means the distribution is close to flat — a deliberate design outcome. Tasks and venues are intentionally long-tailed by comparison — see 3b.
291 distinct skill groups, effective count 40.7. The largest group is Tool Use at 8.9% of clips.
| Skill group | Clips | Share | Distribution |
|---|---|---|---|
| Tool Use | 89 | 8.9% | |
| Assembly | 58 | 5.8% | |
| Cooking | 49 | 4.9% | |
| Cleaning | 48 | 4.8% | |
| Textile Production | 47 | 4.7% | |
| Packaging | 39 | 3.9% | |
| Metalworking | 31 | 3.1% | |
| Repair | 24 | 2.4% | |
| Food Preparation | 22 | 2.2% | |
| Sorting and Packaging | 20 | 2.0% | |
| Garment Care | 13 | 1.3% | |
| Material Handling | 12 | 1.2% |
The distribution is deliberately long-tailed: the top 10 tasks account for 22% of the dataset and the top 50 for 52%, with 282 tasks appearing exactly once. Median clips per task: 1.
| Task | Easy | Medium | Hard | Total | Distribution |
|---|---|---|---|---|---|
| metalworking | 0 | 16 | 39 | 55 | |
| assembling_electronic_components | 3 | 11 | 17 | 31 | |
| weaving_fabric_on_handloom | 0 | 0 | 24 | 24 | |
| packaging_products | 19 | 4 | 0 | 23 | |
| haircut_service | 1 | 0 | 18 | 19 | |
| playing_board_game | 15 | 1 | 0 | 16 | |
| preparing_flatbread | 2 | 12 | 0 | 14 | |
| sorting_and_weighing_food_items | 8 | 5 | 0 | 13 | |
| assembling_product_components | 1 | 12 | 0 | 13 | |
| sorting_and_packaging_electronic_components | 5 | 7 | 0 | 12 | |
| ironing_clothes | 1 | 11 | 0 | 12 | |
| packaging_items_for_shipment | 11 | 1 | 0 | 12 | |
| preparing_food_items | 2 | 9 | 1 | 12 | |
| cooking_food | 0 | 10 | 2 | 12 | |
| repairing_electronic_device | 0 | 0 | 11 | 11 | |
| sorting_and_packaging_items | 2 | 9 | 0 | 11 | |
| cleaning_hotel_room | 1 | 9 | 0 | 10 | |
| processing_customer_order | 1 | 8 | 0 | 9 | |
| managing_sales_and_inventory | 1 | 8 | 0 | 9 | |
| automotive_bodywork | 0 | 1 | 8 | 9 | |
| cleaning_kitchen | 8 | 1 | 0 | 9 | |
| sewing_garments | 0 | 7 | 2 | 9 | |
| packaging_food_products | 4 | 5 | 0 | 9 | |
| using_smartphone | 8 | 0 | 0 | 8 | |
| folding_clothes | 3 | 5 | 0 | 8 |
The red tick marks a 5% / 50-clip ceiling. Only metalworking sits above it.
Every clip in the delivered set passes privacy, usability, orientation, and engagement screening, plus structural validation of its stereo rig data, before inclusion.
Known limitation: subtask timestamps are model-estimated and carry roughly ±5–10s of error on a 150s clip — treat subtask boundaries as approximate, not frame-exact.
Each clip is delivered as an MCAP file alongside a JSON annotation sidecar:
The .annotations.json sidecar carries:
| Field | Description |
|---|---|
| category / skill_group / task_id | Where this clip sits in the taxonomy |
| difficulty_level | Easy / Medium / Hard |
| operator_id | De-identified operator reference, used for the Workers diversity metric |
| clip_usable / pii_risk | Result of the privacy and engagement screens in Quality |
| camera_alignment / two_hands_in_frame | Framing metadata for the manipulation task |
| subtasks[] | Ordered {time_s, description} array — the subtask timeline (see Quality for its accuracy caveat) |
| clip_start_ns / clip_end_ns | Source-recording timestamps this clip was cut from |
This page is a live documentation view of the delivered sample — the frames in Sample clips and the tables above are drawn directly from the real 1,000-clip set.