Household Robot Training Data: Collecting Task Video in Real Homes (2026)

Robots trained in lab kitchens tend to struggle in real ones. How to run household task video collection in actual homes, from camera views and task taxonomy to recruiting households, handling privacy and delivering episodes for imitation learning.

11 min read
Woman washing dishes at a kitchen sink, the kind of everyday task captured as household robot training data

Why household robot training data has to come from real homes

Household robot training data is video and sensor data of people, and sometimes robots, doing domestic work: cooking, washing up, folding laundry, putting things away. Home robot programs need a lot of it, because domestic tasks run long and no two households keep the same objects in the same places.

Most early programs collect in a lab kitchen or a mock apartment. That is fine for getting a pipeline running and poor for generalization. A lab kitchen has one layout and one set of appliances. The lighting is good and the counter is clean, so the policy learns that cup and that drawer handle under that light. Put it in a real apartment with a different tap and a cluttered counter in afternoon glare, and success rates fall. DROID made the case for manipulation by collecting 76,000 trajectories (350 hours) across 564 scenes with 50 collectors in North America, Asia and Europe, instead of in a single lab.

A household task video program should aim for breadth of real homes, with enough consistency in capture that the data still trains well. Getting there is mostly an operations problem. The homes have to be found, and everyone in them has to consent before a camera goes in. Afterwards, hour-long sessions have to be cut into clean episodes.

What household task video programs capture: views and sensors

Household programs usually record the three viewpoints listed below, and each one feeds a different part of the model. Public datasets make it clear why one view is rarely enough.

EPIC-KITCHENS-100 shows what head-mounted capture alone can produce: 100 hours of unscripted kitchen activity from 45 kitchens, annotated into roughly 90,000 action segments. It is the best-known kitchen video dataset and a strong perception benchmark, but it was not designed as a kitchen video dataset for robots: policy learning usually needs wrist views, synchronized multi-camera geometry and, when a robot is in the loop, action streams. Ego4D, with more than 3,670 hours from 923 participants in 74 locations across nine countries, is the reference point for how far egocentric collection can scale geographically.

  • Egocentric (head-mounted)
    • -Head-mounted rigs, Meta Aria glasses or other smart glasses, recorded at up to 4K and 60fps
    • -Shows where the person looks and what the hands do next, the closest human analogue to a humanoid head camera
    • -Carries the most privacy exposure, since it sees everything the participant sees
  • Wrist-mounted
    • -Small cameras on the wrist or a handheld gripper, close to the contact point
    • -Matches the wrist cameras on most manipulation robots, so it transfers well to policy training
    • -Needs careful mounting to limit hand occlusion and motion blur
  • Third-person
    • -One or more fixed cameras covering the work area, time-synchronized with the wearable views
    • -Provides body pose, full-scene context and multi-angle coverage for 3D pose estimation
    • -Useful for annotation and QA even when the model does not train on it directly

Building a household task taxonomy

A taxonomy turns "household tasks" into something a vendor can schedule and a model team can balance. BEHAVIOR-1K, from Stanford, built its benchmark of 1,000 household activities from surveys asking people what they want robots to do for them. It makes a good long list to start from. A real-world collection program needs fewer, better-specified tasks, each broken down into sub-tasks with clear start and end states.

Three levels are usually enough. At the top is the task family (cooking, dishwashing, laundry, tidying, cleaning). Each family breaks into tasks such as loading a dish rack, and each task into sub-tasks or primitives: open cabinet, pick plate, place plate. Annotation and episode segmentation hang on the primitives. Difficulty varies sharply by family, which should shape how hours are allocated.

Household task categories compared: views, sensors and collection difficulty

Task categoryPrimary viewsSensors beyond RGBDifficultyWhat makes it hard
Cooking and food preparationEgocentric, wrist and overhead third-personDepth for pouring and cutting geometry; audio optionalHighLong horizons, liquids and deformable food, heat and knife safety, strong regional variation in dishes and tools
Dishwashing and kitchen cleanupEgocentric and wristDepth; splash protection for cameras near the sinkMedium-highWater, foam and reflective surfaces degrade vision; sink and rack layouts vary widely
Laundry: sorting, loading, foldingThird-person and wrist; egocentric for foldingDepth for cloth geometryHighCloth has near-endless states; front-load and top-load machines, drying racks and dryers all differ
Tidying and object rearrangementEgocentric and third-personDepth; optional room-scale multi-cameraMediumWhere things belong differs by household, so it needs many homes more than many hours
Surface wiping and floor careEgocentric and wristIMU or force/torque where contact pressure mattersLow-mediumCoverage and pressure are hard to judge from video, so success criteria must be defined up front

Step 1: Define the home diversity matrix

Generalization tracks the number of distinct homes far more closely than the hours recorded inside each one. Fifty hours from five homes teaches a model less than fifty hours from fifty homes. Before recruiting, define the variables the dataset must span and a target count for each.

Cultural variation is a real axis in APAC. Vietnamese home kitchens often centre on a gas hob and a rice cooker, with prep at floor level or on a low counter. A Thai kitchen may rely on a mortar and pestle and a wok. Singapore apartments tend to have a compact galley with an induction hob and the washing machine close by. Dishwashers are much less common in many Southeast Asian homes than in North America or Western Europe; handwashing and drying racks dominate instead. A home robot trained on Western kitchen data will first meet chopsticks, woks and wet-market produce in deployment unless the dataset includes them.

  • Home type and layout: high-rise apartment, shophouse, landed house, galley or open-plan kitchen, counter height
  • Appliances: gas or induction hobs, rice cookers, top-load or front-load washers, dishwasher present or not, drying racks or dryers
  • Clutter level: tidy, typical or cluttered, rated on a simple scale at the start of each session
  • Lighting: window orientation, time of day, fluorescent or warm LED
  • Household composition: children, elderly members and pets, which change object sets and safety constraints

Step 2: Recruit homes and participants

Recruiting homes is harder than recruiting people. The whole household has to accept cameras in a private space and a visit schedule, plus rules for other members who may appear or need to stay out of frame. Participants need to be comfortable wearing a head-mounted rig for sessions of 30-90 minutes and following a light protocol without performing for the camera.

Managed programs recruit through existing participant networks and screen households against the diversity matrix. Onboarding is short: a rig fitting and a practice session, plus a briefing on what to do when someone who has not consented walks in. Setup speed also matters, since a field team that needs hours to calibrate cameras in each home caps the program at a handful of homes per week.

Two collection methods apply. Onsite collection sends a trained field team into each home with the full multi-view rig, which suits programs that need synchronized third-person cameras and calibrated geometry. Crowdsourced collection ships simpler wearable kits to vetted participants who record in their own homes to a protocol, which grows home count faster at the cost of fewer views. Many programs use the first for a core dataset and the second for breadth.

Step 3: Set consent and privacy rules for private homes

A private home is the most sensitive place a collection program can operate. Egocentric cameras capture whatever the participant looks at: mail, screens, family photos. They also capture other household members, children included. Consent has to be layered: the participant, every adult in the home who may appear, and guardians for any minors, with a clear rule for what happens when someone who has not consented enters the frame.

Before recording, a room sweep removes or covers documents and screens, and the rig has a pause control. Faces and visible text are blurred automatically before data leaves the collection environment. Storage is access-controlled, with written retention and deletion terms. For collection in Singapore, Malaysia and Thailand, the local PDPA rules set the baseline, and programs serving EU clients typically add GDPR-grade transfer and processing terms. All of this belongs in the contract and the capture protocol from the start.

Step 4: Script sessions and segment them into episodes

Household sessions sit between fully scripted and fully natural. A fully scripted session (pick up the red cup, place it in the sink) gives clean episodes but unnatural behaviour. Leave people entirely to themselves and the behaviour is realistic, but the footage is hard to segment and balance. Most programs give participants a task list with a goal and success condition for each item, such as "wash and rack the lunch dishes", and let them work their own way.

Segmentation turns a 60-minute session into training units. Each episode needs start and end timestamps, a task and sub-task label, a natural-language instruction, a success or failure flag and, where useful, a note on recovery behaviour. For VLA training the language instruction is the conditioning signal, so it cannot be an afterthought. Failed and recovered attempts are worth keeping and labelling, because policies trained only on clean successes handle mistakes badly.

QA at this stage checks that views stay synchronized across each episode and that frames stay sharp during fast hand motion. Task coverage per home is checked against the matrix too.

Step 5: Deliver for imitation learning and VLA pipelines

Fix the delivery format during the pilot. Imitation learning teams commonly want episodes in HDF5, RLDS or LeRobot format with synchronized camera streams, timestamps, calibration files and a language annotation per episode. Open X-Embodiment, which pooled 60 existing datasets from 34 labs into more than one million trajectories across 22 robot embodiments, is a good argument for standardized formats: data built to a common schema can be combined and reused across robots and models.

Human video without robot actions is still useful. Teams pre-train visual representations on it and use it to learn task structure and sub-goal order. It also supplies hand trajectories for retargeting. If the robot has to learn actions directly, go back to the same homes with the same task list and record teleoperated or handheld-gripper demonstrations. Human and robot data then share one taxonomy.

DataX Power runs household programs with the capture modes described above: egocentric rigs including Meta Aria and smart glasses at up to 4K/60fps, third-person arrays with synchronized timestamps for multi-angle 3D pose, and wrist-mounted cameras synchronized with force/torque logging where a task needs it. Collection runs onsite or crowdsourced across participant networks in Vietnam, Thailand, Singapore and Malaysia. Pilot specifications are usually ready within five business days, and programs scale from 100-hour pilots to 50,000-hour production on one contract.

DataX Power runs managed household task video programs for home robot and humanoid teams. We record in real homes across four APAC markets with egocentric, wrist and third-person views, handle consent end to end, and deliver in your training format.

Explore humanoid and home robot data collection

Frequently asked questions about household robot training data

Home robot and humanoid teams tend to ask these when scoping a household collection program.

How many homes does a household robot training data program need?
Home count matters more than hours per home. As a practical estimate, a pilot can show useful signal with 10-20 homes, while a production dataset meant to generalize across a market typically needs 100 or more distinct homes, balanced across layout, appliance and clutter categories. The right number depends on how many task families you cover and how varied your target deployment homes are.
Can we use EPIC-KITCHENS or Ego4D instead of collecting our own home robot dataset?
They are useful for pre-training perception and learning task structure, but they were built for research on human activity rather than for any specific robot. They lack wrist views matched to your hardware and calibrated multi-camera geometry for your setup, and in most cases the task balance and language annotation your policy needs. Check licence terms before any commercial use. Most home robot teams use public data for pre-training and collect their own data for fine-tuning.
Should household task video be recorded in real homes or in a staged apartment?
Both have a role. A staged apartment is useful for early pipeline testing and for tasks that need heavy instrumentation. Real homes supply the diversity that generalization depends on: different layouts and appliances, and real clutter. A common pattern is a small staged set for protocol development, followed by most of the hours in real homes recruited against a diversity matrix.
How is privacy handled when filming inside private homes?
Through layered consent plus technical controls. Every adult who may appear gives consent, and guardians consent for minors. Participants can pause recording at any time, and rooms are prepared before each session to remove documents and screens. Faces and visible text are blurred before delivery, while the contract covers storage, retention and deletion. Local PDPA rules apply in Singapore, Malaysia and Thailand, with GDPR controls added for EU-linked programs.
What does household task video data collection cost?
As a DataX Power practical estimate consistent with our published ranges, egocentric-only household capture typically falls around $80-$200 per captured hour. Programs that add synchronized wrist and third-person cameras or force/torque logging move toward the $200-$450 multi-sensor range. The main additional cost drivers are recruiting across many households and cooking consumables, plus annotation for episode segmentation and language labels.
Data Collection Service

Need the platform layer to make this stick in production? Our Hanoi-based infrastructure team delivers DevOps, FinOps, SecOps, and AI/MLOps for enterprises on AWS, GCP, Azure, and on-premise.

携手打造 下一个里程碑

告诉我们您的挑战 – AI、数据或基础设施。我们将为项目梳理范围,并为您配置合适的团队。