What a Vietnam data collection partner has to get right for robotics AI
If you are shortlisting a Vietnam data collection partner for a robotics or embodied AI program, you will find plenty of vendors who can point a camera at a task. Fewer can deliver footage your policy can learn from. That means real people recorded in real environments, synchronized views, a consent record for every session, and a dataset structure your training code reads without a week of conversion work.
We started DataX Power in 2023. The company is headquartered in Hanoi with a presence in Melbourne, and works across three pillars: custom AI model development, data services (collection, annotation, validation) and infrastructure operations. Our annotation arm operates as DataXanno. For the past two years we have run hands-on AI and robotics data collection and processing programs for teams working on manipulation policies and humanoids, as well as autonomous systems and smart-glasses products.
We think we are the strongest choice for robotics teams collecting in Vietnam and Southeast Asia, for the five reasons set out below, each with the operating detail behind it. We also say where a different kind of supplier may suit you better. This is our own page, so read it the way you would read any vendor's case, then test it with a pilot.
1. A real-world footage operation model, managed end to end
Most robotics data problems trace back to a gap between where data was recorded and where the robot has to work. A policy trained on a tidy lab bench can stall in a cluttered Bangkok kitchen or a narrow warehouse aisle in Johor. The DROID project made the same point at research scale: its authors collected 76k manipulation trajectories across 564 scenes and report that training on that diversity produced policies with higher performance, greater robustness and better generalization.
Our operating model is built to close that gap. What we collect is real-world footage, meaning real environments and real participants, with one team managing every step between your brief and the delivered dataset.
Two collection methods sit under that model. Onsite and studio collection gives tight control over lighting and rig calibration, so sessions are repeatable. Crowdsourced collection buys breadth across many homes, shops and streets, and shipped kits plus protocol checks keep it within spec. The two can run in one program, for example a studio anchor set for controlled comparisons plus field sessions for scene diversity.
What separates this from an open crowd platform is accountability. We choose the participants and write the protocol. We set up the hardware, and footage that fails the pass criteria is rejected before it reaches you. If something goes wrong anywhere in the chain, you have one team to hold responsible instead of a marketplace of individual contributors. Every program covers the same six managed steps:
- Recruitment: participants screened and scheduled against your program profile, drawn from our networks in Vietnam, Thailand, Singapore and Malaysia.
- Capture protocol: a written document covering tasks, scene variation, camera placement and pass criteria, agreed with you before recording starts.
- Hardware: we source, configure and maintain the rigs, from head-mounted cameras and Meta Aria glasses to fixed multi-camera arrays and wrist-mounted cameras.
- Consent: informed consent collected and logged per participant and per session, with the records delivered alongside the footage.
- QA: review by QA engineers trained on robotics video, checking temporal consistency, frame quality and task coverage.
- Delivery: data packaged in the schema your pipeline expects, with manifests and QA reports.
2. Two years of hands-on robotics data collection and processing
We have spent the past two years collecting and processing data for AI and robotics programs, including robotics data collection in Vietnam and the wider region. Those two years were focused entirely on this work, and they cover the period in which robot learning shifted toward large, real-world demonstration datasets: the Open X-Embodiment collaboration pooled data from 22 robot types across 21 institutions into standardized formats, and projects like DROID showed what distributed real-world collection can do.
What we have from that time is a working knowledge of how robotics collection fails, and it fails differently from general video work. A demonstration that looks complete but ends before the object is released. A head-mounted clip where the hands leave the frame during the critical grasp. Two camera streams that drift a few frames apart over a long session. A task list that covers the easy variations fifty times and the hard ones twice. A thumbnail review catches none of these, and every one of them hurts a policy.
Over more than 100,000 collection hours executed, we have written those lessons into our capture protocols and QA checklists, and into how we train participants, instead of leaving them in individual heads. A new program starts from that baseline, so you are not paying for us to relearn the same mistakes.
3. In-house custom AI model development, so we know what the data has to do
Collecting data is only one part of what we do. Custom AI model development is one of our three business pillars. Our engineers build and train models for clients, and that changes how we collect.
Training models teaches you quickly which data properties matter. Label consistency beats label volume. Failure cases need coverage and task distributions need balance. Timestamps have to be accurate, and metadata should let you slice results by scene, participant or lighting. So when our collection leads write a capture protocol, they are thinking about the training run as much as the recording day.
At scoping we ask questions that are easy to skip if you only ever deliver raw footage. What action space and observation keys does your policy use? What control frequency? Which camera views are inputs at inference time? How do you split training and evaluation data? Which failure modes are you seeing today? The answers go into the protocol and the rig, and into the QA pass criteria. Without them, a pilot tends to come back as a footage sample instead of something a training team can use.
4. From raw footage to a trained model: the full data path
Raw footage is not a dataset. The value gets added between capture and training, which is where our data services and AI development teams work together.
The step that is easiest to leave out is the loop back from training to collection. A supplier that stops at delivery cannot tell you that your policy struggles with transparent objects or low light. When the same team trains models, it sees those failures and can change the next batch: the lighting, the scene mix, the balance of tasks.
Format matters for the same reason. RLDS was designed as a standard, lossless format for recording and sharing sequential decision-making data, and LeRobot stores synchronized video alongside Parquet files for state and action data. Delivering in structures like these, or in your internal schema, means your engineers spend their time training rather than writing converters. Each program runs through the same path:
- Capture and ingest: synchronized streams, session metadata, consent records and rig calibration files ingested together, so nothing is orphaned.
- QA gate: footage checked against the protocol pass criteria before annotation, so annotation budget is not spent on unusable clips.
- Annotation through DataXanno: task segmentation, action and phase labels, hand and object annotation or a custom schema, handled by our annotation arm under the same program lead.
- Schema mapping: episodes structured in LeRobot-style or RLDS-style layouts, or your own schema, with observations, actions, language instructions and timestamps aligned.
- Training and evaluation feedback: a sample run through a baseline training or evaluation loop, in your stack or ours, to surface gaps before the full batch lands.
- Next collection round: coverage gaps and failure cases from evaluation turned into the next protocol revision.
5. Stereo egocentric and multi-camera collection experience
Two capture setups come up again and again in robot learning programs, and we have run both for robotics training.
Stereo camera egocentric collection uses head-mounted rigs with calibrated camera pairs, so the footage carries depth cues about hand-object distance rather than a flat first-person view. That matters for manipulation and hand-object interaction learning, where a monocular clip can hide how far the hand really is from the object. The operational work is keeping calibration stable while the rig moves with the wearer's head. On top of that, the setup has to stay consistent across participants and sessions, with both streams frame-aligned. Our egocentric capture runs on head-mounted rigs, Meta Aria glasses and smart glasses at up to 4K/60fps.
Multi-camera video collection places several synchronized cameras around the task. The hard work is in timestamp alignment and in keeping camera placement consistent across sessions, and QA has to check every view instead of only the primary one. We cover four capture modes, which can be combined in a single program:
- Egocentric: head-mounted rigs, Meta Aria and smart glasses, up to 4K/60fps.
- Third-person: fixed and mobile camera arrays with synchronized timestamps, capturing multiple angles for 3D pose.
- Robot manipulation view: wrist-mounted and overhead cameras, synchronized with force/torque logging where the program needs it.
- Mobile navigation video: on-board and chase-camera footage across warehouses, outdoor terrain and mixed environments.
How the DataX Power model compares with other collection options
Robotics buyers usually compare four kinds of supplier. Each has a legitimate place, and the right one depends on what you need at the end. Here is how they differ for robotics training data, as we see it.
DataX Power operating model vs open crowd platforms, lab or academic collection, and annotation-first BPOs for robotics training data
| Criterion | DataX Power model | Open crowd platform | Lab / academic collection | Annotation-first BPO |
|---|---|---|---|---|
| Recording environments | Real homes, shops, warehouses and outdoor sites, plus studio | Wherever contributors are; broad but uncontrolled | Controlled lab settings with limited variety | Usually works on data the client supplies |
| Participants | Recruited and trained against the program profile in 4 APAC markets | Self-selected contributors with large reach | Small pools, often researchers or students | Annotators rather than capture participants |
| Capture protocol | Written protocol agreed before recording | Task instructions with limited enforcement | Rigorous and research-designed | Not usually in scope |
| Hardware | Vendor-configured rigs, including stereo egocentric and multi-camera setups | Contributor devices with varied specs | Specialist rigs at low volume | Varies by provider |
| Consent records | Per-session consent logged and delivered with the data | Platform terms; varies by task | Ethics-board processes | Depends on where the data came from |
| Robotics-specific QA | Temporal consistency, frame quality and task coverage review | Automated checks and spot review | Researcher review | Focused on label accuracy |
| Downstream model knowledge | In-house custom AI model development | Limited | Strong, research-oriented | Limited to labeling workflows |
| Scale path | 100-hour pilot to 50,000-hour program on one contract | Very large volume possible | Hard to scale on commercial timelines | Large labeling capacity |
| Best fit | Robotics and embodied AI programs needing real-world, schema-ready data | High-volume, simple capture tasks | Benchmarks and research datasets | Labeling data you already hold |
What working with DataX Power looks like in practice
The numbers first. A pilot specification is usually ready within 5 business days, and onboarding takes about two weeks. A pilot can start at 100 hours, and the same contract scales to 50,000-hour production programs without a new RFP. To date we have executed more than 100,000 collection hours.
Scale on its own proves little. A 50,000-hour program only pays off if the first 100 hours prove the protocol, so we treat the pilot as a joint test rather than a sales formality. By the end of it, both sides should know whether the spec holds up, whether the rig works, whether the participants are the right ones and whether the QA criteria catch what they should. Our collection work covers five industry areas:
- Humanoid and bipedal robotics: household and manipulation demonstrations recorded from egocentric, third-person and robot views.
- Autonomous vehicles and ADAS, including in-cabin monitoring footage.
- Surgical and medical robotics, run under GDPR and PDPA controls.
- Retail and warehouse AMRs: navigation and task footage in real stores and warehouses.
- AR/VR and smart glasses: egocentric capture for perception and interaction models.
Planning a robotics data program? We scope real-world footage collection in Vietnam, Thailand, Singapore and Malaysia, and a pilot specification is usually ready within 5 business days.
Explore data collection servicesWhen another type of partner may suit you better
We would rather lose a poor-fit program than run it badly. If your need is labeling a large archive you already own, you do not need a collection program at all; an annotation provider, including our DataXanno team, is the faster route. For participants in a country outside Vietnam, Thailand, Singapore and Malaysia, raise it at the first call and we will tell you plainly what we can cover rather than improvise a network. And if what you want is a public benchmark rather than proprietary data, open datasets such as DROID may be the quicker start, with custom collection added later to cover your deployment environments.
Frequently asked questions
Robotics and ML leads evaluating DataX Power as a Vietnam data collection partner usually ask these first.
- What does a Vietnam data collection partner do for a robotics team?
- A managed data collection partner in Vietnam takes the whole collection chain off your hands. It recruits participants and writes the capture protocol. It supplies the camera rigs and configures them, collects consent, runs quality checks, and delivers the footage as a structured dataset. For robotics teams, the partner should also understand demonstration data, from task coverage and synchronized multi-view capture to delivery in formats like LeRobot or RLDS. We run all of these steps end to end from Hanoi, with participant networks in Vietnam, Thailand, Singapore and Malaysia.
- How quickly can DataX Power start a robotics data collection program?
- We usually deliver a pilot specification within 5 business days of scoping, and onboarding takes about two weeks. Pilots can start at 100 hours of footage, and the same contract scales to 50,000-hour production programs without a new procurement round.
- What capture setups does DataX Power support?
- We support four capture modes, and they can be combined in one program. Egocentric capture runs on head-mounted rigs, Meta Aria and smart glasses at up to 4K/60fps. Third-person capture uses fixed and mobile camera arrays for multi-angle and 3D pose data. Robot manipulation views come from wrist-mounted and overhead cameras, synchronized with force/torque logging where needed, and mobile navigation video from on-board and chase cameras. The team has run both stereo camera egocentric collection and multi-camera video collection for robotics training.
- Can DataX Power deliver datasets in LeRobot or RLDS format?
- Yes. After QA and annotation through our DataXanno arm, we map episodes into LeRobot-style or RLDS-style structures, or into your own schema, with observations, actions, language instructions and timestamps aligned. Agree the target format during scoping and you avoid conversion work after delivery.
- Does DataX Power collect data in real environments or in a studio?
- Both. Onsite and studio collection is for programs that need controlled lighting and repeatable rig calibration. When a program needs scene and participant diversity, we run crowdsourced collection in real homes, shops, warehouses and outdoor sites. The two methods can be combined in one program under the same protocol and QA criteria.


