Why DataX Power Is the Best Vietnam Data Collection Partner for Robotics AI

Five reasons robotics and embodied AI teams can rely on DataX Power for real-world footage collection in Vietnam and Southeast Asia, with the operating detail behind each one.

12 min read
Hanoi skyline across West Lake, home base of DataX Power, a Vietnam data collection partner for robotics AI

What a Vietnam data collection partner has to get right for robotics AI

If you are shortlisting a Vietnam data collection partner for a robotics or embodied AI program, you will find plenty of vendors who can point a camera at a task. Fewer can deliver footage your policy can learn from. That means real people recorded in real environments, synchronized views, a consent record for every session, and a dataset structure your training code reads without a week of conversion work.

We started DataX Power in 2023. The company is headquartered in Hanoi with a presence in Melbourne, and works across three pillars: custom AI model development, data services (collection, annotation, validation) and infrastructure operations. Our annotation arm operates as DataXanno. For the past two years we have run hands-on AI and robotics data collection and processing programs for teams working on manipulation policies and humanoids, as well as autonomous systems and smart-glasses products.

We think we are the strongest choice for robotics teams collecting in Vietnam and Southeast Asia, for the five reasons set out below, each with the operating detail behind it. We also say where a different kind of supplier may suit you better. This is our own page, so read it the way you would read any vendor's case, then test it with a pilot.

1. A real-world footage operation model, managed end to end

Most robotics data problems trace back to a gap between where data was recorded and where the robot has to work. A policy trained on a tidy lab bench can stall in a cluttered Bangkok kitchen or a narrow warehouse aisle in Johor. The DROID project made the same point at research scale: its authors collected 76k manipulation trajectories across 564 scenes and report that training on that diversity produced policies with higher performance, greater robustness and better generalization.

Our operating model is built to close that gap. What we collect is real-world footage, meaning real environments and real participants, with one team managing every step between your brief and the delivered dataset.

Two collection methods sit under that model. Onsite and studio collection gives tight control over lighting and rig calibration, so sessions are repeatable. Crowdsourced collection buys breadth across many homes, shops and streets, and shipped kits plus protocol checks keep it within spec. The two can run in one program, for example a studio anchor set for controlled comparisons plus field sessions for scene diversity.

What separates this from an open crowd platform is accountability. We choose the participants and write the protocol. We set up the hardware, and footage that fails the pass criteria is rejected before it reaches you. If something goes wrong anywhere in the chain, you have one team to hold responsible instead of a marketplace of individual contributors. Every program covers the same six managed steps:

  • Recruitment: participants screened and scheduled against your program profile, drawn from our networks in Vietnam, Thailand, Singapore and Malaysia.
  • Capture protocol: a written document covering tasks, scene variation, camera placement and pass criteria, agreed with you before recording starts.
  • Hardware: we source, configure and maintain the rigs, from head-mounted cameras and Meta Aria glasses to fixed multi-camera arrays and wrist-mounted cameras.
  • Consent: informed consent collected and logged per participant and per session, with the records delivered alongside the footage.
  • QA: review by QA engineers trained on robotics video, checking temporal consistency, frame quality and task coverage.
  • Delivery: data packaged in the schema your pipeline expects, with manifests and QA reports.

2. Two years of hands-on robotics data collection and processing

We have spent the past two years collecting and processing data for AI and robotics programs, including robotics data collection in Vietnam and the wider region. Those two years were focused entirely on this work, and they cover the period in which robot learning shifted toward large, real-world demonstration datasets: the Open X-Embodiment collaboration pooled data from 22 robot types across 21 institutions into standardized formats, and projects like DROID showed what distributed real-world collection can do.

What we have from that time is a working knowledge of how robotics collection fails, and it fails differently from general video work. A demonstration that looks complete but ends before the object is released. A head-mounted clip where the hands leave the frame during the critical grasp. Two camera streams that drift a few frames apart over a long session. A task list that covers the easy variations fifty times and the hard ones twice. A thumbnail review catches none of these, and every one of them hurts a policy.

Over more than 100,000 collection hours executed, we have written those lessons into our capture protocols and QA checklists, and into how we train participants, instead of leaving them in individual heads. A new program starts from that baseline, so you are not paying for us to relearn the same mistakes.

3. In-house custom AI model development, so we know what the data has to do

Collecting data is only one part of what we do. Custom AI model development is one of our three business pillars. Our engineers build and train models for clients, and that changes how we collect.

Training models teaches you quickly which data properties matter. Label consistency beats label volume. Failure cases need coverage and task distributions need balance. Timestamps have to be accurate, and metadata should let you slice results by scene, participant or lighting. So when our collection leads write a capture protocol, they are thinking about the training run as much as the recording day.

At scoping we ask questions that are easy to skip if you only ever deliver raw footage. What action space and observation keys does your policy use? What control frequency? Which camera views are inputs at inference time? How do you split training and evaluation data? Which failure modes are you seeing today? The answers go into the protocol and the rig, and into the QA pass criteria. Without them, a pilot tends to come back as a footage sample instead of something a training team can use.

4. From raw footage to a trained model: the full data path

Raw footage is not a dataset. The value gets added between capture and training, which is where our data services and AI development teams work together.

The step that is easiest to leave out is the loop back from training to collection. A supplier that stops at delivery cannot tell you that your policy struggles with transparent objects or low light. When the same team trains models, it sees those failures and can change the next batch: the lighting, the scene mix, the balance of tasks.

Format matters for the same reason. RLDS was designed as a standard, lossless format for recording and sharing sequential decision-making data, and LeRobot stores synchronized video alongside Parquet files for state and action data. Delivering in structures like these, or in your internal schema, means your engineers spend their time training rather than writing converters. Each program runs through the same path:

  • Capture and ingest: synchronized streams, session metadata, consent records and rig calibration files ingested together, so nothing is orphaned.
  • QA gate: footage checked against the protocol pass criteria before annotation, so annotation budget is not spent on unusable clips.
  • Annotation through DataXanno: task segmentation, action and phase labels, hand and object annotation or a custom schema, handled by our annotation arm under the same program lead.
  • Schema mapping: episodes structured in LeRobot-style or RLDS-style layouts, or your own schema, with observations, actions, language instructions and timestamps aligned.
  • Training and evaluation feedback: a sample run through a baseline training or evaluation loop, in your stack or ours, to surface gaps before the full batch lands.
  • Next collection round: coverage gaps and failure cases from evaluation turned into the next protocol revision.

5. Stereo egocentric and multi-camera collection experience

Two capture setups come up again and again in robot learning programs, and we have run both for robotics training.

Stereo camera egocentric collection uses head-mounted rigs with calibrated camera pairs, so the footage carries depth cues about hand-object distance rather than a flat first-person view. That matters for manipulation and hand-object interaction learning, where a monocular clip can hide how far the hand really is from the object. The operational work is keeping calibration stable while the rig moves with the wearer's head. On top of that, the setup has to stay consistent across participants and sessions, with both streams frame-aligned. Our egocentric capture runs on head-mounted rigs, Meta Aria glasses and smart glasses at up to 4K/60fps.

Multi-camera video collection places several synchronized cameras around the task. The hard work is in timestamp alignment and in keeping camera placement consistent across sessions, and QA has to check every view instead of only the primary one. We cover four capture modes, which can be combined in a single program:

  • Egocentric: head-mounted rigs, Meta Aria and smart glasses, up to 4K/60fps.
  • Third-person: fixed and mobile camera arrays with synchronized timestamps, capturing multiple angles for 3D pose.
  • Robot manipulation view: wrist-mounted and overhead cameras, synchronized with force/torque logging where the program needs it.
  • Mobile navigation video: on-board and chase-camera footage across warehouses, outdoor terrain and mixed environments.

How the DataX Power model compares with other collection options

Robotics buyers usually compare four kinds of supplier. Each has a legitimate place, and the right one depends on what you need at the end. Here is how they differ for robotics training data, as we see it.

DataX Power operating model vs open crowd platforms, lab or academic collection, and annotation-first BPOs for robotics training data

CriterionDataX Power modelOpen crowd platformLab / academic collectionAnnotation-first BPO
Recording environmentsReal homes, shops, warehouses and outdoor sites, plus studioWherever contributors are; broad but uncontrolledControlled lab settings with limited varietyUsually works on data the client supplies
ParticipantsRecruited and trained against the program profile in 4 APAC marketsSelf-selected contributors with large reachSmall pools, often researchers or studentsAnnotators rather than capture participants
Capture protocolWritten protocol agreed before recordingTask instructions with limited enforcementRigorous and research-designedNot usually in scope
HardwareVendor-configured rigs, including stereo egocentric and multi-camera setupsContributor devices with varied specsSpecialist rigs at low volumeVaries by provider
Consent recordsPer-session consent logged and delivered with the dataPlatform terms; varies by taskEthics-board processesDepends on where the data came from
Robotics-specific QATemporal consistency, frame quality and task coverage reviewAutomated checks and spot reviewResearcher reviewFocused on label accuracy
Downstream model knowledgeIn-house custom AI model developmentLimitedStrong, research-orientedLimited to labeling workflows
Scale path100-hour pilot to 50,000-hour program on one contractVery large volume possibleHard to scale on commercial timelinesLarge labeling capacity
Best fitRobotics and embodied AI programs needing real-world, schema-ready dataHigh-volume, simple capture tasksBenchmarks and research datasetsLabeling data you already hold

What working with DataX Power looks like in practice

The numbers first. A pilot specification is usually ready within 5 business days, and onboarding takes about two weeks. A pilot can start at 100 hours, and the same contract scales to 50,000-hour production programs without a new RFP. To date we have executed more than 100,000 collection hours.

Scale on its own proves little. A 50,000-hour program only pays off if the first 100 hours prove the protocol, so we treat the pilot as a joint test rather than a sales formality. By the end of it, both sides should know whether the spec holds up, whether the rig works, whether the participants are the right ones and whether the QA criteria catch what they should. Our collection work covers five industry areas:

  • Humanoid and bipedal robotics: household and manipulation demonstrations recorded from egocentric, third-person and robot views.
  • Autonomous vehicles and ADAS, including in-cabin monitoring footage.
  • Surgical and medical robotics, run under GDPR and PDPA controls.
  • Retail and warehouse AMRs: navigation and task footage in real stores and warehouses.
  • AR/VR and smart glasses: egocentric capture for perception and interaction models.

Planning a robotics data program? We scope real-world footage collection in Vietnam, Thailand, Singapore and Malaysia, and a pilot specification is usually ready within 5 business days.

Explore data collection services

When another type of partner may suit you better

We would rather lose a poor-fit program than run it badly. If your need is labeling a large archive you already own, you do not need a collection program at all; an annotation provider, including our DataXanno team, is the faster route. For participants in a country outside Vietnam, Thailand, Singapore and Malaysia, raise it at the first call and we will tell you plainly what we can cover rather than improvise a network. And if what you want is a public benchmark rather than proprietary data, open datasets such as DROID may be the quicker start, with custom collection added later to cover your deployment environments.

Frequently asked questions

Robotics and ML leads evaluating DataX Power as a Vietnam data collection partner usually ask these first.

What does a Vietnam data collection partner do for a robotics team?
A managed data collection partner in Vietnam takes the whole collection chain off your hands. It recruits participants and writes the capture protocol. It supplies the camera rigs and configures them, collects consent, runs quality checks, and delivers the footage as a structured dataset. For robotics teams, the partner should also understand demonstration data, from task coverage and synchronized multi-view capture to delivery in formats like LeRobot or RLDS. We run all of these steps end to end from Hanoi, with participant networks in Vietnam, Thailand, Singapore and Malaysia.
How quickly can DataX Power start a robotics data collection program?
We usually deliver a pilot specification within 5 business days of scoping, and onboarding takes about two weeks. Pilots can start at 100 hours of footage, and the same contract scales to 50,000-hour production programs without a new procurement round.
What capture setups does DataX Power support?
We support four capture modes, and they can be combined in one program. Egocentric capture runs on head-mounted rigs, Meta Aria and smart glasses at up to 4K/60fps. Third-person capture uses fixed and mobile camera arrays for multi-angle and 3D pose data. Robot manipulation views come from wrist-mounted and overhead cameras, synchronized with force/torque logging where needed, and mobile navigation video from on-board and chase cameras. The team has run both stereo camera egocentric collection and multi-camera video collection for robotics training.
Can DataX Power deliver datasets in LeRobot or RLDS format?
Yes. After QA and annotation through our DataXanno arm, we map episodes into LeRobot-style or RLDS-style structures, or into your own schema, with observations, actions, language instructions and timestamps aligned. Agree the target format during scoping and you avoid conversion work after delivery.
Does DataX Power collect data in real environments or in a studio?
Both. Onsite and studio collection is for programs that need controlled lighting and repeatable rig calibration. When a program needs scene and participant diversity, we run crowdsourced collection in real homes, shops, warehouses and outdoor sites. The two methods can be combined in one program under the same protocol and QA criteria.
Data Collection Service

Need the platform layer to make this stick in production? Our Hanoi-based infrastructure team delivers DevOps, FinOps, SecOps, and AI/MLOps for enterprises on AWS, GCP, Azure, and on-premise.

携手打造 下一个里程碑

告诉我们您的挑战 – AI、数据或基础设施。我们将为项目梳理范围,并为您配置合适的团队。