Why vendor selection is the highest-leverage decision in your data program
The difference between a successful AI data program and a failed one is rarely the model architecture, the compute budget, or the annotation tooling. It is almost always the vendor. A vendor who misunderstands your task specification produces data that trains the wrong distribution. A vendor who lacks native speaker depth for your target language produces data that is surface-plausible and substantively wrong. A vendor who runs inadequate QA produces noise at scale.
Vietnam has more data collection and annotation vendors per capita than any other APAC market. The top tier - vendors with documented QA processes, native speaker depth, and production-grade tooling - is a small fraction of the total vendor count. The majority operate at a quality level that is adequate for commodity tasks and inadequate for production AI training data.
This guide gives you the six steps that separate enterprise-grade vendor selection from price-comparison. Each step has a concrete action and a decision criterion.
Step 1: Define your task before you contact any vendor
The single most common failure mode in data program procurement is beginning vendor conversations before the task is fully specified. Vendors who receive a vague brief respond with vague proposals. You cannot compare proposals that are not anchored to the same specification. You cannot evaluate pilot results without a gold standard. And you cannot write a contract that holds a vendor accountable to quality you have not defined.
A complete task specification for a data collection or annotation program includes: the data type (image, video, audio, text, sensor fusion, or multi-modal); the annotation schema with every label class, attribute, and edge case explicitly documented; the success criteria (target IAA score, target F1, acceptable error rate per class); the delivery format (COCO JSON, HDF5, RLDS, WebDataset, plain CSV); and the volume and timeline requirements.
For collection programs specifically, the task specification also defines the collection environment (indoor, outdoor, specific location types), the participant profile (age range, demographics, equipment use), the sensor configuration (camera model, resolution, frame rate, depth sensor, IMU), and the synchronization requirement across modalities.
You do not need to know every detail before starting. But you need to know enough to give a vendor a concrete brief that produces a concrete, comparable response.
Step 2: Set your quality benchmarks before you shop
Quality benchmarks need to be established before vendor conversations begin, not negotiated with vendors during procurement. Vendors set benchmarks to match what they can deliver, not what your model actually needs. If you enter a vendor conversation without quality targets, you will end up with whatever the vendor's default QA standard is - which may be substantially lower than what your training pipeline requires.
The benchmark that matters most for annotation programs is inter-annotator agreement (IAA) on your specific task, measured by Cohen's kappa or Krippendorff's alpha depending on your label type. A minimum acceptable kappa of 0.75 is a reasonable floor for production annotation; 0.80 is the target for programs where annotation quality directly drives model performance on high-stakes tasks.
For collection programs, the relevant benchmarks are: minimum clean-frame rate (percentage of delivered frames that pass motion blur, occlusion, and sensor dropout checks), synchronization tolerance across modalities (for multi-sensor programs), and episode completion rate (percentage of collected demonstrations that meet the task success criteria).
Set these numbers before you brief any vendor. Include them in the RFP. Ask vendors to confirm they can meet them in writing before the pilot begins.
Step 3: Evaluate certifications and compliance posture
For enterprise programs, vendor certifications are a necessary but not sufficient quality signal. ISO 27001 (information security management) is the baseline expectation for any vendor handling personal data or proprietary model training data. ISO 9001 (quality management systems) indicates documented process discipline. SOC 2 Type II is increasingly requested by US-headquartered enterprise clients.
Vietnam's data protection framework (Decree 13/2023/ND-CP) imposes cross-border transfer restrictions for personal data of Vietnamese data subjects. If your program collects data involving Vietnamese individuals - participant video, annotated conversation data, biometric data - the vendor's cross-border transfer protocol needs to be explicitly addressed in the contract.
Practical compliance evaluation: ask the vendor for their data handling agreement template, their sub-processor list, and their breach notification protocol. A vendor who cannot produce these documents within 48 hours either does not have them or is producing them in response to your request - neither is a confident signal for an enterprise engagement.
Step 4: Design a paid pilot before committing to production volume
A structured pilot is the highest-ROI investment in any data program. A well-designed pilot of 200-500 annotation items or 50-100 collection hours will tell you more about a vendor's production quality than any number of reference calls, case studies, or sales deck accuracy claims.
The pilot should be paid at the production rate (not discounted as a loss-leader), run on your actual data or a representative sample, evaluated against your gold standard by your team, and graded on the quality benchmarks you defined in Step 2. If the vendor cannot meet the acceptance criteria on a pilot, they will not meet them at production volume.
The pilot also tests the vendor's project management quality: how they handle questions, how they escalate edge cases, how quickly they turn around revisions, and whether their delivery format matches what your pipeline expects. These operational signals are often more predictive of production outcomes than the quality score on the pilot itself.
Any vendor who resists a structured pilot - who offers only a free sample, who asks to skip straight to production, or who conditions the pilot on a commitment to production volume - is communicating that they are not confident in their own quality under controlled evaluation conditions.
Step 5: Review pricing models and total program cost
Vietnam data collection and annotation pricing comes in three models: per-item (per annotation, per image, per document), per-hour (of annotator time or collection time), and dedicated-team (monthly retainer for a fixed team size). Each model has a different risk profile for the buyer.
Per-item pricing creates incentives for annotators to prioritise speed over quality. It is appropriate for high-volume, low-judgment tasks (bounding box annotation, binary classification) where quality is enforced by automated checking rather than per-item review. For high-judgment annotation tasks - action segmentation, medical imaging, RLHF preference labeling - per-item pricing misaligns incentives.
Per-hour and dedicated-team pricing aligns annotator incentives with quality, because the annotator is paid for time rather than throughput. For specialist annotation programs, these models produce more consistent output. The buyer risk is scope creep: ensure the contract defines productivity expectations (minimum items per hour by task type) alongside the hourly or team rate.
Total program cost includes hidden components that per-item quotes exclude: project management overhead, QA rejection and rework cycles, tooling and setup fees, and the cost of your team's time managing the vendor. Programs that appear cheaper on a per-item basis often cost more in total when rework rates are high and project management overhead is heavy.
DataX Power offers structured pilot programs for data collection and annotation - 50-100 hour collection pilots and 500-item annotation pilots with defined acceptance criteria, dedicated project managers, and production-rate pricing.
Start a pilot programStep 6: Check references for programs of comparable scope
Reference checks for data vendors have a specific failure mode: vendors offer references from their best programs, which may be structurally different from what you need. A vendor with an excellent track record in image classification for a consumer app may have no experience with multi-modal sensor collection for a robotics program. A vendor who has delivered accurately on English annotation may have never run a Vietnamese NLP program at production quality.
Request references specifically for: the same data type as your program, the same language as your program if language-specific, and a comparable volume and timeline. Ask references directly: what was the actual IAA on the production dataset (not the demo), what was the rework rate in the first production batch, and would you engage this vendor again for a program of the same type?
If a vendor cannot provide references that match your program type, treat this as a signal that your program would be their first at that specification - which is a meaningful risk factor, not a disqualifier, but one that warrants pricing in a larger pilot before committing to production volume.
Red flags: what to walk away from
Across the six steps above, certain vendor behaviours consistently predict production failures. Treat these as walk-away signals:
- No IAA reporting: a vendor who cannot produce inter-annotator agreement metrics has no documented quality process. Accuracy claims without IAA are marketing, not measurement.
- Resistance to a structured pilot: any vendor who conditions production engagement on skipping or shortcutting the pilot is not confident in their own quality output.
- Vague answers on annotator recruitment: for language-specific work, ask where annotators are recruited, what their native language is, and what dialect profile the pool covers. Vague answers mean the vendor is sourcing annotators reactively rather than from a pre-qualified pool.
- Per-item pricing for high-judgment tasks: for specialist annotation (medical imaging, manipulation action segmentation, RLHF preference labeling), per-item pricing misaligns incentives. A vendor who insists on per-item pricing for high-judgment work is optimised for throughput, not quality.
- No data handling agreement: if a vendor cannot produce a signed data handling agreement with sub-processor list and breach protocol, they have not built the compliance infrastructure for enterprise data.


