Video annotation vs image annotation - why the distinction matters
Vietnam has a large and competitive data annotation market, but most of that market is image annotation. Video annotation is a distinct discipline requiring temporal consistency, multi-frame object tracking, activity segmentation, and annotation interpolation across variable-length clips. Not all annotation vendors in Vietnam have built these capabilities - and many who list "video annotation" in their service catalog are applying image annotation workflows to individual frames, which produces technically correct but temporally inconsistent results.
For enterprise AI buyers, the difference matters because models trained on temporally inconsistent video annotations learn incorrect motion priors. A tracking model trained on bounding boxes that jump between frames because the annotator re-drew each frame independently will not generalize to smooth motion at inference time. Temporal consistency in video annotation is not a quality preference - it is a correctness requirement for video-based model training.
1. What video annotation services in Vietnam actually cover
Enterprise-grade video annotation services from Vietnam-based vendors cover several distinct task types, each requiring different tooling and annotator training.
Bounding box tracking and temporal interpolation is the most common video annotation type - drawing boxes around objects in keyframes and interpolating box positions between keyframes. Quality depends on the interpolation algorithm, keyframe spacing, and annotator discipline in handling object disappearance, occlusion, and reappearance. A well-specced video annotation vendor should be able to demonstrate their interpolation approach and show IAA metrics on tracking consistency across frames.
Semantic and instance segmentation of video frames produces pixel-level masks for each object in each frame - significantly more annotation effort than bounding boxes, but necessary for scene understanding, depth estimation, and precise object interaction modeling. Vietnam-based vendors with semantic segmentation capability typically use polygon annotation with interpolation rather than per-frame mask drawing.
Action and activity recognition annotation requires temporal segmentation - identifying the start and end frame of each activity event in a clip, and labeling it with an activity class. This requires annotators who understand the activity ontology well enough to make consistent boundary decisions. Action annotation accuracy is highly sensitive to inter-annotator agreement on boundary placement - ambiguous activity boundaries are the primary source of noise in action recognition datasets.
Keypoint and pose annotation tracks human body keypoints across frames - used for pose estimation, action recognition, and human-robot interaction training. Requires annotators trained on the specific keypoint schema and able to handle partial occlusion cases consistently.
2. Quality standards to specify before signing a contract
The most important quality conversation happens before collection begins, not after delivery. Specifying quality standards contractually - rather than evaluating delivered data and arguing about rework - protects both buyer and vendor.
Inter-annotator agreement (IAA) targets should be specified per task type. For bounding box tracking, a minimum IoU of 0.80 across a randomly sampled QA set is a reasonable baseline. For action segmentation, temporal IoU above 0.75 is achievable with a trained team. For keypoint annotation, mean per-joint precision (MPJPE) thresholds should be agreed before annotation begins. Vendors who resist specifying IAA targets are signaling that their workflow does not include systematic quality measurement.
Review and rework procedures should be in the contract. What percentage of delivered annotations will be independently reviewed? What is the threshold that triggers rework? Who pays for rework - the vendor (if quality target was missed) or the buyer (if target was met but buyer changes requirements)? These questions become expensive to resolve retroactively when a 100,000-frame delivery lands below the quality bar.
Annotator consistency should be validated through a paid pilot before committing to a full program. A 500-frame pilot annotated by the production team, reviewed against ground truth created by a domain expert, gives a realistic IAA measurement for the specific task. Pilot IAA is a better predictor of production quality than a vendor's claimed accuracy rate on unrelated past projects.
3. Collection-annotation integration - where programs succeed or fail
The most common failure mode in video data programs is the handoff between collection and annotation. Footage collected without annotation requirements in mind arrives at the annotation team in a format or condition that makes it difficult or impossible to annotate to the required spec. The result is either poor-quality annotations or expensive recollection.
Common handoff failures: footage captured at insufficient frame rate for the tracking density required; lighting conditions that make object boundaries ambiguous; scene composition that places key objects partially outside frame; capture protocol that did not enforce the participant behaviors required for the action recognition ontology; file formats incompatible with the annotation tooling.
The fix is to design collection and annotation as one integrated program, not two sequential projects. The annotation ontology - what objects get labeled, what activity classes exist, what the keypoint schema is - must be finalized before the capture protocol is written. The capture protocol is then designed to produce footage that is annotatable to that ontology.
Vendors who handle both collection and annotation from a single team have a structural advantage here: they can iterate on the protocol based on early annotation feedback, and the annotation team can flag collection issues in the first batch before they propagate through the full program.
4. Pricing for video annotation services in Vietnam
Video annotation pricing in Vietnam varies significantly by task complexity and quality tier.
Bounding box tracking at standard quality (80% IoU target, no temporal consistency review) runs $0.05-0.15 per frame for common object classes in favorable lighting conditions. Semantic segmentation runs $0.30-0.80 per frame depending on scene complexity. Action annotation (temporal segmentation only, no per-frame labeling) runs $0.15-0.40 per clip minute at a standard activity ontology. Keypoint annotation runs $0.20-0.60 per frame per person depending on keypoint count.
These rates assume standard annotation tooling and a trained team. Rates increase for: complex occlusion handling (20-40% premium), non-standard ontologies requiring additional training (project setup fee plus 15-25% rate premium), same-day turnaround (40-60% premium), and domain-specific annotation requiring specialist knowledge.
The lowest-rate vendors in Vietnam are typically image annotation teams applying sequential-frame workflows to video. Their rates are 30-50% below the figures above, but the output requires significant post-processing to achieve temporal consistency. For production training pipelines, the post-processing cost usually eliminates the rate savings.
DataX Power provides integrated video collection and annotation services from Vietnam - with consistent annotation standards across collection and labeling, no handoff failures, and IAA-audited quality on every program.
Learn about our collection and annotation programs

