Dual-Arm Robotics Annotation: Cutting Annotation Effort by up to 60%
40%
reduction in annotation effort
1 hour
annotation time per video minute
3 weeks
to build and test the workflow
USE CASE
Robotic Arm Event Annotation | Action Timing | Motion Descriptions
INDUSTRY
Robotics | Computer Vision
SOLUTION
Data Annotation
The mission: Reduce the work required for dual-arm annotation
A robotics company needed to identify how 2 robotic arms moved objects between containers, including the timing and outcome of each action.
The original scope combined event timestamps with continuous bounding boxes around the arms, objects, and containers. Early testing showed that the boxes required substantial effort beyond the client’s main priorities: event timing and action outcomes.
We proposed a narrower scope that preserved those details while reducing annotation work. Over 3 weeks, we moved from unfamiliar source files to a tested annotation workflow, a sample for client review, and measured delivery benchmarks.
The work: Give each arm its own event sequence
The source data included 3 camera views:
- Left-arm view
- Right-arm view
- Overhead view showing both arms
The overhead view captured the full interaction, including moments when both arms moved at once.
We tracked each arm separately. Every pick-and-place cycle was marked from the start of the grab through placement. Failed and unclear actions received separate labels and short descriptions.
This structure showed which arm performed each action and when it occurred.
Figure 1. Each robotic arm is annotated independently across its pick-and-place cycle. Recreated for illustration.

The challenge: Keep event timing accurate across file conversions
The source footage arrived in MCAP, a format used to store robotics data and synchronized sensor information.
Our available annotation platforms could not open the files directly. Converting them to MP4 made the videos easier to view, but introduced a risk of frame timing shifts. Every event label still had to align with the original footage.
We compared available annotation tools against the project’s tracking, timing, review, and output requirements. None met the full requirement without configuration or file preparation.
Key pressure points:
- Overlapping actions: Both arms could move simultaneously, so each event had to be assigned to the correct arm.
- Full-scope testing: We first completed the timestamp and bounding-box requirements to establish a full-scope baseline and identify where effort could be reduced. This took up to 2.5 hours per minute of video.
- Exception handling: Failed, unclear, and low-visibility actions needed separate rules and short descriptions.
- No delivery baseline: We had no prior figures for annotation time, review effort, platform costs, or staffing needs for this type of footage.
The test: Validate the method on a complete sample
We checked the prepared footage against the source files to confirm that event timing remained aligned.
We selected and configured the closest-fit platform for separate arm records, review, and structured output.
We then annotated a 45-second clip and prepared the corresponding data for client review.
The sample confirmed that the output could capture:
- The arm involved
- The action
- The outcome
- The timestamp
The change: Replace continuous boxes with event-based annotation
Full-scope testing showed that continuous bounding boxes took most of the annotation time.
We removed them and focused the scope on the data needed for review and planning:
- Event timestamp
- Left-arm or right-arm identification
- Action outcome
- Short description for failed or unclear events
This reduced annotation time from up to 2.5 hours to about 1 hour per video minute, a reduction of up to 60%.
Figure 2. Continuous box tracking compared with event-based annotation, recreated for illustration.

The revised workflow captured each action as a structured event record rather than continuously redrawing boxes as the arms and objects moved.
We also measured review effort, platform costs, and staffing needs. This gave the client a clearer basis for estimating larger volumes of footage.
The results: Clearer records for larger dual-arm datasets
For the client, this provided:
- Clearer dual-arm records: Separate records for the left and right arms made overlapping movements easier to distinguish.
- Defined failure handling: Failed and unclear actions were retained in the dataset via dedicated labels and short descriptions.
- Effort aligned to priorities: Annotator time focused on event timing, arm identification, and action outcomes required for client review and planning.
- Stronger delivery estimates: The measured annotation rate gave the client clearer cost and timeline estimates.
This gave the client a defined annotation structure and measured delivery benchmarks before committing larger volumes of robotics footage.
The findings now feed into Chemin's internal robotics tooling, where reducing manual handling and holding timestamps steady across format conversions are core to how we're building the data supply chain for future video projects.
Robotics data labeling built for precise timelines
Establish timing accuracy, review requirements, and delivery benchmarks before committing larger data volumes.
More stories

Turning mission-critical data into waste intelligence
Accelerated waste recognition AI by delivering 1 million high-accuracy, compliance-ready annotations monthly through expert-driven workflows and rapid data turnaround.

Audio AI Annotation for a Premier Gaming Studio: Scaled Output by 294%
Prepared 629.5 hours of conversational audio in 30 days while maintaining 99% accuracy and zero rework.

Gaming Audio Annotation: Structuring Speech Cues at 95%+ Accuracy
Turned emotion, accent, demographic, and speech cues into structured training data for adaptive game AI.