Chemin

Dual-Arm Robotics Annotation: Cutting Annotation Effort by up to 60%

Data Annotation
Delivered clear action records showing what each robotic arm did, when it happened, and whether the movement succeeded.

40%

reduction in annotation effort

1 hour

annotation time per video minute

3 weeks

to build and test the workflow


USE CASE
USE CASE

Robotic Arm Event Annotation | Action Timing | Motion Descriptions

INDUSTRY
INDUSTRY

Robotics | Computer Vision

SOLUTION
SOLUTION

Data Annotation

Dual-Arm Robotics Annotation: Cutting Annotation Effort by up to 60%

The mission: Reduce the work required for dual-arm annotation

A robotics company needed to identify how 2 robotic arms moved objects between containers, including the timing and outcome of each action.

The original scope combined event timestamps with continuous bounding boxes around the arms, objects, and containers. Early testing showed that the boxes required substantial effort beyond the client’s main priorities: event timing and action outcomes.

We proposed a narrower scope that preserved those details while reducing annotation work. Over 3 weeks, we moved from unfamiliar source files to a tested annotation workflow, a sample for client review, and measured delivery benchmarks.

The work: Give each arm its own event sequence

The source data included 3 camera views:

  • Left-arm view
  • Right-arm view
  • Overhead view showing both arms

The overhead view captured the full interaction, including moments when both arms moved at once.

We tracked each arm separately. Every pick-and-place cycle was marked from the start of the grab through placement. Failed and unclear actions received separate labels and short descriptions.

This structure showed which arm performed each action and when it occurred.

Figure 1. Each robotic arm is annotated independently across its pick-and-place cycle. Recreated for illustration.

image.png

The challenge: Keep event timing accurate across file conversions

The source footage arrived in MCAP, a format used to store robotics data and synchronized sensor information.

Our available annotation platforms could not open the files directly. Converting them to MP4 made the videos easier to view, but introduced a risk of frame timing shifts. Every event label still had to align with the original footage.

We compared available annotation tools against the project’s tracking, timing, review, and output requirements. None met the full requirement without configuration or file preparation.

Key pressure points:

  • Overlapping actions: Both arms could move simultaneously, so each event had to be assigned to the correct arm.
  • Full-scope testing: We first completed the timestamp and bounding-box requirements to establish a full-scope baseline and identify where effort could be reduced. This took up to 2.5 hours per minute of video.
  • Exception handling: Failed, unclear, and low-visibility actions needed separate rules and short descriptions.
  • No delivery baseline: We had no prior figures for annotation time, review effort, platform costs, or staffing needs for this type of footage.

The test: Validate the method on a complete sample

We checked the prepared footage against the source files to confirm that event timing remained aligned.

We selected and configured the closest-fit platform for separate arm records, review, and structured output.

We then annotated a 45-second clip and prepared the corresponding data for client review.

The sample confirmed that the output could capture:

  • The arm involved
  • The action
  • The outcome
  • The timestamp

The change: Replace continuous boxes with event-based annotation

Full-scope testing showed that continuous bounding boxes took most of the annotation time. 

We removed them and focused the scope on the data needed for review and planning:

  • Event timestamp
  • Left-arm or right-arm identification
  • Action outcome
  • Short description for failed or unclear events

This reduced annotation time from up to 2.5 hours to about 1 hour per video minute, a reduction of up to 60%.

Figure 2. Continuous box tracking compared with event-based annotation, recreated for illustration.

image.png

The revised workflow captured each action as a structured event record rather than continuously redrawing boxes as the arms and objects moved.

We also measured review effort, platform costs, and staffing needs. This gave the client a clearer basis for estimating larger volumes of footage.

The results: Clearer records for larger dual-arm datasets

For the client, this provided:

  • Clearer dual-arm records: Separate records for the left and right arms made overlapping movements easier to distinguish.
  • Defined failure handling: Failed and unclear actions were retained in the dataset via dedicated labels and short descriptions.
  • Effort aligned to priorities: Annotator time focused on event timing, arm identification, and action outcomes required for client review and planning.
  • Stronger delivery estimates: The measured annotation rate gave the client clearer cost and timeline estimates.

This gave the client a defined annotation structure and measured delivery benchmarks before committing larger volumes of robotics footage.

The findings now feed into Chemin's internal robotics tooling, where reducing manual handling and holding timestamps steady across format conversions are core to how we're building the data supply chain for future video projects.

Robotics data labeling built for precise timelines

Establish timing accuracy, review requirements, and delivery benchmarks before committing larger data volumes.

Share