Welocalize · Language
Project Cursa - Robot Manipulation Video Annotator (V2)
Listed on Welocalize as “Project Cursa - Robot Manipulation Video Annotator (V2)”
What this actually is
You evaluate AI responses in the target language, rate translation quality, write prompts in your native language, and flag errors. Not professional translation work in the traditional sense. The platform title (Project Cursa - Robot Manipulation Video Annotator (V2)) reflects the rate band and the expertise required, not the day-to-day work.
Advertisement
Can you do this on your visa?
F-2 / F-4 / F-5 / F-6: open. E-1 to E-7: needs concurrent-employment permit. D-2 / D-4 students: S-3 permit, 20 hr/week cap. D-10 / D-8: case by case.
Korean tax on USD income
First 5 years in Korea: foreign-source income only taxed if remitted into Korea. After year 5: worldwide income. Full tax guide.
Original posting from Welocalize
About the Role
We are looking for detail-oriented annotators to help label robot manipulation videos for AI training purposes. You'll watch short videos of robots performing manipulation tasks (filmed from three synchronized camera angles) and produce precise, structured, natural-language descriptions of the actions taking place. This work directly supports the development of robotics AI models and requires strong written English, sharp observational skills, and the discipline to follow a detailed style guide consistently.
What You'll Do
- Watch short robot manipulation videos, each filmed from three synchronized camera views (an overhead view and views from each of the robot's two wrist-mounted cameras).
- Break each video into time segments and write clear, natural-language descriptions for each segment.
- Apply labels at three levels of detail for each applicable segment:
- Atomic motion (a few seconds) - a single small movement (e.g., "close fingers around the red handle")
- Skill / subtask (several seconds to ~20 seconds) - a complete, meaningful action (e.g., "pick up the red block by its edge")
- Task / goal (up to ~1 minute) - the overall purpose of a sequence of skills (e.g., "place all blocks in the container")
- Ensure every moment of video is covered by a label at two or more of these levels - no gaps, including idle or pause moments.
- Accurately describe exactly what happens, including when something doesn't go as planned (a dropped object, a failed grasp, a slipped grip). Precision matters more than making the robot look successful.
- Cross-reference all three camera angles: use the overhead view to understand the overall scene and object identity, and the close-up wrist views to confirm exact contact and grasp details.
- Follow a detailed style guide covering vocabulary for actions, spatial relationships, object descriptions, and manner of movement, applying it consistently across many episodes.
- Participate in periodic calibration sessions to align your labeling with the team and the client's reference examples.
What We're Looking For
Required:
- Strong written English - you'll write dozens of short, precise descriptive sentences per video and need to vary your language rather than repeating the same phrases.
- Sharp attention to detail - able to distinguish small differences (a successful grasp vs. a fumble, a push vs. a drag, which specific object part is being touched).
- Comfort following a detailed, structured style guide and applying it consistently, even in ambiguous or edge-case scenarios.
- Basic comfort with spatial/mechanical description (left/right, above/below, naming object parts like handles, lids, or edges).
- Reliable, self-directed work habits - this is often heads-down work with periodic check-ins rather than close supervision.
Nice to Have:
- Prior experience with video annotation, data labeling, transcription, or QA work.
- Familiarity with robotics terminology (grippers, end-effectors, manipulation) - helpful but not necessary, as the style guide is self-contained.
- Experience with annotation tools such as Label Studio.
#L1-CC1
Quoted from Welocalize’s public listing on 2026-10-06. We don’t edit platform copy; honest framing is in the title and the “what this actually is” block above.
Related AI training jobs
Alignerr · Language
Bengali Language Subject Matter Expert – AI Audio Transcription
$15-$35/hr · Remote · USD
Alignerr · Language
Bengali Language Subject Matter Expert – AI Transcription Project Lead
$15-$35/hr · Remote · USD
Alignerr · Language
Bengali Subject Matter Expert – AI Audio Transcription (Remote)
$15-$35/hr · Remote · USD
Alignerr · Language
Bengali Subject Matter Expert – Audio Transcription (AI Training)
$15-$35/hr · Remote · USD
More on this platform
About Welocalize
Language QA at scale. Translation evaluation, speech labeling, multilingual prompt review. Pay is hourly and predictable; project supply is steady for working language pairs.
Is Welocalize legit? Our review
A localization company founded in 1997, not a task marketplace. Scheduled rater work with bounded weekly hours, steadier but lower pay than Outlier or Alignerr ceilings, and a residency catch on Korean-language roles worth checking before you apply.
See all AI training jobs
Browse by category and compare across all eight platforms we cover.