Mercor · Open to all
Applied Engineer, Evaluations Role
Listed on Mercor as “Applied Engineer, Evaluations Role”
What this actually is
You complete short labeling tasks, record everyday activities, or evaluate AI outputs across general knowledge domains. Open to most backgrounds. The platform title (Applied Engineer, Evaluations Role) reflects the rate band and the expertise required, not the day-to-day work.
Advertisement
Can you do this on your visa?
F-2 / F-4 / F-5 / F-6: open. E-1 to E-7: needs concurrent-employment permit. D-2 / D-4 students: S-3 permit, 20 hr/week cap. D-10 / D-8: case by case.
Korean tax on USD income
First 5 years in Korea: foreign-source income only taxed if remitted into Korea. After year 5: worldwide income. Full tax guide.
Original posting from Mercor
Mercor is seeking software engineers to build and refine evaluations. You’ll turn completed pull requests into engineering tasks and use our in-house evaluation framework to multiple frontier coding agents, including Claude Code, Codex and more.
We work alongside an in-house research team and have produced industry leading benchmarks to compare the performance of frontier large language models.
In this role, you will:
- Identify repositories with enough substantive work to support challenging evaluations, and select suitable completed PRs.
- Investigate each problem, its reference solution, and the repository’s architecture, tests, and conventions.
- Write task prompts and build evaluations using automated tests, shell commands, and LLM grading prompts.
- Validate evaluations against reference solutions and deliberately flawed implementations. Find missing checks, incorrect grades, and criteria that unnecessarily constrain how a problem can be solved.
- Run evaluations repeatedly across multiple models and harnesses. Investigate whether failures come from the agent’s solution, the environment, or the grading, and establish that tasks expose meaningful weaknesses in agent performance.
- Refine evaluations through repeated testing and review. Document findings and grading decisions, and work through feedback.
You’ll need:
- Strong programming fundamentals and practical experience working in substantial and complex codebases. Languages we create evals for include TypeScript/JavaScript, Python, Java, Kotlin, Go, Ruby, PHP, C++ and Rust
- The ability to understand unfamiliar code, investigate subtle behavior, and assess whether different implementations solve the same problem correctly.
- Experience writing meaningful tests, including edge cases and regression coverage.
- Attention to detail and patience for repeated investigation and refinement.
- Practical experience using AI coding tools, with the judgment to verify their output and catch mistakes.
- Confidence using Git, test runners, and CI tooling.
- Clear and fluent written and verbal communication with the ability to own a task independently while raising questions when requirements are ambiguous.
Relevant experience can come from open-source, private, or enterprise repositories. Familiarity with a particular language or ecosystem is helpful; the ability to learn the repository and make sound engineering judgments matters more than its popularity or your public contribution history.
We provide onboarding to the evaluation framework and ongoing review feedback. After onboarding, you’ll be expected to own task selection, evaluation development, and iteration without step-by-step direction.
Quoted from Mercor’s public listing on 2026-09-29. We don’t edit platform copy; honest framing is in the title and the “what this actually is” block above.
Related AI training jobs
Mercor · Open to all
Agent Engineer
$100-$500/hr · Remote · USD
Mercor · Open to all
AI Rater Guidelines Writer (Linguist / Instructional Designer)
$45-$65/hr · Remote · USD
Mercor · Open to all
AI Safety Practitioner
$60-$70/hr · Remote · USD
Mercor · Open to all
AI Safety Red Teamer
$70-$84/hr · Remote · USD
More on this platform
About Mercor
AI-interview-based talent network. One application, voice interview with their AI, then matched to projects across coding, research, and specialist work. Pay scales with track and seniority.
Mercor review: AI-interview talent network
4.1/5 on Glassdoor, fastest-growing platform in the category (+509% YoY). What the AI video interview actually asks, real pay across coding/research/medical/legal/finance tracks ($25-$200/hr), and the project-availability problem.
See all AI training jobs
Browse by category and compare across all eight platforms we cover.