All AI training jobs

Mercor · Finance & specialist

SWE-Bench Task Auditor

Listed on Mercor as “SWE-Bench Task Auditor

$70-$90/hrRemoteContractPaid in USD
ShareWhatsAppTelegramEmail

What this actually is

You bring your specialist expertise to AI evaluation. The shape of the work varies but the pattern is the same: review outputs, rate quality, write prompts, flag errors. The platform title (SWE-Bench Task Auditor) reflects the rate band and the expertise required, not the day-to-day work.

Advertisement

Can you do this on your visa?

F-2 / F-4 / F-5 / F-6: open. E-1 to E-7: needs concurrent-employment permit. D-2 / D-4 students: S-3 permit, 20 hr/week cap. D-10 / D-8: case by case.

Korean tax on USD income

First 5 years in Korea: foreign-source income only taxed if remitted into Korea. After year 5: worldwide income. Full tax guide.

Original posting from Mercor

Evaluate the quality, correctness, and reproducibility of software-engineering benchmark tasks used to train and evaluate a frontier AI lab's models. You'll assess repository-level tasks, reference patches, test harnesses, and grading integrity - and provide clear, rubric-based written feedback.

Basic Qualifications

• 3+ years professional software engineering

• Real open-source contribution or maintainer experience (merged PRs, committer / maintainer roles)

• Strong ability to audit reference patches, test runners, and Docker isolation, and to detect answer leakage / reward hacking

• Fluency across common ecosystems (Python and at least one of Java / Go / TypeScript / C++)

Preferred Qualifications

• Familiarity with SWE-Bench (Verified) or similar repository benchmarks

• Maintainer history on major Python OSS (Django, Flask, scikit-learn, sympy, pytest, etc.)

• Prior code-review or task-grading experience

Quoted from Mercor’s public listing on 2026-09-08. We don’t edit platform copy; honest framing is in the title and the “what this actually is” block above.

Apply on Mercor