Mercor · STEM & research
Applied Mathematics Benchmark Specialist - review AI outputs in your specialty
Listed on Mercor as “Applied Mathematics Benchmark Specialist”
What this actually is
You design problems that stump current AI models, evaluate AI reasoning against the correct answer, write rubrics, and provide expert feedback. Often the highest-paid category because the expertise pool is small. The platform title (Applied Mathematics Benchmark Specialist) reflects the rate band and the expertise required, not the day-to-day work.
Advertisement
Can you do this on your visa?
F-2 / F-4 / F-5 / F-6: open. E-1 to E-7: needs concurrent-employment permit. D-2 / D-4 students: S-3 permit, 20 hr/week cap. D-10 / D-8: case by case.
Korean tax on USD income
First 5 years in Korea: foreign-source income only taxed if remitted into Korea. After year 5: worldwide income. Full tax guide.
Original posting from Mercor
Role Overview
We are seeking expert mathematicians to author and review high-quality academic assessment content for an AI research initiative. You will write and verify rigorous multiple-choice questions across core mathematics domains, evaluate solution quality, and help establish gold-standard benchmarks used to advance AI capabilities.
You will be assigned one of two task types:
- Question Authoring - Create original, challenging multiple-choice questions in your area of mathematical expertise, rate their difficulty, and submit them for review.
- Question Verification - Review pre-written questions for accuracy, clarity, and rigor. Edit where needed, rate difficulty, and document any changes made.
Mathematics Domains Covered
Signal Processing, Financial Mathematics & Actuarial Science, Mathematical Economics, Mathematical Modeling of Ecological & Biological Systems, Mathematical Programming & Combinatorial Optimization, Geomathematics & Climate Modeling.
Key Responsibilities
- Author original math questions that test deep conceptual understanding, not surface-level recall
- Ensure questions are unambiguous, self-contained, and precisely defined - all necessary information must be in the problem statement
- Rate each question's difficulty: Medium (intro undergraduate), Hard (advanced undergraduate), or Expert (post-graduate and above)
- Provide 1 correct answer and 9 plausible but subtly incorrect alternatives that challenge expert-level solvers
- Write step-by-step Chain-of-Thought solutions with clear, concise intermediate steps in markdown format
- Supply 1-5 academic references per question from reputable sources (peer-reviewed journals, university repositories)
- For verification tasks: flag issues with clarity, completeness, precision, or solvability and justify any edits made
Ideal Qualifications
- PhD or doctoral candidate in Mathematics, Applied Mathematics, Statistics, or a closely related field
- Master's degree considered for candidates with exceptional depth in a specific subdomain
- Strong command of graduate-level mathematical concepts and formal proof writing
- Experience with rigorous academic problem design or mathematical competition writing is a strong plus
- Excellent written English and ability to express complex ideas clearly and concisely
More About the Opportunity
- Expected commitment: 10+ hours/week
- Asynchronous, fully remote work
Quoted from Mercor’s public listing on 2026-09-08. We don’t edit platform copy; honest framing is in the title and the “what this actually is” block above.
Related AI training jobs
Mercor · STEM & research
Applied Chemistry Benchmark Specialist - review AI outputs in your specialty
$61-$77/hr · Remote · USD
Mercor · STEM & research
Atomistic & Surface Modeling Experts (Computational Materials & Catalysis)
$84/hr · Remote · USD
Mercor · STEM & research
Bilingual Arabic STEM Expert (PhD) — AI Safety - review AI outputs in your specialty
$38-$42/hr · Remote · USD
Mercor · STEM & research
Bilingual Chinese STEM Expert (PhD) — AI Safety - review AI outputs in your specialty
$68-$72/hr · Remote · USD
More on this platform
About Mercor
AI-interview-based talent network. One application, voice interview with their AI, then matched to projects across coding, research, and specialist work. Pay scales with track and seniority.
Mercor review: AI-interview talent network
4.1/5 on Glassdoor, fastest-growing platform in the category (+509% YoY). What the AI video interview actually asks, real pay across coding/research/medical/legal/finance tracks ($25-$200/hr), and the project-availability problem.
See all AI training jobs
Browse by category and compare across all eight platforms we cover.