All AI training jobs

Mercor · STEM & research

Applied Mathematics Benchmark Specialist - review AI outputs in your specialty

Listed on Mercor as “Applied Mathematics Benchmark Specialist

$61-$77/hrRemoteContractPaid in USD
ShareWhatsAppTelegramEmail

What this actually is

You design problems that stump current AI models, evaluate AI reasoning against the correct answer, write rubrics, and provide expert feedback. Often the highest-paid category because the expertise pool is small. The platform title (Applied Mathematics Benchmark Specialist) reflects the rate band and the expertise required, not the day-to-day work.

Advertisement

Can you do this on your visa?

F-2 / F-4 / F-5 / F-6: open. E-1 to E-7: needs concurrent-employment permit. D-2 / D-4 students: S-3 permit, 20 hr/week cap. D-10 / D-8: case by case.

Korean tax on USD income

First 5 years in Korea: foreign-source income only taxed if remitted into Korea. After year 5: worldwide income. Full tax guide.

Original posting from Mercor

Role Overview

We are seeking expert mathematicians to author and review high-quality academic assessment content for an AI research initiative. You will write and verify rigorous multiple-choice questions across core mathematics domains, evaluate solution quality, and help establish gold-standard benchmarks used to advance AI capabilities.

You will be assigned one of two task types:

  • Question Authoring - Create original, challenging multiple-choice questions in your area of mathematical expertise, rate their difficulty, and submit them for review.
  • Question Verification - Review pre-written questions for accuracy, clarity, and rigor. Edit where needed, rate difficulty, and document any changes made.

Mathematics Domains Covered

Signal Processing, Financial Mathematics & Actuarial Science, Mathematical Economics, Mathematical Modeling of Ecological & Biological Systems, Mathematical Programming & Combinatorial Optimization, Geomathematics & Climate Modeling.

Key Responsibilities

  • Author original math questions that test deep conceptual understanding, not surface-level recall
  • Ensure questions are unambiguous, self-contained, and precisely defined - all necessary information must be in the problem statement
  • Rate each question's difficulty: Medium (intro undergraduate), Hard (advanced undergraduate), or Expert (post-graduate and above)
  • Provide 1 correct answer and 9 plausible but subtly incorrect alternatives that challenge expert-level solvers
  • Write step-by-step Chain-of-Thought solutions with clear, concise intermediate steps in markdown format
  • Supply 1-5 academic references per question from reputable sources (peer-reviewed journals, university repositories)
  • For verification tasks: flag issues with clarity, completeness, precision, or solvability and justify any edits made

Ideal Qualifications

  • PhD or doctoral candidate in Mathematics, Applied Mathematics, Statistics, or a closely related field
  • Master's degree considered for candidates with exceptional depth in a specific subdomain
  • Strong command of graduate-level mathematical concepts and formal proof writing
  • Experience with rigorous academic problem design or mathematical competition writing is a strong plus
  • Excellent written English and ability to express complex ideas clearly and concisely

More About the Opportunity

  • Expected commitment: 10+ hours/week
  • Asynchronous, fully remote work

Quoted from Mercor’s public listing on 2026-09-08. We don’t edit platform copy; honest framing is in the title and the “what this actually is” block above.

Apply on Mercor