All AI training jobs

Mercor · Finance & specialist

GPU Kernel Expert - review AI outputs in your specialty

Listed on Mercor as “GPU Kernel Expert

$70-$90/hrRemoteContractPaid in USD
ShareWhatsAppTelegramEmail

What this actually is

You bring your specialist expertise to AI evaluation. The shape of the work varies but the pattern is the same: review outputs, rate quality, write prompts, flag errors. The platform title (GPU Kernel Expert) reflects the rate band and the expertise required, not the day-to-day work.

Advertisement

Can you do this on your visa?

F-2 / F-4 / F-5 / F-6: open. E-1 to E-7: needs concurrent-employment permit. D-2 / D-4 students: S-3 permit, 20 hr/week cap. D-10 / D-8: case by case.

Korean tax on USD income

First 5 years in Korea: foreign-source income only taxed if remitted into Korea. After year 5: worldwide income. Full tax guide.

Original posting from Mercor

Evaluate the quality, correctness, and completeness of GPU/accelerator kernel development tasks used to train and evaluate a frontier AI lab's models. You'll assess numerical correctness, performance-benchmarking fairness, task scoping, and compilation/runtime validity across diverse kernel task types - and provide clear, rubric-based written feedback.

Basic Qualifications

• 3+ years of hands-on experience developing, optimizing, or verifying GPU/accelerator kernels in at least two of: CUDA, Triton, NKI, or Pallas (JAX)

• Strong understanding of numerical-correctness criteria for kernels (absolute/relative/ULP tolerances, reference-implementation selection)

• Demonstrated experience with performance profiling and benchmarking (nsight, ncu, roofline analysis, or framework-native profilers)

• Familiarity with common compilation and runtime failure modes (driver mismatches, OOM, launch-configuration errors, shape/stride mismatches, autotuning failures)

• Experience with at least three kernel task types: generation from specification, translation/lowering across frameworks, migration between hardware targets, debugging, performance optimization, or operator fusion

Preferred Qualifications

• Experience across both NVIDIA GPU (CUDA/Triton) and custom-accelerator (NKI/Pallas/TPU) ecosystems

• Background in compiler engineering, MLIR, or intermediate-representation lowering

• Understanding of memory-hierarchy optimization (shared-memory tiling, register pressure, bank conflicts, coalescing patterns)

• Contributions to kernel libraries (cuBLAS, cuDNN, Triton community kernels, JAX/XLA custom calls)

Quoted from Mercor’s public listing on 2026-09-08. We don’t edit platform copy; honest framing is in the title and the “what this actually is” block above.

Apply on Mercor