GPU Kernel Expert
Mercor · 100% remote · Contract · Posted
- Pay
- $70–90/hr
- Location
- United States
- Languages
- English
- Hours
- Flexible
- Openings
- Not listed
- Level
- Expert
Summary
Evaluate the quality, correctness, and completeness of GPU/accelerator kernel development tasks for training AI models, and provide rubric-based written feedback.
What you'll do
- Evaluate the quality, correctness, and completeness of GPU/accelerator kernel development tasks
- Assess numerical correctness and performance-benchmarking fairness
- Assess task scoping and compilation/runtime validity
- Provide rubric-based written feedback
Requirements
- 3+ years of hands-on GPU/accelerator kernel development
- Experience in at least two of CUDA, Triton, NKI, or Pallas
- Understanding of numerical-correctness criteria (absolute/relative/ULP tolerances)
- Experience with performance profiling and benchmarking (nsight, ncu, roofline)
- Familiarity with compilation and runtime failure modes
- Experience with at least three kernel task types (generation, translation, migration, debugging, performance optimization, operator fusion)
Skills
- CUDA
- Triton
- NKI
- Pallas
- Performance Profiling
- Numerical Correctness
Full description
Evaluate the quality, correctness, and completeness of GPU/accelerator kernel development tasks used to train and evaluate a frontier AI lab's models. You'll assess numerical correctness, performance-benchmarking fairness, task scoping, and compilation/runtime validity across diverse kernel task types — and provide clear, rubric-based written feedback.
Basic Qualifications • 3+ years of hands-on experience developing, optimizing, or verifying GPU/accelerator kernels in at least two of: CUDA, Triton, NKI, or Pallas (JAX) • Strong understanding of numerical-correctness criteria for kernels (absolute/relative/ULP tolerances, reference-implementation selection) • Demonstrated experience with performance profiling and benchmarking (nsight, ncu, roofline analysis, or framework-native profilers) • Familiarity with common compilation and runtime failure modes (driver mismatches, OOM, launch-configuration errors, shape/stride mismatches, autotuning failures) • Experience with at least three kernel task types: generation from specification, translation/lowering across frameworks, migration between hardware targets, debugging, performance optimization, or operator fusion
Preferred Qualifications • Experience across both NVIDIA GPU (CUDA/Triton) and custom-accelerator (NKI/Pallas/TPU) ecosystems • Background in compiler engineering, MLIR, or intermediate-representation lowering • Understanding of memory-hierarchy optimization (shared-memory tiling, register pressure, bank conflicts, coalescing patterns) • Contributions to kernel libraries (cuBLAS, cuDNN, Triton community kernels, JAX/XLA custom calls)
Location: open to applicants in United States.
Similar jobs
Staff Software Engineer, Mobile (Android, Kotlin, Full Stack)AndroidKotlinJetpackiOS/SwiftBackend ServicesAPIs+5$50–90/hr
🇺🇸 US
Machine Learning Engineers: Scenario Building for Reinforcement LearningReinforcement LearningSimulation DesignRL FrameworksScenario BuildingPlatform Interfaces+4$90/hr
🌍 Worldwide
AI Developer Trace Task AuditorCursorGitHub CopilotClaude CodeDebuggingFull-Stack DevelopmentBackend Systems+6$70–90/hr
🇺🇸 US
SWE-Bench Task AuditorPythonJavaGoTypeScriptC++SWE-Bench+5$70–90/hr
🇺🇸 US