Trainium (NKI) Kernel Expert
Mercor · 100% remote · Contract · Posted
- Pay
- $70–90/hr
- Location
- United States
- Languages
- English
- Hours
- Flexible
- Openings
- Not listed
- Level
- Expert
Summary
Evaluate the quality, correctness, and hardware-appropriateness of Neuron Kernel Interface (NKI) development tasks, assess CUDA-to-NKI migration fidelity, Trainium performance optimization, and numerical correctness, and provide rubric-based written feedback.
What you'll do
- Evaluate NKI development tasks for quality, correctness, and hardware-appropriateness
- Assess CUDA-to-NKI migration fidelity
- Assess Trainium-specific performance optimization quality
- Assess cross-platform numerical correctness standards
- Provide clear, rubric-based written feedback
Requirements
- Have 2+ years of hands-on experience developing or optimizing kernels using NKI targeting AWS Trainium/Inferentia2
- Understand NKI-specific development patterns: tile-based computation, SBUF/PSUM/HBM memory-hierarchy management, partition-dimension constraints, DMA orchestration
- Demonstrate experience assessing CUDA-to-NKI migration quality
- Be familiar with Trainium-specific performance profiling (NeuronCore pipeline utilization, tensor-engine throughput, memory-bandwidth bottlenecks)
- Have experience defining or evaluating cross-platform numerical-correctness standards (GPU vs Trainium accumulation order, rounding behavior, mixed-precision semantics)
Skills
- NKI
- AWS Trainium
- CUDA
- Triton
- Performance Profiling
- Numerical Correctness
- Mixed Precision
- Neuron SDK
Full description
Evaluate the quality, correctness, and hardware-appropriateness of Neuron Kernel Interface (NKI) development tasks used to train and evaluate a frontier AI lab's models. You'll assess CUDA→NKI migration fidelity, Trainium-specific performance-optimization quality, and cross-platform numerical-correctness standards — and provide clear, rubric-based written feedback.
Basic Qualifications • 2+ years of hands-on experience developing or optimizing kernels using the Neuron Kernel Interface (NKI) targeting AWS Trainium/Inferentia2 hardware • Strong understanding of NKI-specific development patterns: tile-based computation, SBUF/PSUM/HBM memory-hierarchy management, partition-dimension constraints, and DMA orchestration • Demonstrated experience assessing CUDA→NKI migration quality • Familiarity with Trainium-specific performance profiling (NeuronCore pipeline utilization, tensor-engine throughput, memory-bandwidth bottlenecks) • Experience defining or evaluating cross-platform numerical-correctness standards (GPU vs Trainium accumulation order, rounding behavior, mixed-precision semantics)
Preferred Qualifications • Direct experience with AWS Neuron SDK, Neuron Compiler internals, or contributions to NKI kernel libraries • Prior CUDA or Triton kernel development • Familiarity with Trainium hardware specifications (NeuronCore-v2 architecture, on-chip SRAM topology, supported data types: FP32/BF16/FP8/INT8) • Experience benchmarking ML training workloads on Trn1/Trn2 instances
Location: open to applicants in United States.
Similar jobs
Staff Software Engineer, Mobile (Android, Kotlin, Full Stack)AndroidKotlinJetpackiOS/SwiftBackend ServicesAPIs+5$50–90/hr
🇺🇸 US
Machine Learning Engineers: Scenario Building for Reinforcement LearningReinforcement LearningSimulation DesignRL FrameworksScenario BuildingPlatform Interfaces+4$90/hr
🌍 Worldwide
AI Developer Trace Task AuditorCursorGitHub CopilotClaude CodeDebuggingFull-Stack DevelopmentBackend Systems+6$70–90/hr
🇺🇸 US
SWE-Bench Task AuditorPythonJavaGoTypeScriptC++SWE-Bench+5$70–90/hr
🇺🇸 US