Skip to content
AI Trainer Jobs
Companies

Trainium (NKI) Kernel Expert

Mercor · 100% remote · Contract · Posted

Pay
$70–90/hr
Location
United States
Languages
English
Hours
Flexible
Openings
Not listed
Level
Expert
Apply for this position

Summary

Evaluate the quality, correctness, and hardware-appropriateness of Neuron Kernel Interface (NKI) development tasks, assess CUDA-to-NKI migration fidelity, Trainium performance optimization, and numerical correctness, and provide rubric-based written feedback.

What you'll do

  • Evaluate NKI development tasks for quality, correctness, and hardware-appropriateness
  • Assess CUDA-to-NKI migration fidelity
  • Assess Trainium-specific performance optimization quality
  • Assess cross-platform numerical correctness standards
  • Provide clear, rubric-based written feedback

Requirements

  • Have 2+ years of hands-on experience developing or optimizing kernels using NKI targeting AWS Trainium/Inferentia2
  • Understand NKI-specific development patterns: tile-based computation, SBUF/PSUM/HBM memory-hierarchy management, partition-dimension constraints, DMA orchestration
  • Demonstrate experience assessing CUDA-to-NKI migration quality
  • Be familiar with Trainium-specific performance profiling (NeuronCore pipeline utilization, tensor-engine throughput, memory-bandwidth bottlenecks)
  • Have experience defining or evaluating cross-platform numerical-correctness standards (GPU vs Trainium accumulation order, rounding behavior, mixed-precision semantics)

Skills

  • NKI
  • AWS Trainium
  • CUDA
  • Triton
  • Performance Profiling
  • Numerical Correctness
  • Mixed Precision
  • Neuron SDK

Full description

Evaluate the quality, correctness, and hardware-appropriateness of Neuron Kernel Interface (NKI) development tasks used to train and evaluate a frontier AI lab's models. You'll assess CUDA→NKI migration fidelity, Trainium-specific performance-optimization quality, and cross-platform numerical-correctness standards — and provide clear, rubric-based written feedback.

Basic Qualifications • 2+ years of hands-on experience developing or optimizing kernels using the Neuron Kernel Interface (NKI) targeting AWS Trainium/Inferentia2 hardware • Strong understanding of NKI-specific development patterns: tile-based computation, SBUF/PSUM/HBM memory-hierarchy management, partition-dimension constraints, and DMA orchestration • Demonstrated experience assessing CUDA→NKI migration quality • Familiarity with Trainium-specific performance profiling (NeuronCore pipeline utilization, tensor-engine throughput, memory-bandwidth bottlenecks) • Experience defining or evaluating cross-platform numerical-correctness standards (GPU vs Trainium accumulation order, rounding behavior, mixed-precision semantics)

Preferred Qualifications • Direct experience with AWS Neuron SDK, Neuron Compiler internals, or contributions to NKI kernel libraries • Prior CUDA or Triton kernel development • Familiarity with Trainium hardware specifications (NeuronCore-v2 architecture, on-chip SRAM topology, supported data types: FP32/BF16/FP8/INT8) • Experience benchmarking ML training workloads on Trn1/Trn2 instances

Location: open to applicants in United States.

Similar jobs

View all jobs