AWS Trainium / NKI Kernel Expert
Anyone AI · 100% remote · Contract · Posted
- Pay
- $65/hr
- Location
- Latin America and Europe
- Languages
- English
- Hours
- Part-time, project-based consulting
- Openings
- Not listed
- Level
- Experienced
Summary
Evaluate technical tasks for NKI kernel development on AWS Trainium, reviewing correctness, performance, and idiomatic implementations. Assess CUDA to NKI migrations and provide quality feedback.
What you'll do
- Review NKI kernel correctness and Trainium-specific development patterns
- Evaluate CUDA to NKI kernel migrations
- Assess Trainium performance optimization and benchmarking
- Check memory management across SBUF, PSUM, and HBM
- Verify tile-based computation and DMA scheduling
- Compare numerical correctness across CUDA/Triton and NKI
- Provide technical feedback and quality assessment
Requirements
- Have 2+ years of NKI kernel development or optimization
- Experience with AWS Trainium or Inferentia2 hardware
- Understand tile-based computation and partition dimension constraints
- Know SBUF/PSUM/HBM memory hierarchy and DMA orchestration
- Ability to profile and optimize workloads on Trainium
- Evaluate CUDA to NKI migrations and numerical differences
- Provide clear written feedback on complex implementations
Skills
- NKI
- AWS Trainium
- CUDA
- Triton
- Neuron SDK
- Kernel Optimization
- Memory Hierarchy
- Benchmarking
Full description
Anyone AI is recruiting experienced AWS Trainium / Neuron Kernel Interface (NKI) engineers for a specialized project focused on evaluating and improving kernel development tasks for AI workloads.
We’re looking for engineers with hands-on experience building or optimizing NKI kernels on AWS Trainium or Inferentia2 hardware who understand how Trainium’s architecture differs from traditional GPU programming.
What You’ll Work On
You’ll review and evaluate technical tasks involving:
NKI kernel correctness and Trainium-specific development patterns
CUDA → NKI kernel migrations
Trainium performance optimization and benchmarking
Memory management across SBUF, PSUM, and HBM
Tile-based computation and DMA scheduling
Cross-platform numerical correctness between CUDA/Triton and NKI
Trainium-specific performance bottlenecks and optimization opportunities
Technical feedback and quality assessment of kernel implementations
The work involves determining whether implementations are not only technically correct, but also idiomatic and optimized for Trainium hardware rather than simply translated from GPU-based approaches.
What We’re Looking For
2+ years of hands-on experience developing or optimizing kernels with the Neuron Kernel Interface (NKI)
Experience working with AWS Trainium and/or Inferentia2
Strong understanding of:
Tile-based computation
SBUF / PSUM / HBM memory hierarchy
Partition dimension constraints
DMA orchestration
Trainium-specific optimization techniques
Ability to evaluate CUDA → NKI migrations
Experience profiling and optimizing workloads on Trainium
Understanding of numerical differences across GPU and Trainium backends
Strong ability to analyze complex technical implementations and provide clear written feedback
Nice to Have
Experience with the AWS Neuron SDK or Neuron Compiler
CUDA or Triton kernel development experience
Knowledge of NeuronCore-v2 architecture
Experience with FP32, BF16, FP8, and INT8 workloads
Experience benchmarking workloads on Trn1 or Trn2 instances
Familiarity with nki.language, @nki.jit, or XLA custom calls
Experience with technical evaluation, AI/ML data projects, RLHF, or rubric-based assessment
Engagement
Work Type: Remote
Engagement: Part-time, project-based consulting
Focus: AWS Trainium / NKI kernel engineering and technical evaluation
This is a strong fit for engineers who have worked deeply with AWS Trainium infrastructure and low-level ML kernel optimization and are interested in applying that expertise to technically challenging AI projects.
Location: open to applicants in Argentina, Brazil, Chile, Colombia, Ecuador, Mexico, Portugal, Spain, Uruguay.
Similar jobs
Software Engineers: Paid Code Review for AI Agent EvaluationSoftware EngineeringTest SuitesEvaluation HarnessesCode ReviewBackend DevelopmentFull-Stack Development+7$65/hr
🌍 Worldwide
Competitive CoderBasics C++Competitive Programming+1$45–65/hr
🌍 Worldwide
Senior Software Engineer – Open Source & SWE-Bench EvaluationSoftware EngineeringOpen SourceUnit TestingGitGitHubDebugging+7$65/hr
🌍 Latin America and Europe
Machine Learning Engineer – ML Evaluation & Experiment DesignMachine LearningExperiment DesignModel EvaluationData QualityStatistical TestingHyperparameter Tuning+7$65/hr
🌍 Latin America and Europe