GPU Kernel Engineer – CUDA, Triton & Accelerator Performance
Anyone AI · 100% remote · Contract · Posted
- Pay
- $65/hr
- Location
- Latin America and Europe
- Languages
- English
- Hours
- Part-time, project-based consulting
- Openings
- Not listed
- Level
- Experienced
Summary
Review, debug, and evaluate high-performance GPU and accelerator kernels for AI workloads, assessing correctness, performance, and optimization.
What you'll do
- Review GPU and accelerator kernel implementations for correctness
- Compare outputs against reference implementations and evaluate numerical tolerance thresholds
- Review kernel benchmarks and determine whether comparisons are fair
- Identify performance bottlenecks and optimization opportunities
- Assess whether performance targets are realistic given hardware limits
- Review kernel translations and hardware migrations
- Identify compilation, driver, memory, shape, and runtime issues
- Provide clear, actionable technical feedback
Requirements
- Have 3+ years of hands-on experience developing, optimizing, or debugging GPU or accelerator kernels
- Proficiency in at least two of CUDA, Triton, NKI/AWS Neuron, or Pallas/JAX
- Understand GPU performance optimization and profiling tools such as Nsight, NCU, or roofline analysis
- Understand memory bandwidth, compute throughput, GPU occupancy, shared memory, register pressure, memory coalescing, and bank conflicts
- Understand floating-point numerical correctness and tolerance thresholds
- Debug kernel compilation and runtime issues
- Distinguish software defects, environment problems, and genuine optimization challenges
Skills
- CUDA
- Triton
- NKI
- Pallas
- JAX
- Nsight
- GPU Performance Optimization
- Kernel Profiling
Full description
Anyone AI is recruiting experienced GPU Kernel Engineers for a specialized project focused on reviewing, debugging, and evaluating high-performance compute kernels used in AI workloads.
We’re looking for engineers with hands-on experience writing and optimizing kernels across frameworks such as CUDA, Triton, NKI, or Pallas, with a strong understanding of numerical correctness, GPU performance, memory optimization, and benchmarking.
What You’ll Work On
You’ll work with GPU and accelerator kernel tasks involving:
Kernel implementation and debugging
CUDA and Triton optimization
Translation between kernel frameworks
Hardware migration
Operator fusion
Performance profiling and benchmarking
Numerical correctness verification
Compilation and runtime debugging
Memory hierarchy optimization
Kernel-level AI workload performance
You’ll assess whether implementations are technically correct, efficiently designed, reproducible, and appropriately optimized for the target hardware.
What We’re Looking For
3+ years of hands-on experience developing, optimizing, or debugging GPU or accelerator kernels
Strong experience with at least two of the following:
CUDA
Triton
NKI / AWS Neuron
Pallas / JAX
Strong understanding of GPU performance optimization
Experience with kernel profiling tools such as Nsight, NCU, roofline analysis, or framework-native profilers
Understanding of:
Memory bandwidth
Compute throughput
GPU occupancy
Shared memory
Register pressure
Memory coalescing
Bank conflicts
Strong understanding of floating-point numerical correctness and tolerance thresholds
Experience debugging kernel compilation and runtime issues
Ability to distinguish software defects, environment problems, and genuine optimization challenges
Relevant Experience
Candidates should have experience with several of the following types of work:
Writing kernels from technical specifications
Translating kernels between CUDA, Triton, or other frameworks
Migrating kernels across hardware platforms
Debugging incorrect kernel implementations
Optimizing kernel performance
Fusing multiple operations into optimized kernels
Nice to Have
Experience across both NVIDIA GPU and custom accelerator ecosystems
Experience with AWS Trainium, TPU, JAX, or other accelerators
Compiler engineering experience
Familiarity with MLIR, XLA, or intermediate representation lowering
Contributions to GPU or ML kernel libraries
Experience with cuBLAS, cuDNN, Triton community kernels, or JAX/XLA custom calls
Experience with AI model evaluation, RLHF, or technical benchmark development
What You’ll Be Responsible For
Reviewing GPU and accelerator kernel implementations for correctness
Comparing outputs against reference implementations
Evaluating numerical tolerance thresholds
Reviewing kernel benchmarks and determining whether comparisons are fair
Identifying performance bottlenecks and optimization opportunities
Assessing whether performance targets are realistic given hardware limits
Reviewing kernel translations and hardware migrations
Identifying compilation, driver, memory, shape, and runtime issues
Determining whether technical tasks are genuinely difficult or incorrectly configured
Providing clear, actionable technical feedback
Engagement
Work Type: Remote
Engagement: Part-time, project-based consulting
Focus: GPU kernels, performance engineering, debugging, and technical evaluation
This role is ideal for engineers who enjoy working close to the hardware, optimizing GPU workloads, debugging low-level performance issues, and pushing AI compute systems toward their performance limits.
Location: open to applicants in Argentina, Brazil, Chile, Colombia, Ecuador, Mexico, Portugal, Spain, Uruguay.
Similar jobs
Software Engineers: Paid Code Review for AI Agent EvaluationSoftware EngineeringTest SuitesEvaluation HarnessesCode ReviewBackend DevelopmentFull-Stack Development+7$65/hr
🌍 Worldwide
Competitive CoderBasics C++Competitive Programming+1$45–65/hr
🌍 Worldwide
AWS Trainium / NKI Kernel ExpertNKIAWS TrainiumCUDATritonNeuron SDKKernel Optimization+5$65/hr
🌍 Latin America and Europe
Senior Software Engineer – Open Source & SWE-Bench EvaluationSoftware EngineeringOpen SourceUnit TestingGitGitHubDebugging+7$65/hr
🌍 Latin America and Europe