AI Developer Trace Task Auditor
Mercor · 100% remote · Contract · Posted
- Pay
- $70–90/hr
- Location
- United States
- Languages
- English
- Hours
- Flexible
- Openings
- Not listed
- Level
- Experienced
Summary
Evaluate the quality and correctness of AI-assisted software development traces, assessing end-to-end coding sessions for correctness, workflow soundness, and reasoning, and provide rubric-based written feedback.
What you'll do
- Evaluate quality and correctness of AI-assisted software development traces
- Assess end-to-end coding sessions produced with AI-assisted developer tools
- Judge correctness, workflow soundness, and reasoning
- Provide clear, rubric-based written feedback
Requirements
- Have 3+ years of professional software development experience
- Use AI-assisted coding tools and agentic or spec-driven workflows
- Read and debug code across full-stack or backend systems
- Evaluate multi-step coding trajectories for correctness and best practice
Skills
- Cursor
- GitHub Copilot
- Claude Code
- Debugging
- Full-Stack Development
- Backend Systems
- Agentic Workflows
- Spec-Driven Development
Full description
Evaluate the quality and correctness of AI-assisted software-development traces used to train and evaluate a frontier AI lab's models. You'll assess end-to-end coding sessions produced with AI-assisted developer tools — judging correctness, workflow soundness, and reasoning — and provide clear, rubric-based written feedback.
Basic Qualifications • 3+ years professional software development • Hands-on experience with AI-assisted coding tools and agentic / spec-driven workflows (Cursor, GitHub Copilot, Claude Code, or similar) • Strong code-reading and debugging skills across full-stack or backend systems • Ability to evaluate multi-step coding trajectories for correctness and best practice
Preferred Qualifications • Experience with Kiro or Amazon CodeCatalyst • Prior work evaluating or grading AI-generated code • Contributions to developer tooling
Location: open to applicants in United States.
Similar jobs
Staff Software Engineer, Mobile (Android, Kotlin, Full Stack)AndroidKotlinJetpackiOS/SwiftBackend ServicesAPIs+5$50–90/hr
🇺🇸 US
Machine Learning Engineers: Scenario Building for Reinforcement LearningReinforcement LearningSimulation DesignRL FrameworksScenario BuildingPlatform Interfaces+4$90/hr
🌍 Worldwide
SWE-Bench Task AuditorPythonJavaGoTypeScriptC++SWE-Bench+5$70–90/hr
🇺🇸 US
ML Challenge Task AuditorPyTorchTensorFlowScikit-LearnXGBoostExperiment DesignModel Selection+6$70–90/hr
🇺🇸 US