SWE-Bench Task Auditor
Mercor · 100% remote · Contract · Posted
- Pay
- $70–90/hr
- Location
- United States
- Languages
- English
- Hours
- Flexible
- Openings
- Not listed
- Level
- Experienced
Summary
Evaluate software-engineering benchmark tasks for quality, correctness, and reproducibility. Provide rubric-based feedback on repository-level tasks, reference patches, and test harnesses.
What you'll do
- Evaluate the quality, correctness, and reproducibility of software-engineering benchmark tasks
- Assess repository-level tasks, reference patches, test harnesses, and grading integrity
- Provide clear, rubric-based written feedback
Requirements
- Have 3+ years of professional software engineering experience
- Provide real open-source contributions or maintainer experience (merged PRs, committer, maintainer)
- Audit reference patches, test runners, and Docker isolation
- Detect answer leakage and reward hacking
- Show fluency in Python and at least one of Java, Go, TypeScript, or C++
Skills
- Python
- Java
- Go
- TypeScript
- C++
- SWE-Bench
- Docker
- Code Review
Full description
Evaluate the quality, correctness, and reproducibility of software-engineering benchmark tasks used to train and evaluate a frontier AI lab's models. You'll assess repository-level tasks, reference patches, test harnesses, and grading integrity — and provide clear, rubric-based written feedback.
Basic Qualifications • 3+ years professional software engineering • Real open-source contribution or maintainer experience (merged PRs, committer / maintainer roles) • Strong ability to audit reference patches, test runners, and Docker isolation, and to detect answer leakage / reward hacking • Fluency across common ecosystems (Python and at least one of Java / Go / TypeScript / C++)
Preferred Qualifications • Familiarity with SWE-Bench (Verified) or similar repository benchmarks • Maintainer history on major Python OSS (Django, Flask, scikit-learn, sympy, pytest, etc.) • Prior code-review or task-grading experience
Location: open to applicants in United States.
Similar jobs
Staff Software Engineer, Mobile (Android, Kotlin, Full Stack)AndroidKotlinJetpackiOS/SwiftBackend ServicesAPIs+5$50–90/hr
🇺🇸 US
Machine Learning Engineers: Scenario Building for Reinforcement LearningReinforcement LearningSimulation DesignRL FrameworksScenario BuildingPlatform Interfaces+4$90/hr
🌍 Worldwide
AI Developer Trace Task AuditorCursorGitHub CopilotClaude CodeDebuggingFull-Stack DevelopmentBackend Systems+6$70–90/hr
🇺🇸 US
ML Challenge Task AuditorPyTorchTensorFlowScikit-LearnXGBoostExperiment DesignModel Selection+6$70–90/hr
🇺🇸 US