ML Challenge Task Auditor
Mercor · 100% remote · Contract · Posted
- Pay
- $70–90/hr
- Location
- United States
- Languages
- English
- Hours
- Flexible
- Openings
- Not listed
- Level
- Experienced
Summary
Evaluate applied ML tasks for quality and methodological rigor, and provide rubric-based feedback on experiment design and evaluation methodology.
What you'll do
- Evaluate quality, correctness, and methodological rigor of applied ML tasks
- Assess experiment design, model-selection reasoning, and evaluation methodology
- Provide rubric-based written feedback
Requirements
- 3+ years of applied/experimental ML experience
- Strong grasp of data-quality rigor including leakage detection and metric gaming
- Proficiency with PyTorch, TensorFlow, scikit-learn, XGBoost
- Ability to critique ML claims against evidence and reproduce results
Skills
- PyTorch
- TensorFlow
- Scikit-Learn
- XGBoost
- Experiment Design
- Model Selection
- Evaluation Methodology
- Data-Quality Rigor
Full description
Evaluate the quality, correctness, and methodological rigor of applied machine-learning tasks used to train and evaluate a frontier AI lab's models. You'll assess experiment design, model-selection reasoning, and evaluation methodology — and provide clear, rubric-based written feedback.
Basic Qualifications • 3+ years hands-on applied/experimental ML (experiment design, model selection, hyperparameter tuning, evaluation methodology) • Strong grasp of data-quality rigor: leakage detection, metric gaming, and train/test/CV hygiene • Proficiency with standard ML frameworks (PyTorch, TensorFlow, scikit-learn, XGBoost) • Ability to critique ML claims against evidence and reproduce results
Preferred Qualifications • Competition / benchmark experience (e.g., Kaggle) • Graduate research or publication record in applied ML • Prior task-grading or peer-review experience
Note: this role evaluates applied/experimental ML rigor — it is not an LLM-application-building or MLOps role.
Location: open to applicants in United States.
Similar jobs
Staff Software Engineer, Mobile (Android, Kotlin, Full Stack)AndroidKotlinJetpackiOS/SwiftBackend ServicesAPIs+5$50–90/hr
🇺🇸 US
Machine Learning Engineers: Scenario Building for Reinforcement LearningReinforcement LearningSimulation DesignRL FrameworksScenario BuildingPlatform Interfaces+4$90/hr
🌍 Worldwide
AI Developer Trace Task AuditorCursorGitHub CopilotClaude CodeDebuggingFull-Stack DevelopmentBackend Systems+6$70–90/hr
🇺🇸 US
SWE-Bench Task AuditorPythonJavaGoTypeScriptC++SWE-Bench+5$70–90/hr
🇺🇸 US