South Asian Software Engineers: Coding Tasks for AI Evaluation
Terac · 100% remote · Contract · Posted
- Pay
- $40/hr
- Location
- South Asia
- Languages
- English
- Hours
- Flexible
- Openings
- Not listed
- Level
- Experienced
Summary
Review proposed coding tasks and environments during an AI-moderated session, providing technical feedback on evaluation harnesses and suggesting improvements to test suites.
What you'll do
- Review the structure, difficulty, and realism of programming challenges.
- Provide technical feedback on evaluation harnesses.
- Suggest improvements to test suites.
- Ensure tasks reflect real-world engineering requirements.
Requirements
- Have experience as a software engineer.
- Possess a strong background in building and reviewing complex codebases.
- Be familiar with evaluation harnesses.
- Have hands-on experience verifying realistic programming tasks.
Skills
- Software Engineering
- Code Review
- Test Automation
- Evaluation Harnesses
Full description
What We're Researching
We're running a paid study on the effectiveness of coding evaluation harnesses used to test AI agents. Our team is building a comprehensive suite of programming environments designed to measure agent performance accurately. This work ensures that AI coding assistants are evaluated against realistic, high-quality development scenarios.
How It Works
You will walk through a series of proposed coding tasks and environments during an AI-moderated session. We will ask you to review the structure, difficulty, and realism of these programming challenges. You will provide technical feedback on the evaluation harnesses and suggest improvements to the test suites. The conversation will focus on ensuring these tasks accurately reflect real-world engineering requirements.
Who This Is For
We are hiring experienced software engineers based in South Asia who have a strong background in building and reviewing complex codebases. We welcome backend developers, test automation engineers, and full-stack engineers who understand evaluation harnesses. Ideal candidates have hands-on experience verifying realistic programming tasks in professional environments.
Similar jobs
Brazil & Argentina Software Engineers: Coding Tasks for AI EvaluationSoftware EngineeringCode ReviewEvaluation HarnessesAI AgentsQuality Assurance Automation+4$40/hr
🌍 Brazil and Argentina
Part-Time Full-Stack Engineers: Farm-Data Platform DevelopmentFull-Stack DevelopmentMachine LearningData InfrastructureAgricultural Technology+3$39/hr
🌍 Worldwide
Senior AI TrainerVideo AnnotationData LabelingVideo Annotation Tools+2$14–36/hr
🌍 Northern America and Germany
AI Safety Experts — English & ThaiAdversarial MLPrompt InjectionRed TeamingCybersecurityConversational AI TestingRLHF/DPO Attacks+5$24–35/hr
🌍 Worldwide