Brazil & Argentina Software Engineers: Coding Tasks for AI Evaluation
Terac · 100% remote · Contract · Posted
- Pay
- $40/hr
- Location
- Brazil and Argentina
- Languages
- English
- Hours
- Flexible
- Openings
- Not listed
- Level
- Experienced
Summary
Review proposed coding tasks and verify evaluation harness logic for AI agents. Provide technical feedback on task structure and realism during a screen-shared session.
What you'll do
- Review proposed coding tasks
- Assess suitability for AI agent testing
- Verify evaluation harness logic
- Provide feedback on task realism
- Share screen to walk through environments
Requirements
- Be based in Brazil or Argentina
- Have experience building or reviewing evaluation harnesses
- Have strong software engineering background
- Be able to discuss technical details
Skills
- Software Engineering
- Code Review
- Evaluation Harnesses
- AI Agents
- Quality Assurance Automation
Full description
What We're Researching
We're running a paid study on the effectiveness of coding tasks used to evaluate AI agents. Our team is building a comprehensive suite of programming environments designed to test complex software capabilities. We want to ensure these evaluation frameworks are realistic, accurate, and properly calibrated.
How It Works
You will review a series of proposed coding tasks and assess their suitability for testing AI agents. We will ask you to verify the logic of the evaluation harnesses and provide feedback on their realistic application. You will share your screen to walk through the environments and point out potential flaws or improvements. The session involves a mix of code review, technical discussion, and direct feedback on the task structures.
Who This Is For
We are looking for software engineers based in Brazil and Argentina with strong technical backgrounds. You should have direct experience building, reviewing, or testing evaluation harnesses and programming tasks. We welcome backend developers, full-stack engineers, and quality assurance automation specialists.
Similar jobs
South Asian Software Engineers: Coding Tasks for AI EvaluationSoftware EngineeringCode ReviewTest AutomationEvaluation Harnesses+3$40/hr
🌍 South Asia
Part-Time Full-Stack Engineers: Farm-Data Platform DevelopmentFull-Stack DevelopmentMachine LearningData InfrastructureAgricultural Technology+3$39/hr
🌍 Worldwide
Senior AI TrainerVideo AnnotationData LabelingVideo Annotation Tools+2$14–36/hr
🌍 Northern America and Germany
AI Safety Experts — English & ThaiAdversarial MLPrompt InjectionRed TeamingCybersecurityConversational AI TestingRLHF/DPO Attacks+5$24–35/hr
🌍 Worldwide