Skip to content
AI Trainer Jobs
Companies

Senior Software Engineer – Open Source & SWE-Bench Evaluation

Anyone AI · 100% remote · Contract · Posted

Pay
$65/hr
Location
Latin America and Europe
Languages
English
Hours
Part-time, project-based consulting
Openings
Not listed
Level
Experienced
Apply for this position

Summary

Review and evaluate real-world software engineering tasks from open-source repositories, assessing technical soundness, reproducibility, test coverage, and clarity. Provide recommendations on whether tasks should be accepted, improved, or excluded.

What you'll do

  • Review coding tasks derived from real GitHub issues and pull requests
  • Assess problem statements and success criteria for clarity and completeness
  • Evaluate unit tests for correctness, coverage, and robustness
  • Identify flaky tests, missing dependencies, version conflicts, and environment issues
  • Determine whether tasks can be reliably reproduced across environments
  • Assess the real-world difficulty and complexity of each task
  • Provide clear recommendations on task acceptance, improvement, or exclusion

Requirements

  • Have 3+ years of professional software engineering experience
  • Work with large, multi-file codebases
  • Experience reviewing pull requests and debugging issues
  • Strong understanding of unit testing and test coverage
  • Ability to evaluate tests for correctness without restricting implementation
  • Experience with dependency management, environment setup, and reproducibility
  • Strong understanding of Git and GitHub workflows
  • Ability to analyze complex technical problems and provide clear written feedback

Skills

  • Software Engineering
  • Open Source
  • Unit Testing
  • Git
  • GitHub
  • Debugging
  • Dependency Management
  • Code Review

Full description

Anyone AI is recruiting experienced Software Engineers for a specialized project focused on reviewing and evaluating real-world software engineering tasks derived from open-source repositories.

The work involves assessing whether coding tasks based on real GitHub issues and pull requests are technically sound, reproducible, appropriately tested, and representative of the kinds of problems professional software engineers solve every day.

What You’ll Work On

You’ll review software engineering tasks involving:

  • Real-world bug fixes and feature implementations

  • Open-source repositories and pull requests

  • Unit tests and test coverage

  • Repository setup and dependency management

  • Reproducibility and environment configuration

  • Task difficulty and complexity

  • Multi-file and cross-module code changes

  • Technical feedback and quality assessment

You’ll determine whether tasks are clearly specified, technically solvable, supported by sufficient tests, and free from issues such as flaky tests, missing dependencies, ambiguous requirements, or environment-specific behavior.

What We’re Looking For

  • 3+ years of professional software engineering experience

  • Strong experience working with large, multi-file codebases

  • Experience reviewing pull requests, debugging issues, and maintaining production code

  • Strong understanding of unit testing and test coverage

  • Ability to evaluate whether tests correctly validate a solution without unnecessarily restricting implementation approaches

  • Experience with dependency management, environment setup, and reproducibility

  • Strong understanding of Git and GitHub-based development workflows

  • Ability to analyze complex technical problems and provide clear written feedback

Nice to Have

  • Contributions to or maintenance of open-source projects

  • Experience with SWE-Bench, SWE-Bench Verified, or similar coding benchmarks

  • Experience with major Python open-source projects such as Django, Flask, scikit-learn, SymPy, matplotlib, requests, or pytest

  • Experience with Docker, CI/CD, pip, conda, or dependency pinning

  • Knowledge of test fixtures, test isolation, or property-based testing

  • Experience designing technical assessments or reviewing coding challenges

  • Experience with AI/ML evaluation, data curation, RLHF, or benchmark development

What You’ll Be Responsible For

  • Reviewing coding tasks derived from real GitHub issues and pull requests

  • Assessing whether problem statements and success criteria are clear and complete

  • Evaluating unit tests for correctness, coverage, and robustness

  • Identifying flaky tests, missing dependencies, version conflicts, and environment issues

  • Determining whether tasks can be reliably reproduced across environments

  • Assessing the real-world difficulty and complexity of each task

  • Providing clear recommendations on whether tasks should be accepted, improved, or excluded

Engagement

Work Type: Remote
Engagement: Part-time, project-based consulting
Focus: Software engineering, open-source code review, testing, and technical evaluation

This role is a strong fit for experienced engineers who enjoy debugging complex codebases, reviewing pull requests, working with open-source software, and evaluating what makes a software engineering problem well designed.

Location: open to applicants in Argentina, Brazil, Chile, Colombia, Ecuador, Mexico, Portugal, Spain, Uruguay.

Similar jobs

View all jobs