Skip to content
AI Trainer Jobs
Companies

Agent Engineer

Mercor · 100% remote · Contract · Posted

Pay
$100–500/task
Location
Worldwide
Languages
English
Hours
Flexible
Openings
Not listed
Level
Expert
Apply for this position

Summary

Build and operate production LLM agents, focusing on reliability, evaluation, and adoption. Share war stories and tradeoffs from real systems.

What you'll do

  • Build and operate LLM agents in production
  • Evaluate agent performance after changes
  • Run scheduled and long-running agent jobs
  • Integrate agents with company data, tools, and MCP surfaces
  • Monitor agent adoption and usage

Requirements

  • Have shipped a production agent that real users depended on
  • Know how to evaluate whether an agent improved after a change
  • Have experience running agents that outlive a single request
  • Understand how agents are adopted and used inside an organization

Skills

  • LLM Agents
  • Agent Evaluation
  • Production Systems
  • Cloud Sandboxes
  • MCP
  • Company Data Integration

Full description

We are looking for engineers who build and operate LLM agents in production, and who have real visibility into how agents are actually used inside a company.

You have probably:

  • Shipped an agent that real users depended on, and been on the hook when it broke.

  • Figured out how to tell whether an agent got better or worse after a change.

  • Run agents that outlive a single request: scheduled jobs, long-running work, cloud sandboxes.

  • Watched your org build an internal assistant, and seen who adopted it and who quietly did not.

We are especially interested in the layers most people do not talk about: internal monoagents wired into company data, shared company memory, reusable skills and playbooks, the tool and MCP surfaces agents call, and how anyone sees what agents did and what they cost.

Applying starts with a short conversational AI interview. No coding, no take-home. We want to hear how you actually think about agent reliability, evaluation, and adoption, and the tradeoffs you have made in real systems. Bring war stories. The messier and more specific, the better.

If that screen stands out, we will invite you to a live 30 minute conversation with our team. We pay $100 to $500 for that conversation, paid on completion of the call, with the amount depending on depth of experience.

If that sounds like you, apply and complete the screen. We review every submission.

Similar jobs

View all jobs