AI Evaluators: Assessing a Shopping Assistant
Terac · 100% remote · Contract · Posted
- Pay
- $30/hr
- Location
- Worldwide
- Languages
- English
- Hours
- 20+ hours per week
- Openings
- Not listed
- Level
- Experienced
Summary
Review user interaction traces with a shopping assistant to pinpoint failures and create structured rubrics for evaluating response quality.
What you'll do
- Review real interaction traces between users and the shopping assistant
- Identify failures, logical errors, or unhelpful product recommendations
- Create structured rubrics and verifiers to judge response quality
Requirements
- Have prior experience in prompt engineering, complex data annotation, or software testing
- Analyze text interactions deeply
- Build structured evaluation frameworks from scratch
- Possess a strong eye for detail
Skills
- Prompt Engineering
- Data Annotation
- Software Testing
- E-Commerce
- Quality Assurance
Full description
What We're Researching
We're hiring AI evaluators to assess the accuracy and helpfulness of a new digital shopping assistant. This project focuses on understanding how well the system handles real-world e-commerce queries and where it falls short in its logic. Your analysis will directly feed into improving the underlying model and its response quality.
How It Works
You will review real interaction traces between users and the shopping assistant within our custom platform. As you analyze these conversations, you will pinpoint specific failures, logical errors, or unhelpful product recommendations. From there, you will create structured rubrics and verifiers to consistently judge future response quality. This is an ongoing remote engagement requiring 20+ hours per week.
Who This Is For
This opportunity is ideal for quality assurance specialists, AI data evaluators, and e-commerce professionals with a strong eye for detail. We welcome applicants with prior experience in prompt engineering, complex data annotation, or software testing. You should be comfortable analyzing text interactions deeply and building structured evaluation frameworks from scratch.
Similar jobs
Software Engineer, Full Stack — IndiaPythonJavaRustC#C++TypeScript+5$25–30/hr
🇮🇳 India
AI Safety Experts — English & ThaiAdversarial MLPrompt InjectionRed TeamingCybersecurityConversational AI TestingRLHF/DPO Attacks+5$24–35/hr
🌍 Worldwide
AI Safety Experts — English & IndonesianRed TeamingAdversarial MLPrompt InjectionCybersecurityPenetration TestingConversational AI+4$17–25/hr
🌍 Worldwide
AI Safety Experts — English & VietnameseAI Red TeamingPrompt InjectionCybersecurityAdversarial Machine LearningConversational AI Testing+4$17–25/hr
🌍 Worldwide