An older role: first posted 39d ago, and AI expert network still listed it when we checked today. Newer roles tend to fill faster. See the jobs hiring now.
What the work is
Evaluate the quality and correctness of AI-assisted software-development traces used to train and evaluate a frontier AI lab's models. You'll assess end-to-end coding sessions produced with AI-assisted developer tools — judging correctness, workflow soundness, and reasoning — and provide clear, rubric-based written feedback.
Basic Qualifications
- 3+ years professional software development
- Hands-on experience with AI-assisted coding tools and agentic / spec-driven workflows (Cursor, GitHub Copilot, Claude Code, or similar)
- Strong code-reading and debugging skills across full-stack or backend systems
- Ability to evaluate multi-step coding trajectories for correctness and best practice
Preferred Qualifications
- Experience with Kiro or Amazon CodeCatalyst
- Prior work evaluating or grading AI-generated code
- Contributions to developer tooling
Pay
$70–90/hr, fully remote.