Skip to content
Clera logo

Research Engineer, Benchmarks

OCaml, Linux, Python +2 Tags

**About The Role** Join a small, technically elite team --- including International Olympiad medalists and published AI researchers --- building high-quality benchmarks to evaluate frontier AI agents on realistic, domain-specific workflows. As a **Research Engineer, Benchmarks** , you'll own the design and implementation of evaluations that frontier labs and enterprise customers trust. This role is central to ensuring our benchmarks are rigorous, credible, and tightly aligned with real-world agent performance. This is an on-site role based in **San Francisco, CA** . Visa sponsorship is available. What You'll Do * Design, implement, and own the quality of internal benchmarks for evaluating frontier agents on domain-specific tasks. * Partner with subject-matter experts to define realistic workflows and tasks for domain-specific evaluations. * Build reliable infrastructure to run models and agents against benchmark tasks at scale. * Develop metrics and analyses that measure benchmark difficulty, reliability, and failure modes. * Validate that benchmark performance correlates with real-world evaluations, customer needs, and frontier lab expectations. * Write clear documentation and benchmark reports that make results legible and credible to technical audiences. What We're Looking For **Required** * 2--4 years of experience in software engineering, ML engineering, or research roles. * Strong proficiency in Python, Docker, and Linux environments. * Experience building environments, evaluations, or benchmarks for AI systems. * Published research or technical writing on topics such as public benchmarks, model failure modes, or evaluation methodology. * Deep understanding of what makes a benchmark realistic, reliable, and practically useful. * Curiosity and genuine ability to understand how real-world workflows operate across diverse domains. * Strong attention to detail --- a habit of spotting subtle inconsistencies and edge cases in task design. * Ability to reason from first principles about task design, scoring, and failure modes. * Comfort thriving in unstructured problem spaces and working independently in fast-paced, early-stage environments. * Excellent communication skills for collaborating across time zones and with technical teams. **Compensation \& Benefits** * Salary: $150,000 -- $250,000 USD annually, depending on experience. * Visa sponsorship available. * Opportunity for significant early-stage equity and career growth within a high-impact, research-driven team. Location This is a full-time, **on-site** position in **San Francisco, CA** . Candidates must be willing and able to work in-office.