B

AI Evaluation Infrastructure Engineer

icon building Company : Block
icon briefcase Job Type : Full Time
icon remote-alt Remote / Work from Home

Job Description - AI Evaluation Infrastructure Engineer

It all started with an idea at Block in 2013. Initially built to take the pain out of peer-to-peer payments, Cash App has gone from a simple product with a single purpose to a dynamic ecosystem, developing unique financial products, including Afterpay/Clearpay, to provide a better way to send, spend, invest, borrow and save to our 50+ million monthly active customers. We want to redefine the world’s relationship with money to make it more relatable, instantly available, and universally accessible.

Today, Cash App has thousands of employees working globally across office and remote locations, with a culture geared toward innovation, collaboration and impact. We’ve been a distributed team since day one, and many of our roles can be done remotely from the countries where Cash App operates. No matter the location, we tailor our experience to ensure our employees are creative, productive, and happy.


The Role


We build AI products, and the quality of our evaluations sets the ceiling for how good those products can be. The speed of our evaluations determines how quickly we can improve them.


We are looking for an engineer to build the infrastructure and tooling that make high-quality AI evaluation possible at Block's scale. You will help teams understand whether a model or product change is actually better, whether a result is statistically meaningful, and whether offline evaluation is predicting what happens with real users.


Our evaluation approach combines offline evals that encode our definition of a good response, online evals that show how people actually respond, and a feedback loop that keeps the two converging. Your work will turn that approach into systems that product teams can use quickly, reliably, and with confidence.


This is a high-impact, early-stage area with broad surface area. You will help decide what to build first, then build the platform that helps teams ship better AI products faster.


You Will



  • Build an execution engine that can score candidate versions against task sets in minutes, not hours.

  • Create task set tooling that samples from production logs and validates tasks before they are admitted into an evaluation set.

  • Build grader infrastructure across ground truth checks, rubrics, and LLM-as-judge approaches.

  • Develop tooling that helps human reviewers calibrate judges, measure judge-to-human agreement, and monitor drift over time.

  • Build leaderboards and reporting systems that include sample size, confidence intervals, and run-to-run variance, so teams can distinguish real improvements from noise.

  • Support in-product side-by-side serving, feedback capture, and implicit signal extraction from real conversations.

  • Build the loop that compares offline scores with online outcomes, identifies eval sets that have stopped predicting reality, and helps teams improve them.

  • Partner with product, engineering, data, and ML teams to make evaluation workflows fast enough and trustworthy enough to become part of everyday development.


You Have



  • Experience building production platforms or infrastructure, including distributed batch execution, data pipelines, or systems that process production logs.

  • Strong statistical literacy, including comfort with confidence intervals, variance, power, and multiple comparisons.

  • The judgment to identify when a result is meaningful and when it is noise.

  • Experience evaluating LLM or ML systems, or deep systems engineering experience with a strong interest in AI evaluation.

  • Product instinct for internal tools. You understand that leaderboards, annotation tools, and workflows only matter if teams actually use them.

  • A bias toward building reliable, observable systems that other engineers can trust.

  • Strong collaboration skills and the ability to work across ambiguous product, data, and engineering problems.


Success in your first year



  • Product teams can stand up credible evals for new AI surfaces in days.

  • Eval results gate CI and run quickly enough that engineers do not route around them.

  • Judge-to-human agreement is measured, published, and monitored for drift.

  • Launch decisions are not made on results that are within statistical noise.

  • At least one case is documented where online reality disagreed with offline scores, and the evaluation was improved as a result.


Why this matters


AI product development moves quickly, but speed only helps when teams can trust the signal they are using to make decisions. This role will build the systems that make those signals faster, more accurate, and more actionable. Your work will directly influence how Block evaluates, improves, and ships AI products.

Original job AI Evaluation Infrastructure Engineer posted on GrabJobs ©. To flag any issues with this job please use the Report Job button on GrabJobs.
Share Job
Share Job

Similar AI Evaluation Infrastructure Engineer Jobs in the US

GrabJobs is the no1 job portal in the US, connecting you to thousands of jobs fast! Find the best jobs in the US, apply in 1 click and get a job today!

Mobile Apps

Copyright © 2026 Grabjobs Pte.Ltd. All Rights Reserved.