Want to learn evals? buildevals.com →

// CAREERS

Build the loop with us.

RunAgain is the self-improving loop for AI agents: trace what happened, evaluate what changed, replay it exactly as it ran, and turn every issue into a test. We are early, the surface area is enormous, and everyone here owns their work end to end.

// OPEN ROLES
Every open role lives on Wellfound.

Roles, levels and compensation are kept up to date there. Apply in one click.

View open roles →

What you would work on

01
Ingestion and traces

OpenTelemetry pipelines that swallow every span, tool call and token from production agents, then make them searchable, diffable and replayable at volume.

02
Evals and simulation

Mocked environments, deterministic replays, LLM-as-judge and rubric scoring, dataset regression: the machinery that says whether a change actually made the agent better.

03
Product surface

The app people live in all day: dashboards, trace views, the SDKs and the developer experience around them. Fast, dense, and pleasant to use.

Nothing open that fits?

Write anyway. Tell us what you would build and point at something you have shipped: tamas@runagain.ai. Good people get remembered when a role opens.

wellfound.com/company/runagain-ai →