# RunAgain: the self-improving loop for AI agents

> Trace. Experiment. Test. Evaluate. Improve. Monitor. Run again.

RunAgain (runagain.ai) is an observability, testing and evaluation platform for AI agents, currently in private beta. Book a demo: https://cal.com/tamas-szuromi/30min · Get in touch: tamas@runagain.ai

## The feedback loop

1. **Trace**: every span, tool call and token, captured live.
2. **Experiment**: fork prompts, models and tools behind flags.
3. **Test**: deterministic runs against mocks and sims.
4. **Evaluate**: LLM-judge, rubric, assertion and dataset scoring.
5. **Improve**: diff regressions, fix the weak spans, ship.
6. **Run again**: the loop closes; every pass raises the score.

## Works with

Vercel AI SDK, Vercel Eve, Claude Agent SDK, MCP, REST API, and most agent SDKs (OpenAI Agents SDK, LangChain, LangGraph, LlamaIndex, CrewAI, AutoGen, Mastra, Pydantic AI, Semantic Kernel, smolagents, DSPy, Haystack, Google ADK, Cloudflare Agents, Agno, any OpenTelemetry SDK).

## Explore

- [Solutions](https://runagain.ai/solutions.md): nine capability pages
- [Use cases](https://runagain.ai/use-cases.md): twelve task-level guides
- [Eval catalog](https://runagain.ai/evals.md): all 47 supported evals, scorers and judges
- [For agents](https://runagain.ai/for-agents.md): what you can do here as an agent
- [Docs](https://docs.runagain.ai/)

---

Markdown mirror of https://runagain.ai/ for agents and LLMs. Append .md to any runagain.ai page URL for its markdown twin. Overview: https://runagain.ai/llms.txt · For agents: https://runagain.ai/for-agents.md
