Agent Development Toolkit | Arthur

Stop Guessing. Start Shipping Agents.

An open-source toolkit for building, testing, and monitoring AI agents in production.

How it works

One workflow for the whole agent lifecycle.

Step 1

Manage

Keep prompts versioned, tagged, and promotable across environments. Roll back in seconds when something regresses.

Step 2

Experiment

Test prompt changes, model swaps, and RAG configs against real data before anything ships. Know what changed and why it mattered.

Step 3

Monitor

Trace every agent run end to end. Catch hallucinations, failures, and drift in production before your users do.

Get Started Today

Ship Reliable AI Agents. Fast.

Prompts that behave like code.

Most teams treat prompts like config files — unversioned, untracked, and painful to roll back. One bad change can quietly break production.

Test changes before they reach users.

Swapping a model or tweaking a prompt is a gamble without structured tests. Most teams ship first and find out what broke second.

See exactly what your agent did.

When an agent fails, you shouldn't have to piece together logs and hope for the best.

Know before your users do.

Quality problems in production are invisible until someone complains. By then it's too late.

Works with what you already have.

You shouldn't have to rebuild your stack to get observability.

Agent Framework Eval Platform Arthur Engine + Toolkit
Build and run agents
Prompt versioning & management Basic
Structured A/B experiments
Real-time guardrails (hallucination, PII, injection)
End-to-end trace debugging
Traditional ML model eval
Self-hosted / open-source Varies SaaS MIT licensed

How It Fits Into the Arthur Engine

The Agent Toolkit is part of the Arthur Engine — Arthur's free, open-source AI evaluation and monitoring platform. The Engine provides the foundation: real-time guardrails, LLM eval infrastructure, and flexible deployment. The Toolkit builds on top of that with the full agent development workflow.

Works with every model and framework

Ready to turn your AI into real-world impact?

We'll help you move from pilots and prototypes to production-grade applications, with evaluation every step of the way.