Agent reliability & control plane

Ship AI agents like you ship production software.

Plan, trace, approve, evaluate and roll back every agent run from one control plane — with policy gates, durable traces and multi-provider failover already wired in.

  • Docker compose in one command
  • Approval-gated tool calls
  • Automatic provider failover
99.9%Target uptime
<8sMedian run
2LLM providers
100%Runs traced
Platform

Everything you need to operate agents

From the first prompt to the production rollout — controls, evidence and observability in a single pane of glass.

Approval gates

Human-in-the-loop checkpoints pause every risky side effect until someone signs off.

Durable traces

Each run persists planner, tool and synthesis steps with timing and token usage.

Evaluation suites

Regression datasets score every agent version so changes ship with evidence.

Knowledge base

Postgres full-text retrieval grounds answers in your own documents and tickets.

Versioned agents

Draft, publish and roll back agent versions with changelogs and production pins.

Multi-provider LLMs

Route across Ollama and OmniRoute with automatic retries and failover built in.

Explore

See it before you sign up

Documentation, a 3D architecture map and a live run playground.

Product tour

See the control plane at work

Three surfaces carry most of the weight: a durable trace for every run, a human gate in front of risky side effects, and evaluation scores that decide what ships.

COMPLETEDtriage-agentrun_8f31c2 · staging · v4
4.2 s1,284 tok3 tools
  1. planClassify request → choose tools → draft execution plan
    120 ms214 tok
  2. search_knowledge3 chunks · incident-runbook.md, sso-faq.md
    840 ms402 tok
  3. ticket:writegated by policy “high-risk write” → human approved
    +2m 04s—
  4. synthesizeFinal answer with cited evidence · confidence 0.91
    1.1 s668 tok

Every step persists input, output, latency and token usage — open any run to replay it, line by line.

How it works

From prompt to evidence in four steps

No bespoke glue code: define an agent, queue a run, let policies decide what needs a human, then read the evidence.

  1. 1

    Define an agent

    Instructions, tool allowlist and knowledge scope. Save a draft, iterate freely.

  2. 2

    Trigger a run

    From the Run Lab, CI, a webhook or the API. Runs are queued as durable jobs.

  3. 3

    Gate the side effects

    Policies pause risky tool calls until a human approves — or reject them outright.

  4. 4

    Ship with evidence

    Every step, token and decision is stored, scored and replayable against evals.

POST /api/control/runs
curl -X POST http://localhost:4001/api/control/runs \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer rsk_live_..." \
  -d '{
    "agentId": "clx_agent_42",
    "prompt": "Investigate the authentication incident and use internal evidence.",
    "environment": "staging",
    "trigger": "ci"
  }'

One endpoint away

Everything the dashboard does is available over HTTP with a workspace API key: create runs, read traces, decide approvals, push knowledge and trigger evaluations from CI.

  • Session cookies or rsk_… bearer keys
  • Tenant-scoped by default
  • Full reference in the docs
Open the API reference
Get started

Your control plane is one click away

Create a workspace, provision the sample agents and run your first traced, approval-gated agent run in minutes.