Trust infrastructure for autonomous AI agents

Connect agents to tools.Safely and reliably.

Find an MCP server and review its automatic pre-check before connecting your account. Then discover useful agents with AI, approve their actions, and follow every run.

Already running an agent? Observe an MCP workflow. Recommended onboarding: run the customer-controlled adapter in passive observe mode, learn from counterfactual findings, then choose when to enforce.

  • Apache-2.0
  • Model-agnostic
  • MCP-aware
  • Fail-closed controls
decision.assessassurance active
intentissue customer refund
decisionapprove $750 refund
evidenceorder verified · 8y tenure
policy>$500 requires approval
CHALLENGE

Decision is plausible. Manager approval is still required.

decision_tracetrace_01JAA4…A91C

For MCP developers · Preview

Building an MCP server?

Check protocol discovery, tool descriptions, and declared schemas. Get actionable findings with versioned spec guidance, without signing in. Publish your report to make those observations available in AgentAction discovery.

Anonymous metadata inspection. No credentials or tool execution. Runtime behavior remains untested.

Check your MCP server

01 / The gap

Useful agents need more than a connection.

Know whether your agent chose the right action, whether it was allowed, and whether it achieved the job. AgentAction connects policies, approvals, and outcome evidence—without inspecting hidden chain-of-thought.

02 / Beyond IAM

Access is necessary. It is not assurance.

AgentAction complements identity providers, policy engines, and API gateways. It adds the decision and evidence controls autonomous systems require at the moment of action.

Traditional IAMAutonomous agent systems
User identityAgent identity and workload context
Role permissionsContextual, action-specific authorization
API accessAutonomous decision assurance
Activity logsVerifiable receipts and observations
Periodic auditsContinuous, versioned assurance

03 / The trust lifecycle

Assure the decision before authorizing the action.

Before an agent acts, assess whether the decision is justified. When it acts, ensure the action is authorized. Afterward, preserve proof of what happened and evaluate the result.

  1. 01

    Declare intent

    The agent declares the goal, proposed action, and relevant context.

  2. 02

    Assure the decision

    Decision evidence is checked for justification, uncertainty, and safer alternatives.

  3. 03

    Enforce policy

    Identity, policy, approvals, and durable state return allow, deny, or challenge.

  4. 04

    Execute

    The provider verifies the exact-action authority and applies its own rules.

  5. 05

    Preserve evidence

    Receipts and verified observations record what was authorized and what occurred.

  6. 06

    Evaluate continuously

    Immutable assessments feed assurance signals across runs, profiles, and versions.

Trusted action boundary
intent → assessed → authorized → executed → evidenced → evaluated

04 / The platform

Build, evaluate, and govern your agents.

Turn a job into an agent you can put to work. Start with a recipe, evaluate it against your intent, then bring decision assurance, action authorization, and ongoing visibility into every run.

01Available now

Agent Creation & Evaluation

Find servers by capability, review public authentication and tool risk signals, and build a supervised agent. Test it against your intent and track outcomes across runs.

Discover → inspect → connect → supervise

Public findings are not safety certification; authenticated tools may remain unseen. OAuth discovery is available; OAuth login is not yet supported.

Open My agents
02Available now

Decision Assurance

Assess the declared basis for a consequential choice: policy factors, alternatives, assumptions, uncertainty, and supporting evidence.

Decision evidence → allow · deny · challenge

Uses normalized decision evidence—not private chain-of-thought or hidden model reasoning.

03Available now

Action Authorization

Gate the exact tool call against policy and prior state, then issue action-bound authority that providers can verify.

Tool call → policy → signed receipt

Deploy locally today or connect the decision and evidence service behind a gateway.

05 / Deploy the boundary

One governed endpoint for enterprise AI.

Route model and tool traffic through company policy. Use the least expensive qualified model, stop unsafe or duplicate actions, and preserve evidence from selection through execution.

Explore the Gateway
01

Route intelligently

Select an approved, cost-effective model for the task, risk, privacy boundary, and company policy.

02

Apply corporate policy

Control models, providers, tools, destinations, sensitive data, budgets, and versioned system context.

03

Execute once

Deduplicate consequential actions and replay the original provider result for an identical safe retry.

04

Prove what happened

Correlate routing, authorization, approval, execution, replay, observation, and assessment evidence.

Available now

Action decisions, approvals, data-flow controls, safe replay, signed receipts, and MCP reference integration.

Product direction

Production gateway packaging, risk-aware inference routing, managed company context, and broader protocol coverage.

06 / Trust model

The agent never becomes its own authority.

Security facts come from authenticated systems and durable state, not from conversation text or agent-editable memory.

SignalPostureBoundary rule
Agent outputUntrusted proposalNever treated as authority by itself
Identity and job contextVerified inputDerived by the trusted runtime or gateway
Policy and prior stateEnforcement inputHeld outside prompts and agent-editable memory
Authorization receiptPortable evidenceBound to the exact action, audience, and decision
Execution and outcomeIndependent evidenceKept distinct from the authorization decision

07 / Inspect the evidence

See how agent outcomes hold up across runs.

The read-only console keeps unlike intent-profile versions separate and makes failed, partial, indeterminate, and low-confidence outcomes visible instead of blending them into one score.

AgentAction public observability console showing synthetic support-refund outcome, constraint, confidence, execution, and data-quality metrics
Public demo · synthetic fixturesExplore the live console
Public · synthetic data

Explore the observability console

Inspect Fleet Overview, filter finalized Jobs, and open a deterministic evidence timeline without signing in. The demo is isolated from production services and contains no customer data.

  • Profile-scoped outcome and constraint rollups
  • Finalized Jobs with confidence and discipline signals
  • Job detail with immutable digests and evidence timeline
Open the public demo
Operator access

Keep real tenant evidence protected

The production operator console remains behind Cloudflare Access. Identity establishes the tenant boundary before the console can reach the private, read-only gateway service.

Available now: profile-scoped intent and outcome observability. Roadmap: richer OpenTelemetry correlation across runtime boundaries, retries, provider execution, observations, and assessments.

08 / Proof, not promises

What exists—and what comes next.

AgentAction labels experimental work and roadmap items plainly. The public repository, runnable examples, fixtures, and tests are the source of truth.

Available now

Runtime action control

Local and hosted checks for exact tool calls, approvals, amount caps, budgets, circuit breakers, sensitive-data movement, idempotency, and replay protection.

Available now

Provider-verifiable authority

Signed, action-bound authorization receipts, public verification keys, provider middleware, contracts, fixtures, and negative conformance cases.

Available now

Execution assurance

Linked execution receipts, immutable evidence snapshots, verified observations, versioned intent contracts, and outcome assessments.

Available now

Observability console

Explore synthetic fleet rollups, finalized jobs, constraint outcomes, evidence confidence, and deterministic evidence timelines in the public read-only demo.

Open the public console demo

Roadmap

Causal observability

Richer OpenTelemetry correlation across runs, boundary decisions, tool calls, retries, provider execution, observations, and assessment evidence.

Roadmap

Monotonic task authority

Task-scoped capability state that can remove incompatible authority after protected events without silently expanding what an agent may do.

External proof target

Independent interoperability

Two independent providers passing the same public action-authorization cases without project-specific runtime coordination.

09 / Recommended onboarding

Available now

Observe first. Enforce when ready.

Run the customer-controlled MCP adapter beside an existing workflow. In observe mode, it forwards every MCP call unchanged while recording the counterfactual allow, deny, or challenge decision and actionable findings in process-local shadow state.

Run the observer quick start

This reference adapter is a quick onboarding and integration path, not yet a production-complete MCP gateway. When the findings look right, switch the same deployment from observe toenforce deliberately.

Need an in-process boundary instead? Embed the TypeScript guard

customer environmentmode: observe
MCP clientObserver adapterMCP server
gateway_outcome
forwarded
counterfactual_decision
deny
finding
idempotency key missing
downstream_result
returned unchanged
after validationobserve → enforce

10 / One boundary, four entry points

Meet the project where you build.

01

Agent developers

Wrap consequential tool calls with a small authorization boundary inside an existing agent loop.

02

Platform and gateway teams

Apply consistent action policy before forwarding MCP tools/call or other privileged operations.

03

API and SaaS providers

Verify action-specific enterprise authority before mutation, then apply provider business authorization.

04

Security and assurance teams

Connect proposals, approvals, decisions, execution, observations, and assessments without trusting the model as the record of truth.

11 / Open standards posture

Build interoperability before vocabulary.

AgentAction reuses established identity, policy, transport, signing, and provenance work where it fits. The project contributes mappings, negative fixtures, reference verifiers, conformance cases, and narrowly scoped experimental profiles. Its community drafts are not adopted standards or external certifications.

Agent recipes

Start with a job worth handing off.

Choose a recipe built around real MCP servers. Adapt it to your needs, test it in your environment, and follow its outcomes in AgentAction.

Browse all recipes

Each starter includes setup guidance and synthetic test cases. Validate your agent with connected services before putting it to work.

12 / Enterprise transition blueprint

Move from AI assistance to bounded autonomy.

Do not decide whether an agent is simply “autonomous.” Decide which action classes it has earned the right to perform, under which conditions, and with what evidence.

Advance one bounded workflow at a time. Every stage changes the operating mode only after its exit evidence is available.

  1. 01Frame

    Frame the job

    Define one bounded workflow, its owner, value, risk, consequential actions, and human baseline.

    Evidence must answer

    Is the job, owner, value, risk, and human baseline clear?

  2. 02Prove offline

    Prove behavior

    Test historical, edge, adversarial, and failure cases before the agent touches production traffic.

    Evidence must answer

    Does the agent meet quality and safety thresholds on representative cases?

  3. 03Shadow

    Shadow production

    Record counterfactual decisions on representative traffic without duplicating or changing side effects.

    Evidence must answer

    What would happen on real traffic if policy were enforced?

  4. 04Supervise

    Supervise actions

    Bind approval and short-lived authority to the exact actor, job, tool, resource, and payload.

    Evidence must answer

    Can exact actions execute safely with approval and recovery controls?

  5. 05Bound and scale

    Scale the proven envelope

    Automate only earned action classes; keep exceptions, missing context, and high-impact work supervised.

    Evidence must answer

    Which actions can run autonomously without exceeding the risk envelope?

Referenced practitioner field guide · 9 pages

Take the complete transition blueprint into your planning session.

Includes the evidence stack, evaluation scorecard, 90-day launch plan, three applied workflow playbooks, and reference anchors.

Start a transition assessment

Tell us about one consequential workflow.

We will reply with a practical first take on its current stage, evidence gaps, and next step.

Sent privately to info@agentaction.dev. Do not include credentials, secrets, or sensitive production data.