VAIL — Agent Assurance Engine

Peace of mind that your agents are working.

Your business now runs on models, APIs, and agents operated by someone else, and changing without notice. The Agent Assurance Engine watches your endpoints continuously from your network, confirms they're working the way you need them to, tells you the moment something changes, and what to do next.

In the background, never in the way.

The Agent Assurance Engine is a lightweight service that runs on your terms. It only needs permission to reach the same AI services your company already uses and it does not access sensitive production traffic. Your IP stays yours, guaranteed.

YOUR NETWORK · INSIDE THE FIREWALL Agents & apps Codex · Claude Code · OpenClaw AI Gateway / Router Databricks · Vercel · Cloudflare Agent Assurance Engine Your telemetry & SIEM OTel · Prometheus · Splunk · Datadog findings & alerts steer routing & policy steer agents PRODUCTION ENDPOINTS api.openai.com/v1 OpenAI · Anthropic · Google api.fireworks.ai/inference/v1 Fireworks · Together AI · Bedrock api.salesforce.com/einstein Salesforce · ServiceNow · Zendesk vllm-gpu01.internal.corp:8000 vLLM · SGLang · Ollama production traffic no customer data
Running in minutes
A single container for Docker or Kubernetes, pointed at the services and agents you want watched. Prefer zero upkeep? VAIL runs it for you as a managed service.
Inside your walls
Checks come from your own network, under the security rules you already enforce. Nothing new is exposed to the internet, and no customer data is ever read.
Works with your tools
Findings land in the dashboards and alerting your teams already watch — Datadog, Grafana, Splunk, Slack, PagerDuty — not another console someone has to remember to check.

Out of the request path, but not out of the loop. What the Engine finds — a swapped model, a drifting agent — can feed straight into the routing decisions and governance policies your AI gateway or router enforces, so traffic shifts and access tightens the moment something changes.

When something changes, the right person already knows — and knows what to do.

AI platform lead · Platform engineering
A provider quietly swaps a model

The Engine surfaces An identity check on a production endpoint stops matching the approved baseline. A PagerDuty alert names the endpoint, when it changed, and how far it has moved.

The action The platform lead pins traffic to a verified fallback through the gateway, re-runs evals against the new model, and re-approves it — before customers notice anything.

Engineering lead · Agent & app teams
Tool calling starts to slip

The Engine surfaces Consistency probes show tool-call success down 5% over seven hours — on the provider's side, not in your code. The trend lands on the team's Datadog dashboard.

The action The lead skips the internal fire drill, files a provider ticket with the evidence attached, and shifts the affected agent to a backup model in the meantime.

AI governance lead · Security & compliance
A coding agent drifts off course

The Engine surfaces Trajectory tracking shows an agent's behavior trending away from its baseline. A Slack alert connects the drift to the model change that started it.

The action The governance lead tightens the agent's access policy and quarantines the endpoint pending re-approval — with the whole timeline already logged for auditors.

Four questions, answered around the clock.

For every model, endpoint, and agent on the watchlist, the Engine keeps answering the questions your teams would otherwise have to take on faith — and delivers each answer to the team that needs it.

/ 01 · Identity Security & Compliance

Which model is really answering?

An endpoint tells you a model's name — it doesn't prove it. The Engine regularly fingerprints whatever is actually answering each service and confirms it's the model you approved, whether it's a big-name API, a hosted open model, or AI built into a vendor's product. And when auditors or regulators ask, you have the records to show for it, the kind of evidence frameworks like the EU AI Act expect.

Answers Is this the model we certified? · Which model processed regulated data?
/ 02 · Stability Security & MLOps

Has it changed since you approved it?

Models get updated, downgraded, or replaced behind the same name and URL. The Engine compares each endpoint against its own history and records exactly when every shift happened, so you know which version was serving your customers on any given day, and when it's time to re-test. No more "it feels different lately" with nothing to point to.

Answers Did the endpoint change? · When, and for how long?
/ 03 · Consistency Productivity & Debugging

Is it behaving the way your agents expect?

When an AI workflow starts failing, the first question is always "is it us or them?" The Engine tracks the specific behaviors agents rely on: using tools correctly, returning well-formed responses, and following your instructions. Your team can tell in minutes whether a problem lives in your code or on the provider's side, instead of losing days to debugging the wrong thing.

Answers Is the endpoint behaving consistently for agents? · Us or them?
/ 04 · Trajectory Productivity & Optimization

Are your agents still on course?

Agents evolve: their instructions, memory, and tools change over time, and their behavior shifts with them. The Engine keeps a running picture of how each agent is behaving, so a slow drift off course shows up as a trend you can act on early, not a surprise you discover after it's already caused damage.

Answers Are our agents still on course? · What changed their behavior?

The same capabilities, already proving themselves.

You don't have to take the Engine's approach on faith. The kinds of information it collects are already on display — in public, and with partners — so you can judge them for yourself before any conversation.

Cisco partnership

Provenance Explorer

VAIL provides the model identity and similarity matching behind Cisco's Provenance Explorer — a resource for compliance teams evaluating models, powered by the Cisco AI Security Framework. It's the same matching the Engine uses to confirm which model is really answering an endpoint, and how close it is to the one you approved.

Live · Updated hourly

Stability Arena

Our public dashboard demonstrating how the Engine measures endpoint stability — including quiet changes and provider-to-provider differences in behavior that ripple into application and agent workflows. What the Arena does in public, the Engine does for your services, privately, inside your own network.

Open Stability Arena
Peer-reviewed research

Built on published methods

The Engine's checks aren't a black box either. The methods for verifying model identity and detecting change, and for tracking agent behavior over time, are published, peer-reviewed research — presented at ACM CAIS 2026 and ICML 2026.

Endpoint stability paper Agent trajectories paper

Stop wondering whether your AI changed. Know.

Tell us which models, providers, and agents your business depends on. We'll show you what the Engine would watch from day one — and how the answers show up in your own dashboards.

Nothing to install in your apps. No customer traffic touched. Checks run from inside your own network — or from VAIL's, if you prefer the managed service.