A card is a list of claims.
An A2A agent publishes an Agent Card that names the skills it offers. The specification defines how to describe a skill. It has no mechanism for checking that the agent performs it.
Suncly reads an A2A Agent Card, tests every declared skill many times in a sandbox, judges the results, and hands your reviewers signed evidence to approve on.
capabilities.streamingproduction endpoint (sandbox only)An A2A agent publishes an Agent Card that names the skills it offers. The specification defines how to describe a skill. It has no mechanism for checking that the agent performs it.
Agent Cards can be signed. A signature shows the card was not altered, and who signed it. It says nothing about how the agent behaves.
Platform and security teams read the card, send a few prompts, and sign off. Suncly is built for those teams.
Point it at an Agent Card and it does five things.
CI, a schedule, a card change, or a person.
If the hash is new, a model drafts test cases for every declared skill, and a human approves them as a contract.
Each test case runs many times, within a cost budget.
As an A2A client, capturing every message. It is the only component that holds your credentials.
Deterministic checks first; a pinned model only where those cannot decide. The verdict is pass, fail or inconclusive, and inconclusive never counts as a pass.
It applies your policy for the agent's risk level, decides approve, flag for human review, or block, and signs the attestation.
A registry status, a CI gate, and an evidence report that states what was NOT tested.
Each component has one responsibility, and a list of things it must never do. A fault can make Suncly approve less, never more.
Turns the Agent Card into a versioned contract: one or more test cases per declared skill. A model drafts them; a human approves them. Approved contracts are immutable.
neverMarks a contract approved without a human.
Expands an approved contract into runs: every test case times the number of repetitions. Owns concurrency, retries, timeouts and the cost budget per attestation.
neverEnqueues a run once the budget is reached.
The only component that holds your credentials and calls the agent. Acts as an A2A client in its own container, network-limited to the target, and redacts every transcript before it leaves.
neverLets a secret leave, or calls a production endpoint.
Layer 1 is deterministic: valid schema, final task state, required fields, latency limit. Layer 2 uses a pinned model with a fixed rubric, only for criteria Layer 1 cannot decide.
neverCounts inconclusive as a pass.
Append-only. Per run: the full redacted transcript, verdict, timings and cost. Each attestation is signed and references the card hash and contract version.
neverUpdates or deletes a record. Corrections are new records.
Aggregates results per test case and applies your thresholds for the agent's risk level. Outcome: approve, flag or block. Signs the attestation.
neverApproves automatically when a human is required.
Registry status, CI gate and evidence report. They translate results into other systems and never decide anything.
Suncly canonicalizes and hashes the card, then drafts at least one test case for every declared skill. Nothing runs until a human approves the contract. Approved contracts are immutable: any edit is a new version, and a changed card is a new contract.
Every test case runs repeatedly against a sandbox or dry-run endpoint, so nothing real is booked, paid or deleted. Each run has a deterministic key, so retries never double count, and the attestation stops when it reaches its cost cap.
Alongside skill tests, the contract carries probes: behaviour the card does not declare, injected instructions, and failure conditions. They go through the same human approval and the same sandbox rule.
Valid schema, final task state, required fields and the latency limit are decided without a model. Where a criterion needs one, a model pinned by version applies a fixed rubric and stores its rationale. Inconclusive is a third verdict, and it is never a pass.
Every run keeps its transcript, verdict, timings and cost. The attestation is signed over the card hash, the contract version, the aggregated results and the hash of every transcript. The report states what was not tested.
The policy engine applies your thresholds per risk level and writes a decision. Adapters push it where approval happens: a status in your agent registry, pass or fail in your CI pipeline, and a report for the reviewer.
Every agent has a risk level. The default policy ties it to how an attestation is approved. Thresholds are configured per customer and per risk level; Suncly ships no numbers of its own.
| Risk | Example | Approval |
|---|---|---|
| low | read-only lookup | automatic on pass |
| medium | writes to internal systems | automatic on pass, human on any drop |
| high | payments, personal data | human sign-off every time |
The results meet your policy for the agent's risk level. No human is required.
A human must decide. Their resolution is a second decision record; the first is never edited.
The results do not meet the policy.
An attestation that ran out of budget, or whose card changed mid-run, gets no decision at all. No decision is never an approval.
Each one is a decision record. Changing one means changing the schema first.
Every run has a deterministic key. Retries never double count, so the signed results describe runs that happened the way they say.
Records are never edited. Corrections are new records, so a signature and the approval behind it can be trusted later.
Transcripts are redacted before storage. The Runner is the only component with your credentials, which is what lets it run inside your network later.
An attestation measures the agent, not changes in the judge. A new judge model is a deliberate configuration change, never drift.
Runs are test cases times repetitions, so costs multiply. The component that enqueues runs is the one that stops them. No surprise bills.
Probes deliberately inject instructions and failure conditions. Against production that could book, pay or delete something real.
An approval on partial evidence misleads if the gaps are invisible. Suncly does not produce a score that would hide them.
Suncly does not ask you to describe your agent twice. It reads the card you already publish under the A2A protocol, version 1.0, and tests what is in it.
// from docs/API.md (proposed response shape), shortened
{
"id": "3c4d5e6f-7a8b-4c9d-8e0f-1a2b3c4d5e6f",
"trigger": "ci",
"status": "completed",
"budget_limit": 25.0,
"cost_total": 3.42,
"signature": "<signature>",
"signing_key_id": "<signing_key_id>",
"results": [
{ "skill_id": "order-status", "kind": "skill",
"pass": 49, "fail": 0, "inconclusive": 1 },
{ "skill_id": "order-status", "kind": "skill",
"pass": 50, "fail": 0, "inconclusive": 0 }
],
"decisions": [
{ "outcome": "flag", "decided_by": "policy",
"policy_version": "<policy_version>" }
]
}The CLI arrives first and is enough to run a pilot by hand. The HTTP API follows and calls the same library.
$ suncly attest <card-url> --runs 50POST /attestations
{
"agent_id": "6f1c2d3e-…",
"card_url": "https://agent.example.com/.well-known/agent-card.json",
"runs": 50,
"trigger": "ci",
"budget_limit": 25.0
}Suncly is pre-prototype: the architecture is written down, the package skeleton exists, and the build order is fixed. Stage 1 is enough to run a first pilot by hand.
Runner, deterministic judge, CLI and a file report.
Drafted test cases with the human approval step.
Cloudflare D1 tables, R2 transcripts, and attestation signing.
Layer 2 with a pinned model, plus undeclared, injection and failure probes.
The four endpoints, the policy engine and the CI gate.
Approval status written into your agent registry.
Suncly is not a registry, a gateway, an identity system, a monitoring platform, a universal score, or a payments product. It writes into the registry you have and gates the pipeline you run.
Eight documents describe the system. Where they and the schema disagree, the schema wins. They ship with early access.
The architecture schema. The source of truth.
Each component's responsibility, inputs, outputs, what it must never do, and how it fails safely.
The seven entities, with fields, keys, enums and an ER diagram.
One attestation from trigger to result, with every failure path.
The CLI command and the four HTTP endpoints, with examples.
Risk levels, approval rules, and when a human is required.
The seven non-negotiable rules, as decision records.
Six build stages, each with a definition of done.
Suncly is in early access. If your team approves A2A agents by hand today, leave your email and we will write back. No list size, no countdown, just a reply from the team in Tallinn.
An attestation tool for A2A agents. It reads an agent's Agent Card, tests every declared skill many times in a sandbox, judges the results, stores signed evidence, and outputs an approval decision for your registry or CI pipeline.
Under the A2A protocol, an agent publishes a JSON document, normally at /.well-known/agent-card.json, that names it and lists the skills it offers. Suncly treats that list as the claims to test.
A card signature shows the card was not altered and who signed it. It says nothing about how the agent behaves. Suncly tests the behaviour.
No. Tests hit a sandbox or dry-run endpoint, so nothing real is booked, paid or deleted. The report says so.
A model drafts at least one test case per declared skill. A human approves them as a contract before anything runs. Approved contracts are immutable; an edit is a new version.
Pass, fail or inconclusive. Deterministic checks decide first; a pinned model with a fixed rubric decides only what they cannot. Inconclusive is never counted as a pass.
Approve, flag for human review, or block, under your policy for the agent's risk level. A high-risk agent is never approved without a human.
No. The outcome is a decision under one customer's policy, not a universal score. A score would hide what was not tested.
Only to the Runner, the one component that calls the agent. It runs in its own container, limited to the target, and redacts transcripts before they are stored. It is designed to run inside your network later.
Each attestation has a budget. The Orchestrator stops enqueuing runs when the cost reaches it, the attestation ends as failed, and no decision is made.
You bring your own model keys for drafting test cases and for the model-based judge. The judge model is pinned by version so results do not drift.
It is pre-prototype and in early access. The architecture is documented and the build order is fixed; stage 1 is enough to run a first pilot by hand. Leave your email and we will write to you.