Skip to content
Early access · for platform and security teams

Does your agent dowhat its card says?

Suncly reads an A2A Agent Card, tests every declared skill many times in a sandbox, judges the results, and hands your reviewers signed evidence to approve on.

attestation 3c4d5e6f… · trigger ci · completed
Signed · signing_key_id
Order Status Agentowner example-teamrisk low
agent.example.com/.well-known/agent-card.json
contract v1 · approved by reviewer@example.com
Test casePassFailInconclusive
  • order-status · 9b0c1d2e
    kind skill
    4901
  • order-status · 2d3e4f5a
    kind skill
    5000
  1. flagdecided_by policy09:42
  2. approvedecided_by reviewer@example.com11:05
Budget
3.42 of 25.0
Not testedcapabilities.streamingproduction endpoint (sandbox only)
The gap

The protocol describes. Nothing verifies.

  • A card is a list of claims.

    An A2A agent publishes an Agent Card that names the skills it offers. The specification defines how to describe a skill. It has no mechanism for checking that the agent performs it.

  • A signature is not a test.

    Agent Cards can be signed. A signature shows the card was not altered, and who signed it. It says nothing about how the agent behaves.

  • So approval is done by hand.

    Platform and security teams read the card, send a few prompts, and sign off. Suncly is built for those teams.

What Suncly does

A roadworthiness inspection for AI agents.

Point it at an Agent Card and it does five things.

  1. 01Reads the agent's Agent Card.
  2. 02Generates test cases for each declared skill.
  3. 03Runs each test many times against the agent, in a sandbox.
  4. 04Judges the results and stores signed evidence.
  5. 05Outputs an approval decision for your agent registry or CI pipeline.
How it works

Seven steps from trigger to signed decision.

  1. 1

    A trigger arrives.

    CI, a schedule, a card change, or a person.

  2. 2

    Suncly fetches the Agent Card and hashes it.

    If the hash is new, a model drafts test cases for every declared skill, and a human approves them as a contract.

  3. 3

    The Orchestrator expands the contract into runs.

    Each test case runs many times, within a cost budget.

  4. 4

    The Runner calls the agent's sandbox endpoint.

    As an A2A client, capturing every message. It is the only component that holds your credentials.

  5. 5

    The Judge scores each run.

    Deterministic checks first; a pinned model only where those cannot decide. The verdict is pass, fail or inconclusive, and inconclusive never counts as a pass.

  6. 6

    The Policy engine decides and signs.

    It applies your policy for the agent's risk level, decides approve, flag for human review, or block, and signs the attestation.

  7. 7

    Adapters publish the result.

    A registry status, a CI gate, and an evidence report that states what was NOT tested.

terminal
$
    Architecture

    Six components in the core. Three thin adapters outside it.

    Each component has one responsibility, and a list of things it must never do. A fault can make Suncly approve less, never more.

    System layout: Agent Card and trigger enter the Suncly core, where contract builder, orchestrator and runner feed judge, evidence store and policy engine; the runner alone talks to the agent endpoint; outputs are registry status, CI gate and evidence reportAgent Card/.well-known/agent-card.jsonTriggerci · schedule · card change · manualAgent endpointsandbox or dry-runSUNCLY COREContract buildermodel drafts · human approvesOrchestratorruns · retries · budgetRunnerisolated · holds credentialsJudgedeterministic · pinned modelEvidence storeappend-only · signedPolicy engineapprove · flag · blockRegistry status · CI gate · Evidence reportadapters: thin, outside the core, never decide
    • Contract builder

      Turns the Agent Card into a versioned contract: one or more test cases per declared skill. A model drafts them; a human approves them. Approved contracts are immutable.

      neverMarks a contract approved without a human.

    • Orchestrator

      Expands an approved contract into runs: every test case times the number of repetitions. Owns concurrency, retries, timeouts and the cost budget per attestation.

      neverEnqueues a run once the budget is reached.

    • Runner

      The only component that holds your credentials and calls the agent. Acts as an A2A client in its own container, network-limited to the target, and redacts every transcript before it leaves.

      neverLets a secret leave, or calls a production endpoint.

    • Judge

      Layer 1 is deterministic: valid schema, final task state, required fields, latency limit. Layer 2 uses a pinned model with a fixed rubric, only for criteria Layer 1 cannot decide.

      neverCounts inconclusive as a pass.

    • Evidence store

      Append-only. Per run: the full redacted transcript, verdict, timings and cost. Each attestation is signed and references the card hash and contract version.

      neverUpdates or deletes a record. Corrections are new records.

    • Policy engine

      Aggregates results per test case and applies your thresholds for the agent's risk level. Outcome: approve, flag or block. Signs the attestation.

      neverApproves automatically when a human is required.

    Adapters

    Registry status, CI gate and evidence report. They translate results into other systems and never decide anything.

    • Registry status
    • CI gate
    • Evidence report

    Product

    01Contract

    Tests drafted by a model. Approved by a person.

    Suncly canonicalizes and hashes the card, then drafts at least one test case for every declared skill. Nothing runs until a human approves the contract. Approved contracts are immutable: any edit is a new version, and a changed card is a new contract.

    An Agent Card is hashed, a model drafts test cases per skill, a human approves them, and the contract becomes an immutable versionagent-card.jsonskills[0]skills[1]skills[2]card_hash 7f3a…c1draft · by modeltest_case · skills[0]test_case · skills[1]test_case · skills[2]probe_injectionprobe_failurehuman approvescontractv1immutablea changed card is a new hash, a new draft, and a new approval
    02Runs

    Many times, in a sandbox, inside a budget.

    Every test case runs repeatedly against a sandbox or dry-run endpoint, so nothing real is booked, paid or deleted. Each run has a deterministic key, so retries never double count, and the attestation stops when it reaches its cost cap.

    Each test case is run many times against a sandbox endpoint, every run has a deterministic key, and the attestation stops at its budgettest cases × repetitionsrun key = (attestation, test_case, attempt)sandbox / dry-run endpointnothing real is booked, paid or deletedbudget_limit 25.0cost_total 3.42stops here · no decisionretries reuse the run key, so nothing is counted twice
    03Probes

    Beyond what the card declares.

    Alongside skill tests, the contract carries probes: behaviour the card does not declare, injected instructions, and failure conditions. They go through the same human approval and the same sandbox rule.

    Three probe kinds, undeclared behaviour, injected instructions and failure conditions, are sent to the agent inside the sandbox, after the same human approval as skill testsprobe_undeclaredbehaviour the card does not declareprobe_injectioninjected instructionsprobe_failurefailure conditionssandboxsame human approvaltest_case.kind: skill · probe_undeclared · probe_injection · probe_failure
    04Judge

    Deterministic first. A pinned model only when needed.

    Valid schema, final task state, required fields and the latency limit are decided without a model. Where a criterion needs one, a model pinned by version applies a fixed rubric and stores its rationale. Inconclusive is a third verdict, and it is never a pass.

    A transcript passes through deterministic Layer 1 checks, then a pinned model with a fixed rubric only where needed, producing a pass, fail or inconclusive verdicttranscriptredactedLayer 1 · deterministicvalid schemafinal task staterequired fieldslatency limitLayer 2 · pinned modelfixed rubric · rationale storedverdict per runpassfailinconclusivenever counted as a pass
    05Evidence

    Signed, append-only, and honest about gaps.

    Every run keeps its transcript, verdict, timings and cost. The attestation is signed over the card hash, the contract version, the aggregated results and the hash of every transcript. The report states what was not tested.

    An append-only evidence ledger of runs and decisions, a signature over card hash, contract version, results and transcript hashes, and a statement of what was not testedevidence store · append-onlyrun · attempt 1passrun · attempt 2passrun · attempt 3inconclusivedecision · policyflagdecision · reviewerapproverecords are never edited · corrections are new recordssignedcard_hashcontract versionaggregated resultstranscript hashesdecision · policy_versionNOT TESTEDcapabilities.streaming · production endpoint
    06Decision

    Approve, flag, or block. Then gate on it.

    The policy engine applies your thresholds per risk level and writes a decision. Adapters push it where approval happens: a status in your agent registry, pass or fail in your CI pipeline, and a report for the reviewer.

    The policy engine combines the agent's risk level with aggregated results into approve, flag or block, and adapters push the decision to a registry, a CI pipeline and a reportagent.risk_levellowautomatic on passmediumhuman on any drophighhuman every timePolicy engineyour thresholdsper risk levelapproveflagblock→ human decidesRegistry statuswritten by adapterCI gatepass or failEvidence reportstates what was not tested
    Approval policy

    Risk decides how much a human is in the loop.

    Every agent has a risk level. The default policy ties it to how an attestation is approved. Thresholds are configured per customer and per risk level; Suncly ships no numbers of its own.

    RiskExampleApproval
    lowread-only lookupautomatic on pass
    mediumwrites to internal systemsautomatic on pass, human on any drop
    highpayments, personal datahuman sign-off every time

    A human is always required for

    • the first approval of a contract
    • new or changed skills
    • borderline or dropping results
    • every attestation of a high-risk agent
    • approve

      The results meet your policy for the agent's risk level. No human is required.

    • flag

      A human must decide. Their resolution is a second decision record; the first is never edited.

    • block

      The results do not meet the policy.

    An attestation that ran out of budget, or whose card changed mid-run, gets no decision at all. No decision is never an approval.

    Non-negotiable

    Seven rules the system is built around.

    Each one is a decision record. Changing one means changing the schema first.

    1. DR-001

      Idempotent runs.

      Every run has a deterministic key. Retries never double count, so the signed results describe runs that happened the way they say.

    2. DR-002

      Evidence is immutable.

      Records are never edited. Corrections are new records, so a signature and the approval behind it can be trusted later.

    3. DR-003

      Secrets never leave the Runner.

      Transcripts are redacted before storage. The Runner is the only component with your credentials, which is what lets it run inside your network later.

    4. DR-004

      The judge model is pinned.

      An attestation measures the agent, not changes in the judge. A new judge model is a deliberate configuration change, never drift.

    5. DR-005

      Budget caps live in the Orchestrator.

      Runs are test cases times repetitions, so costs multiply. The component that enqueues runs is the one that stops them. No surprise bills.

    6. DR-006

      Tests hit a sandbox or dry-run endpoint.

      Probes deliberately inject instructions and failure conditions. Against production that could book, pay or delete something real.

    7. DR-007

      Reports state what was NOT tested.

      An approval on partial evidence misleads if the gaps are invisible. Suncly does not produce a score that would hide them.

    Open standard

    Reads the A2A Agent Card. Emits a signed attestation.

    Suncly does not ask you to describe your agent twice. It reads the card you already publish under the A2A protocol, version 1.0, and tests what is in it.

    What Suncly reads

    • /.well-known/agent-card.json
      where A2A agents publish their card
    • skills[]
      id · name · description · tags · examples. The claims under test.
    • supportedInterfaces[]
      url · protocolBinding · protocolVersion. The Runner picks the first it supports.
    • capabilities · securitySchemes
      what the card declares beyond skills, and how to authenticate
    • signatures
      integrity and signer only. Not behaviour.

    What Suncly emits

    • results[]
      per test case: pass, fail and inconclusive counts
    • decisions[]
      approve · flag · block, with policy_version and who decided
    • signature · signing_key_id
      over card_hash, contract version, results and transcript hashes
    • evidence report
      with a statement of what was NOT tested
    GET /attestations/3c4d5e6f…
    // from docs/API.md (proposed response shape), shortened
    {
      "id": "3c4d5e6f-7a8b-4c9d-8e0f-1a2b3c4d5e6f",
      "trigger": "ci",
      "status": "completed",
      "budget_limit": 25.0,
      "cost_total": 3.42,
      "signature": "<signature>",
      "signing_key_id": "<signing_key_id>",
      "results": [
        { "skill_id": "order-status", "kind": "skill",
          "pass": 49, "fail": 0, "inconclusive": 1 },
        { "skill_id": "order-status", "kind": "skill",
          "pass": 50, "fail": 0, "inconclusive": 0 }
      ],
      "decisions": [
        { "outcome": "flag", "decided_by": "policy",
          "policy_version": "<policy_version>" }
      ]
    }
    Interfaces

    One command. Four endpoints. The same core library.

    The CLI arrives first and is enough to run a pilot by hand. The HTTP API follows and calls the same library.

    CLI
    $ suncly attest <card-url> --runs 50
    <card-url>
    the URL of the agent's Agent Card
    --runs
    repetitions per test case
    • Stage 1Fetches the card, runs each test case --runs times against a sandbox or dry-run endpoint, judges each run deterministically, and writes a file report that states what was NOT tested.
    • From stage 2Runs only an approved contract. If the card's hash is new, a draft contract is created and must be approved first.
    HTTP API
    • POST/attestationsstart an attestation
    • GET/attestations/{id}status and results
    • POST/contracts/{id}/approvehuman approval
    • GET/agents/{id}/evidenceevidence history
    POST /attestations
    {
      "agent_id": "6f1c2d3e-…",
      "card_url": "https://agent.example.com/.well-known/agent-card.json",
      "runs": 50,
      "trigger": "ci",
      "budget_limit": 25.0
    }
    Where we are

    Six stages. The first one runs a pilot.

    Suncly is pre-prototype: the architecture is written down, the package skeleton exists, and the build order is fixed. Stage 1 is enough to run a first pilot by hand.

    1. 1first pilot

      Core library

      Runner, deterministic judge, CLI and a file report.

    2. 2

      Contract builder

      Drafted test cases with the human approval step.

    3. 3

      Evidence store

      Cloudflare D1 tables, R2 transcripts, and attestation signing.

    4. 4

      Model judge and probes

      Layer 2 with a pinned model, plus undeclared, injection and failure probes.

    5. 5

      API, policy, CI

      The four endpoints, the policy engine and the CI gate.

    6. 6

      Registry adapters

      Approval status written into your agent registry.

    Stack

    Language
    Python, for the most mature A2A SDK
    Storage
    Cloudflare D1 for tables, R2 for transcripts
    Signing
    asymmetric signatures over the canonicalized attestation
    Models
    you bring your own keys; the judge model is pinned by version

    Not on the roadmap

    Suncly is not a registry, a gateway, an identity system, a monitoring platform, a universal score, or a payments product. It writes into the registry you have and gates the pipeline you run.

    Run the first pilot with us

    Run the first pilot with us.

    Suncly is in early access. If your team approves A2A agents by hand today, leave your email and we will write back. No list size, no countdown, just a reply from the team in Tallinn.

    FAQ

    Questions.

    What is Suncly?

    An attestation tool for A2A agents. It reads an agent's Agent Card, tests every declared skill many times in a sandbox, judges the results, stores signed evidence, and outputs an approval decision for your registry or CI pipeline.

    What is an Agent Card?

    Under the A2A protocol, an agent publishes a JSON document, normally at /.well-known/agent-card.json, that names it and lists the skills it offers. Suncly treats that list as the claims to test.

    Why is a signed card not enough?

    A card signature shows the card was not altered and who signed it. It says nothing about how the agent behaves. Suncly tests the behaviour.

    Does Suncly call my production agent?

    No. Tests hit a sandbox or dry-run endpoint, so nothing real is booked, paid or deleted. The report says so.

    Who writes the tests?

    A model drafts at least one test case per declared skill. A human approves them as a contract before anything runs. Approved contracts are immutable; an edit is a new version.

    What can a run score?

    Pass, fail or inconclusive. Deterministic checks decide first; a pinned model with a fixed rubric decides only what they cannot. Inconclusive is never counted as a pass.

    What decision comes out?

    Approve, flag for human review, or block, under your policy for the agent's risk level. A high-risk agent is never approved without a human.

    Does it produce a score I can compare agents on?

    No. The outcome is a decision under one customer's policy, not a universal score. A score would hide what was not tested.

    Where do my credentials go?

    Only to the Runner, the one component that calls the agent. It runs in its own container, limited to the target, and redacts transcripts before they are stored. It is designed to run inside your network later.

    What stops a runaway bill?

    Each attestation has a budget. The Orchestrator stops enqueuing runs when the cost reaches it, the attestation ends as failed, and no decision is made.

    Which models does it use?

    You bring your own model keys for drafting test cases and for the model-based judge. The judge model is pinned by version so results do not drift.

    Is Suncly available today?

    It is pre-prototype and in early access. The architecture is documented and the build order is fixed; stage 1 is enough to run a first pilot by hand. Leave your email and we will write to you.