Checking identity is not enough.
An authenticated agent can still act against the merchant's interest. Knowing who it is without checking its track record does not guarantee the operation ends well.
The score measures readiness, not identity.The Trust Score is an operational reliability indicator. It helps AI assistants understand whether your store has current information, clear policies, and enough controls to serve a purchase.
This page describes how it is calculated. What that number decides — visibility in agent listings and, below a certain cut-off, automatic suspension of the store — is published with its exact thresholds and your GDPR Article 22 rights in Privacy.
It is a prioritisation signal with context: not a seal, not a customer review, not a guarantee of outcome.
An agent decides in milliseconds, without seeing your site, reading your reviews or having room to interpret. It needs to know up front which actions it may attempt, which conditions will apply and what proof will remain afterwards. The Trust Score is that upfront reading: the same one for the merchant who publishes it and the agent who queries it.
An authenticated agent can still act against the merchant's interest. Knowing who it is without checking its track record does not guarantee the operation ends well.
The score measures readiness, not identity.Stars and opinions describe a past human experience, not the technical ability to answer a delegated action today.
A structured, versioned, queryable signal.The merchant assumes the agent will behave; the agent assumes the merchant will respond. Both find out in the same transaction.
One reference both read the same way.What could not be observed is not counted: it is marked not applicable and excluded from the calculation. The actual figure and the quality of each pillar travel in the signed response, with their date; that is why this page publishes no number, which would age the moment it is written.
merchant_reliability_baseagentic_readinessagentic_evidenceThe response is signed: the EdDSA signature and the key id travel in the headers (x-jws-signature, x-jws-kid), alongside the methodology version.
Open the demo store's signed responseThe weight of each pillar and the quality of each signal travel inside the response. There is no hidden weighting.
The score is calculated during the diagnosis and reads backwards: every point you miss names a concrete capability that is not operational yet.
You learn which control and evidence signals exist in your systems today and which ones you assumed existed. The score does not reward declared intentions.
When an agent compares stores that solve the same need, it picks by complete and current signals. Being measured and signed is the condition for entering that comparison.
The score does not win a dispute, but it records which controls were active and with what provenance at the time of the operation.
Every tool call costs time, tokens and one chance to fail in front of the user. The Trust Score is queried during planning, before the first attempt is spent.
Among several stores selling the same thing, you start with the one whose signals are most complete and recent, not with the first in the index.
Fewer retries after an unexpected rejection.Thresholds, human review and supported surfaces are visible before attempting the action, not after the error.
Planning with the condition known.The response arrives signed and versioned, so your agent's decision is explainable to the user who delegated it.
An auditable choice, not an opaque heuristic.There are not two scores or two methodologies. There is one signed response and two legitimate ways to use it, and that is what makes it a reference.
It does not certify the merchant, does not guarantee a transaction goes well and does not replace the evidence of a specific operation. It is a prioritisation signal with visible provenance and sample. Presenting it as anything more would make it useless for both sides.
Review limits and evidenceGET /trust-scoreThe pillars that depend on real operations switch on once there is enough sample in your pilot.
No. Every pillar is checked against the system, not against what a manifest declares. Declaring a capability that does not answer does not move the score; at best it leaves it marked not applicable.
With a small sample, Bayesian shrinkage pulls the result towards the mean instead of rewarding a few favourable cases. The score rises as real operations accumulate, not all at once.
No. It is a prioritisation signal — catalogue, policies, checkout, evidence — not a guarantee of outcome. The proof of a specific operation is its receipt, not the score.
Signals decay over time, so an old capture weighs less even if nothing changes. Every signed response states its date and methodology version so you can decide whether it still holds.
Only by reading the quality of each pillar. Two stores with the same figure may have different pillars measured, and that is explicit in the response.
The assessment is scoped to one operation, some systems and one metric. You leave with a figure, its provenance and the list of what is worth solving first.