Published by ApexClaw Research. Engine ARB/1.1. 49 checks, 9 weighted pillars, every finding carrying the exact string observed on the target so any result can be reproduced or disputed by inspection.
ARB/1.1 measures whether this one public page is legible and independently verifiable to an autonomous agent: structured data, machine-readable dates, off-site corroboration, and the page's public transport and security surface. It does not measure the organisation's actual security posture, the quality of its product, or its reputation -- those require access, judgement, and time this engine does not have and will not simulate. Its security_surface pillar reads public response headers and TLS parameters observed on one fetch; it is NOT a penetration test and must never be read as one.
Rubric authored and reviewed by ApexClaw Research. Engine ARB/1.1, published 2026-08-26, revised whenever a check or weight changes -- full history in Versioning below.
Agent readiness is the degree to which a website can be fetched, parsed, cited and transacted with by an autonomous agent rather than a human reader.
Search ranking assumes a human clicks a result. An agent never clicks: it retrieves, extracts an answer, and either cites the source or omits it. Omission is a retrieval failure, not a ranking one.
Each check returns one of three verdicts: VERIFIED, UNVERIFIED, or UNOBSERVABLE. There is no partial credit -- a check either verifies the fact it checks for, or it does not. An UNOBSERVABLE check (the evidence needed to decide was absent, e.g. a fetch failed or a third-party API was unreachable) is removed from the denominator rather than counted as a failure, so nothing is inferred from missing information.
| Pillar | Weight | Checks |
|---|---|---|
| Machine Access | 14% | 6 |
| Structured Data | 14% | 6 |
| Answer Extractability | 13% | 5 |
| Evidence & Citability | 11% | 5 |
| Entity Clarity | 10% | 6 |
| Agent Interaction Surface | 10% | 4 |
| Governance & Trust | 8% | 3 |
| Off-site Corroboration | 10% | 5 |
| Transport & Security Surface | 10% | 9 |
The percentile compares a score against 30 real sites measured with this same engine. Median 59.8.
The cohort is 30 sites, scored on an unrecorded date. Lowest 3.2, median 59.8, highest 77.7. Every member and its score is listed in the machine version.
A deterministic check reads a structural fact: a header, a status code, a JSON-LD type, an element count. A heuristic check infers from a text pattern and can miss a true signal.
| ID | Check | Method | Weight |
|---|---|---|---|
| Machine Access | |||
ma.reachable | Page returns 200 to an agent user-agent | DETERMINISTIC | 4 |
ma.robots | robots.txt present | DETERMINISTIC | 2 |
ma.ai_crawlers | AI crawlers not blocked in robots.txt | DETERMINISTIC | 4 |
ma.llms_txt | /llms.txt published | DETERMINISTIC | 3 |
ma.sitemap | XML sitemap resolves | DETERMINISTIC | 2 |
ma.no_js_content | Primary content present without JavaScript | DETERMINISTIC | 3 |
| Structured Data | |||
sd.any | JSON-LD structured data present | DETERMINISTIC | 4 |
sd.org | Organization / Product entity declared | DETERMINISTIC | 4 |
sd.faq | FAQPage or QAPage markup | DETERMINISTIC | 3 |
sd.breadcrumb | BreadcrumbList markup | DETERMINISTIC | 2 |
sd.article | Article / TechArticle / Dataset for content pages | DETERMINISTIC | 2 |
sd.sameas | sameAs identity links in JSON-LD | DETERMINISTIC | 3 |
| Answer Extractability | |||
ex.h1 | Exactly one H1 | DETERMINISTIC | 2 |
ex.question_heads | Question-form headings | DETERMINISTIC | 4 |
ex.direct_answers | Short direct-answer paragraphs under headings | DETERMINISTIC | 4 |
ex.structured_blocks | Tables or lists carry comparable facts | DETERMINISTIC | 3 |
ex.depth | Sufficient topical depth | DETERMINISTIC | 3 |
| Evidence & Citability | |||
ev.dated | Content carries a machine-readable date | DETERMINISTIC | 4 |
ev.author | Named author or reviewer | HEURISTIC | 3 |
ev.citations | Outbound citations to independent sources | DETERMINISTIC | 3 |
ev.canonical | Canonical URL declared | DETERMINISTIC | 2 |
ev.specifics | Quantified, checkable claims | HEURISTIC | 2 |
| Entity Clarity | |||
en.title | Descriptive <title> | DETERMINISTIC | 2 |
en.description | Meta description present | DETERMINISTIC | 2 |
en.opengraph | OpenGraph metadata | DETERMINISTIC | 2 |
en.about_contact | About and Contact reachable from this page | DETERMINISTIC | 3 |
en.lang | Language declared | DETERMINISTIC | 1 |
en.https | Served over HTTPS | DETERMINISTIC | 2 |
| Agent Interaction Surface | |||
ag.forms | Forms are machine-fillable (named fields) | DETERMINISTIC | 3 |
ag.booking | Booking endpoint linked | DETERMINISTIC | 3 |
ag.api | Machine endpoint or API discoverable | HEURISTIC | 3 |
ag.contact_ld | Contact details in structured data | DETERMINISTIC | 3 |
| Governance & Trust | |||
gv.security_txt | /.well-known/security.txt published | DETERMINISTIC | 2 |
gv.policies | Privacy and Terms linked | DETERMINISTIC | 2 |
gv.ai_policy | Stated policy for AI/agent access | HEURISTIC | 3 |
| Off-site Corroboration | |||
offsite.sameas_resolve | sameAs targets resolve and are independent | DETERMINISTIC | 3 |
offsite.citations_resolve | Citations resolve across distinct domains | DETERMINISTIC | 3 |
offsite.security_contact | security.txt contact is reachable | DETERMINISTIC | 2 |
offsite.wikidata | Entity present in Wikidata | DETERMINISTIC | 2 |
offsite.claims_attributable | Quantified claims are attributable to a source | HEURISTIC | 2 |
| Transport & Security Surface | |||
sec.tls_version | Modern TLS version negotiated | DETERMINISTIC | 3 |
sec.cert_valid | TLS certificate chain is valid | DETERMINISTIC | 3 |
sec.hsts | HSTS present with a sane max-age | DETERMINISTIC | 2 |
sec.csp | CSP present and not trivially permissive | DETERMINISTIC | 3 |
sec.xcto | X-Content-Type-Options: nosniff | DETERMINISTIC | 1 |
sec.referrer_policy | Referrer-Policy declared | DETERMINISTIC | 1 |
sec.permissions_policy | Permissions-Policy declared | DETERMINISTIC | 1 |
sec.mixed_content | No mixed content on an HTTPS page | DETERMINISTIC | 2 |
sec.server_banner | Server banner does not leak a version | DETERMINISTIC | 1 |
A run scores one URL and the site-wide files at its root: robots.txt, llms.txt, sitemap.xml and the well-known paths. A finding of absent means absent from that page, not from the site.
Every finding carries its evidence string. A dispute is resolved by fetching the same URL and comparing, not by argument. Re-running after a fix shows the delta attributable to that single change.
The rubric is versioned. Weights and checks may change; a published score names the engine version that produced it, and an earlier version's published score is never restated under a later version's rubric. Changes are dated in the machine version.
| Version | Date | Change |
|---|---|---|
ARB/1.0 | 2026-08-24 | First published rubric: 36 checks, seven pillars, PASS/PARTIAL/FAIL/UNKNOWN verdicts, not-evaluable checks excluded from the denominator. |
ARB/1.1 | 2026-08-26 | Three-state rubric: VERIFIED/UNVERIFIED/UNOBSERVABLE replaces PASS/PARTIAL/FAIL/UNKNOWN (no partial credit). Two pillars added -- off-site corroboration and transport/security surface -- with the seven original pillar weights scaled down proportionally to make room. 49 checks, nine pillars. Every report now carries an explicit scope statement stating the security pillar is not a penetration test. |
A benchmark that cannot be reproduced is an opinion. Publishing the checks, the weights and the cohort is what makes a result arguable, and a result that can be argued with is one that can be cited.
The alternative, a hidden score, forces a reader to trust the scorer. That works for an established authority and fails for a new one. So the whole method is on this page, the machine-readable version is one request away, and any figure quoted from either carries the engine version that produced it. A competitor can implement the same rubric and get the same numbers; that is the point rather than a risk. Standards earn their position by being adopted, and nothing is adopted that cannot first be inspected.
ARB/1.1 measures whether this one public page is legible and independently verifiable to an autonomous agent: structured data, machine-readable dates, off-site corroboration, and the page's public transport and security surface. It does not measure the organisation's actual security posture, the quality of its product, or its reputation -- those require access, judgement, and time this engine does not have and will not simulate. Its security_surface pillar reads public response headers and TLS parameters observed on one fetch; it is NOT a penetration test and must never be read as one.
Four limits follow from that, and each is stated on the face of every report rather than buried here. A run covers one URL plus the files at the domain root, so an absent finding means absent from that page. 5 checks infer from a text pattern instead of reading a structural fact, and those are tagged as heuristic wherever they appear. A check that cannot be evaluated at all is UNOBSERVABLE and is removed from the denominator instead of being scored as a failure, because counting an unknown as a failure would penalise a site for the benchmark's own blind spot. The security_surface pillar reads response headers and TLS parameters observed on one fetch -- it is NOT a penetration test and must never be read as one.
Score any public URL at the audit form. The free result names the score, the pillar breakdown and nine findings; the full report adds the remaining checks, the ranked remediation queue and a JSON export.
Contact hello@apexclawai.com.
About ApexClaw · Contact · Privacy · Terms · Agent access policy