Skip to content

Capability Claims: Evidence with an Expiry

Two Outcomes report successful work and the same favorable verdict. One came from a constrained local Environment with a large sample, low cost, and an expiry next month; the other used different tools, fewer Cases, higher cost, and a shorter evidence window. Their shared verdict—or one impressive number—cannot choose a universal winner. The useful result is an exact statement of what each subject revision demonstrated under named conditions.

Capability claims are Designed; no claim or routing path runs. State of Work records maturity; ADR 34 owns the claim law.

Shape the supported claim

A claim derives only from a healthy, version-pinned TrialSuite@1 and its retained Outcomes. This page defines no health algorithm. The evidence must survive its declared controls, harness and Evaluator health, leakage checks, drift review, and required receipts.

The claim pins exact Animator, model, adapter, toolset, and configuration revisions. It names the task class, Cases, Rubric, Evaluators, and Environment; retains sample count, controls, verdict distribution, uncertainty, and noise; and records cost and latency per admitted success, including failed attempts. Creation time, expiry, and evidence references close the support boundary.

Accuracy, latency, VRAM, cost, restraint, and tool behavior remain separate axes. The consumer decides which dimensions matter for its present request. Riddle does not collapse them into one scalar.

Expire without rewriting history

Missing, unhealthy, or expired evidence yields no active claim. A changed subject, toolset, configuration, Rubric, Evaluator, or Environment requires new Outcomes and a new claim; no silent refresh inherits the earlier conclusion. The stale claim remains historical evidence without steering current decisions.

Later loss, leakage, Evaluator drift, or contradiction invalidates only claims whose complete support closure has broken. Independently supported claims remain intact. Recovery means a new bounded trial and a new claim, while causal uncertainty can move through Returning findings.

Missing or stale evidence preserves the Dispatcher's documented fallback. It creates no synthetic intelligence floor and supplies no reason to rank untested subjects.

Let each consumer decide

A fresh admitted claim may inform Dispatcher only after Ward, compatibility, availability, privacy, and authority have produced an eligible set. Dispatcher retains selection and lease ownership. Its current readiness order prefers open admission, then active over inactive and warm over merely active, with stable name and capability key tie-breakers; it does not optimize quality or price. No evaluation-driven ordering is delivered.

For Toll, retained cost and latency evidence may inform spending policy. It neither authorizes payment nor supplies a budget, signature, or settlement. A cheaper local success can be relevant without making all Portal use wasteful.

Soulforge may compare baseline evidence, sealed-holdout Outcomes, expected change, and regressions for a frozen candidate. Passing Riddle establishes eligibility only. The candidate handoff keeps that eligibility separate from the externally owned Promotion Decision; a claim cannot admit corpus, register a candidate, promote, route, or grant capability.

A capability claim remains what its evidence can carry: versioned, scoped, expiring, and useful only inside the consuming owner's decision.