Skip to content

34. Evaluation

Context

A capable-looking answer can fail where its tools, authority, hardware, dialect, and consumer matter. Riddle turns the question of whether a spirit bears its name into versioned, attributable evidence rather than benchmark theatre or self-report.

Decision

Riddle is LychD's singular evaluation jurisdiction: it defines Cases and Trial Suites, captures observations, applies versioned Rubrics, reports uncertainty, and returns bounded findings. It does not execute unsafe payloads or own Tomb; select Animator/capability; authorize spend, publication, repair, or deployment; define Persona; mutate Pattern, Composition, or artifact; admit training; or promote a Soulstone. Execution, Dispatcher, Toll, Spellweaver, Mirror, Smith, Soulforge, and HitL keep those effects. Evaluation is evidence offered to policy, never policy disguised as a score.

Delivery boundary

Riddle is Designed. There is no harness, maintained Trial Suite, evaluator store, capability matrix, benchmark history, Altar route, or Dispatcher update driven by evaluation. State of Work owns delivery.

Trial contract

Record Required content
Case input/fixtures, expected and forbidden behavior, oracle, effect class, stop
Trial Suite (TrialSuite@1) versioned Cases/controls, order, repetitions, aggregation
Rubric criteria, verdict vocabulary, thresholds, missing-evidence policy, revision
Evaluator kind/identity/revision, independence, calibration, limitations
Environment subject/prompt/tool/dependency/hardware/harness/state/budget/policy revisions
Outcome observations, measures, verdicts, uncertainty, errors, cost, latency, evidence

Changed subject, prompt, schema, Rubric, Evaluator, or Environment creates a new Outcome; it cannot rewrite an earlier result. Libraries implement this port but do not own evidence or routing. Riddle separates observed exit/files/rows/tool requests/admitted effects/resource measures/provider receipts from a criterion match, quality grade, attribution claim, or judge score. Model self-report is output under test. Missing evidence stays missing; mechanically observable receipts outrank textual similarity. Shadow may isolate candidates and Tomb may execute them, but only contribute typed observations.

trial status: completed | subject_error | harness_error | evaluator_error | blocked
claim verdict: PASS | PARTIAL | FAIL | CONTRADICTED | UNKNOWN | DISPUTED

Unavailable dependency is not subject failure; a hidden validator precondition is a harness/state contract candidate. Refusal on an impossible task is not success unless matched solvable controls show action when action is possible.

Adversarial evidence and calibration

Sphinx Cases pressure a boundary: forbidden authoritative requests, false premises, demanded certainty amid missing evidence, contradictions/impossible completion, repeated nudges, recoverable distortion, and attempted tool/memory/identity/completion claims beyond supplied evidence. Matched positive and negative controls record pressure round/order, recovery, over-refusal, truthful non-completion, and downstream contamination separately—not a mutable “integrity” scalar. Magus dialect perturbations are still test data with provenance, scope, and release rules.

Trial Suites predeclare repetitions/stops and retain distributions, order, applicable seeds, blocked/error trials, and exclusions. Non-deterministic Evaluators are calibrated against labelled controls and known ambiguity; qualitative work records independent agreement/disagreement where warranted. An LLM judge is a declared, bounded Evaluator: its prompt, revision, inputs, lineage, calibration, and conflicts belong to Environment; hidden chain-of-thought is never required. Sealed Cases and holdouts defend against tuning; leakage, duplicates, unstable harnesses, and evaluator drift invalidate only the claims they undermine.

Capability claims and routing

Riddle may derive a scoped claim from a healthy Trial Suite, pinning Animator/model/adapter/tool/config revisions; task class, Cases, Rubric, Evaluators, Environment; sample/controls/distribution, uncertainty/noise; cost/latency per admitted success including failures; and creation/expiry/evidence references. There is no universal rank: accuracy, latency, VRAM, cost, restraint, and tool behavior are distinct policy-valued axes. Dispatcher may consume fresh admitted claims only after Ward, compatibility, availability, privacy, and authority construct an eligible set. Missing/stale evidence preserves its documented fallback, never an invented intelligence floor. Toll may use the same measures in spend policy without making one local or frontier win universal routing authority.

Evaluation before and after training

Soulforge pins any proposal to expected change, baseline Outcomes, holdout evidence, and unacceptable regressions. Post-training work uses that contract or makes every change visible; training-facing improvement cannot promote. Riddle returns evidence, neither selects corpus nor registers model; passing an identity/behavior Trial Suite grants no Persona, Sigil, tool, or privileged route.

Returning findings across a Composition Suite

An exact, version-pinned Composition Suite may return a consumer consequence as evidence without reverse execution. It must retain member Composition/Pattern revisions, handoffs, failing observation, Rubric/Evaluator/Environment/verdict/ uncertainty, and declared artifact/evidence dependencies. Composition Suites coordinate applications but do not merge rows, secrets, Sigils, approvals, policies, or effect authority.

Inert record Law
CompositionSuiteFindingSet@1 binds Composition Suite/Rubric, subjects, Environment, observations/measures, Evaluator, verdicts, uncertainty
AttributionCandidate@1 possible boundary, supporting/conflicting evidence, rivals, uncertainty; never causal certainty
InvalidationSet@1 claims whose support fails and claims with intact closure
CorrectionRequest@1 bounded owner delta, preserved constraints, evidence, scope, repair budget

These grant no authority, spend, publication, deletion, training, promotion, or mutation. Spellweaver may admit a new forward Invocation under ordinary policy/HitL; old Runs/Outcomes remain lineage. Riddle walks declared dependencies backwards to the smallest supported cut, not nearest producer. Reuse requires matching complete input closure, artifact revisions, Rubric, Evaluator, relevant Environment, and evidence contract. A failing consumer does not condemn shared artifacts. Missing lineage, flakiness, contagion, capture, or rival explanations yield UNKNOWN/DISPUTED, then at most a broader bounded trial—not reconstructed history or convenient blame.

Consequences

Accepted

Claims are reproducible and scoped; receipts, judgment, and absence remain distinct; pressure measures restraint without rewarding blanket refusal; consumers retain their own authority.

Cost

Cases, controls, environments, calibration, and evidence maintenance cost ongoing human, hardware, and provider effort; observability may leave attribution disputed and claims stale.

Acceptance evidence

Riddle remains Designed until one versioned Trial Suite with controls distinguishes subject/harness/ evaluator failure, reproduces an Outcome, calibrates each non-deterministic Evaluator, preserves raw evidence/uncertainty, and proves routing/repair consumers reapply their own policy. State of Work, not this ADR alone, records promotion.