Evidence matrix · observations, not a reliability rate

Claims stop where the receipts stop.

SkipHow's outcome-first orchestration separates two kinds of proof. Package checks can prove files and invariants. Only controlled model runs can support behavior claims.

What controlled runs have shown.

Observation What the run established Boundary
Fully specified request Work proceeded without asking the owner to choose engineering mechanics. Observed in isolated fixtures; not a general selection rate.
Open product choice The owner received the visible product choice and a recommendation before dependent work was written. Some boundary cases varied by host and are documented in the full record.
Intermittent failure A flaky test was diagnosed at its cause without retrying, skipping, or weakening the assertion. One controlled problem class, not a general debugging score.
False-positive fix A plausible change was rejected after its test also passed against unfixed code. Shows the evidence discipline can fire; not that it always does.
Multi-outcome plan A larger build was split into independently verifiable units with only a genuine dependency edge. The split was observed; concurrent delegated execution was not.

Hold the environment fixed. Change only the package.

A useful receipt uses a throwaway fixture repository, a fresh host session carrying only the candidate package and built-ins, the same prompt on both sides of a change, and a control run that proves isolation.

The run must execute through the host's own permissions and budget controls. Isolation is confirmed from the session transcript, not by asking the model what it can see. A failing case runs before the candidate as well as after it.

Deterministic checks are necessary, not behavioral evidence

They verify the package layout, reachable resources, metadata, versions, links, portability, and site structure. They do not start a model and cannot prove how a model will act.

What SkipHow does not claim.

Comparative advantage
No paired product benchmark shows SkipHow is faster, cheaper, more reliable, or better than a base agent or another framework.
Delegation
No controlled pass has demonstrated concurrent lanes, isolated worktrees, or separately integrated delegated units.
Automatic selection
No general rate has been measured for Claude Code or Codex selecting the skill when it is not named.
Production delivery
Real production and public-delivery actions are outside the retained controlled evidence.
Public adoption
Independent case studies and broad public usage remain limited. Popularity is not used as a proxy for correctness.

The concise page is not the source ledger.

The repository's evidence document records host differences, regressions, negative controls, unresolved behavior, and the exact limits on every summary above. Decision history records why rules were kept, changed, or rejected.