PAILEvidence Runtime
Release evidence

Claims with a workload. And a boundary.

01

Labeled evidence-control campaign

A deliberately messy record fixture measured useful answers, wrong releases, refusals, and conflict handling. No external LLM was used.

MESSY RECORDS4,980Across 3,000 generated business cases.
ANSWERABLE COVERAGE98%147 of 150 answerable questions released the expected fact.
WRONG RELEASES0No wrong answer on this fixture; three safe false refusals remained.
CORRECT REFUSAL100%150 absent answers and 130 conflicts withheld.

The comparison baseline was deliberately naive and always answered. Its failure rate is an unsupported-commitment baseline, not a measured universal LLM hallucination rate.

02

Payment and POS safety campaign

Terminal CSV, switch JSONL, settlement rows, and operations logs were joined by explicit transaction evidence.

CROSS-SOURCE RECORDS4,027Generated records across four source shapes.
GUARDED OUTCOMES140/140100 answerable, 20 conflicts, and 20 absent identifiers.
INSTRUCTION QUARANTINE7/7Unsafe message fields withheld while safe facts remained.
LOCAL P50 / P950.99 / 1.30sOne Mac, sequential single-process run.

This is not a PhonePe workload, payment-network certification, hardware acceptance test, or customer result.

03

Current public release checks

The site and Worker contract are tested separately from the private rule implementation.

PUBLIC CONTRACT42Automated JavaScript checks passing in this release tree.
PRIVATE GATEWAY4Real PAIL ingest, query, tenant isolation, duplicate-name preservation, and deletion checks.
CONFLICT TO LLM0Top-level conflict response is withheld even on HTTP 200.
TRIAL LIMIT10The eleventh protected query is rejected before runtime work.
04

Behavior claim ledger

A pass establishes only the named boundary.

ClaimMeasured behaviorDoes not prove
Exact-anchor isolationThe requested transaction excludes neighboring entities.Broad semantic retrieval quality.
Conflict withholdingRequired identity fields disagree and no owner is chosen.Resolution of the source conflict.
Instruction quarantineInstruction-like evidence fields are withheld.Coverage against every prompt attack.
Proof-preserving reductionSelected evidence retains identifiers and provenance.One universal token-reduction percentage.
Point-in-time scopeEvidence after an explicit cutoff is excluded.Complete event-sourcing semantics.
05

What remains unproven

These require customer data, external infrastructure, or independent review.

CUSTOMER ACCURACY

Held-out domain questions

Measure against source-of-truth labels from a real workflow.

INDEPENDENTLY UNPROVEN
PRODUCTION RELIABILITY

Concurrency and recovery

Run multi-user soak, backup, restore, failover, and deletion drills.

NO PRODUCTION SLA
SECURITY ASSURANCE

External assessment

Complete penetration testing, managed-key review, and customer identity integration.

NOT CERTIFIED
BUSINESS IMPACT

Operational outcome

Measure prevented wrong releases, review time, and workflow cost.

NEEDS DESIGN PARTNER
The next meaningful proof

Run held-out customer questions against the current baseline.

Compare wrong joins, unsupported answers, retained answer coverage, latency, and context volume.