experiment / evidence
Browser-localMethod & privacy
BEFORE YOU CALL IT A WIN

Does the evidence support the next step?

Turn experiment exports into a decision you can trace. See who counted, what broke, and how much uncertainty remains.

A fictional support-assistant pilot. Real calculations. One click to inspect.

No file upload · No sign-in · No analytics
ANALYSIS NOTE EE / 001
A lift is a question.
Evidence is the answer.
01 ASSIGNMENTS
02 EXPOSURES
03 OUTCOME EVENTS
Inspect the whole storyIntegrity · Uncertainty · Guardrails
01

Bring the raw evidence

Three linked CSV or NDJSON exports. Explicit columns and event definitions.

02

Challenge the apparent lift

Preserve assignments. Check missing follow-up, allocation and adverse outcomes.

03

Keep a reproducible record

Download the definitions, source fingerprints and calculations. Reopen to rerun.

Already have a replay pack?

Reopen it locally and recalculate from its original inputs.

KNOW THE BOUNDARIES

Useful evidence.
Honest limits.

What this analysis assumes

One prespecified binary outcome, one independent randomized user per unit, two arms and a fixed final horizon. Absence of an event means no occurrence only when the analyst declares complete coverage. Continuous, ratio and clustered metrics are unsupported. No adjustment for peeking, specification search or interference. Funnels and exposure views are exploratory.

How the review gates work

Blocking integrity problems, allocation mismatch (SRM p < .001) or an observed adverse increase above tolerance → Investigate. Missing declarations, immature windows, fewer than 50 mature users per arm, absent guardrail or uncertain estimates → Insufficient evidence. Otherwise, the primary lower 95% interval must clear the meaningful threshold, Fisher p must be < .05, and the guardrail upper two-sided 95% endpoint must stay within tolerance. All conditions are required. Fifty is an operational gate, not a power calculation.

Rates: Wilson 95%. Treatment-minus-control difference: Newcombe 95%, without continuity correction. Fisher exact: two-sided probability ordering on [[control outcomes, control non-outcomes], [treatment outcomes, treatment non-outcomes]]. Methods can disagree; review requires both. No corrections for multiple candidate outcomes or repeated looks.

SciPy Fisher reference · statsmodels interval reference

Privacy, storage & reproducibility

Imported content is processed in your browser. The app intentionally makes no data requests or application-storage writes. Cloudflare receives normal page-request metadata. Deliberate downloads persist on your device. Reset releases application references; it does not guarantee secure erasure from browser memory or disk.

The full replay pack contains raw input, mapping and specification fingerprints, methods, audit rules and results. Reopening recalculates from raw evidence and verifies fingerprints. Hashes detect changes, not authenticity. Use anonymized exports; inspect downloads before sharing. Reloading the page loses the current workspace.

Experiment Evidence / independent portfolio toolEngine 1.0.0 · Binary outcomes only