Launch evidence guides
How to Evaluate Launch Evidence: A Practical Pilot Scorecard
Published (UTC)
A small launch-evidence pilot should answer a practical question: does this information improve a decision your team already makes?
Start there, rather than with a request for a general demonstration of accuracy. A bounded historical replay can reveal field gaps, review effort, and useful context in a particular workflow. It cannot, by itself, establish global reliability or a delivery guarantee.
This guide provides a proposed buyer-owned scorecard, including an explicitly synthetic row to show the arithmetic. No customer pilot, benchmark, or buyer interview is represented as completed. Thresholds below are examples to replace before a real evaluation.
Name one decision the pilot could improve
Choose a task narrow enough to run twice. For a newsroom, it might be: prepare a source-linked background brief for each eligible launch using a defined evidence standard. For a research team, it might be: decide whether an event record contains enough provenance to enter a review queue.
State what success would let the team do differently. “Explore the dashboard” is an activity. “Produce a traceable brief within our review budget, with fewer unresolved source gaps than our existing process” identifies a decision and the evidence needed to evaluate it.
Keep the proposed use bounded. The public API guide excludes operational navigation, collision avoidance, and emergency use. A retrospective briefing evaluation does not establish suitability for those decisions.
LaunchDetect’s pilot page describes agreeing on a historical replay, comparing it with an existing baseline, and separately scoping a live evaluation if the replay shows value. Scope, measures, timing, and pricing are agreed before the work. The scorecard here adds a method for making that discussion concrete; it introduces no commercial commitment.
Freeze the sample and baseline before seeing results
Write down the eligible events, period, region if relevant, inclusion rule, and exclusions. Use a defined event list appropriate to the question. Selecting only records that already have convenient evidence makes the replay easier but changes what its results mean.
If the sample comes from the public feed, describe it as the selected available records. The API guide documents a bounded recent feed and summaries that omit raw imagery. Those constraints matter when choosing the sample; the feed alone cannot define all events that should have been detected.
Keep an exclusion log. Record the event, the rule applied, the reason, and who applied it. Do not silently remove difficult cases after scoring begins. Report an unusable or missing record according to the prespecified missing-data rule.
Define the baseline as a repeatable task with named sources and an output standard. Fix the allowed information and time budget. Use the same question in both runs. When a person performs both, record order and familiarity because the first run can teach them facts that make the second easier. A larger evaluation may need stronger controls; a small replay should at least expose this limitation.
Use measures with visible denominators
The following scorecard is a proposed starting point. An owner and threshold should be assigned before the replay. Avoid a single weighted score that lets a fast workflow compensate for a critical evidence error.
| Measure | Definition and denominator | Evidence to retain | Example decision rule |
|---|---|---|---|
| Traceable completed briefs | Briefs meeting every required source criterion / all eligible assigned events | Briefs, source checklist, exclusion log | Agree the acceptable fraction before replay |
| Critical claim errors | Count of unsupported or materially wrong claims under a predefined rubric | Claim-level review decisions and corrections | Hold any expansion until critical errors are resolved |
| Analyst effort | Active review minutes per assigned event, including failed or incomplete attempts | Start/stop log and interruption rule | Compare paired tasks; agree a review budget |
| Missing required fields | Missing required values / expected required field opportunities | Field dictionary and captured records | Do not turn missing values into zeros |
| Timeliness | Named endpoint interval for each case with compatible evidence | Publication and receipt history, clock notes | No timing pass when an endpoint is unknown |
| Integration completion | Records transformed and accepted under the agreed contract / eligible attempted records | Validation logs and rejection reasons | Track failures and corrections explicitly |
These measures address different questions. An accurately sourced brief can be too costly to prepare. A quick import can produce an unusable record. Report both outcomes rather than averaging away the distinction.
Separate timeliness from timestamp availability
Choose the timing claim before choosing a subtraction. A scheduled launch time, recorded detection time, feed publication time, and receiver’s retrieval time describe different events. The published field definitions explain those distinctions for the public API.
A later refresh of an old detection does not reveal its original publication delay. To evaluate delivery, retain a suitable first-publication record and receiver logs, or describe the bounds your polling observations actually establish. If those endpoints are absent, mark the metric unscorable for that case and report how many cases were affected.
Do not silently calculate the metric only on convenient successes. A table showing “7 of 10 cases had comparable endpoints; 3 were unscorable” is more useful than a median that conceals the missing cases. The values in this sentence are illustrative, not observed pilot results.
Replay only the information available at the step
Historical records can contain later corrections and enrichment. A replay using today’s complete record may answer “Can we prepare a useful retrospective brief?” while failing to answer “What could the desk have known at that historical moment?”
Choose which question you are testing. For a time-sensitive replay, freeze versions and release information according to an evidenced availability sequence. If the historical sequence cannot be reconstructed, label that limitation and narrow the question. Do not present a retrospective reconstruction as a live alert test.
Capture the output, effort, missing information, source versions, and rationale for each decision. Record disagreements between reviewers rather than forcing agreement simply to fill the scorecard. Leave review status pending until a review has actually occurred.
A synthetic worked row
Suppose a fictional newsroom preselects ten historical events. Its hypothetical rule is that at least eight must produce a brief meeting all required source criteria, with no unresolved critical errors. Missing records stay in the denominator.
In a made-up replay, seven briefs meet the criteria, two lack a required source, and one cannot be completed. The completion measure is 7/10, or 70%. It misses the hypothetical eight-of-ten threshold. The outcome is “revise the scope or stop,” even if the seven completed briefs were quick to prepare.
This row illustrates a decision rule, not LaunchDetect performance. The useful habit is deciding the treatment of all three incomplete outcomes before the replay, so the result does not depend on changing the denominator afterward.
Agree what happens next
Write the stop rules alongside the success rules. Stop or pause when an essential source cannot be obtained, a critical error remains unresolved, the review budget is exceeded, or a required timing endpoint cannot be evidenced. Specify whether the result ends the evaluation or calls for a narrower question.
If the replay supports moving forward, separately agree a live period, event scope, delivery route, support expectations, ownership, and stop criteria. A positive historical result does not automatically establish those terms.
Use the blank scorecard to record the decision, sample, baseline, measure, denominator, evidence, threshold, exclusions, result, and stop rule. Then discuss a scoped pilot with one decision and a proposed baseline in hand. That makes a small evaluation useful whether the answer is proceed, revise, or stop.