An automatic check failed
5 of 34 calls the rule could decide (14.7%). First flagged 2026-09-18T20:06:11Z, latest 2026-09-19T01:18:14Z (in this range).
A starting point for a guided review of this local demonstration app: six journeys, each showing real computed content from a page that already exists (the same range and filter choices carry across this page) — not a plain list of links. Every section states its own source — real recorded calls, a synthetic simulation, or the local weather pilot — and they are never blended into one figure.
One real demo call made to the Flight Service phone line, opened to everything the archive kept for it — the current time range chosen below applies everywhere on this page.
Every recorded call with a usable start time, as of Tue 29 Sep 2026 14:14 UTC (Tue 29 Sep 2026 09:14 CDT).
| Call | When (UTC) | Outcome | Recording | Transcript | Weather | Findings |
|---|---|---|---|---|---|---|
| #58 | Mon 28 Sep 2026 15:36 UTC | Completed | Recording | 25 turns | Stored | 0 |
| #57 | Sun 27 Sep 2026 20:21 UTC | Completed | Recording | 27 turns | None | 0 |
| #56 | Wed 23 Sep 2026 18:05 UTC | Completed | Recording | 45 turns | Stored | 0 |
| #55 | Wed 23 Sep 2026 13:10 UTC | Not connected | None | None | None | 0 |
Whether a briefing covered what it needed to — a small, fixed set of seven local checks of what was recorded (text and fields). No recording is listened to and no model judges a call; a call a check cannot decide is "unknown", never counted as fine.
Found by a small fixed set of local checks of what was recorded (text and fields). No recording was listened to and no model judged a call. A call a check cannot decide is "unknown", and is never counted as fine. Counted over 58 calls in All recorded — never over the synthetic examples in the separate synthetic lab, which are never combined with these. Review status: none is recorded here; the human review of real calls stays on the archive's own Review & Learn board above, and these checks never close or reopen anything.
5 of 34 calls the rule could decide (14.7%). First flagged 2026-09-18T20:06:11Z, latest 2026-09-19T01:18:14Z (in this range).
3 of 34 calls the rule could decide (8.8%). First flagged 2026-09-18T20:06:11Z, latest 2026-09-20T06:47:37Z (in this range).
3 of 4 calls the rule could decide (a share is shown from 30 calls). First flagged 2026-09-18T22:37:31Z, latest 2026-09-19T01:18:14Z (in this range).
1 of 16 calls the rule could decide (a share is shown from 30 calls). First flagged 2026-09-19T01:18:14Z, latest 2026-09-19T01:18:14Z (in this range).
Examples: call
Open the full Briefing quality tab · or see the synthetic lab's own briefing-coverage check instead (a different mechanism — which lookups ran per section, not these same seven checks)
Every recommended action is rule-based or a person's own recorded finding — never a language model's judgement — and carries its own evidence. A recommendation is a hypothesis, never a decision made for a person.
Read-only mirror of the archive's own Review & Learn board. Nothing here can be accepted, fixed, verified or dismissed — that stays in the archive's own tool. "Implemented, verification not recorded" means implemented, not independently checked; a "Marked verified in archive" record is shown as the archive marks it, and its evidence is not included here — this page never claims to have checked it again. A reference or a history note may be shown truncated.
8 of 8 finding(s) shown.
Plain historical rates for the real recorded calls — never a calibrated per-call probability, never a causal claim.
Two clearly specified questions about a future demo call. For each, the forecast is the observed historical rate — the baseline expectation for the next comparable call — and beside it is a 95% range for how much that underlying rate could reasonably differ from what was observed: a statement about the rate's own uncertainty, not about the odds of any one specific future call. Below 5 eligible calls, only counts are shown: too little evidence for either.
| Question | Observed and forecast |
|---|---|
| Will it connect? (of calls with a known connection status) | 34 of 58 — 58.6% (95% range for the underlying rate: 45.8%–70.4%) |
| If it completes, will it run longer than 5 minutes? (of completed calls with a recorded length) | 15 of 33 — 45.5% (95% range for the underlying rate: 29.8%–62%) |
1 connected call(s) did not complete (still going, or failed — the archive does not distinguish these here). They count toward the connection forecast above (they did connect); this dashboard does not offer a separate completion forecast from them, because a clean rate would need to guess which side each belongs on.
Call length is how long a completed call lasted. It is not how fast the service answered, and it is not linked to any timing record — see below.
The same small, fixed, local checks the Improvement queue uses, over the calls in this range: what was flagged most, each with a recommended action and the calls it was flagged on. No recording is listened to and no model judges a call; a call a check cannot decide is left out of that check's own rate, never counted as passing.
5 of 34 (14.7%) the check could decide, flagged this. First flagged 2026-09-18T20:06:11Z, latest 2026-09-19T01:18:14Z.
Recommended: Open each call's automatic-check results and address what failed.
3 of 34 (8.8%) the check could decide, flagged this. First flagged 2026-09-18T20:06:11Z, latest 2026-09-20T06:47:37Z.
Recommended: Review these calls' transcripts and confirm the required notice is being given consistently.
3 of 4 (a share is shown from 30 calls; counts only below that) the check could decide, flagged this. First flagged 2026-09-18T22:37:31Z, latest 2026-09-19T01:18:14Z.
Recommended: Verify or dismiss the open or implemented review finding on these calls, in the archive's own Review & Learn tool.
Every check, and the rest of the flagged calls: Improvement queue.
A separate, bounded, local pilot: real public weather from NOAA's Aviation Weather Center at 3 real airports — a modest, explicitly NOT nationwide footprint — a labelled hypothetical flight compared against that same real weather, and wholly fabricated stress-test cases. Three provenance modes, never blended.
The synthetic request lab's own simulation, walked through as one story: load and timing, three forecasting methods compared on a chronological holdout, a historical replay, and a simulated human approval that can be rolled back. Never a measurement of the real service.
Currently open for a decision, highest priority first (the same items the Improvement queue and Overview show):
Synthetic stand-in — from the synthetic lab's own Improvement queue, never the recorded-calls archive.
26 synthetic requests finished with under 1.5 seconds of the 15-second limit left; 1 of them skipped a lookup, so part of what the briefing checks was not assessed; the rest of that section may be complete. The service stops starting new lookups when time is nearly gone, so a briefing that close can leave something a pilot needs unchecked.
Automated tests check the wording the service generates for speech. Nothing on this dashboard checks what a caller hears.
This is a local demonstration app for review. It carries no endorsement from the FAA, DOT, or any other authority, and nothing here is a production-readiness or compliance claim.
What would measure it: A formal review by the relevant authority, using its own process.
Every recommended action shown here comes from a small, fixed, local rule or a person's own recorded finding — never a language model's judgement of a call or a record. Where a check is rules-based, this page and the page it links to both say so.
What would measure it: A labelled AI-assisted review, which none of these pages claim to be.
Each journey above states its own evidence tier (real recorded calls, a synthetic simulation, or a local weather pilot) and its own stated uncertainty. They are never combined into one score, and none of them is a validated accuracy, acoustic, live-latency or compliance measurement.
What would measure it: A calibration study specific to one method, over real records, which is not built here.