KBM Nexus Flight Service
Operations analysis
Mixed — each section below states its own source
Engineering details: shown, hide them

Review guide

A starting point for a guided review of this local demonstration app: six journeys, each showing real computed content from a page that already exists (the same range and filter choices carry across this page) — not a plain list of links. Every section states its own source — real recorded calls, a synthetic simulation, or the local weather pilot — and they are never blended into one figure.

This page reuses, it does not recompute. Every figure and recommendation below comes from the same function the destination tab itself calls, over the same current range; nothing here is a second calculation that could disagree with it. Open the dashboard directly if you would rather explore without a guided path.

A recorded call, in full

Recorded demo calls

One real demo call made to the Flight Service phone line, opened to everything the archive kept for it — the current time range chosen below applies everywhere on this page.

Sat 26 Sep 2026 14:53 UTC (Sat 26 Sep 2026 09:53 CDT) to Tue 29 Sep 2026 14:53 UTC (Tue 29 Sep 2026 09:53 CDT), by the server's clock right now. The boundary is inclusive: a call exactly at the start of the window is included, one second earlier is not.

Not counted here: 56 older than this window.

The most recent recorded calls
CallWhen (UTC)OutcomeRecordingTranscriptWeatherFindings
#58Mon 28 Sep 2026 15:36 UTCCompletedRecording25 turnsStored0
#57Sun 27 Sep 2026 20:21 UTCCompletedRecording27 turnsNone0

See all 2 calls in this range, each opening to its recording, transcript, captured flight details and stored weather briefing

Briefing completeness: seven rule-based checks

Recorded demo calls

Whether a briefing covered what it needed to — a small, fixed set of seven local checks of what was recorded (text and fields). No recording is listened to and no model judges a call; a call a check cannot decide is "unknown", never counted as fine.

Found by a small fixed set of local checks of what was recorded (text and fields). No recording was listened to and no model judged a call. A call a check cannot decide is "unknown", and is never counted as fine. Counted over 2 calls in Last 72 hours — never over the synthetic examples in the separate synthetic lab, which are never combined with these. Review status: none is recorded here; the human review of real calls stays on the archive's own Review & Learn board above, and these checks never close or reopen anything.

No check flagged a call in this range.

Open the full Briefing quality tab · or see the synthetic lab's own briefing-coverage check instead (a different mechanism — which lookups ran per section, not these same seven checks)

Improvement queue and Review & Learn

Recorded demo calls

Every recommended action is rule-based or a person's own recorded finding — never a language model's judgement — and carries its own evidence. A recommendation is a hypothesis, never a decision made for a person.

Read-only mirror of the archive's own Review & Learn board. Nothing here can be accepted, fixed, verified or dismissed — that stays in the archive's own tool. "Implemented, verification not recorded" means implemented, not independently checked; a "Marked verified in archive" record is shown as the archive marks it, and its evidence is not included here — this page never claims to have checked it again. A reference or a history note may be shown truncated.

Clear filters

0 of 0 finding(s) shown.

No finding matches this filter.

Open the full Improvement queue (recorded)

Predictive evaluation of the recorded calls

Recorded demo calls Predicted

Plain historical rates for the real recorded calls — never a calibrated per-call probability, never a causal claim.

Forecast for the next comparable call

Recorded demo calls

Two clearly specified questions about a future demo call. For each, the forecast is the observed historical rate — the baseline expectation for the next comparable call — and beside it is a 95% range for how much that underlying rate could reasonably differ from what was observed: a statement about the rate's own uncertainty, not about the odds of any one specific future call. Below 5 eligible calls, only counts are shown: too little evidence for either.

Forecast for the next comparable demo call, over the 2 recorded calls in Last 72 hours
QuestionObserved and forecast
Will it connect? (of calls with a known connection status)2 of 2 — insufficient evidence for a forecast (fewer than 5 eligible calls)
If it completes, will it run longer than 5 minutes? (of completed calls with a recorded length)1 of 2 — insufficient evidence for a forecast (fewer than 5 eligible calls)

Call length is how long a completed call lasted. It is not how fast the service answered, and it is not linked to any timing record — see below.

Assumes a future demo call is comparable to the ones counted here. Not validated against outcomes this estimate did not see, and not a backtested forecast: this dataset does not record when each field became known during a call, so there is no way to check what it would have said at the time. Does not adjust for repeated calls from the same caller — the archive does label callers (see the Calls tab); this estimate simply does not use that label — or for a change to the service between calls; either can make the calls counted less independent and less comparable than a plain historical rate assumes.

Recurring patterns in this range

Recorded demo calls

The same small, fixed, local checks the Improvement queue uses, over the calls in this range: what was flagged most, each with a recommended action and the calls it was flagged on. No recording is listened to and no model judges a call; a call a check cannot decide is left out of that check's own rate, never counted as passing.

No check flagged a call in this range.

Every check, and the rest of the flagged calls: Improvement queue.

Open the full Predictive Analysis tab

Airport weather and simulated flight impact

Simulated

A separate, bounded, local pilot: real public weather from NOAA's Aviation Weather Center at 3 real airports — a modest, explicitly NOT nationwide footprint — a labelled hypothetical flight compared against that same real weather, and wholly fabricated stress-test cases. Three provenance modes, never blended.

  1. Ceiling & Visibility Outlook — the real TAF at the cutoff, real later observations as the outcome, and — for a past example — how a simple "hold the last observation" persistence baseline compares against what was actually observed

The synthetic engine, end to end

Synthetic

The synthetic request lab's own simulation, walked through as one story: load and timing, three forecasting methods compared on a chronological holdout, a historical replay, and a simulated human approval that can be rolled back. Never a measurement of the real service.

Currently open for a decision, highest priority first (the same items the Improvement queue and Overview show):

Synthetic stand-in — from the synthetic lab's own Improvement queue, never the recorded-calls archive.

High Synthetic example

Some briefings finished with almost no time left

26 synthetic requests finished with under 1.5 seconds of the 15-second limit left; 1 of them skipped a lookup, so part of what the briefing checks was not assessed; the rest of that section may be complete. The service stops starting new lookups when time is nearly gone, so a briefing that close can leave something a pilot needs unchecked.

Recommended action (rule-based)
Hypothesis to test: fetching the shared feeds ahead of time, so requests do not wait for them, leaves more time for the sections that matter. This data does not show that it would.
How we would know (uncertainty)
Run the pre-fetch on alternate days, or on a random half of requests, for long enough to see a few dozen near-limit briefings in each group, then compare how many finish near the limit.
High Synthetic example

Voice quality is not measured

Automated tests check the wording the service generates for speech. Nothing on this dashboard checks what a caller hears.

Recommended action (rule-based)
Start a weekly listening sample of real recordings. Listen for the altimeter setting and the winds first.
How we would know (uncertainty)
Done when a person has listened to a sample and recorded, for each call, whether the altimeter setting and winds were heard exactly as written.
  1. System performance (synthetic) — timing and budget behaviour across the synthetic window
  2. Trends (synthetic) — how the synthetic figures move over the window
  3. Predictive Analysis (synthetic) — three simple methods compared on a chronological holdout (never data the method could have seen), a historical replay, and — under "Decision: a person approves" — a simulated human-in-the-loop approval and rollback: no sign-in, no live model, and a log of decisions that is not proof anything helped
  4. Data readiness — the dataset-readiness gate behind these figures: what passed structural validation and what did not, and why predictive readiness and accuracy are stated as not assessed here

What this guide is not

Not measured

An endorsement or certification Not measured

This is a local demonstration app for review. It carries no endorsement from the FAA, DOT, or any other authority, and nothing here is a production-readiness or compliance claim.

What would measure it: A formal review by the relevant authority, using its own process.

An AI-generated recommendation Not measured

Every recommended action shown here comes from a small, fixed, local rule or a person's own recorded finding — never a language model's judgement of a call or a record. Where a check is rules-based, this page and the page it links to both say so.

What would measure it: A labelled AI-assisted review, which none of these pages claim to be.

A single accuracy figure for the whole app Not measured

Each journey above states its own evidence tier (real recorded calls, a synthetic simulation, or a local weather pilot) and its own stated uncertainty. They are never combined into one score, and none of them is a validated accuracy, acoustic, live-latency or compliance measurement.

What would measure it: A calibration study specific to one method, over real records, which is not built here.