KBM Nexus Flight Service
Operations analysis
Mixed — each section below states its own source
Engineering details: hidden, show them

Review guide

A starting point for a guided review of this local demonstration app: six journeys, each showing real computed content from a page that already exists (the same range and filter choices carry across this page) — not a plain list of links. Every section states its own source — real recorded calls, a synthetic simulation, or the local weather pilot — and they are never blended into one figure.

This page reuses, it does not recompute. Every figure and recommendation below comes from the same function the destination tab itself calls, over the same current range; nothing here is a second calculation that could disagree with it. Open the dashboard directly if you would rather explore without a guided path.

A recorded call, in full

Recorded demo calls

One real demo call made to the Flight Service phone line, opened to everything the archive kept for it — the current time range chosen below applies everywhere on this page.

Every recorded call with a usable start time, as of Tue 29 Sep 2026 14:54 UTC (Tue 29 Sep 2026 09:54 CDT).

The most recent recorded calls
CallWhen (UTC)OutcomeRecordingTranscriptWeatherFindings
#58Mon 28 Sep 2026 15:36 UTCCompletedRecording25 turnsStored0
#57Sun 27 Sep 2026 20:21 UTCCompletedRecording27 turnsNone0
#56Wed 23 Sep 2026 18:05 UTCCompletedRecording45 turnsStored0
#55Wed 23 Sep 2026 13:10 UTCNot connectedNoneNoneNone0

See all 58 calls in this range, each opening to its recording, transcript, captured flight details and stored weather briefing

Briefing completeness: seven rule-based checks

Recorded demo calls

Whether a briefing covered what it needed to — a small, fixed set of seven local checks of what was recorded (text and fields). No recording is listened to and no model judges a call; a call a check cannot decide is "unknown", never counted as fine.

Found by a small fixed set of local checks of what was recorded (text and fields). No recording was listened to and no model judged a call. A call a check cannot decide is "unknown", and is never counted as fine. Counted over 58 calls in All recorded — never over the synthetic examples in the separate synthetic lab, which are never combined with these. Review status: none is recorded here; the human review of real calls stays on the archive's own Review & Learn board above, and these checks never close or reopen anything.

M-AUTOMATIC_CHECKS Flagged on 5 calls in this range

An automatic check failed

5 of 34 calls the rule could decide (14.7%). First flagged 2026-09-18T20:06:11Z, latest 2026-09-19T01:18:14Z (in this range).

Examples: call · call · call · call · call

M-ADVISORY_NOTICE Flagged on 3 calls in this range

The agent spoke but did not give the notice

3 of 34 calls the rule could decide (8.8%). First flagged 2026-09-18T20:06:11Z, latest 2026-09-20T06:47:37Z (in this range).

Examples: call · call · call

M-FINDINGS_VERIFIED Flagged on 3 calls in this range

A review finding is open, or implemented and not yet verified

3 of 4 calls the rule could decide (a share is shown from 30 calls). First flagged 2026-09-18T22:37:31Z, latest 2026-09-19T01:18:14Z (in this range).

Examples: call · call · call

M-WEATHER_LOOKUP Flagged on 1 call in this range

The weather lookup returned an error

1 of 16 calls the rule could decide (a share is shown from 30 calls). First flagged 2026-09-19T01:18:14Z, latest 2026-09-19T01:18:14Z (in this range).

Examples: call

Open the full Briefing quality tab · or see the synthetic lab's own briefing-coverage check instead (a different mechanism — which lookups ran per section, not these same seven checks)

Improvement queue and Review & Learn

Recorded demo calls

Every recommended action is rule-based or a person's own recorded finding — never a language model's judgement — and carries its own evidence. A recommendation is a hypothesis, never a decision made for a person.

Read-only mirror of the archive's own Review & Learn board. Nothing here can be accepted, fixed, verified or dismissed — that stays in the archive's own tool. "Implemented, verification not recorded" means implemented, not independently checked; a "Marked verified in archive" record is shown as the archive marks it, and its evidence is not included here — this page never claims to have checked it again. A reference or a history note may be shown truncated.

Clear filters

8 of 8 finding(s) shown.

Implemented, verification not recorded phrasing medium call 28
What went wrong
Mike: read field elevation and density altitude, not how high the density altitude is above field elevation.
What was expected
field elevation X feet, density altitude Y feet
Must say
field elevation · density altitude
Must not say
above field elevation
Reference recorded with the status
flight-service 9262aed
History (1)
  • Sat 19 Sep 2026 03:24 UTC — Implemented, verification not recorded: Mike's request on the call
Implemented, verification not recorded platform high call 28
What went wrong
The caller's 'Yes' abandoned the weather lookup while it was running ('Tool execution was abandoned due to user input'); Albert had to call again.
What was expected
The lookup finishes even if the caller speaks.
Reference recorded with the status
voice platform tool setting interruption_mode = disable_during_tool (19 Sep)
History (1)
  • Sat 19 Sep 2026 03:24 UTC — Implemented, verification not recorded: config change, no briefing-text test
Not a problem in the archive other low call 28
What went wrong
Automatic check failed on the freezing-level G-AIRMET ('No SIGMETs or airmets on route' followed by the freezing level).
What was expected
Reference recorded with the status
evaluation criterion clarified 19 Sep
History (1)
  • Sat 19 Sep 2026 03:24 UTC — Not a problem in the archive: dismissed: the check was wrong
Not a problem in the archive other low call 27
What went wrong
Automatic check failed on the freezing-level G-AIRMET ('No SIGMETs or airmets on route' followed by the freezing level).
What was expected
Reference recorded with the status
evaluation criterion clarified 19 Sep
History (1)
  • Sat 19 Sep 2026 03:24 UTC — Not a problem in the archive: dismissed: the check was wrong
Implemented, verification not recorded flow medium call 26
What went wrong
The greeting was cut off by 'Mm-hmm', and Albert repeated the advisory statement three times before taking the flight.
What was expected
Greeting plays once, uninterrupted; acknowledgements don't interrupt; advisory said once.
Reference recorded with the status
agent config 19 Sep: greeting not interruptible, acknowledgement ignore terms, …
History (1)
  • Sat 19 Sep 2026 03:24 UTC — Implemented, verification not recorded: config change
Implemented, verification not recorded readback medium call 26
What went wrong
DD-175 blocks not taken: alternate, ETE to alternate, unit/home base; pilot rank handled as 'not part of a filed plan'.
What was expected
Take every DD-175 block and read it back in block order.
Reference recorded with the status
agent prompt 19 Sep: DD Form 175 section (blocks 1-16 + unit)
History (1)
  • Sat 19 Sep 2026 03:24 UTC — Implemented, verification not recorded: verified on Mike's 01:18Z call
Implemented, verification not recorded wrong value high call 24
What went wrong
Arrival forecast wind was read from the TAF period that had already ended: 190 at 6, while the FM190000 group (150 at 3) governed the ETA of 0045Z. Ties in the +/-1 h window went to the first period scanned.
What was expected
Forecast for McAlester at the arrival time: wind one five zero at three.
Must say
one five zero at three
Must not say
one nine zero at six
Reference recorded with the status
flight-service 8b8e750 (TAF window tie-break)
History (1)
  • Sat 19 Sep 2026 03:24 UTC — Implemented, verification not recorded: from the 18 Sep audit
Implemented, verification not recorded missing item medium call 24
What went wrong
Freezing level was not stated, and the automatic check flagged 'No SIGMETs or airmets' against the freezing-level G-AIRMET.
What was expected
State the freezing level after the adverse conditions.
Must say
Freezing level
Reference recorded with the status
flight-service 8b8e750; evaluation criterion clarified 19 Sep
History (1)
  • Sat 19 Sep 2026 03:24 UTC — Implemented, verification not recorded: from the 18 Sep audit

Open the full Improvement queue (recorded)

Predictive evaluation of the recorded calls

Recorded demo calls Predicted

Plain historical rates for the real recorded calls — never a calibrated per-call probability, never a causal claim.

Forecast for the next comparable call

Recorded demo calls

Two clearly specified questions about a future demo call. For each, the forecast is the observed historical rate — the baseline expectation for the next comparable call — and beside it is a 95% range for how much that underlying rate could reasonably differ from what was observed: a statement about the rate's own uncertainty, not about the odds of any one specific future call. Below 5 eligible calls, only counts are shown: too little evidence for either.

Forecast for the next comparable demo call, over the 58 recorded calls in All recorded
QuestionObserved and forecast
Will it connect? (of calls with a known connection status)34 of 58 — 58.6% (95% range for the underlying rate: 45.8%–70.4%)
If it completes, will it run longer than 5 minutes? (of completed calls with a recorded length)15 of 33 — 45.5% (95% range for the underlying rate: 29.8%–62%)

1 connected call(s) did not complete (still going, or failed — the archive does not distinguish these here). They count toward the connection forecast above (they did connect); this dashboard does not offer a separate completion forecast from them, because a clean rate would need to guess which side each belongs on.

Call length is how long a completed call lasted. It is not how fast the service answered, and it is not linked to any timing record — see below.

Assumes a future demo call is comparable to the ones counted here. Not validated against outcomes this estimate did not see, and not a backtested forecast: this dataset does not record when each field became known during a call, so there is no way to check what it would have said at the time. Does not adjust for repeated calls from the same caller — the archive does label callers (see the Calls tab); this estimate simply does not use that label — or for a change to the service between calls; either can make the calls counted less independent and less comparable than a plain historical rate assumes.

Recurring patterns in this range

Recorded demo calls

The same small, fixed, local checks the Improvement queue uses, over the calls in this range: what was flagged most, each with a recommended action and the calls it was flagged on. No recording is listened to and no model judges a call; a call a check cannot decide is left out of that check's own rate, never counted as passing.

Recurring in this range

An automatic check failed

5 of 34 (14.7%) the check could decide, flagged this. First flagged 2026-09-18T20:06:11Z, latest 2026-09-19T01:18:14Z.

Recommended: Open each call's automatic-check results and address what failed.

Evidence: call · call · call · call · call

Recurring in this range

The agent spoke but did not give the notice

3 of 34 (8.8%) the check could decide, flagged this. First flagged 2026-09-18T20:06:11Z, latest 2026-09-20T06:47:37Z.

Recommended: Review these calls' transcripts and confirm the required notice is being given consistently.

Evidence: call · call · call

Recurring in this range

A review finding is open, or implemented and not yet verified

3 of 4 (a share is shown from 30 calls; counts only below that) the check could decide, flagged this. First flagged 2026-09-18T22:37:31Z, latest 2026-09-19T01:18:14Z.

Recommended: Verify or dismiss the open or implemented review finding on these calls, in the archive's own Review & Learn tool.

Evidence: call · call · call

Every check, and the rest of the flagged calls: Improvement queue.

Open the full Predictive Analysis tab

Airport weather and simulated flight impact

Simulated

A separate, bounded, local pilot: real public weather from NOAA's Aviation Weather Center at 3 real airports — a modest, explicitly NOT nationwide footprint — a labelled hypothetical flight compared against that same real weather, and wholly fabricated stress-test cases. Three provenance modes, never blended.

  1. Ceiling & Visibility Outlook — the real TAF at the cutoff, real later observations as the outcome, and — for a past example — how a simple "hold the last observation" persistence baseline compares against what was actually observed

The synthetic engine, end to end

Synthetic

The synthetic request lab's own simulation, walked through as one story: load and timing, three forecasting methods compared on a chronological holdout, a historical replay, and a simulated human approval that can be rolled back. Never a measurement of the real service.

Currently open for a decision, highest priority first (the same items the Improvement queue and Overview show):

Synthetic stand-in — from the synthetic lab's own Improvement queue, never the recorded-calls archive.

High Synthetic example

Some briefings finished with almost no time left

26 synthetic requests finished with under 1.5 seconds of the 15-second limit left; 1 of them skipped a lookup, so part of what the briefing checks was not assessed; the rest of that section may be complete. The service stops starting new lookups when time is nearly gone, so a briefing that close can leave something a pilot needs unchecked.

Recommended action (rule-based)
Hypothesis to test: fetching the shared feeds ahead of time, so requests do not wait for them, leaves more time for the sections that matter. This data does not show that it would.
How we would know (uncertainty)
Run the pre-fetch on alternate days, or on a random half of requests, for long enough to see a few dozen near-limit briefings in each group, then compare how many finish near the limit.
High Synthetic example

Voice quality is not measured

Automated tests check the wording the service generates for speech. Nothing on this dashboard checks what a caller hears.

Recommended action (rule-based)
Start a weekly listening sample of real recordings. Listen for the altimeter setting and the winds first.
How we would know (uncertainty)
Done when a person has listened to a sample and recorded, for each call, whether the altimeter setting and winds were heard exactly as written.
  1. System performance (synthetic) — timing and budget behaviour across the synthetic window
  2. Trends (synthetic) — how the synthetic figures move over the window
  3. Predictive Analysis (synthetic) — three simple methods compared on a chronological holdout (never data the method could have seen), a historical replay, and — under "Decision: a person approves" — a simulated human-in-the-loop approval and rollback: no sign-in, no live model, and a log of decisions that is not proof anything helped
  4. Data readiness — the dataset-readiness gate behind these figures: what passed structural validation and what did not, and why predictive readiness and accuracy are stated as not assessed here

What this guide is not

Not measured

An endorsement or certification Not measured

This is a local demonstration app for review. It carries no endorsement from the FAA, DOT, or any other authority, and nothing here is a production-readiness or compliance claim.

What would measure it: A formal review by the relevant authority, using its own process.

An AI-generated recommendation Not measured

Every recommended action shown here comes from a small, fixed, local rule or a person's own recorded finding — never a language model's judgement of a call or a record. Where a check is rules-based, this page and the page it links to both say so.

What would measure it: A labelled AI-assisted review, which none of these pages claim to be.

A single accuracy figure for the whole app Not measured

Each journey above states its own evidence tier (real recorded calls, a synthetic simulation, or a local weather pilot) and its own stated uncertainty. They are never combined into one score, and none of them is a validated accuracy, acoustic, live-latency or compliance measurement.

What would measure it: A calibration study specific to one method, over real records, which is not built here.