KBM Nexus Flight Service
Operations analysis

Improvement queue

Findings and recommended actions, each with its evidence. A recommendation is a hypothesis to test: a pattern in data, or a log that a change was made, does not show that the change helped. "How we would know" says what a fair test would need.

This page always uses the whole synthetic window: about 76 hours, Mon 05 Jan 00:01 UTC to Thu 08 Jan 04:25 UTC (1,200 synthetic requests), whatever time range is chosen on other pages. Dates are synthetic.

Two kinds of item. A synthetic example is a finding in this synthetic data. A known gap is something found by reading the service's source; it is not from any data, and reading the repository does not verify what is running live.
ShowAll items · 12Synthetic examples · 6Known gaps · 6
Real evidence from the recorded demo calls is on this tab too. Show the Review & Learn findings and the recorded checks instead. They are never mixed into the synthetic items below.

Items

Synthetic Known gap Simulated

12 shown, in priority order. Decisions on these items are simulated.

Q-01 High priority Synthetic example System speed Open

Some briefings finished with almost no time left

Why it matters
26 synthetic requests finished with under 1.5 seconds of the 15-second limit left; 1 of them skipped a lookup, so part of what the briefing checks was not assessed; the rest of that section may be complete. The service stops starting new lookups when time is nearly gone, so a briefing that close can leave something a pilot needs unchecked.
Evidence

Examples below are picked (those that skipped a lookup first, then the slowest), not sampled. They show what the cases look like, not how common they are. In total: 26 requests that finished with under 1.5 seconds left, or skipped a lookup.

See it on the System performance page

Recommended action (a hypothesis)
Hypothesis to test: fetching the shared feeds ahead of time, so requests do not wait for them, leaves more time for the sections that matter. This data does not show that it would.
How we would know
Run the pre-fetch on alternate days, or on a random half of requests, for long enough to see a few dozen near-limit briefings in each group, then compare how many finish near the limit. Things getting better over time would not show that the pre-fetch helped.
Q-09 High priority Known gap Voice quality Open

Voice quality is not measured

Why it matters
Automated tests check the wording the service generates for speech. Nothing on this dashboard checks what a caller hears. Whether the altimeter setting and winds were spoken as written can only be found out by listening.
Evidence

From source review Not measured by this dashboard. The real service archives call recordings; this dashboard reads none.

Read from the repository, not verified against the live service. No data is behind this item.

See the Voice quality page

Recommended action
Start a weekly listening sample of real recordings. Listen for the altimeter setting and the winds first.
How we would know
Done when a person has listened to a sample and recorded, for each call, whether the altimeter setting and winds were heard exactly as written.
Q-10 High priority Known gap Briefing quality Open

NOTAMs and special use airspace are not live in any briefing

Why it matters
Every briefing carries placeholder text for NOTAMs and says special use airspace was not checked. Both are things a pilot needs. The briefing marks them as not live, so the gap is disclosed, but it is still a gap.
Evidence

From source review Source review of the request handler.

Read from the repository, not verified against the live service. No data is behind this item.

See the Briefing quality page

Recommended action
Needs official data access that is outside this dashboard. Until then the briefing must keep saying they were not checked.
How we would know
Done when both come from a live source and a sample of briefings is checked against the official version.
Q-11 High priority Known gap Data gaps Open

This dashboard is not connected to any real request timing

Why it matters
It reads only the synthetic set, so nothing on it says anything about the live service. Whether request timing is switched on anywhere was not checked when this was built.
Evidence

From source review This dashboard reads synthetic data only. Whether request timing is enabled on the live service was not checked.

Read from the repository, not verified against the live service. No data is behind this item.

See the System performance page

Recommended action
Find out whether timing capture is enabled. If it is not, enabling it needs the owner's approval and a key kept apart from the call-record key. Then connect this dashboard to the stored records.
How we would know
Done when real request records exist and this dashboard reads them in place of the synthetic set.
Q-02 Medium priority Synthetic example System speed Open

Briefings that needed many new lookups were slow more often

Why it matters
When 3 or more stored-copy checks found no copy, 42.5% of 501 briefings ran over 8 seconds. When every check found a stored copy, 0.0% of 356 briefings ran over 8 seconds. That is an association in synthetic data. It does not show what would happen if lookups were changed.
Evidence

Examples below are picked (slowest first), not sampled. They show what the cases look like, not how common they are. In total: 213 slow briefings in which 3 or more checks found no stored copy.

See it on the System performance page

Recommended action (a hypothesis)
Hypothesis to test: starting independent lookups together instead of one group after another, or fetching the shared feeds ahead of time, would leave fewer briefings waiting on many new lookups. Both would work within the freshness limits the service already keeps to. Loosening those limits is a safety setting and is not a speed fix.
How we would know
A comparison needs some requests to use the change and others not, at the same time of day, plus a check that no briefing used data older than the current limits allow. Comparing this week with last week would not do.
Q-04 Medium priority Synthetic example Data gaps Open

Some requests left no timing record

Why it matters
54 synthetic requests have no timing record, so how long they took, and how they ended, is unknown. It is not zero, and they are left out of every time figure. Because this data is synthetic we can see that 17 of them ran long and were cut off. The real service may not be able to tell.
Evidence

Examples below are picked (in time order), not sampled. They show what the cases look like, not how common they are. In total: 54 requests with no timing record.

See it on the System performance page

Recommended action (a hypothesis)
Count requests separately from their timing records, so a missing record shows up as a gap (see Q-08).
How we would know
Done when the number of requests counted separately equals the number of stored records plus the number reported missing, over a full day.
Q-06 Medium priority Synthetic example Predictions Open

A simple rule scored better than two baselines on past synthetic requests

Why it matters
On 397 held-out synthetic requests the cache-state rule had a lower error score than both baselines. The synthetic data was built with that effect in it, so this shows the checking works, not that the rule works on the real service.
Evidence

Examples below are picked (highest rule score first), not sampled. They show what the cases look like, not how common they are. In total: 86 of the 397 scored requests that ran long.

See it on the Predictive Analysis page

Recommended action (a hypothesis)
Review the historical trial run on the Predictions page. Approval there is simulated. A real decision needs real timing records (see Q-11).
How we would know
A better error score on past requests does not show that acting on a prediction helps. That needs a comparison where some requests get the action and others do not.
Q-07 Medium priority Known gap Data gaps Open

A timing record cannot be matched to the call it came from

Why it matters
A timing record holds no call identifier; that is deliberate, to keep identifiers out of the timing data. The phone agent does pass a conversation identifier to the service for some requests, but it is not stored with the timing. Until a match exists there is no dependable way to say which call a slow request belonged to, and no count of calls can be worked out from requests.
Evidence

From source review Source review of the timing record design and the call log (docs/predictive/source-field-inventory.md).

Read from the repository, not verified against the live service. No data is behind this item.

See the Calls page

Recommended action
Decide whether a pseudonymous call reference may be stored with each timing record. That touches what is recorded about callers and needs the owner's approval.
How we would know
Done when a sample of real timing records can each be matched to exactly one call, and the match can be checked in both directions.
Q-08 Medium priority Known gap Data gaps Open

Requests are not counted separately from their timing records

Why it matters
A request that is cut off before its record is written, or whose record cannot be written, may leave nothing behind, so it could not be told from a request that never happened. How often that happens is not known. If it does happen, time figures built from records alone are biased towards the requests that finished.
Evidence

From source review Source review of the request handler and the timing record design.

Read from the repository, not verified against the live service. No data is behind this item.

See the System performance page

Recommended action
Add a small counter that is updated at the start of each request, independently of the timing record.
How we would know
Done when, over a day, requests counted at the start equal records stored plus records reported as dropped.
Q-03 Low priority Synthetic example System speed Open

Slow briefings were more common in the quiet hours

Why it matters
In the hours with fewer requests (00:00–12:00, 18:00–20:00, 23:00–24:00 UTC), 34.3% of 356 briefings ran over 8 seconds, against 17.5% of 549 briefings in the busier hours. That is an association. One possible reason is that stored copies expire between requests when few come in; this data does not show it.
Evidence

Examples below are picked (slowest first), not sampled. They show what the cases look like, not how common they are. In total: 122 slow briefings in the quiet hours.

See it on the Trends page

Recommended action (a hypothesis)
Hypothesis to test: a scheduled pre-fetch through the quiet hours would keep the stored copies current for the first request that arrives.
How we would know
The same test as Q-01, compared within the quiet hours only, on days with and without the pre-fetch. Comparing quiet hours with busy hours would only repeat this association.
Q-05 Low priority Synthetic example Briefing quality Open

Some briefing requests were refused as invalid

Why it matters
32 synthetic requests were refused before any lookup. The data does not say what made them invalid. One possible cause is an airport identifier the service did not recognise; nothing here shows that.
Evidence

Examples below are picked (in time order), not sampled. They show what the cases look like, not how common they are. In total: 32 requests refused as invalid.

Recommended action (a hypothesis)
Look at what the refused requests had in common before deciding whether the request, or the way it was collected, needs to change.
How we would know
Needs the reason for each refusal, which this data does not carry. Real requests would need it recorded before this can be judged.
Q-12 Low priority Known gap Data gaps Open

Call records keep the time of the latest refresh, not the first save

Why it matters
A prediction may only use what was known at the time. A call record is rewritten with a new "refreshed at" time each time it is refreshed, so when a fact was first known cannot be shown.
Evidence

From source review Source review of the call log (docs/predictive/source-field-inventory.md).

Read from the repository, not verified against the live service. No data is behind this item.

See the Calls page

Recommended action
Keep the first-saved time and never change it.
How we would know
Done when every call record has a first-saved time that later refreshes leave alone.

Decision log

Simulated
Decisions taken in this simulation, in order
StepItemActionResultWhy
1Q-06acceptRefusedthe evidence has to be reviewed first
2Q-03reopenRefusedit is already open
3Q-09acceptRefusedthe evidence has to be reviewed first
4Q-04acceptRefusedthe evidence has to be reviewed first

The log is the list of decisions in this page's address, replayed from the start on every load. Anyone can edit the address, so it simulates a log that is only added to. It is not a tamper-proof audit record.

Accepting an item here records a simulated decision only. It changes nothing in the service, and a log of decisions does not show that any change helped.
Engineering details: how items are found

Synthetic examples are computed from the synthetic requests over the whole window. Each has a count, a list of example request IDs and the rule for how the examples were picked. If the data holds none of a kind, the item does not appear.

  • Q-01: requests with budget.out_of_time_at_end or a skipped lookup.
  • Q-02: slow briefings with three or more stored-copy checks that found no copy; shown only as a comparison when there are briefings to compare with.
  • Q-03: slow briefings in hours quieter than usual, when that share is higher than in busier hours (at least 30 briefings in each).
  • Q-04: requests with no timing record.
  • Q-05: requests refused with HTTP 400.
  • Q-06: shown only when every approval condition is met on the Predictions page; the examples are requests that ran long, highest rule score first.
  • Q-07 to Q-12: known gaps, written from source review and listed whatever the data says.