KBM Nexus Flight Service
Operations analysis

Improvement queue

Findings and recommended actions, each with its evidence. A recommendation is a hypothesis to test: a pattern in data, or a log that a change was made, does not show that the change helped. "How we would know" says what a fair test would need.

This page always uses the whole synthetic window: about 76 hours, Mon 05 Jan 00:01 UTC to Thu 08 Jan 04:25 UTC (1,200 synthetic requests), whatever time range is chosen on other pages. Dates are synthetic.

Two kinds of item. A synthetic example is a finding in this synthetic data. A known gap is something found by reading the service's source; it is not from any data, and reading the repository does not verify what is running live.
ShowAll items · 12Synthetic examples · 6Known gaps · 6
Real evidence from the recorded demo calls is on this tab too. Show the Review & Learn findings and the recorded checks instead. They are never mixed into the synthetic items below.

Items

Synthetic Simulated

6 shown, in priority order. Decisions on these items are simulated.

Q-01 High priority Synthetic example System speed Open

Some briefings finished with almost no time left

Why it matters
26 synthetic requests finished with under 1.5 seconds of the 15-second limit left; 1 of them skipped a lookup, so part of what the briefing checks was not assessed; the rest of that section may be complete. The service stops starting new lookups when time is nearly gone, so a briefing that close can leave something a pilot needs unchecked.
Evidence

Examples below are picked (those that skipped a lookup first, then the slowest), not sampled. They show what the cases look like, not how common they are. In total: 26 requests that finished with under 1.5 seconds left, or skipped a lookup.

See it on the System performance page

Recommended action (a hypothesis)
Hypothesis to test: fetching the shared feeds ahead of time, so requests do not wait for them, leaves more time for the sections that matter. This data does not show that it would.
How we would know
Run the pre-fetch on alternate days, or on a random half of requests, for long enough to see a few dozen near-limit briefings in each group, then compare how many finish near the limit. Things getting better over time would not show that the pre-fetch helped.
Q-02 Medium priority Synthetic example System speed Open

Briefings that needed many new lookups were slow more often

Why it matters
When 3 or more stored-copy checks found no copy, 42.5% of 501 briefings ran over 8 seconds. When every check found a stored copy, 0.0% of 356 briefings ran over 8 seconds. That is an association in synthetic data. It does not show what would happen if lookups were changed.
Evidence

Examples below are picked (slowest first), not sampled. They show what the cases look like, not how common they are. In total: 213 slow briefings in which 3 or more checks found no stored copy.

See it on the System performance page

Recommended action (a hypothesis)
Hypothesis to test: starting independent lookups together instead of one group after another, or fetching the shared feeds ahead of time, would leave fewer briefings waiting on many new lookups. Both would work within the freshness limits the service already keeps to. Loosening those limits is a safety setting and is not a speed fix.
How we would know
A comparison needs some requests to use the change and others not, at the same time of day, plus a check that no briefing used data older than the current limits allow. Comparing this week with last week would not do.
Q-04 Medium priority Synthetic example Data gaps Open

Some requests left no timing record

Why it matters
54 synthetic requests have no timing record, so how long they took, and how they ended, is unknown. It is not zero, and they are left out of every time figure. Because this data is synthetic we can see that 17 of them ran long and were cut off. The real service may not be able to tell.
Evidence

Examples below are picked (in time order), not sampled. They show what the cases look like, not how common they are. In total: 54 requests with no timing record.

See it on the System performance page

Recommended action (a hypothesis)
Count requests separately from their timing records, so a missing record shows up as a gap (see Q-08).
How we would know
Done when the number of requests counted separately equals the number of stored records plus the number reported missing, over a full day.
Q-06 Medium priority Synthetic example Predictions Open

A simple rule scored better than two baselines on past synthetic requests

Why it matters
On 397 held-out synthetic requests the cache-state rule had a lower error score than both baselines. The synthetic data was built with that effect in it, so this shows the checking works, not that the rule works on the real service.
Evidence

Examples below are picked (highest rule score first), not sampled. They show what the cases look like, not how common they are. In total: 86 of the 397 scored requests that ran long.

See it on the Predictive Analysis page

Recommended action (a hypothesis)
Review the historical trial run on the Predictions page. Approval there is simulated. A real decision needs real timing records (see Q-11).
How we would know
A better error score on past requests does not show that acting on a prediction helps. That needs a comparison where some requests get the action and others do not.
Q-03 Low priority Synthetic example System speed Open

Slow briefings were more common in the quiet hours

Why it matters
In the hours with fewer requests (00:00–12:00, 18:00–20:00, 23:00–24:00 UTC), 34.3% of 356 briefings ran over 8 seconds, against 17.5% of 549 briefings in the busier hours. That is an association. One possible reason is that stored copies expire between requests when few come in; this data does not show it.
Evidence

Examples below are picked (slowest first), not sampled. They show what the cases look like, not how common they are. In total: 122 slow briefings in the quiet hours.

See it on the Trends page

Recommended action (a hypothesis)
Hypothesis to test: a scheduled pre-fetch through the quiet hours would keep the stored copies current for the first request that arrives.
How we would know
The same test as Q-01, compared within the quiet hours only, on days with and without the pre-fetch. Comparing quiet hours with busy hours would only repeat this association.
Q-05 Low priority Synthetic example Briefing quality Open

Some briefing requests were refused as invalid

Why it matters
32 synthetic requests were refused before any lookup. The data does not say what made them invalid. One possible cause is an airport identifier the service did not recognise; nothing here shows that.
Evidence

Examples below are picked (in time order), not sampled. They show what the cases look like, not how common they are. In total: 32 requests refused as invalid.

Recommended action (a hypothesis)
Look at what the refused requests had in common before deciding whether the request, or the way it was collected, needs to change.
How we would know
Needs the reason for each refusal, which this data does not carry. Real requests would need it recorded before this can be judged.

Decision log

Simulated
Decisions taken in this simulation, in order
StepItemActionResultWhy
No decision yet. Try accepting an item first: it is refused until the evidence has been reviewed.

The log is the list of decisions in this page's address, replayed from the start on every load. Anyone can edit the address, so it simulates a log that is only added to. It is not a tamper-proof audit record.

Accepting an item here records a simulated decision only. It changes nothing in the service, and a log of decisions does not show that any change helped.
Engineering details: how items are found

Synthetic examples are computed from the synthetic requests over the whole window. Each has a count, a list of example request IDs and the rule for how the examples were picked. If the data holds none of a kind, the item does not appear.

  • Q-01: requests with budget.out_of_time_at_end or a skipped lookup.
  • Q-02: slow briefings with three or more stored-copy checks that found no copy; shown only as a comparison when there are briefings to compare with.
  • Q-03: slow briefings in hours quieter than usual, when that share is higher than in busier hours (at least 30 briefings in each).
  • Q-04: requests with no timing record.
  • Q-05: requests refused with HTTP 400.
  • Q-06: shown only when every approval condition is met on the Predictions page; the examples are requests that ran long, highest rule score first.
  • Q-07 to Q-12: known gaps, written from source review and listed whatever the data says.