Synthetic demonstration. Every number and example on these pages comes from a generator, not from the live service. A separate, real source exists — 58 recorded demo calls (current, kept up to date on a schedule) — kept apart from these figures: the two are never counted together.
Show the recorded demo calls instead. Timing covers the briefing service's own processing only: it leaves out the telephone connection, speech recognition and speech synthesis. A synthetic request is one ask of the service. It is never a call.
Improvement queue
Findings and recommended actions, each with its evidence. A recommendation is a hypothesis to test: a pattern in data, or a log that a change was made, does not show that the change helped. "How we would know" says what a fair test would need.
This page always uses the whole synthetic window: about 76 hours, Mon 05 Jan 00:01 UTC to Thu 08 Jan 04:25 UTC (1,200 synthetic requests), whatever time range is chosen on other pages. Dates are synthetic.
Two kinds of item. A synthetic example is a finding in this synthetic data. A known gap is something found by reading the service's source; it is not from any data, and reading the repository does not verify what is running live.
Items
Synthetic Known gap Simulated12 shown, in priority order. Decisions on these items are simulated.
Q-01 High priority Synthetic example System speed Open
Some briefings finished with almost no time left
- Why it matters
- 26 synthetic requests finished with under 1.5 seconds of the 15-second limit left; 1 of them skipped a lookup, so part of what the briefing checks was not assessed; the rest of that section may be complete. The service stops starting new lookups when time is nearly gone, so a briefing that close can leave something a pilot needs unchecked.
- Evidence
Examples below are picked (those that skipped a lookup first, then the slowest), not sampled. They show what the cases look like, not how common they are. In total: 26 requests that finished with under 1.5 seconds left, or skipped a lookup.
See it on the System performance page
- Recommended action (a hypothesis)
- Hypothesis to test: fetching the shared feeds ahead of time, so requests do not wait for them, leaves more time for the sections that matter. This data does not show that it would.
- How we would know
- Run the pre-fetch on alternate days, or on a random half of requests, for long enough to see a few dozen near-limit briefings in each group, then compare how many finish near the limit. Things getting better over time would not show that the pre-fetch helped.
Q-09 High priority Known gap Voice quality Open
Voice quality is not measured
- Why it matters
- Automated tests check the wording the service generates for speech. Nothing on this dashboard checks what a caller hears. Whether the altimeter setting and winds were spoken as written can only be found out by listening.
- Evidence
From source review Not measured by this dashboard. The real service archives call recordings; this dashboard reads none.
Read from the repository, not verified against the live service. No data is behind this item.
See the Voice quality page
- Recommended action
- Start a weekly listening sample of real recordings. Listen for the altimeter setting and the winds first.
- How we would know
- Done when a person has listened to a sample and recorded, for each call, whether the altimeter setting and winds were heard exactly as written.
Q-10 High priority Known gap Briefing quality Open
NOTAMs and special use airspace are not live in any briefing
- Why it matters
- Every briefing carries placeholder text for NOTAMs and says special use airspace was not checked. Both are things a pilot needs. The briefing marks them as not live, so the gap is disclosed, but it is still a gap.
- Evidence
From source review Source review of the request handler.
Read from the repository, not verified against the live service. No data is behind this item.
See the Briefing quality page
- Recommended action
- Needs official data access that is outside this dashboard. Until then the briefing must keep saying they were not checked.
- How we would know
- Done when both come from a live source and a sample of briefings is checked against the official version.
Q-11 High priority Known gap Data gaps Open
This dashboard is not connected to any real request timing
- Why it matters
- It reads only the synthetic set, so nothing on it says anything about the live service. Whether request timing is switched on anywhere was not checked when this was built.
- Evidence
From source review This dashboard reads synthetic data only. Whether request timing is enabled on the live service was not checked.
Read from the repository, not verified against the live service. No data is behind this item.
See the System performance page
- Recommended action
- Find out whether timing capture is enabled. If it is not, enabling it needs the owner's approval and a key kept apart from the call-record key. Then connect this dashboard to the stored records.
- How we would know
- Done when real request records exist and this dashboard reads them in place of the synthetic set.
Q-02 Medium priority Synthetic example System speed Open
Briefings that needed many new lookups were slow more often
- Why it matters
- When 3 or more stored-copy checks found no copy, 42.5% of 501 briefings ran over 8 seconds. When every check found a stored copy, 0.0% of 356 briefings ran over 8 seconds. That is an association in synthetic data. It does not show what would happen if lookups were changed.
- Evidence
Examples below are picked (slowest first), not sampled. They show what the cases look like, not how common they are. In total: 213 slow briefings in which 3 or more checks found no stored copy.
See it on the System performance page
- Recommended action (a hypothesis)
- Hypothesis to test: starting independent lookups together instead of one group after another, or fetching the shared feeds ahead of time, would leave fewer briefings waiting on many new lookups. Both would work within the freshness limits the service already keeps to. Loosening those limits is a safety setting and is not a speed fix.
- How we would know
- A comparison needs some requests to use the change and others not, at the same time of day, plus a check that no briefing used data older than the current limits allow. Comparing this week with last week would not do.
Q-04 Medium priority Synthetic example Data gaps Open
Some requests left no timing record
- Why it matters
- 54 synthetic requests have no timing record, so how long they took, and how they ended, is unknown. It is not zero, and they are left out of every time figure. Because this data is synthetic we can see that 17 of them ran long and were cut off. The real service may not be able to tell.
- Evidence
Examples below are picked (in time order), not sampled. They show what the cases look like, not how common they are. In total: 54 requests with no timing record.
See it on the System performance page
- Recommended action (a hypothesis)
- Count requests separately from their timing records, so a missing record shows up as a gap (see Q-08).
- How we would know
- Done when the number of requests counted separately equals the number of stored records plus the number reported missing, over a full day.
Q-06 Medium priority Synthetic example Predictions Dismissed
A simple rule scored better than two baselines on past synthetic requests
- Why it matters
- On 397 held-out synthetic requests the cache-state rule had a lower error score than both baselines. The synthetic data was built with that effect in it, so this shows the checking works, not that the rule works on the real service.
- Evidence
Examples below are picked (highest rule score first), not sampled. They show what the cases look like, not how common they are. In total: 86 of the 397 scored requests that ran long.
See it on the Predictive Analysis page
- Recommended action (a hypothesis)
- Review the historical trial run on the Predictions page. Approval there is simulated. A real decision needs real timing records (see Q-11).
- How we would know
- A better error score on past requests does not show that acting on a prediction helps. That needs a comparison where some requests get the action and others do not.
Q-07 Medium priority Known gap Data gaps Open
A timing record cannot be matched to the call it came from
- Why it matters
- A timing record holds no call identifier; that is deliberate, to keep identifiers out of the timing data. The phone agent does pass a conversation identifier to the service for some requests, but it is not stored with the timing. Until a match exists there is no dependable way to say which call a slow request belonged to, and no count of calls can be worked out from requests.
- Evidence
From source review Source review of the timing record design and the call log (docs/predictive/source-field-inventory.md).
Read from the repository, not verified against the live service. No data is behind this item.
See the Calls page
- Recommended action
- Decide whether a pseudonymous call reference may be stored with each timing record. That touches what is recorded about callers and needs the owner's approval.
- How we would know
- Done when a sample of real timing records can each be matched to exactly one call, and the match can be checked in both directions.
Q-08 Medium priority Known gap Data gaps Open
Requests are not counted separately from their timing records
- Why it matters
- A request that is cut off before its record is written, or whose record cannot be written, may leave nothing behind, so it could not be told from a request that never happened. How often that happens is not known. If it does happen, time figures built from records alone are biased towards the requests that finished.
- Evidence
From source review Source review of the request handler and the timing record design.
Read from the repository, not verified against the live service. No data is behind this item.
See the System performance page
- Recommended action
- Add a small counter that is updated at the start of each request, independently of the timing record.
- How we would know
- Done when, over a day, requests counted at the start equal records stored plus records reported as dropped.
Q-03 Low priority Synthetic example System speed Open
Slow briefings were more common in the quiet hours
- Why it matters
- In the hours with fewer requests (00:00–12:00, 18:00–20:00, 23:00–24:00 UTC), 34.3% of 356 briefings ran over 8 seconds, against 17.5% of 549 briefings in the busier hours. That is an association. One possible reason is that stored copies expire between requests when few come in; this data does not show it.
- Evidence
Examples below are picked (slowest first), not sampled. They show what the cases look like, not how common they are. In total: 122 slow briefings in the quiet hours.
See it on the Trends page
- Recommended action (a hypothesis)
- Hypothesis to test: a scheduled pre-fetch through the quiet hours would keep the stored copies current for the first request that arrives.
- How we would know
- The same test as Q-01, compared within the quiet hours only, on days with and without the pre-fetch. Comparing quiet hours with busy hours would only repeat this association.
Q-05 Low priority Synthetic example Briefing quality Open
Some briefing requests were refused as invalid
- Why it matters
- 32 synthetic requests were refused before any lookup. The data does not say what made them invalid. One possible cause is an airport identifier the service did not recognise; nothing here shows that.
- Evidence
Examples below are picked (in time order), not sampled. They show what the cases look like, not how common they are. In total: 32 requests refused as invalid.
- Recommended action (a hypothesis)
- Look at what the refused requests had in common before deciding whether the request, or the way it was collected, needs to change.
- How we would know
- Needs the reason for each refusal, which this data does not carry. Real requests would need it recorded before this can be judged.
Q-12 Low priority Known gap Data gaps Accepted
Call records keep the time of the latest refresh, not the first save
- Why it matters
- A prediction may only use what was known at the time. A call record is rewritten with a new "refreshed at" time each time it is refreshed, so when a fact was first known cannot be shown.
- Evidence
From source review Source review of the call log (docs/predictive/source-field-inventory.md).
Read from the repository, not verified against the live service. No data is behind this item.
See the Calls page
- Recommended action
- Keep the first-saved time and never change it.
- How we would know
- Done when every call record has a first-saved time that later refreshes leave alone.
Decision log
SimulatedThe log is the list of decisions in this page's address, replayed from the start on every load. Anyone can edit the address, so it simulates a log that is only added to. It is not a tamper-proof audit record.
Accepting an item here records a simulated decision only. It changes nothing in the service, and a log of decisions does not show that any change helped.
Engineering details: how items are found
Synthetic examples are computed from the synthetic requests over the whole window. Each has a count, a list of example request IDs and the rule for how the examples were picked. If the data holds none of a kind, the item does not appear.
- Q-01: requests with budget.out_of_time_at_end or a skipped lookup.
- Q-02: slow briefings with three or more stored-copy checks that found no copy; shown only as a comparison when there are briefings to compare with.
- Q-03: slow briefings in hours quieter than usual, when that share is higher than in busier hours (at least 30 briefings in each).
- Q-04: requests with no timing record.
- Q-05: requests refused with HTTP 400.
- Q-06: shown only when every approval condition is met on the Predictions page; the examples are requests that ran long, highest rule score first.
- Q-07 to Q-12: known gaps, written from source review and listed whatever the data says.