Overview
How the briefing service is doing across the synthetic demonstration. A briefing is the aviation weather and flight information the service prepares for a pilot's flight. The five sections below answer, in order: what happened, what needs attention, why, what may happen next, and what to do about it.
Showing 59 synthetic requests, Wed 07 Jan 22:25 UTC to Thu 08 Jan 04:25 UTC. Dates are synthetic. Trends, Predictions and the Improvement queue always use the whole window.
This synthetic window
What is real, and what is not
- Synthetic: every number and example above. They come from a generator with a cache effect built in, so a rule that uses cache state does well here by construction.
- Known gaps: 6 gaps in the real service are listed in the Improvement queue. They come from reading the service's source, not from data, and are not verified against the live service.
- Not measured: calls, voice quality, whether briefing content was right, what callers heard or did next, and anything about the real service's speed. The Recorded demo calls tab is a separate snapshot of real demo calls. None of the synthetic figures use it or describe it, and it measures none of these things either.
What happened?
In this window the service handled 59 synthetic requests: 48 briefing requests, 1 alternate-airport checks, 5 flight plans, and 5 whose type was not recorded because their timing record is missing.
Each request is counted once, under the most serious thing that happened to it: a skipped lookup outranks being slow. Slow requests in total: 11, which includes 0 that also had a lookup skipped.
What needs attention?
The items below are the highest-priority findings still open in the synthetic data. They are examples, in priority order.
Some briefings finished with almost no time left
26 synthetic requests finished with under 1.5 seconds of the 15-second limit left; 1 of them skipped a lookup, so part of what the briefing checks was not assessed; the rest of that section may be complete. The service stops starting new lookups when time is nearly gone, so a briefing that close can leave something a pilot needs unchecked.
4 examples of 26 requests that finished with under 1.5 seconds left, or skipped a lookup. See the whole item
Briefings that needed many new lookups were slow more often
When 3 or more stored-copy checks found no copy, 42.5% of 501 briefings ran over 8 seconds. When every check found a stored copy, 0.0% of 356 briefings ran over 8 seconds.
4 examples of 213 slow briefings in which 3 or more checks found no stored copy. See the whole item
Some requests left no timing record
54 synthetic requests have no timing record, so how long they took, and how they ended, is unknown. It is not zero, and they are left out of every time figure.
4 examples of 54 requests with no timing record. See the whole item
Why?
Two patterns in this data. Both are associations: they show what went with what, not what caused what.
Slow briefings, by how many stored-copy checks found no copy
Slow briefings, by how busy the hour was
Which hours count as busier or quieter is decided over the whole synthetic window. The briefings counted are the ones in the time range chosen above. A group with fewer than 10 briefings is not compared.
What may happen next?
A simple rule looks at each request one second in and flags the ones likely to run past 8 seconds. It is a table of counts, not machine learning, and it is one of three methods compared. On the 397 synthetic requests it was tested on, its error score was 0.152, against 0.173 for a rolling history and 0.170 for one fixed guess (lower is better). That shows the checking works. It says nothing about how the real service will behave.
Among the situations the rule tells apart, the one with the highest share of long requests in the training period was Standard briefing · at the 1-second mark: no stored-copy check done yet (92 of 327), and the lowest was Alternate-airport check · at the 1-second mark: no stored-copy check done yet (0 of 34).
In this window requests were busiest at 12:00–18:00, 20:00–23:00 UTC. That is a description of the past, not a forecast.
In use in this simulation: the rolling history baseline. Nothing is served from any prediction and nothing is trained live. Open Predictions.
What action is recommended?
The highest-priority items still open, from the Improvement queue. A recommendation is a hypothesis to test. A pattern in this data, or a log that a change was made, does not show that the change helps.
Some briefings finished with almost no time left
- Recommended action
- Hypothesis to test: fetching the shared feeds ahead of time, so requests do not wait for them, leaves more time for the sections that matter. This data does not show that it would.
- How we would know
- Run the pre-fetch on alternate days, or on a random half of requests, for long enough to see a few dozen near-limit briefings in each group, then compare how many finish near the limit.
Voice quality is not measured
- Recommended action
- Start a weekly listening sample of real recordings. Listen for the altimeter setting and the winds first.
- How we would know
- Done when a person has listened to a sample and recorded, for each call, whether the altimeter setting and winds were heard exactly as written.
NOTAMs and special use airspace are not live in any briefing
- Recommended action
- Needs official data access that is outside this dashboard. Until then the briefing must keep saying they were not checked.
- How we would know
- Done when both come from a live source and a sample of briefings is checked against the official version.
Engineering details: how these numbers are made
| Dataset ID | syn-feab7bc56e79 |
|---|---|
| Generator version | pa-proto-0.1 |
| Content SHA-256 of the events | 08a17326968ac351d80464c4fed4464b61837f679542b925fa076e488bf5de8e |
| Requests / with a timing record / without | 1,200 / 1,146 / 54 |
| Without a record, by what the synthetic world knows | {"missing_dropped":37,"missing_killed":17} (known only because the data is synthetic) |
| Prediction sets | {"train":581,"pre_fit_unlabelled":4,"holdout":397,"completed_before_decision":164,"no_record":54} |
Preparation time is the request's own processing time on the service host (in the records, server.script_end_ms). Percentiles are nearest-rank over briefings and alternate checks that have a recorded time; a request with no recorded time is left out of every time figure and counted on its own. A briefing is "slow" when that time is over 8,000 ms. "Finished near the time limit" means under 1,500 ms of the 15,000 ms budget was left.
The events use the real record shape (fs.telemetry.request.v0), read through the same field registry and point-in-time reader real records would use. Predictions and promotion come from pa_run() and pa_replay() unchanged.