Manoeuvre Detection
Manoeuvre detection calibrated against objects that physically cannot manoeuvre
August 2026
The screener I built before this one could be graded. A US government agency published the correct answers and I compared mine against them line by line. Nobody publishes the correct answers for this one, and that absence is the shape of the problem rather than an inconvenience around it.
About 35,000 orbiting objects are tracked, and the US military publishes a periodic estimate of where each one is through Space-Track. No estimate carries an error bar. Nothing says whether today's is good to ten metres or ten kilometres. So the claim worth making — this orbit changed by more than measurement error explains — has nothing to appeal to. Assume a noise model instead and the first person who disputes the assumption has disputed the result, when the entire point was to survive being disputed.
The way out is to measure the noise. The catalogue holds 12,045 objects that certainly did not manoeuvre because they physically cannot: a fragment of a shattered rocket has no engine and never will. I sampled 348 of them, put them through exactly the processing a suspect would get, and collected 2.1 million disagreements between each estimate and the next. That spread is the noise, measured. A suspect is ranked against it: out of all the innocent behaviour on record, how much looked at least this extreme?
Whether that number means anything needs manoeuvres somebody else recorded. Seven satellites in the International DORIS Service publish their operators' thruster logs. This is what the scoring program prints, unedited, against 416 of those burns:
=== 416 real manoeuvres scored, against 157 in the fixed window ===
sat window burns sigma@1h excl det S det RSW det family_wise FA S
null = oracle
h2a 2012-2015 24 0.007176 17 0% 0% 0% 0.0%
ja2 2017-2020 99 0.001017 17 62% 63% 40% 0.0%
ja3 2024-2026 50 0.000918 17 100% 100% 100% 1.5%
s3a 2023-2026 72 0.002592 10 50% 50% 47% 1.0%
s3b 2023-2026 84 0.003741 10 12% 17% 42% 0.0%
s6a 2020-2023 20 0.001887 17 65% 65% 65% 0.5%
sp5 2013-2016 67 0.013060 17 100% 66% 97% 0.5%
null = naive
h2a 2012-2015 24 0.007176 17 0% 0% 0% 0.0%
ja2 2017-2020 99 0.001017 17 9% 9% 7% 0.0%
ja3 2024-2026 50 0.000918 17 10% 0% 8% 0.0%
s3a 2023-2026 72 0.002592 10 1% 3% 4% 0.0%
s3b 2023-2026 84 0.003741 10 0% 1% 1% 0.0%
s6a 2020-2023 20 0.001887 17 15% 0% 0% 0.0%
sp5 2013-2016 67 0.013060 17 19% 19% 6% 0.0%
oracle burn-weighted detection: S 57% RSW 53% family_wise 57% | false alarms S 0.50%
naive burn-weighted detection: S 7% RSW 6% family_wise 5% | false alarms S 0.00%
The 157 in the header is what an earlier run scored, when all seven were judged over one fixed two-year window that three of their logs end before. sigma@1h is the satellite's own noise level, excl how many days around each burn were kept out of the baseline, the three det columns three ways of choosing which direction to look in, and FA the false-alarm rate on days the log says nothing happened. The two halves are the same satellites, data and statistic. They differ only in what the baseline of innocent behaviour — the block's null — was built from: oracle uses the published log to find the quiet stretches, naive does not.
The false-alarm column is the claim that held. One satellite overshoots the 1% budget, ja3 at 1.5%; one sits exactly on it; the average across the seven is 0.50%, over fourteen years.
Detection is the claim that did not hold. With the operator's log, 57% of real burns are caught. Without it, 7%. The method works on operators who publish when they fired their thrusters, and does not work on the ones you would most want to watch.
Why one satellite returns nothing
Detection in that table runs from 100% to zero, and the spread is not noise. HY-2A is the zero.
Its burns are the smallest in the set by a distance. The median is 0.66 centimetres per second and the largest in three years is 0.8, where every other satellite here pushes at least 1.8 metres per second at some point. Its baseline is also the widest: the score a burn has to beat runs to 255, where no other satellite reaches 30. The smallest burns in the set have to clear the highest bar in the set.
What decides detection is therefore not a sensitivity figure belonging to the method. It is the burn measured against that object's own noise, and those two properties predict the zero without running anything. Reporting nothing there is the correct answer. Adjusting until it reported something is the operation the next section measures.
The repair that makes it worse
Judging a satellite against its own history has an obvious hole: a satellite that manoeuvres has manoeuvres in its history, so the yardstick contains the thing being measured. The direction of that error saves it: a past burn makes the baseline wider, so anything measured against it looks less unusual. The error runs toward missing real manoeuvres, never toward accusing an innocent operator.
The tempting repair is to clean the baseline: discard the most extreme parts of a history on the grounds that they were probably burns. This is what that does, where K is how many real manoeuvres the history already holds.
False-alarm rate at p<0.01 after ejecting the top q% of the null.
Nominal is 1%. Anything above it is a false-accusation route.
eject q K=0 K=5
0% 0.52% 0.07%
1% 1.22% 0.15%
2% 2.27% 0.28%
5% 5.40% 0.61%
10% 10.46% 1.16%
20% 20.42% 9.00%
Discarding the top 1% more than doubles the false-alarm rate, past its budget; discarding 20% multiplies it fortyfold. The extremes being discarded are mostly not manoeuvres — this data is naturally wild and its rare large values are ordinary.
The right-hand column is the trap. Someone holding a contaminated history sees 0.07%, reads it as needlessly cautious, cleans, and walks the rate back toward 1% and straight past. At 10% it reads 1.16%: apparently perfect calibration, arrived at by luck, with nothing in the numbers marking it as the place to stop. So cleaning is refused unless it rests on evidence from outside the scores, and can never be applied backwards to a case already judged. The rule lives in the code, not a document.
What it doesn't do
It is poor at seeing a satellite tilt its orbit. A forward push changes the orbit's size and the error grows every hour — 260.7 km after a day for one metre per second. A sideways push tilts the plane and the error only wobbles: 0.381 km at the same point. Under the statistic all my earlier results used, a sideways burn scored at the control rate at every magnitude tested, up to 100 m/s. Reading the three directions separately takes that to 50.4%. Half, not most.
It does not accuse anyone. "This object's momentum changed in a way that drag, sunlight and data artefacts do not explain" is a measurement. "This operator deliberately manoeuvred" is a different claim needing different evidence. Sunlight alone pushes a light, high-area object for weeks in a way that mimics a thruster, and gas venting from a dead satellite is a real push that was nobody's decision.
One route past HY-2A's zero looked promising and was wrong. Its 24 burns are regular, one every 46 days or so, so instead of asking whether a single moment is anomalous you can ask whether an object's noisy days clump together. Manoeuvring satellites did look clumpier than debris — until the separation turned out to be a single step change. HY-2A's second half sits 3.9 times above its first, and measuring each day against its neighbours rather than against the whole record drops its clumpiness from 4.84 to 1.06, under the debris median of 1.15. Nothing was built on it.
Every real manoeuvre it has been scored against is one mission class: altimetry satellites in high, quiet orbits, on chemical propulsion. None is on electric propulsion or in geostationary orbit. On invented burns, holding tracking cadence fixed, a 0.3 m/s nudge is caught 32% of the time below 900 km against 82% above 1,400 km, and no real satellite has been checked down there.