Retune the Kalman process noise, and fuse the two thermometers

Two changes to the same signal path, one large and one small.

The large one: all three filters were tuned to track one to three decades
faster than their signals move. In a still room the temperature filter
reported a median rate of 12.4 C/h while the air moved 0.4 C/h, and it
overshot a real -36 C/h event by 77%. Sweeping q against the RMSE of the
reported rate versus the true rate, using noise measured on the board
(temperature 0.088 C, pressure 0.022 hPa, humidity 0.40 %):

   temperature   6.45 -> 0.37 C/h RMSE     2e-6 -> 1e-9
   pressure      2.15 -> 0.24 hPa/h RMSE   1e-5 -> 1e-8
   humidity     27.94 -> 3.55 %/h RMSE     5e-5 -> 2e-8

Tracking does not suffer. Lag against a genuine 2 C/h ramp is 0.003 C at both
the old and new values, and the peak response to a five-minute event moves
closer to the truth rather than further from it, because the overshoot goes
away. What is given up is response to sub-minute transients, which for a
station forecasting fifteen minutes to a day ahead is noise to reject.

This matters most for pressure, whose tendency drives the precipitation
forecast, and which was the worst tuned of the three.

config.yaml shadowed kalman_q_temp, so editing the dataclass alone changed
nothing. All six values are now listed there with that hazard spelled out,
because a silent shadow cost real time here.

The small one: temp_raw was the plain average of two thermometers whose
white-noise sds differ by 7x (LPS25HB 0.007 C, HTS221 0.049 C), which throws
the quiet one away. Inverse-variance weighting cuts the raw noise 3.5x.

The trap is that the chips do not agree. They sit at different distances from
the SoC and stand about 1.3 C apart, so weighting by variance alone drags
temp_raw 0.48 C onto the LPS25HB, which after the 1.55x gain of the inverse
compensator is 0.75 C of silent bias on every reading, since k was fitted
against the mean of the two. The gradient is therefore tracked and removed
before weighting and only the deviations are fused: measured mean shift
0.0001 C, noise still 3.5x lower. The tracked gradient is retained because it
is a second observation of self-heating.

Also corrected: the earlier claim that the HTS221 was the quieter channel was
wrong, taken from twelve samples at a cadence slow enough that real drift
dominated. At 0.5 s over 120 samples the LPS25HB is quieter by 7x and takes
98% of the weight.
This commit is contained in:
2026-08-19 19:14:46 +01:00
parent e3176e29c9
commit 40f934901d
4 changed files with 205 additions and 6 deletions
+85
View File
@@ -248,3 +248,88 @@ def test_forecast_head_migrates_state_from_before_the_setpoint_member():
# discover a migration bug.
assert back.member_mae.size == len(MEMBERS)
back.learn(np.zeros(4), 20.0, 20.5, 0.1, 0.2) # must not raise
# ------------------------------------------------- dual-thermometer fusion
def _bare_board():
from ashvale.sensors import SD_HTS221, SD_LPS25HB, SenseBoard, _ChannelNoise
b = SenseBoard.__new__(SenseBoard)
b._noise_h = _ChannelNoise(SD_HTS221)
b._noise_p = _ChannelNoise(SD_LPS25HB)
b._gradient = None
b._gradient_lam = 0.9967
return b
def _two_channels(n=4000, seed=5):
from ashvale.sensors import K_HTS221, K_LPS25HB
rng = np.random.default_rng(seed)
cpu = 43.0 + 0.5 * np.sin(np.arange(n) / 500.0)
th = (24.0 + K_HTS221 * cpu) / (1 + K_HTS221) + 0.049 * rng.normal(size=n)
tp = (24.0 + K_LPS25HB * cpu) / (1 + K_LPS25HB) + 0.007 * rng.normal(size=n)
return th, tp
def test_fusion_does_not_move_the_mean():
"""The whole point of removing the gradient first.
The two chips stand about 1.3 C apart, so weighting them by variance drags
temp_raw onto the quieter one. k was fitted against the mean of the two, and
after the 1.55x gain of the inverse model that shift becomes about a degree
of silent bias on every reading downstream.
"""
th, tp = _two_channels()
board = _bare_board()
fused = np.array([board._fuse(th[i], tp[i])[0] for i in range(th.size)])
avg = (th + tp) / 2.0
w = slice(1000, None)
assert abs(fused[w].mean() - avg[w].mean()) < 0.01, "fusion shifted the calibration"
def test_fusion_is_quieter_than_the_average():
th, tp = _two_channels()
board = _bare_board()
fused = np.array([board._fuse(th[i], tp[i])[0] for i in range(th.size)])
avg = (th + tp) / 2.0
w = slice(1000, None)
def wn(x):
return np.std(np.diff(x)) / np.sqrt(2)
assert wn(fused[w]) < wn(avg[w]) / 2.0, "fusion did not halve the noise"
def test_fusion_survives_one_dead_channel():
board = _bare_board()
value, var = board._fuse(float("nan"), 29.5)
assert value == 29.5, "a dead HTS221 must not poison the reading"
value, var = board._fuse(30.5, float("nan"))
assert value == 30.5
value, var = board._fuse(float("nan"), float("nan"))
assert not np.isfinite(value)
def test_kalman_rate_is_physical_in_a_still_room():
"""The tuning failure this guards against.
On a real station the temperature filter reported a median rate of
12.4 C/h while the room moved 0.37 C/h. Process noise was set to track
perhaps a hundred times faster than any of these signals actually move.
"""
from ashvale.config import CONFIG
from ashvale.estimation import KalmanCV
dt = CONFIG.sensor.sample_period_s
rng = np.random.default_rng(3)
n = 6000
truth = 24.0 + 0.4 * np.arange(n) * dt / 3600.0 # a real 0.4 C/h drift
z = truth + 0.0877 * rng.normal(size=n) # measured input noise
kf = KalmanCV(CONFIG.sensor.kalman_q_temp, CONFIG.sensor.kalman_r_temp)
rates = [kf.update(z[i], dt)[1] * 3600.0 for i in range(n)]
settled = np.abs(np.array(rates[600:]))
assert np.median(settled) < 3.0, (
f"median |rate| {np.median(settled):.1f} C/h in a room drifting 0.4 C/h")
assert np.percentile(settled, 95) < 10.0