mirror of
https://github.com/lynchaos/ashvale-station.git
synced 2026-09-12 12:47:49 +00:00
18b9a29ffa1fc2557302aefedbe30bda3526dd9b
11
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
18b9a29ffa |
Earn the blend weights and the intervals from forecasts, not from refits
fit() called learn(), and learn() updated three things: the RLS, the conformal
calibrator and the Hedge weights. Only the first belongs to a refit. The comment
above that loop already said so, and was wrong about what the code did.
Measured on 8.2 days of the live station, the 15 minute head had taken 977,078
Hedge updates from 758 distinct supervised pairs, a factor of 1,289, and the
1 day head 296,715 from 12 pairs, a factor of 24,726. A refit is not an outcome.
It is the same week of weather being read again, once every seven minutes.
Hedge is multiplicative, so an edge far too small to be real compounds to
certainty: twelve of twelve temperature and humidity heads had collapsed onto
climatology at a weight of 0.991 or above, while their own member_mae said the
members were within a few percent of each other. The ACI integrator moves by
gamma per observation, so it had likewise pinned against its clips, leaving the
6 hour temperature band (1.571 C) narrower than the 3 hour one (2.258 C), and
pressure at 1 day covering 3 of 7 with alpha jammed at the 0.005 floor.
Three changes, because fixing only the first would freeze the weights forever:
- fit() calls refit_step(), which touches the regression and nothing else.
The climatology and setpoint members were evaluated in that loop purely to
feed the Hedge update, so fit() no longer needs a climatology or a
setpoint_fn at all.
- verify() feeds observe_outcome() with the member predictions the forecast
was actually blended from. These are now written to the forecasts table at
issue time, because the learned member cannot be recovered afterwards: the
RLS has moved on.
- a matured forecast teaches exactly once. It stays readable for an hour so
the scorecard can aggregate a rolling window, which meant verify() was
feeding the calibrator the same outcome about twelve times.
The Hedge weights additionally decline an outcome that overlaps the last one
they took, which is the stride rule from fit() applied on the scoring side.
Forecasts are issued every retrain tick, so at the 1 day horizon roughly two
hundred a day resolve against very nearly the same outcome. The conformal window
absorbs that, a quantile over duplicated scores being merely overconfident about
its sample size, but exponentiated gradient cannot.
Walk-forward over the full 8.2 day record, against the current code:
mean MAE 0.856 (0.938 over the second half alone)
heads improved 17/18
beats persistence 7/18 -> 12/18
coverage |dev from .90| 0.188 -> 0.060, second half 0.112 -> 0.041
Decimating the conformal feed as well was measured and rejected. It reads better
(coverage |dev| 0.023) and is not: three heads fall below MIN_SCORES, drop to
1.645*sigma, and "cover" with a median band of +/- 107% relative humidity. At
h/4 and h/8 it never starves and lands within noise of not decimating at all, so
the simpler rule wins. Honest regressions: humidity at 12 hours is 19% worse,
and pressure past 6 hours is still under-covered, because at 1 day the point
forecast is genuinely poor and ACI can only widen so far.
A state file from before this change has its weights, member_mae, n_scored and
alpha reset on load. They are products of the replay, they are not evidence, and
they do not decay on their own: Hedge needs about twenty independent outcomes to
climb off its 1e-4 floor and the 1 d head sees one a day. The conformal scores
are kept, being residuals of roughly the right size, and the window refreshes
within about two days.
Schema migration verified against a pristine copy of the live database: 1,437
forecast rows preserved, five columns added, idempotent across restarts.
|
||
|
|
9a1033973d |
Treat the board being moved as a regime change, from the fused IMU attitude
The accelerometer and gyroscope were logged and never used. They measure nothing about weather, but they do measure the one thing about this station that nothing else can see: whether the sensor is still where it was. Measured over four and a half days on the real station, four genuine movements each stepped the temperature by a median of 1.02 C, against an ordinary fifteen minute change of 0.107 C with a 95th percentile of 0.841. A move therefore lands past the 95th percentile of normal variation. The heads carry about 55 hours of memory, so an undeclared move contaminates two days of training with a discontinuity they will try to fit rather than ignore. This now gets the same treatment set_environment gives a window being opened, because it is the same event: the coupling between the sensor and what it is measuring changed, and nothing in the data says so. Three choices in here were made by measurement, and the obvious one was wrong. Raw accelerometer looks like the natural input and is not. Over the same record a gravity-vector detector fires 112 times against this one's 4, because RTIMULib's gyro fusion removes exactly the desk vibration a bare accelerometer picks up. The fused pitch and roll have a p99 sample-to-sample noise of 0.0001 degrees, so a one degree trigger carries four decades of headroom. Yaw and compass are excluded. They are the only attitude outputs that depend on the magnetometer, and indoors the magnetometer is measuring the building. RTIMULib restarts its fusion from a default attitude when SenseHat is reconstructed, which put an 18 degree step in the record on every one of this station's seven service restarts. Without a settle window every deploy would queue a retrain. 300 seconds rather than 180: one artifact appeared three minutes after a restart, still converging. Replayed against the full record the detector finds 4 genuine movements and leaks 0 artifacts. |
||
|
|
3dd45f7ebf |
Joystick labelling, a clock guard, and throttle logging
Three things the hardware offers that the code ignored. The joystick has never had a line of code. Left records a dry label, right a wet one, middle cycles the LED scene, and a full-panel flash acknowledges the press because a headless box gives no other sign and a button you cannot tell worked gets pressed twice. Precipitation is the weakest head in the bank and strong labels are its binding constraint: this station has 80 of them against thousands of proxy ones, entirely because the only label control lives in a web page, and a web page is not where anyone is standing when it starts raining. The board has no RTC, so a power cut without a network gives a clock somewhere in 1970 on the next boot. Solar elevation, the diurnal harmonics and a sample's position on the 5-minute grid then all lie with complete confidence, and unlike a gap in the record the damage cannot be identified afterwards. train() now refuses a clock below 2025 or one that has stepped behind the newest stored row, and logs the refusal rather than training on fiction. Undervoltage and thermal capping both shift the SoC temperature, which is the regressor in the self-heating compensation, so a weak power supply presents as an unexplained temperature bias rather than as anything resembling a power problem. get_throttled is now sampled hourly and logged when set. Measured and deliberately not done: colour features. r, g and b are logged and 74% of rows carry usable colour, but adding blue/red, green/red and saturation made MAE 1.50% worse and helped in only 13 of 72 cases. Three more regressors on a 33-feature model whose longest horizon trains on 13 independent pairs is straightforwardly overfitting. That also prompted a sweep of the RLS prior and forgetting factor in both directions; delta = 100 with lambda = 0.9985 is a local optimum on both axes, so neither moved. |
||
|
|
bda42a0468 |
Log both Sense HAT thermometers, and migrate schemas that predate them
The board carries two independent thermometers and the code averaged them into temp_raw without ever recording either. Measured over 12 samples on a real station: HTS221 30.973 C at sd 0.060, LPS25HB 29.810 C at sd 0.443, a standing gradient of 1.163 C with the SoC at 44.55 C. Two things follow from that and neither is possible without the raw channels. A plain average of a quiet sensor and one seven times noisier lands at sd 0.223 where inverse-variance weighting reaches 0.060, and the gradient between two chips at different distances from the SoC is a second observation of self-heating that could identify the compensator's k with no reference thermometer. Both need history, and history cannot be backfilled, so the columns land on their own ahead of the work that consumes them. CREATE TABLE IF NOT EXISTS is a no-op against a table that already exists, so adding to COLUMNS would have reached a fresh install and silently missed every station already running, then surfaced as an OperationalError inside insert_telemetry. That sits on the sample loop, so it takes a station down rather than leaving a gap. Store now reconciles the table against COLUMNS on open, which makes every future column addition safe rather than just this one. The simulator gains the same two channels, with couplings solved so their forward models average to exactly the k = 0.55 the compensator is tuned against. Aggregate behaviour is unchanged; only the per-channel detail is new. Simulated temp_raw noise does rise from 0.05 to 0.223, which is not a regression but the end of an over-optimistic figure: it was modelling the quiet sensor and calling it the average. |
||
|
|
485affe956 |
Fix the runaway forecasts: refits accumulated, and annual terms fitted too early
Reported from a real station after 1.5 days: a six hour temperature forecast of 53 C in a 24 C room, and 9 C at one day, both carrying a plus or minus of 0.43. Confidently wrong is the one failure this project is supposed to refuse. Root cause. fit() replayed history into the live RLS on every retrain tick and never reset, so 453 grid rows had produced 64,676 updates in a day and a half. RLS with forgetting reads every update as fresh evidence, so the model believed it had a hundred times the data it had: P collapsed, in-sample error looked excellent, and the weights drifted without bound in directions the data never excited. Measured: cond(P) 3.1e9 and ||theta|| 1680 against a median |theta| of 1.67. A refit now starts from the prior, which makes retraining idempotent. Across 25 refits on the real data ||theta|| holds at 11.35, drifting 0.03, where before it grew without limit. The two largest weights were sin_doy and cos_doy at +1174 and +1191. Annual harmonics were in the design matrix from the first sample, where they are near-constant, near-collinear with each other and with the bias, and a rank-deficient regressor is what RLS answers with enormous cancelling weights. They are now held at zero until the record spans the same 120 days the climatology fit already requires, because a day and a half of data says nothing whatsoever about the season. Also raised the standardiser's variance floor from 1e-8, which only caught a bit-exactly constant column, to 1e-3. A feature that merely barely moves was being divided by its own noise. The conformal calibrators and Hedge weights are deliberately not reset by a refit: those are earned from scored forecasts, not from this regression. Backtest unchanged within noise, coverage still 89 to 91 across all 18 heads. Four regression tests added, including that refitting the same history twice must give the same model. |
||
|
|
4cca40388f |
Heated environment: a thermostat member in the forecast ensemble
A room held at a setpoint is a different process from one left to drift. It is
a closed loop, and persistence, the baseline everything here is scored against,
is the wrong statement about it: the truth is not that it stays where it is, it
is that it returns to the setpoint.
So site.heating adds a fourth ensemble member, first order because that is what
a controlled system is:
dT_set(h) = (T_set - T_now) * (1 - exp(-h / tau))
Humidity follows and is the part that is easy to get wrong. Heating adds no
moisture, so vapour pressure is conserved and not relative humidity:
RH(h) = RH_now * es(T_now) / es(T_now + dT_set(h))
Warm the air and RH falls although nothing was dried, which is why a heated
house in winter is dry. The test asserts the dew point is unchanged to 1e-6.
Pressure gets zero: a thermostat cannot move the synoptic field.
Offered, not imposed. Hedge scores this member on realised error like any
other, so a wrong tau or a stale setpoint costs accuracy and gets down-weighted
rather than quietly biasing every forecast. Verified: on history with no
heating the ensemble assigned it weight 0.000. With heating off it returns zero
and is identical to persistence.
Going from three members to four means old saved heads must migrate.
from_dict reinitialises weights and member_mae. I missed member_mae first time
and it did not fail on load, it failed later inside learn() on a broadcast
error, which is a much worse place to find out; the migration test now covers
both and calls learn() to prove it.
Settings tab gains the toggle, setpoint and time constant. Turning heating on
or off is treated as a regime change like a door: discontinuity marker plus a
queued retrain.
|
||
|
|
50b29f8077 |
Readout scene, environment regime tracking, and a Kalman cadence bug in recompute
recompute replayed the Kalman over stored rows at their own spacing while q stays tuned for the live 2 s cadence. Q scales with dt^3, so at the 30 s persist interval the process noise was 3375x too large and the filter tracked noise instead of smoothing: it wrote indoor temperature rates of +/-20 C/h into the history. This is the exact trap DESIGN.md section 2 documents for simulate.py, which does scale q, and I walked into it anyway. Now rescaled per step, because tiering means the stored cadence is not constant. Mean |rate| on the real board dropped to 2.73 C/h; what remains above 10 is the filter's warm-up transient in the first four samples, which is honest. Readout scene puts the actual numbers between the animations: temperature, humidity, sea-level pressure and the signed three hour forecast, each in its channel colour, scrolling. Text is drawn whole-pixel on purpose. Everything else here is sub-pixel and that is what makes it look good, but splitting a 3 px glyph across two columns halves its peak and smears it illegible. Crisp beats smooth when the thing has to be read. site.environment and site.enclosure record where the sensor lives and what has changed around it, with POST /api/environment to change them at runtime. This is not cosmetic: closing a door changes how strongly the sensor couples to outside, which is a regime change in the process the heads are fitting, and at lambda 0.9985 they carry about 55 hours of memory. Left alone they keep predicting the old room for two days. Page-Hinkley would notice eventually but needs matured forecasts to do it, which at the long horizons is the same two days. So the endpoint marks a discontinuity and queues a retrain. |
||
|
|
98210bff8f |
Six enhancements: recompute, markers, vendoring, tests, nerd stats, DS18B20
1. POST /api/recompute re-derives every compensated column from the untouched raw values, removing the step a calibration otherwise leaves through the history. Possible because temp_raw, cpu_temp and hum are never overwritten. Idempotent by construction and tested per row: 0 of 6051 rows change on a second run. 6069 rows in 0.25 s here, so a few seconds on the Pi. 2. Calibration now emits a 'discontinuity' event alongside the calibration log, so downstream views can find the boundary without parsing prose. 3. Vendored Tailwind, Chart.js, hammer, the zoom plugin, KaTeX with its 20 woff2 faces, and both Google fonts into ashvale/static, served by the station. 1.4 MB. Verified with every non-localhost request aborted in the browser: zero external requests, equations still render, fonts still load. The dashboard no longer needs internet. 4. 54 pytest cases over the pure numerics: physics closed forms and round trips, both compensator inverse properties, the Kalman covariance invariants and NIS consistency, the RLS trace cap under a deliberately unexcited regressor, conformal coverage, and the Zambretti ordering. Wired into CI after the seed step so the recompute cases have history. Writing them caught my own sign error on the conformal update: a hit raises alpha and narrows the band, which reads backwards until you follow it through. 5. Stats for Nerds gains the condition number of each head's covariance, a standardised innovation histogram per Kalman filter from a bounded 600 sample ring buffer, and a reliability strip of realised against nominal coverage. All arithmetic on data already in memory. 6. OutdoorProbe reads a DS18B20 over the kernel 1-Wire driver, no new dependency. Polled on its own slower cadence because the sensor blocks for up to 750 ms during conversion, which would eat a third of the 2 s sample budget. Rejects the 85000 power-on sentinel and out-of-range values, and reports age so a dead probe cannot masquerade as fresh. |
||
|
|
e27a4b41c8 |
Outlook to Live, humidity calibration, Models tab rebuilt without scrollers
Seven day outlook moves from History to Live, which now runs four rows. Conditions ahead tightened so the Live column no longer needs a scroller. Adds HumidityCompensator: an additive RH offset estimated by one-step RLS from a trusted hygrometer, clamped to +/-35%, persisted, exposed at POST /api/calibrate/humidity and on the renamed Models and calibration tab. It also implements the psychrometric term (RH moved from element temperature onto air temperature via conserved vapour pressure) but leaves it OFF by default. The thermal argument predicts a hot element reads low; measured against a reference hygrometer this board read 75.4% where the truth was 50.4%, so it reads HIGH and that correction would push it the wrong way. When the flag is enabled, simulate.py applies the exact inverse, per the simulator/compensator trap in DESIGN.md section 2. Models pane rebuilt: the scorecard is one column per target so all 18 heads are visible, and no panel on the tab uses an internal scroller. Verified in Chromium at 1600x900: Live, History and Models all report zero scrollbars, zero clipping, no page scroll, zero console errors. Backtest is numerically identical to the previous commit, confirming the humidity work is a no-op while the flag is off. |
||
|
|
dbbea1a95f | chore: enforce ruff config and fix lint findings | ||
|
|
06ce53bc44 | Initial release: Ashvale Station 1.0.0 |