Next Actions
Immediate Priority
- Run
./scripts/run-real-transfer-trial.shafter each INGV refresh. The 2026-07-29 rerun scored held-out balanced accuracy0.693435versus the historical spatial-rate baseline0.686013, with precision0.315353and recall0.783505. VLF and astronomy were missing-mask features, so this remains an experimental seismic-history baseline, not evidence of multimodal prediction. - Use
./scripts/evaluate-piezo-group-holdout.shas the primary synthetic stability check. The latest 27-run evaluation averages0.578712balanced accuracy, with positive recall0.166667--0.916667and negative recall0.148148--0.914286; it fails the stability gate. - Keep random-init piezo/VLF-only as the leading controlled Transformer architecture. It averages
0.619033within episodes, but has not passed unseen-episode stability; direct and summary branches remain disabled by default. - Keep label-free real VLF pretraining as the default real-data path while supervised VLF-aligned labels remain one-class or sparse; require reconstruction to beat both zero and last-patch baselines.
- Continue periodic INGV refresh and prospective relabeling; the current catalog-coverage guard prevents false maturity. Region-level tables remain one-class, so use
./scripts/prepare-italy-spatial-model-inputs.shfor the fixed-cell baseline until more temporal coverage arrives. - Repeat
./scripts/evaluate-italy-spatial-baseline.shafter each refresh; its--group-by-timesplit keeps all cells from one VLF window in the same partition. The 2026-07-29 rerun scored calibrated balanced accuracy0.655320, precision0.259861, and recall0.666667on 1,064 held-out rows; keep it separate from the fixed-cell transfer trial, which uses a different contract. - Do not promote the spatial model from the current cell holdout: only 5 of 19 cells contain positive labels, so leave-one-cell-out evaluation is mostly one-class. Accumulate more time coverage before using this as a transfer test.
- Treat the permutation result as a stop signal for interpretation: the five timestamp-shuffled controls averaged
0.679362, above the real-order score0.655320. Do not tune the model against this table until the live capture history is substantially longer. - Follow Synthetic Event Alignment Strategies: first compare real and synthetic event-process statistics, then calibrate time/rate, magnitude, and spatial density before another transfer-model sweep.
- Review
data/derived/reports/italy_event_catalog_alignment.json; the first comparison shows synthetic magnitudes are far too high and synthetic episodes occupy fewer cells, so do not use raw synthetic magnitudes for transfer without calibration. - Compare calibrated catalogs using
data/derived/reports/italy_event_catalog_alignment_calibrated.json; rate and magnitude alignment improve, but spatial support remains incomplete. Use the new per-event spatial weights only in a weighted-training diagnostic until a source-observation model is validated. - Use the matching central-Italy catalog for central-Italy simulation profiles. The current raw seed-40 rate is about 16 times too high; do not select an extractor from a five-event sample. Generate and combine several independent episodes before judging temporal or spatial alignment.
- Use the three-episode combined calibrated profile as the current alignment candidate:
data/derived/reports/central_italy_event_catalog_alignment_combined_spatial.json. Its sample-matched nearest-neighbour distance is now3.469km; validate this improvement with more episodes and held-out spatial cells before promoting it. - Do not promote the 40,000-step seed-
4500profile. Its tuned ten-event catalog is too small and has poor magnitude and clustering alignment; generate several independent long episodes and require a minimum event count before recalibrating. - Use explicit simulation coverage duration in future catalog comparisons. The corrected five-episode report is
data/derived/reports/central_italy_event_catalog_alignment_5episode_duration_final.json; its rate ratio is0.924, but held-out episode and cell validation is still required. - Treat
./scripts/evaluate-italy-synthetic-episode-alignment.shas a regime-stability diagnostic only after all inputs share one simulator profile. The current0.278--14.990spread is configuration drift: regenerate seeds under the current source-localized refill profile before judging seed sensitivity. - Use the matched current-profile baseline at
data/derived/reports/central_italy_matched_3episode_q09996_w240_final.jsonfor the next alignment/model smoke tests. Keep the extractor settings (q=0.9996, window240, no event cap) explicit. - Use the per-episode reports under
data/derived/reports/central_italy_matched_q09996_episode_alignment/as a seed-stability gate. Rate is stable, but raw sample-matched clustering spans23.5--42.9 km; do not tune spatial transport against the aggregate alone. - Treat
data/derived/models/central_italy_matched_3episode_transfer_suite.jsonas a matched-model smoke baseline only. Its historical-rate control beats synthetic transfer and all thresholds are recall-driven; do not interpret the high balanced accuracy as skill. - Use
data/derived/models/central_italy_matched_3episode_target_calibration.jsonfor precision-aware comparisons. Report rate-matched precision and recall alongside balanced accuracy; do not select a threshold from the final holdout. - Use
data/derived/models/central_italy_matched_3episode_target_calibration_rolling_controls.jsonas the current precision gate: transfer rate-matched precision0.214is below the historical-rate control at0.228across four rolling folds. Improve features or training before changing threshold policy again. - Follow Feature And Training Options: multiscale seismic/neighbour features are now implemented behind
FEATURE_MODE=multiscale; the first MLP transfer check is negative, so retain compact features as the baseline and test the multiscale path only with more data and stronger calibration. - Use
data/derived/models/central_italy_transformer_sweep_long/summary.jsonas the current five-seed fixed-split Transformer reference. The longer run still favours piezo/VLF-only, but do not select it until the episode-held-out range improves. - Use
data/derived/models/central_italy_transformer_episode_holdout/as the current nine-fold multi-task diagnostic. Mean calibrated balanced accuracy is0.5508, with a0.3788--0.7009range; improve regime robustness before adding more capacity. - The identical occurrence-only control scores
0.5533, slightly above multi-task0.5508; keep occurrence-only as default. Only run count/energy loss-weight sweeps if a representation diagnostic justifies the added objectives, with normalization and threshold selection training-only. - Test domain-robust features with
FEATURE_MODE=relative: use causal local/neighbour activity and magnitude/energy relative to the preceding Italy-wide baseline. Require improvement over the compact model on rolling folds and at least five of nine episode folds before retaining it. - Expand synthetic domain randomization across source schedules, deposition rates, thresholds, erosion, and sensor corruption. Select settings using training episodes only; reserve complete episodes for final evaluation.
- Require cross-regime consensus and calibrated uncertainty for any future event list. Abstain when ensemble disagreement or domain distance is high, and compare predicted rates with the historical INGV rate.
- Keep synthetic pretraining and self-supervised representation learning separate from supervised evidence. Once real labels contain both classes, use chronological real validation with seismic-only, VLF-only, astronomy-only, full, and shuffled-modality controls.
- Treat
data/derived/models/domain_randomized_transformer_episode_holdout/as a failed but reusable stress-test baseline: mean calibrated balanced accuracy0.5096. Before adding more regimes, test target alignment and regime-relative normalization on these same 12 folds. - Treat
data/derived/models/domain_randomized_transformer_episode_holdout_per_window/as the normalization control: mean0.4909, below global normalization. Do not add model capacity yet; inspect event extraction thresholds, horizon labels, and causal regime-relative targets on the same folds. - Add a train-only target audit for each held-out episode: event count, positive rate, event timing, and source/profile metadata. Require comparable target support before interpreting model scores across regimes.
- Obtain one exact ISEE CDF file URL from the archive, record its station/date/units/use policy, and store the unchanged file under
data/raw/vlf/japan/. Do not mark the source usable until the pull is reproducibly nonempty. - After installing
cdflib, runINPUT=data/raw/vlf/japan/<file>.cdf ./scripts/normalize-japan-vlf-cdf.sh; inspect epoch variables and channel units before building features. - For the first ISEE sample, retain the CDF metadata JSON as the source contract. Confirm whether the archive variable is a scalar trace or a time-frequency array before adding a feature adapter; do not flatten a spectrum without preserving its time and frequency axes.
- The native CDF spectrum adapter is implemented in
extract-japan-vlf-cdf-features.sh. Validate its band definitions against additional Moshiri months, preserveresearch_use_only, and compare features with Cumiana only in explicitly declared scientific experiments. - Use
build-japan-vlf-cdf-window-features.shagainst Japan seismic windows after the Japan USGS history is populated; require nonempty overlap and retain Japan-only research reports before any cross-region representation experiment. - Completed the initial Japan temporal extension through 2025-07-15: 319 normalized events, 26 mature windows, and one overlapping window for each Moshiri sample. Keep this as an ingestion gate only; it is not sufficient model coverage.
- Use
./scripts/process-japan-vlf-manifest.shas the standard Japan preprocessing entry point. Add more manifest rows only after station/date/permission metadata are recorded, then rerun the workflow and audit overlap before model training. - Install and monitor
elfquake-japan-vlf.timeras a separate Japan research-data collector. Confirm the archive's publication delay and adjustLOOKBACK_MONTHSorMAX_FILESonly after checking storage and overlap growth. - Run
./scripts/build-japan-vlf-cdf-dataset.shafter each refresh to produce one combined Japan VLF row per seismic window; use this artifact as the input to the Japan design-matrix join. - The Japan refresh now rebuilds the combined CDF windows and model-input table automatically when
WINDOWSis set. The current table has 78 target windows but 75 missing VLF rows; acquire matching CDF dates before training or cross-region comparison. - Use
data/derived/models/japan_vlf_model_input.m5.csvas the current Japan smoke-training table only. Its comparable M5.0 target has49/29positive/negative windows; eight now contain observed VLF, including three negative and five later/test-era windows. The refreshed 80/20 tabular run scored0.375calibrated balanced accuracy and failed negative recall; do not tune on this sample. - Run
./scripts/probe-japan-target-thresholds.shafter each Japan refresh. Select thresholds using class balance and VLF overlap; do not optimize the threshold against model scores while only three VLF windows are observed. - Run the Japan VLF sequence smoke test with the capture-specific manifests listed in
data/derived/models/japan_moshiri_sequences/manifests.txt; treat it as an interface check only until the number of VLF-observed target windows grows substantially. - Require more than one observed Japan VLF capture in both chronological train and test periods before evaluating the Japan Transformer. The absolute coverage blocker is cleared, but seven windows remain insufficient for a reliable score and most target rows are still masked.
- Use
data/derived/models/mixed_source_transformer_fixture.alignment.jsonand Transformer Fixture as the current cross-source inventory. Implement the common window builder only after preserving domain-specific time scales and missing-modality masks. - Use
./scripts/build-common-transformer-fixture.shto refresh the 5,546-row mixed window fixture. Treat itsready_for_smoke_trainingstatus as a tabular/tensor interface gate only; build continuous per-dataset sequence inputs before patch-Transformer training. - Use
./scripts/materialize-common-transformer-sequences.shfollowed by a two-epochsequence_common_multimodalCPU smoke run. The current run completes across 48 manifests but scores0.489320calibrated balanced accuracy, below the majority baseline; retain it as an interface/masking test only and do not tune against it. - Run
./scripts/audit-common-transformer-alignment.shbefore any mixed-source model comparison. The current audit finds 5,301 Italy seismic/VLF/astronomy rows, 8 Japan seismic/observed-VLF rows, and 167 fully co-observed synthetic sensor rows; acquire more temporally matched Japan VLF before evaluating Japan transfer. - Use
./scripts/compare-japan-synthetic-shapes.shand Japan And Synthetic Shape Comparison as the current signal-shape gate. The first run finds Japan VLF low-band power0.771versus synthetic piezo0.361, and Japan seismic event energy is much sparser and heavier-tailed than the synthetic catalog. Tune rate/clustering and the causal piezo envelope before another transfer sweep. - Run
./scripts/evaluate-piezo-japan-shape-variants.sh. Retain a slow-envelope setting only if it moves the piezo PSD and low-band ratio toward the Japan feature trace across multiple seeds without creating an artificial trend. - Run
./scripts/tune-japan-avalanche-events.shover multiple seeds and time-held-out episodes. Treat the current 25-event candidate as a rate-calibration control only; improve clustering without relying on a global event cap. - Use
data/derived/reports/japan-avalanche-policy-seeds/summary.csvas the current multi-seed gate. The fixed policy is rate-stable enough for a control but fails clustering, so do not promote it as the simulator default. - Add a configurable clustered-loading/relaxation regime to the sandpile simulation, preserving localized source locations and deterministic seeds. The first per-source persistence smoke run was stable but did not improve clustering; do not promote it.
- Retain
SOURCE_REGIME_DECAY=0.2,SOURCE_REGIME_BOOST=0.8, andTARGET_FILL_REGIME_FLOOR=0.25only as a negative control. The 5,000-step probe remained stable but failed event-shape alignment. - Refactor the synthetic event aggregation boundary or add a longer-lived stress-release state. Compare event inter-arrival, burst-run, and PSD metrics without changing source coordinates or injecting independent events.
- Use
data/derived/reports/avalanche-burst-extractor/summary.csvas the current burst-extraction diagnostic. Thedecay99_gap120candidate improves rate, PSD slope, and tails without a global event cap, but remains a single-seed control. - Use
data/derived/reports/avalanche-burst-seeds/summary.csvas the multi-seed diagnostic. Rate and tail behavior are promising, but burst clustering and PSD sign are not stable; do not use it for model fixtures yet. - Add train-only burst-threshold calibration: estimate the baseline-score threshold from training episodes, apply it unchanged to held-out seeds, and report event rate, burst runs, tails, and PSD without per-test-episode retuning.
- Use
data/derived/reports/avalanche-burst-train-test/summary.csvas the leakage-safe gate. The fixed threshold fails held-out rate transfer (0.101–0.170versus0.084), so the extractor is not ready for fixtures. - Retain the relative-baseline normalization as a negative control. It does not improve held-out rate transfer, so do not use it for fixtures.
- Use the bounded source stress reservoir as the next simulator candidate. The 500-step probe is stable and produces localized releases, but it has not yet passed event-shape evaluation.
- Retain the tested stress-reservoir parameters as a stable negative control. They produce releases without safety failure but only one extracted event over 5,000 steps.
- Do not increase stress release mass blindly. The per-source cooldown control (
SOURCE_STRESS_RELEASE_COOLDOWN_STEPS=120) is stable but still yields one extracted direct event over 5,000 steps. Redesign the stress-release coupling or event representation so localized releases create distinct, short-lived global activity episodes, then repeat the causal burst and safety gates. - Use the optional
*.source_stress.csvoutput to compare source-local stress-release pulses with nearby avalanche activity using causal, per-regime normalization. The output is implemented and preserves source coordinates; evaluate it against the existing direct signal rather than replacing it silently. - Build the release-aware diagnostic and report lead/lag, local activity, and event-shape metrics. Promote only if it improves held-out multi-seed alignment without safety failures or independent event injection.
- Run
scripts/analyze-source-stress-alignment.shover tagged episodes. The first 1,000-step cooldown result has 938 release rows, only5.9%with positive local excess activity, local/global excess-AUC ratio0.037, and median local peak lag96steps. This does not support a source-local precursor effect yet. - Repeat the release-aware diagnostic over multiple seeds and radii. Completed on seeds
40–43: positive local excess was2.2%at radius 16,6.8%at radius 32, and23.5%at radius 64; local/global excess-AUC ratios were0.010,0.043, and0.188, with median lags near 90–102 steps. This is weak and spatially broad, so the stress reservoir remains a negative control. - Improve the spatial avalanche activity representation before further simulator tuning. Compact
*.avalanche_regions.csvoutput is now available as a configurable regional grid, preserving locality without per-cell output. Use it to rerun the source-stress diagnostic and compare against the existing global signal. - Add a region-aware source-stress diagnostic using the regional table. Report activity in the release region versus matched non-release regions, with time-held-out seeds.
- Track the Japan VLF archive until 28 July data are available. The first pre-event check covers 13 daily Moshiri CDFs from 15–27 July; July 26 is the highest robust-deviation day, but earlier elevated days prevent a precursor claim. Re-run with hourly event-day/post-event coverage and matched controls.
- Checked the official ISEE/ERGSC July 2026 Moshiri archive on 2026-07-30. Twenty additional hourly CDF URLs are listed for 2026-07-27 04:00--23:00 UTC; the newest file was downloaded and decoded successfully (
8,646rows,ch1/ch2spectrogram features). No July 28--30 Moshiri CDFs were listed, and the Kagoshima July directory returned no CDF filenames. Japan data remain research-use-only. - Checked the new Moshiri sample against the 14-day pre-event baseline. The 2026-07-27 23:00 UTC score is elevated (
2.322) but is below earlier 15 July (2.657) and 16 July (2.345) scores. Treat it as a candidate anomaly only; obtain 28 July event-day/post-event files and matched controls before any precursor interpretation. - Added the MkDocs Material public site scaffold with curated navigation, status homepage, blog support, first journal post, local authoring guidance, and GitHub Pages Actions deployment. Install
requirements-site.txtand runmkdocs build --strictbefore enabling Pages. - Refreshed the Italy transfer-trial artifact and held-out map on 2026-07-30. The result remains
0.693435balanced accuracy and0.315353precision, with 4 actual and 8 predicted held-out events; this is unchanged experimental baseline behavior, not evidence of prediction skill. The Pages build now has zero broken internal links and uses full Git history for RSS dates. - Reran the synthetic piezo group holdout on 2026-07-30. The 27-run CPU evaluation scored
0.578712mean balanced accuracy and failed the recall-stability gate; retain the current architecture as a diagnostic and investigate episode/regime drift before another capacity sweep. - Reran the synthetic event-list drift audit on 2026-07-30. The temporal split has
31.9%positive training rows versus92.8%positive test rows, a0.609rate shift; episode totals alone conceal this failure. Do not interpret the Transformer holdout score until the target horizon/rate drift is reduced or explicitly modeled. - Keep the Italy transfer-trial artifact at
data/derived/models/real_transfer_trial/report.jsonas the current chronological baseline. Compare future multimodal runs against its confusion matrix and historical-rate control, with missing-modality masks reported explicitly. - Compare the two Italy baselines using identical time ranges and target contracts before interpreting any apparent improvement. The mismatch is documented in Italy Baseline Comparison; do not attribute the transfer-trial difference to VLF or astronomy until the spatial feature and label definitions are matched.
- Build one weekly fixed-cell evaluation table and run matched seismic-history-only, VLF-mask-only, full multimodal, and synthetic-pretraining ablations with the same chronological folds and training-only threshold calibration. Completed in
data/derived/models/matched_italy_ablation_suite/summary.json; full/pretrained scored0.693435, seismic-history-only/real-only0.687797, and VLF-mask-only0.601370. - Completed the matched suite across seeds
42,99, and123. Full/pretrained averaged0.690230balanced accuracy (0.684675--0.693435) versus0.671673without pretraining; seismic-history/pretrained averaged0.685992, and seismic+VLF-mask/pretrained0.689776. VLF-mask-only stayed near0.601, confirming that the constant holdout mask cannot identify VLF utility. - Build a real-VLF holdout with observed variation before making a modality claim. A new Cumiana image was captured at
2026-07-29T08:45Zand integrated; the descriptive association table now has four VLF-observed weeks, but the new target is pending. The separate M2.5 central-Italy table has 229/50 positive/negative rows, but its chronological test has only one negative and its0.990909score is majority-class driven. Keep it as a data-shape artifact, not model evidence. - Extend Cumiana capture coverage until at least three positive and three negative VLF-observed target weeks exist in the same threshold and horizon contract. Then run matched seismic-only, VLF-image-only, and full multimodal time-held-out baselines with training-only threshold calibration.
- Completed the self-supervised VLF refresh: 280 rows, 257 causal windows, and 41 exploratory alerts at score
>=0.8. Re-label the new2026-07-29T08:51Zwindow after its horizon matures, keeping the anomaly threshold and model checkpoint fixed for the audit. - Added
./scripts/render-vlf-anomaly-chart.shanddocs/images/anomaly.png. The renderer breaks lines across observation gaps, preventing sparse captures from appearing as continuous trends. - Added the dated current-window audit
data/derived/reports/italy_data_coverage_20260729.json: 280 INGV events, 284 VLF metadata records, 257 anomaly windows, and 4 VLF/seismic overlap weeks. Use this alongside the longer historical coverage report to avoid mixing time ranges. - Added capture continuity monitoring with
./scripts/report-vlf-capture-gaps.sh. The current report finds 280 captures and 8 gaps over one hour, including a 308-hour gap before the July 29 capture; use it to verify the systemd collector is producing sustained VLF coverage before interpreting future anomaly/earthquake overlap. - Hardened
deploy/systemd/elfquake.servicewith unconditional restart and no start-limit lockout. Install the updated unit and monitor its journal; rerun the gap report after several capture intervals to confirm continuity improves. - Fixed the namespace failure caused by the
/home/danny/github -> /chalet/githubsymlink: the service now uses/chalet/github/elfquakefor its working directory, Python path, documentation path, and writable data path.systemd-analyze verifypasses. - The post-fix journal showed the unit was still using
/usr/bin/python3, which lackednumba; it entered a restart loop. PointedExecStartat/chalet/github/elfquake/.venv/bin/pythonand added a five-restart-in-five-minutes guard. Reinstall the unit, reset its failed state, and verify one successful capture. - Configured Numba’s
UserProvidedCacheLocatorwith writable cache directory/chalet/github/elfquake/data/derived/numba_cache, resolving theno locator availablefailure caused by systemd filesystem protection. Reinstall and restart the unit, then verify a capture in the journal. - Made Numba caching explicitly configurable through
ELFQUAKE_NUMBA_CACHE; the systemd service disables disk caching while ordinary simulation runs keep it enabled. Service output now confirms the import-time cache failure is resolved. - Leave the collector running and rerun
./scripts/report-vlf-capture-gaps.shafter several 30-minute intervals. Confirm that new captures are being added and that no multi-hour gaps recur before using future VLF anomaly windows in analysis. - Verified
elfquake-prospective.serviceand its timer after installing the real/chaletpaths, project venv, and service-mode Numba settings; the manual update completed without errors. - The collector has since added a
2026-07-29T11:15ZCumiana capture. Current prospective summaries have 280 rows and one pending target each; continue monitoring the 30-minute cadence and rerun the gap report before the next matured-label refresh. - Ran the refreshed real-VLF versus synthetic-piezo embedding probe. The closest 25% synthetic windows improved centroid distance
3.045->2.531and nearest distance2.186->1.404, but synthetic reconstruction MSE remained11.604versus real0.562. Treat inlier filtering as diagnostic only; next improve simulator signal-shape statistics before model training. - Ran the 20,000-step piezo shape sweep.
gain_burstis the leading single-run candidate (1.512centroid,1.028nearest distance), withfast_burstclose behind; validate both across multiple seeds and compare against the current profile before changing defaults. - Completed the multi-seed variant check with
./scripts/evaluate-piezo-vlf-variant-seeds.sh.gain_burstled narrowly on the fixed alignment seed; the independent model-seed rerun gives means of1.7488centroid /1.0537nearest forgain_burst,1.7558/1.0636forfast_burst, and1.7490/1.2575for current. The preference is weak because per-run ranges are wide; do not promote a transform yet. - Completed the Cross-Region Generative Smoke Test: synthetic masked pretraining, Japan ISEE self-supervised continuation, chronological Italy fine-tuning on seismic + Italy VLF + astronomy, and continuous coordinate/magnitude event-list output. All qualifying events are retained in a two-slot target tensor rather than reducing each cell to its strongest event. The CPU run used 5,263 mature Italy rows, held out 1,023 rows, trained on 45 Japan windows, and emitted 5 continuous event rows in
docs/images/cross-region-generative-smoke.png. It is a pipeline sanity check only; Japan remains research-use-only and no prediction skill is demonstrated. - Repeat the cross-region smoke run with at least three seeds. Measure coordinate error, in-Italy rate, spatial dispersion, and duplicate-location rate; compare event count against the historical spatial-rate baseline before treating the output as more than an interface artifact.
- Use
./scripts/trial-weekly-event-forecast.shas the current end-to-end event-list contract smoke test, not as a validated predictor. - Use
./scripts/balance-italy-synthetic-episode-rates.shonly as an auditable training/observation-model diagnostic. It can thin overactive episodes, but it must not synthesize events for underactive episodes; the matched rerun is preferred.
Modeling
- Run
./scripts/run-transfer-experiments.shafter each real-data refresh. It compares historical rate, real-only random initialization, synthetic transfer, rolling-origin folds, and a train-only grid selection before one final holdout evaluation. The default synthetic corpus now includes four long episodes; add more 20,000-step episodes before treating transfer changes as stable. - Generate more independent warmed episodes and rerun leave-one-episode-out evaluation; nine episodes are not enough to estimate regime robustness tightly.
- Combine at least five scope-matched long episodes, apply calibration using training dates only, and report rate, magnitude, inter-event, sample-matched clustering, and occupancy metrics together.
- Keep the five-episode candidate as a diagnostic benchmark, not a training default, until its corrected rate and spatial metrics survive held-out episode and cell checks.
- Add regime-conditioned reporting or a mixture-of-regimes simulation before another global calibration pass; the current episode rate spread is too large to treat one global thinning factor as a physical correction.
- Investigate the remaining matched rate and clustering differences using the
4600--4800source/loading trajectories; change simulation dynamics only if a stable cause is found. - Do not add the default piezo potential channel to model training yet. Its spatial average failed a nine-episode causal lead-time check; event-nearest diagnostics are positive but use future event locations and are not valid inputs.
- Calibrate weekly event counts against historical INGV
>M2rates before trusting any neural score scale. - Compare every weekly forecast run with
./scripts/compare-weekly-forecasts.shand track Stage 1/Stage 2 pass/fail status. - Keep direct avalanche-derived seismic features separate from piezo/VLF-like features; use ablations to test their contribution independently.
Data
- Keep accumulating Cumiana VLF image captures and refreshing image features.
- Refresh prospective INGV labels as target windows mature; train supervised real models only after one table has both positive and negative labels.
- Validate Abelian Cumiana live/archive audio only if a reproducible nonempty pull is found; current probes returned zero usable bytes.
- Extend historical INGV backfill earlier than 2024 only if weekly baseline calibration needs longer seasonal coverage.
- Repeat mixed real/synthetic VLF alignment after new Cumiana captures; require improvements over centroid and random controls before relying on inlier selection.
- Keep event-count, energy, and spatial-occupancy targets alongside binary occurrence; do not make one thresholded event label carry all timing, magnitude, and location information.
- Fit magnitude calibration on the real training period only, then compare calibrated and uncalibrated synthetic catalogs before adding temporal-rate or spatial-density transforms.
- Treat rate thinning as an observation model, not a simulation fix; retain the raw event catalog and test whether spatial reweighting improves cell occupancy without moving localized source events.
- Add a joint alignment score with minimum sample gates: rate ratio, magnitude distance, inter-event distance, nearest-neighbour distance, and spatial occupancy must be reported together.
- Preserve the combined-episode time offsets and calibration metadata in every synthetic training artifact; do not collapse episodes back onto their shared demonstration clock.
- Use sample-size-matched nearest-neighbour statistics for catalog clustering. Do not compare a 32-event synthetic catalog directly against all 594 real events.
Simulation
- Run
./scripts/run-longer-synthetic-transformer-batch.shwhen CPU time is available, validate drift, then rerun./scripts/evaluate-piezo-group-holdout.shagainst the larger episode set. - Keep
damage_totalas a validated synthetic precursor diagnostic, not a default Transformer feature. A matched nine-fold screen regressed from0.599648without damage channels to0.586848with them. - Keep the duration-aligned
SOURCE_COUNT=64, refill470, removal interval20, andq=0.998/window=120profile as a valid synthetic target baseline (47.0%positives, temporal drift0.182). It has no confirmed piezo lead and is not a precursor-training profile. - The first two-stage mature-weakness profile failed its nine-episode causal confirmation despite stable target drift. Do not tune its scalar parameters immediately or train a model. Document a stronger physical mechanism proposal, such as a spatially propagating rupture/nucleation state, before another synthetic dynamics run.
- Compare future episode-batch h6 drift against the current scaled
WARMUP_STEPS=3000delta0.187025. - Revisit structured initial fill only with delayed bottom-layer removal; the first fill probe drifted at
0.307937. - Tune the piezo/VLF mapping only from
*.piezo.csvand compare against Cumiana VLF shape reports.
Maintenance
- Keep docs concise: one current source doc, one simulation doc, one modeling doc, one operations/steps doc, and one report.
- Split
tests/test_acquisition_scaffold.pyby subsystem if test maintenance starts slowing changes. - Add chunked sandpile snapshot storage only if larger pretraining runs outgrow current
.npysanity snapshots. - Keep optional dependencies CPU-compatible on this system; do not add GPU-only paths.
Recent Completed
- Added
run-transfer-experiment-suiteand./scripts/run-transfer-experiments.sh. The suite performs matched real-only versus synthetic-pretrained ablations, four rolling-origin folds, and a train-only threshold/grid selection followed by final holdout evaluation. - Generated a 20,000-step CPU episode at seed
4300and stacked its dense synthetic records with seeds40--42. Synthetic episodes are offset in synthetic time during training because their source timestamps intentionally share one demonstration origin; real timestamps are unchanged. - The expanded transfer corpus contains 79,976 synthetic records and 190 weekly spatial samples. Transfer remains below the historical spatial-rate baseline on precision and is not predictive evidence.
- Added modality-specific Transformer patch adapters, elapsed-time inputs, separate observed/corruption/padding masks, masked reconstruction, modality dropout, frozen probes, compatible checkpoint transfer, and missing-modality checks.
- Added stable name-derived Transformer initialization. Adding or reordering unused modality adapters no longer changes shared weights or advances the global PyTorch RNG.
- Reran seven transfer regimes under controlled initialization. Random-init piezo/VLF-only is strongest at mean
0.619033; synthetic pretraining falls to0.534272, so current self-supervision has no demonstrated downstream gain. - Reran late gated fusion under the same control. Random anchored full and direct-only variants reach
0.609241and0.606385, but neither beats the piezo/VLF anchor and direct improves when disabled. - Added and ran nine-episode by three-seed leave-one-episode-out evaluation. Mean balanced accuracy is
0.578712, the worst fold is0.275641, and 14 of 27 folds meet both recall floors. - Added a fixed three-seed probability ensemble with training-only threshold calibration. It raises mean unseen-episode balanced accuracy to
0.632634, but only 6 of 9 episodes meet both recall floors; a fixed threshold passes the same six. - Diagnosed episode instability: the same piezo features change class-effect sign between trajectories, while h6 labels commonly form six-row positive runs against only 12 minutes of model context.
- Tested h3 labels, h6 with 60-minute context, and opt-in spatial sensor aggregates. Their ensemble means are
0.580991,0.512103, and0.544096; all trail the h6/12-minute mean-only baseline. - Added
compare-piezo-group-holdoutsand./scripts/summarize-piezo-group-holdouts.sh; zero of four current variants passes both the mean and recall-stability gates. - Added synthetic event-list target generation from avalanche event CSVs, including future count, occurrence, magnitude, centroid, and time-to-first-event fields.
- Added dependency-light synthetic event-list heads for occurrence, count, max magnitude, and centroid. The h6 balanced split reaches balanced accuracy
0.887566, count MAE0.506783, and centroid median error145.585806 km; the temporal split still fails with balanced accuracy0.500000. - Added drift diagnostics, episode annotation, and a validation wrapper. Current h6 temporal split has train positive rate
0.318653and test positive rate0.927835; episode-balanced validation reaches balanced accuracy0.878079. - Added
./scripts/run-synthetic-episode-batch.shfor shorter stationarity-tuned synthetic episodes with localized sources preserved. - Ran two stationarity profiles. The first six-episode 5000-step batch still drifted badly (
0.652766positive-rate delta). The aggressive three-episode 3000-step probe fixed drift (0.065608delta, warningok) but is too small for model selection. - Added structured initial fill and ran a three-seed probe. It started loaded and stayed near mean height
3.2, but h6 drift was0.307937, worse than the no-prefill aggressive profile. - Added unrecorded sandpile warm-up. The first three-seed
WARMUP_STEPS=1000probe improved h6 drift to0.017989with warningok. - Scaled
WARMUP_STEPS=1000to nine episodes; drift failed again (0.294146delta), showing the small result did not scale. - Tested
WARMUP_STEPS=3000on three seeds. It passed h6 drift (0.048677, warningok) and the temporal event-list smoke model reached balanced accuracy0.674342; this is the new episode-batch default but still needs scaling. - Scaled
WARMUP_STEPS=3000to nine episodes. The run stayed drift-ok (0.187025delta) with 396 labeled rows, but temporal model balanced accuracy was only0.468045. - Tested denser direct avalanche extraction on the same nine episodes.
_q099_w60_m10saturated h6 labels (0.828283positive rate), while_q0995_w120_m5improved class balance (0.512626) but did not beat the sparse default model checks. - Added deterministic feature-bag ensembles to the dependency-light event-list occurrence head. The default script now uses 8 members with 50% feature bags, improving the scaled sparse temporal/balanced/episode-balanced checks to
0.498120,0.616473, and0.607551. - Added richer synthetic event-list targets for event-rate, log magnitude energy, early/middle/final horizon counts, peak timing, event duration, and spatial spread.
- Added regression shape heads for the richer event-list targets and report metrics for each head. On the balanced split, event-rate MAE is
0.068936, early/middle/final count MAE is0.180167/0.246829/0.194170, and spatial-spread MAE is20.317859 km. - Added an optional boosted-stump occurrence head. It improves the balanced engineering split to
0.629173balanced accuracy, but fails the chronological split at0.347744; keep the feature-bag logistic ensemble as the default. - Added
./scripts/probe-synthetic-event-list-models.shandsummarize-synthetic-event-list-probesto make horizon, burn-in, feature-cap, ensemble-size, boosted-stump, and balanced-control checks reproducible. - Smoke-tested the probe harness with
HORIZONS=3,BURN_IN_FRACTIONS=0,RUN_BALANCED_CONTROLS=0, andEPOCHS=20; it wrote comparable summaries and confirmed the harness works end to end. - Ran the full synthetic event-list probe. The best drift-ok chronological result was h6 with a 32-feature cap at balanced accuracy
0.511278; the best drift-ok balanced control was h12 boosted stumps at0.651543. - Updated probe summaries so model rows include matching target drift warning and positive-rate delta, making drift-ok temporal results easy to separate from suspect runs.
- Added
build-synthetic-lagged-contextand./scripts/probe-synthetic-event-list-lagged-context.shto test explicit recent-history features without target leakage. - Ran lagged-context h6 probes. The best drift-ok chronological result was the 256-feature cap at balanced accuracy
0.587093, just below the synthetic gate; all features regressed to0.530702. - Added
train-synthetic-event-list-sequence-headand./scripts/train-synthetic-event-list-sequence-head.sh. The h6 lookback-12 seed-42 sequence head reaches calibrated balanced accuracy0.609649, but seed/config checks range from0.467419to0.603383, so it needs stability work before promotion. - Added
./scripts/sweep-synthetic-event-list-sequence-head.shandsummarize-synthetic-event-list-sequence-heads. The first 18-run sweep selected lookback12, dropout0.1as the best mean config: mean0.600459, min0.576441, max0.645363. - Added
ensemble-synthetic-event-list-sequence-headsand./scripts/ensemble-synthetic-event-list-sequence-head.sh. The three-seed ensemble scored0.591479; pairwise42+99scored0.644110, showing useful but selection-sensitive variance reduction. - Added validation-selected and early-stopped sequence-head controls. Validation-selected lookback-12/dropout-0.1 averaged
0.539265; early-stopped dropout-0.1 averaged0.494570, so neither should be promoted. - Added
prepare-transformer-target-input,./scripts/train-synthetic-event-list-patch-transformer.sh, and./scripts/sweep-synthetic-event-list-patch-transformer.shto train the patch Transformer directly on the richer h6 event-list target table. - Added causal multiscale transfer features, ran the MLP baseline and a four-configuration CPU patch-Transformer sweep, and recorded the negative MLP comparison and piezo/VLF-only Transformer result in the model docs.
- Extended the patch-Transformer sweep to five seeds (
7,17,42,99,123) and 36 epochs across 20 runs. Piezo/VLF-only remained the strongest mean ablation at calibrated balanced accuracy0.5620; this remains synthetic-only evidence. - Added optional Transformer regression heads and ran a nine-fold leave-one-episode-out multi-task check. Occurrence balanced accuracy averaged
0.5508; count and log-energy MAE averaged0.4234and2.7124. - Added causal
FEATURE_MODE=relativefeatures and ran the matched transfer probe. Rolling balanced accuracy was0.907935, below the multiscale control at0.912232; retain the compact baseline and treat relative features as an available diagnostic. - Compared multi-task and occurrence-only Transformer heads on identical episode folds. Occurrence-only was marginally better (
0.5533versus0.5508), so auxiliary targets remain opt-in. - Added
run-domain-randomized-synthetic-batch.shwith explicit baseline, slow-fill, and fast-localized profiles. Generated 12 episodes and ran complete-episode holdouts; the model averaged0.5096, failing the regime-invariance gate but establishing a reproducible stress test. - Added opt-in per-window sequence normalization and ran the same 12-fold control. It scored
0.4909, so it is retained only as a negative normalization control. - Updated the sandpile determinism test for the current damage-reporting schema. The focused simulation test and the 138-test acquisition/feature suite now pass.
- Ran a four-config h6 patch-Transformer sweep on the warmed nine-episode synthetic data. Piezo/VLF-only was the strongest ablation in every config; best short-run calibrated balanced accuracy was
0.608629with lookback12, patch3, dropout0.1. - Added
./scripts/run-longer-synthetic-transformer-batch.shas the repeatable CPU-only route for larger warmed synthetic batches before Transformer retuning. - Added and ran
./scripts/trial-weekly-event-forecast.sh; the current2026-07-08trial emits 25 capped>M2event-coordinate rows for2026-07-08to2026-07-15. - Added and ran
./scripts/learned-weekly-event-forecast.sh; it trains a synthetic-window logistic scorer and emits the same weekly event-list CSV contract. - Added learned-scorer metadata to the forecast report without changing the CSV event-row contract.
- Added
docs/success-criteria.mdwith staged scaffold, synthetic utility, real readiness, and prediction-claim gates. - Added and ran
./scripts/compare-weekly-forecasts.sh; Stage 1 event-contract criteria pass, Stage 2 synthetic-model criteria fail. - Refreshed INGV prospective labels and rebuilt real model inputs; both scopes now have 69 labeled rows but remain class-blocked.
- Added and ran
./scripts/trial-forecast-map.sh, renderingdata/derived/maps/mag_gt2_weekly_trial_forecast_map.pngfrom the trial forecast CSV. - Added
docs/forecast-interface.mdto define the stable weekly event-list output contract for trial and future learned scorers. - Added
docs/output-example.mdwith the top three highest-magnitude trial rows and nearest mapped places. - Added self-supervised real VLF pretraining and label-free anomaly scoring as the default real-data development path while labels are sparse.
- Built current real VLF-aligned all-Italy and central-Italy model inputs; both remain class-blocked for supervised real training.
- Extended INGV historical backfill from
2024-01-01through2026-07-07, producing historical seismic baseline windows. - Added synthetic aligned sequence/tensor paths, GRU and patch-Transformer smoke models, missing-modality checks, and model-run summaries.
- Added direct avalanche-derived event extraction, synthetic event maps, and separate piezo/VLF-like signal outputs.
- Added repo-local Codex skills for source ingest, data refresh, simulation, and synthetic modeling workflows.
- Added the research-only Japan model-input builder and registered the
japan_vlffeature group, ablation, tensor-spec modality, and missing-VLF quality mask. The current 26-row table has only two observed VLF windows and remains one-class at M3.0. - Extended Japan seismic coverage through 2026-07-08, rebuilt 78 aligned weekly windows, matched all three Moshiri CDF samples, and ran the first M5.0 CPU tabular smoke baseline. The target has
49/29positive/negative windows, but VLF is present in only three; no predictive utility is demonstrated. - Added the Japan CDF sequence adapter, stable dataset IDs, capture-specific sequence manifests, and
sequence_japan_vlf_only. The capture-safe Transformer fixture rejects evaluation when held-out windows have no Japan VLF coverage instead of crossing multi-month gaps. - Added three bounded June 2026 Moshiri CDF captures from the official archive, bringing observed Japan VLF windows from three to five, with three later/test-era windows. Rebuilt the research-only combined windows and model input, and preserved multi-file source provenance for overlapping windows.
- Rebuilt six capture-specific Japan sequence manifests and ran the guarded Japan patch-Transformer smoke test. It returned
insufficient_split_rowsbecause the current M3 target is still one-class; no model score was reported. - Added four more bounded June 2026 Moshiri CDF captures. The Japan pipeline now has ten validated CDF samples, seven observed weekly windows, and five later/test-era windows; 71 of 78 target windows remain VLF-missing.
- Rebuilt the comparable M5.0 Japan model input over 2025-01 to 2026-07: 78 rows with
49/29positive/negative targets and seven observed VLF windows. All observed windows are positive, so supervised VLF evaluation remains blocked. - Added a January 15, 2025 Moshiri CDF from a known negative M5 target week. The refreshed table now has eight observed windows, three negative; the CPU tabular smoke baseline scored
0.375calibrated balanced accuracy and remains a diagnostic only. - Generated the first mixed-source alignment manifest covering synthetic, Italy, and Japan artifacts. It is an auditable provenance fixture, not yet a trainable mixed-domain batch; the remaining gaps are documented in
docs/model-fixture.md. - Built the common window fixture and materialized 48 dataset/modality sequence manifests with explicit zero-filled missing channels and masks. The CPU patch-Transformer smoke run completed end to end, but its calibrated balanced accuracy was
0.489320, so it remains an interface check rather than model evidence. - Added the common-fixture alignment audit and corrected zero-valued missing-count handling. The audit now separates nominal interval overlap from actual same-row observations and reports the current Italy, Japan, and synthetic co-observation gates.
- Refreshed the common cross-source fixture after Japan acquisition: 11 dataset IDs and 88 missing-aware sequence manifests are now represented, with eight Japan seismic/VLF co-observations.
- Added the Japan CDF versus avalanche shape comparison. The report covers time-domain tails, burstiness, autocorrelation, and PSD shape, and identifies low-frequency piezo mismatch plus synthetic event-rate/tail mismatch as the next simulation targets.
- Fixed prospective-label maturity so a target is only labeled when the event catalog covers its complete horizon. The incremental refresh now rebuilds existing candidate rows, uses stable current catalog paths, and caches unchanged VLF image features.
- Added causal pre-relaxation piezo-to-post-relaxation avalanche lead-time analysis. Old release signals are same-step effects; a new stored-potential sensor did not pass the spatially averaged nine-episode test. Event-nearest diagnostic support is not usable for prediction because it relies on future event coordinates.
- Fixed mountain target refill so
TARGET_FILL_MODE=sourcesdeposits refill mass at persistent source locations instead of uniformly random cells. The episode batch now uses this localized loading mode by default. - Added causal
top_kandtop_k_risesensor pooling to the lead-time analyzer. Both fail on the nine-episode localized default, so dynamic pooling alone does not recover the oracle diagnostic. - Made
SOURCE_COUNTexplicit in the episode-batch runner and derived its end time from simulated steps, removing 34 hours of false empty padding after 3,000-step runs. - Confirmed the corrected 64-source target baseline over nine episodes: 387 labeled rows,
182/205positive/negative, temporal positive-rate delta0.182, and h6 temporal balanced accuracy0.518750for the simple event-list model. - Rejected the three-episode potential-lead result: its 15--30 and 180--360 step effects did not survive nine-episode causal confirmation across 40 events.
- Added modular pre-relaxation spatial-state metrics: near-critical contact count, coherence, and weighted stress. Their encouraging three-episode screen failed nine-episode causal confirmation across 39 events, so they remain diagnostics only.
- Added opt-in delayed local failure through
DamageConfig: near-critical cells accumulate damage, damage lowers only their local relaxation threshold, and toppling resets it. The damage-enabled nine-episode run has 387 labeled rows,197/190positives/negatives, drift0.245581, and a confirmed pre-relaxationdamage_totallead at 5--15 steps (AUC 0.652315, positive in 6/9 episodes). - Added named piezo-channel exclusions to group holdout for matched feature ablations. On the damage profile, the single-seed nine-fold Transformer screen is lower with damage channels (
0.586848) than without (0.599648); do not promote them yet. - Added 15-minute synthetic step targets sampled every 5 minutes to match the damage lead. The matched short-horizon nine-fold screen is also lower with damage (
0.529700) than without (0.530826), so a generic Transformer does not yet exploit the causal state. - Tested dedicated damage-only patch-Transformer branches at 12- and 24-minute lookback. They reach
0.500824and0.504695; isolating or extending context does not recover predictive utility.
Japan parallel data path
- Run
./scripts/backfill-japan-history.shand verify nonempty USGS raw and normalized outputs. - Identify one reproducible current passive broadband ELF/VLF Japan sample and add it to
data/raw/vlf/japan/manifest.csv; prioritize ISEE Moshiri or Kagoshima over WALDO. - Compare Japan and Italy source coverage before any cross-region model training.
- Use the confirmed ISEE permission to obtain one recent Moshiri or Kagoshima digital sample, then build the Japan VLF adapter for its native CDF format.
-
Keep WALDO out of the main acquisition schedule; revisit it only for a defined historical case study or optional self-supervised pretraining corpus.
-
Refreshed Italy data through 2026-07-16: 67 new INGV events were pulled, one new Cumiana
last_E_VLFimage was captured, and the prospective tables now contain 279 rows with 277 mature rows. Both all-Italy (277/0) and central-Italy (0/277) remain class-blocked. - Rebuilt the real VLF sequence and model inputs: 279 image rows and 256 anomaly windows now extend through 2026-07-16. The label-free smoke forecast remains a novelty artifact, not a seismic prediction.
- Audited the mirrored all-Italy/central-Italy label counts. Equal row counts are correct because both scopes use the same VLF anchors; the one-class labels are target saturation, not a region-filter bug. All-Italy is
277/0at M3+ and central Italy is0/277; at M2.5+ central Italy is228/49but all-Italy remains277/0. - Added
./scripts/report-italy-data-coverage.sh. The latest report contains 4,836 INGV events, 283 Cumiana capture metadata records, 256 VLF anomaly windows, and only two weeks with both VLF and seismic observations. This is descriptive coverage evidence, not an association result. - Added
./scripts/analyze-italy-vlf-event-association.sh. The first refreshed permutation-controlled association remainsinsufficient_controls: three VLF-observed weeks provide only one M2.5+ event week and two controls.
Italy coverage diagnostics
- Run
./scripts/report-italy-data-coverage.shafter each refresh. It reports INGV event coverage, Cumiana capture coverage, label-free anomaly coverage, and descriptive weekly overlap. - Treat anomaly/event overlap as exploratory only until enough mature windows contain both positive and negative targets.
- Replace binary all-Italy targets with fixed spatial-cell targets or count regression; use central-Italy M2.5+ only as a temporary exploratory control.
The fixed-cell implementation is available in data/derived/multimodal/all_italy.spatial_vlf_image_windows.labeled.csv and is prepared by ./scripts/prepare-italy-spatial-model-inputs.sh. The current smoke artifact has 5,301 rows across 19 cells, with 812 positive, 4,451 negative, and 38 pending labels. This fixes target saturation but does not fix the short time coverage or establish predictive skill.
The first grouped-time logistic smoke baseline reached calibrated balanced accuracy 0.655320 for the all-feature ablation. The seismic-only and VLF-only ablations collapsed to balanced accuracy 0.5 under their calibrated thresholds. These figures are a single short-window diagnostic and are not evidence that either modality predicts earthquakes.
The first 19-cell leave-one-cell-out probe is stored under data/derived/models/all_italy_spatial_cell_holdouts_v2. Only 5 cells have positive test labels; the other 14 folds are one-class. The valid folds range from 0.146597 to 0.855263 calibrated balanced accuracy, with mean 0.333370 across all folds. This instability and class sparsity block meaningful spatial transfer evaluation.
The timestamp-permutation null control is stored under data/derived/models/all_italy_spatial_permutation_controls. It preserves each timestamp's complete spatial label pattern but shuffles those patterns across time. Five controls scored 0.643309--0.709108, mean 0.679362; all five matched or exceeded the real-order 0.655320. The current multimodal score therefore has no demonstrated temporal signal.