Analysis Report
Date: 2026-07-29
Progress So Far
The ideal outcome for ELFQuake would be to identify strong, consequential earthquakes in Italy before they happen, giving a useful estimate of when and where they will occur and how large they may be, while keeping false alarms low enough for the result to be actionable. That is the bottom-line standard against which the project should be judged.
Progress so far is mainly in building the machinery needed to test that goal. ELFQuake now has a working live-data pipeline for Italy, regularly collecting INGV earthquake events, Cumiana VLF spectrogram images, and available astronomy and space-weather measurements. It retains the original source files, aligns observations by time, trains baseline and Transformer models, tests the contribution of each modality, and can emit a demonstration seven-day event list and map. Historical INGV data provides seismic context, while the avalanche simulation supplies controlled synthetic seismic-like and piezo/VLF-like signals for model development. Self-supervised learning also lets the Transformer learn from live VLF observations before enough real earthquake labels have accumulated.
Compared with the ideal of predicting strong events, the present system is still at an early feasibility stage. It has not predicted a strong earthquake on held-out real data, and no multimodal model has yet beaten seismic-only and historical-rate baselines over a substantial real period. The overlapping live VLF and seismic record is short, and strong earthquakes are rare enough that a credible test will require much longer coverage. Current weekly forecasts demonstrate the intended software and output format only; they are not actionable predictions. The hypothesis remains open and testable, but useful strong-event prediction is not currently demonstrated or operationally viable.
Decision Summary
The table below assesses practical project utility, not predictive success. “Yes (probably)” means the work is already useful infrastructure or a reproducible measurement. “Maybe (further research needed)” means it is a plausible research direction without sufficient evidence. “No (probably)” means the current form should not be used as a claim, feature, or operational forecast.
| Experiment | Utility | Comment |
|---|---|---|
| Italy-scoped INGV acquisition and normalization | Yes (probably) | Reproducible seismic records with timestamps, provenance, and uncertainty fields. |
| Cumiana live VLF image capture and feature extraction | Yes (probably) | Provides a repeatable natural-radio observation stream, although it is image-derived rather than raw waveform data. |
| Astronomy and space-weather ingestion | Maybe (further research needed) | The interfaces and captures work, but there is no demonstrated incremental value. |
| Self-supervised learning on real VLF images | Maybe (further research needed) | Useful for representation learning while labels are sparse; reconstruction does not establish earthquake relevance. |
| Sandpile simulation as synthetic data generation | Yes (probably) | Useful for controlled software, sensor, target, and ablation experiments, subject to regime and realism limits. |
| Piezo-like signal as a VLF analogue | Maybe (further research needed) | Shape statistics can be tuned toward VLF, but no causal precursor value has been demonstrated. |
| Direct avalanche signal as a seismic analogue | Maybe (further research needed) | Event extraction and spatial rendering are useful, but synthetic events remain difficult to align with real seismic sparsity and timing. |
| One-to-one synthetic-to-INGV event matching | No (probably) | The two catalogs are generated by different processes; matching individual events would imply unsupported precision. |
| Multi-resolution count, energy, and spatial-occupancy alignment | Maybe (further research needed) | Retains more information than binary occurrence and is the recommended next alignment path. |
| Catalog-level synthetic calibration | Maybe (further research needed) | Training-only calibration of rate, magnitude, timing, and spatial density could make synthetic pretraining more representative. |
| Three-episode central-Italy calibrated profile | Maybe (further research needed) | Rate, magnitude, and median inter-event time are promising; spatial clustering remains too different. |
| Five-episode central-Italy calibrated profile | Maybe (further research needed) | Explicit episode duration brings the rate to 0.924x; magnitude and sample-matched spatial clustering are promising, but held-out validation is still required. |
| Matched current-profile three-episode baseline | Maybe (further research needed) | Regenerated seeds have stable heights and, with q=0.9996, produce rate 0.847x, magnitude W 0.085, and four-cell coverage; the sample remains small. |
| Training-only rate-matched precision calibration | Maybe (further research needed) | Transfer precision improves from 0.163 to 0.263 at recall 0.333; the historical-rate control reaches 0.269, so this reduces false positives without demonstrating incremental value. |
| 40,000-step central-Italy episode and extraction tuning | Maybe (further research needed) | The longer run is stable and its signal-shape score improves slightly, but the tuned catalog has only 10 events and remains poorly matched in magnitude and clustering. |
| Synthetic Transformer and sequence-model experiments | Maybe (further research needed) | Useful for interface and training checks; leave-one-episode-out stability is not yet adequate. |
| Synthetic-to-real transfer | No (probably) | Current transfer is below or comparable to historical-rate controls and is not predictive evidence. |
| Fixed spatial-cell target redesign | Yes (probably) | Removes all-Italy target saturation and creates a valid spatial evaluation contract. |
| Grouped-time spatial baseline | Maybe (further research needed) | It is a useful leakage-controlled diagnostic, but the 0.655320 score is not above its null controls. |
| Leave-one-cell-out evaluation | No (probably) | Currently blocked by only 5 of 19 cells containing positive labels. |
| Timestamp-permutation null control | Yes (probably) | It is an important reproducible falsification check; all five null runs matched or exceeded the real-order score. |
| Seven-day event-list forecasts and maps | Yes (probably) | Useful demonstration and integration outputs only; not actionable predictions. |
| Useful strong-earthquake prediction at present | No (probably) | No held-out real-data evidence supports this capability. |
The prospective-label pipeline only marks a target mature when the fetched INGV catalog covers its complete horizon. The 2026-07-29 refresh fetched 123 new INGV rows for 29 June--29 July and rebuilt 280 image-aligned rows. A new Cumiana image was captured at 2026-07-29T08:45Z; its seven-day target is correctly pending. Among the 279 mature rows, all-Italy remains 279/0 positive/negative and central Italy 1/278. These region-level tables are unsuitable for a credible supervised holdout.
Latest Italy coverage check
The long coverage report combines 4,836 historical normalized INGV events through 2026-07-07 with 284 Cumiana capture metadata records through 2026-07-29. A separate current-window audit, data/derived/reports/italy_data_coverage_20260729.json, uses the same 1 June--29 July event file as the association test: 280 INGV events, 257 self-supervised anomaly windows, and 4 weeks with both VLF and seismic observations out of 9 weeks. The newest anomaly score is 0.960976, and 41 windows exceed the exploratory 0.8 alert threshold. These are novelty measurements, not earthquake predictions. The new capture adds a fourth VLF-observed week to the descriptive anomaly table, although its target remains pending. The machine-readable outputs are data/derived/reports/italy_data_coverage.json, data/derived/reports/italy_data_coverage_20260729.json, and data/derived/models/self_supervised/real_vlf_anomaly_scores.csv.

The chart leaves gaps where no VLF windows were observed; it does not interpolate between the July 6, July 11, July 16, and July 29 captures.
A permutation-controlled association check gives the same conclusion. Using the refreshed 29 June--29 July event file, there are four VLF-observed weeks. At M2.5+ all four are event weeks; at M3+ there are two event weeks and two controls. The report still returns insufficient_controls rather than a p-value because at least three of each are required. The observed M3+ anomaly difference of -0.3563 is not evidence of a precursor. The refreshed diagnostic is data/derived/reports/italy_vlf_event_association_20260729.json.
Label-scope audit
The apparently mirrored all-Italy/central-Italy counts are not a row-generation error. Both tables are anchored on the same 279 Cumiana windows, so equal row counts are expected. Under the current M3 target, all-Italy is saturated (279/0 positive/negative), while central Italy has only 1/278 positive/negative. At M2.5+, central Italy becomes usable for a limited exploratory split (229/50 positive/negative), but all-Italy remains saturated. The all-Italy table must therefore move to fixed spatial-cell targets or event-count regression; it must not be used as a binary supervised table in its current form.
The fixed-cell replacement is now available. ./scripts/build-italy-spatial-vlf-targets.sh expands 279 VLF windows across 19 Italy cells to 5,301 rows: 812 positive, 4,451 negative, and 38 pending. ./scripts/prepare-italy-spatial-model-inputs.sh builds its tensor, aligned-window, and readiness artifacts. This corrects the target geometry, but does not create additional independent time periods, so it remains a smoke baseline rather than evidence of predictive skill.
As a separate coverage experiment, a central-Italy M2.5 table was built at data/derived/multimodal/central_italy.m2_5.prospective_vlf_image_windows.labeled.csv. It contains 229 positive and 50 negative rows with varying real Cumiana image features. The chronological 80/20 PyTorch tabular smoke run scored 0.990909 balanced accuracy, but its test period contains only one negative row and the majority-class baseline also scores 0.5 balanced accuracy. The result is therefore not useful evidence; independent VLF coverage across more target weeks is still required.
The leakage-controlled spatial smoke evaluation is ./scripts/evaluate-italy-spatial-baseline.sh; it keeps every cell from a timestamp together. Its first grouped-time run used 4,199 training rows and 1,064 test rows. The all-feature logistic ablation reached calibrated balanced accuracy 0.655320; seismic-only and VLF-only both reached 0.5. With only 279 capture times and a short overlap, this is an engineering baseline, not useful earthquake prediction.
The leave-one-cell-out control is ./scripts/evaluate-italy-spatial-cell-holdouts.sh. Across 19 cells, only 5 had positive labels; the remaining folds were one-class. Calibrated all-feature balanced accuracy ranged from 0.146597 to 0.855263 on the mixed folds and averaged 0.333370 across all folds. This shows that the earlier grouped-time score does not yet demonstrate spatial transfer; more observations are required.
The timestamp-permutation control is ./scripts/evaluate-italy-spatial-permutation-controls.sh. It preserves each timestamp's spatial event pattern and shuffles complete patterns across time. Five null runs scored 0.643309--0.709108, mean 0.679362, while the real-order run scored 0.655320. Since every null run was at least as strong, the real-order score provides no evidence of a temporal VLF or multimodal signal.
The first catalog-alignment implementation is now available through ./scripts/compare-italy-event-catalogs.sh and ./scripts/calibrate-italy-synthetic-magnitudes.sh. Against the current all-Italy INGV catalog, the real magnitude median is 2.30 while the three raw synthetic episodes are approximately 4.65--4.73; synthetic event rates are about twice the real daily rate, and each synthetic episode occupies 13 fixed cells versus 63 cells in the real catalog. This confirms that raw synthetic event marks are not yet aligned. The derived magnitude calibration is an observation-model diagnostic only and must be fitted on training dates before any transfer experiment.
The combined calibration is available through ./scripts/calibrate-italy-synthetic-catalog.sh. For seed 40, deterministic rate thinning reduced the rate ratio from 1.998 to 1.143; magnitude Wasserstein distance fell from 2.316 to 0.146, and the magnitude KS statistic fell from 0.999 to 0.281. The derived rows also carry per-event cell-rate importance weights, while coordinates remain unchanged. Spatial support remained incomplete, falling from 13 to 11 occupied cells versus 63 in the real catalog. Rate and magnitude alignment are therefore useful preprocessing checks, not evidence that the simulated process is physically aligned.
The scope-matched central-Italy comparison is data/derived/reports/central_italy_event_catalog_alignment.json. Raw seed 40 has a rate ratio of 16.242 against the central-Italy INGV catalog. Event-extraction tuning selected q=0.975, local-max window 60, and max_events=5; the resulting five-event catalog has a rate ratio of 0.709, but magnitude Wasserstein distance 2.536 and nearest-neighbour Wasserstein distance 77.318. The sample is too small for a reliable distribution comparison. Good synthetic-to-real event alignment has therefore not yet been reached; the next valid test is a multi-episode, scope-matched calibration with minimum sample thresholds.
The first multi-episode candidate is now available in data/derived/reports/central_italy_event_catalog_alignment_combined_spatial.json. Three independent episodes were offset by 21 days, rate-thinned, magnitude-calibrated, and then spatially transported as an explicit observation model. It reaches rate ratio 0.956, magnitude Wasserstein distance 0.082, and median inter-event time 23.6 hours versus 21.6 hours for real central Italy. Because the catalogs have different sizes, clustering is evaluated with repeated 32-event real subsamples: the untransported comparison is 21.570 km and the spatially transported comparison is 3.469 km. It covers 3 of 4 coarse cells. This is a substantial clustering improvement, but it still needs validation across more episodes and held-out spatial cells.
A 40,000-step episode at seed 4500 was then run to test whether duration alone improves the direct avalanche catalog. The simulation remained stable, with mean height near 8.17 and bounded maximum height near 114. The default extractor still returned only five events. A tuning grid selected q=0.975, a 60-step local-max window, and at most 10 events; its signal-shape score (0.177394) was slightly better than the earlier long-run profile, but the resulting catalog had rate ratio 0.606, magnitude Wasserstein distance 2.501, and sample-matched nearest-neighbour distance 91.279 km. The longer run therefore validates computational stability, not event alignment. More independent episodes or a better event-process mechanism are needed before using this profile for transfer.
A five-episode central-Italy catalog was then assembled from the default extraction outputs for seeds 40, 41, 42, 4300, and 4500, with episode clocks offset before calibration. An initial comparison incorrectly inferred duration from the first and last emitted events, producing an inflated rate ratio of 1.625. The calibration interface now accepts explicit coverage duration; using the known 60-second timestep and 111.78-day combined coverage gives 67 retained events, rate ratio 0.924, magnitude Wasserstein distance 0.046, inter-event Wasserstein distance 22.885 hours, sample-matched nearest-neighbour distance 1.600 km, and occupancy in 3 of 4 coarse cells. This is the strongest current catalog-level diagnostic, but it still requires held-out episode and cell validation.
The new per-episode diagnostic initially appeared to show distinct event-rate regimes: seeds 40--42 produce rate ratios of 14.435, 14.324, and 14.990, while seeds 4300 and 4500 produce 0.555 and 0.278. A configuration audit explains the difference. The older 40--42 runs have mean height near 146, no target refill, and approximately 64,000 units removed in the final step; the newer 4300 and 4500 profile has mean height near 8, continuous source-localized target refill, and approximately 24,000 units removed. These are different simulation profiles, not a clean seed-only comparison. Global thinning can make the combined rate look realistic, but it cannot replace rerunning a matched corpus. Reports are stored under data/derived/reports/central_italy_episode_alignment/.
An explicit episode-rate balancing diagnostic now deterministically thins the overactive episodes to approximately the real central-Italy rate while preserving underactive episodes unchanged. It produces 9, 9, 9, 5, 5 events for the five episodes. With a 14-day offset, followed by magnitude and spatial calibration, the aggregate reaches rate ratio 0.681, magnitude Wasserstein distance 0.089, inter-event distance 13.718 hours, and sample-matched clustering distance 3.067 km. The rate remains below one because underactive episodes cannot be filled without inventing events. This is suitable for a weighted or balanced training diagnostic, not evidence that the simulator itself is corrected.
The configuration-matched regeneration is the better correction. Three fresh 20,000-step episodes (4600, 4700, 4800) use the current source-localized refill profile and have final mean heights 8.35, 8.52, and 8.18. The default q=0.999/five-event cap is too sparse, while q=0.99 produces 154--162 events per episode. An intermediate q=0.9996, 240-step local-maximum profile produces 7--8 events per episode. The resulting 23-event catalog, with magnitude and spatial calibration, reaches rate ratio 0.847, magnitude Wasserstein distance 0.085, inter-event distance 10.415 hours, sample-matched clustering distance 7.637 km, and all 4 coarse central-Italy cells. This is the current matched synthetic baseline, not evidence of predictive value.
The matched corpus was also used in the transfer smoke suite at data/derived/models/central_italy_matched_3episode_transfer_suite.json. On the 26-week real holdout, the historical spatial-rate baseline reached balanced accuracy 0.934238 and precision 0.192308; real seismic random initialization reached 0.919624 and 0.163043; synthetic-pretrained transfer also reached 0.919624 and 0.163043. All models had 15 positive holdout samples and thresholds selected for near-perfect recall. These scores are therefore an imbalance/control diagnostic, not useful prediction; the historical baseline remains preferable in this trial.
Target calibration was extended with two training-only alternatives. A precision-oriented threshold maximizes training precision subject to a recall floor; a rate-matched threshold selects the score quantile implied by the training positive rate. On the matched holdout, the rate-matched historical baseline has precision 0.263 and recall 0.333, while synthetic transfer has precision 0.263 and recall 0.333 in the current run; the precision-floor variant gives transfer precision 0.220 and recall 0.867. The apparent precision improvement is therefore a useful false-positive control, not evidence that synthetic features add value. Outputs are in data/derived/models/central_italy_matched_3episode_target_calibration.json.
The calibration was then checked over four chronological rolling folds with the historical-rate control included. Rate-matched synthetic transfer precision averages 0.214 with recall 0.254, while the corresponding historical-rate control averages 0.228 precision. The final-holdout precision increase is therefore not robust across time. Keep rate-matched calibration available for honest low-volume forecasts, but do not claim a model improvement until it beats the historical control over rolling folds.
Per-episode checks under the same extractor give raw rate ratios of 0.888, 0.777, and 0.888 for seeds 4600, 4700, and 4800. Their untransported sample-matched clustering distances are 23.5, 38.5, and 42.9 km, so spatial clustering remains seed-sensitive even after configuration matching. The aggregate spatial result should therefore be used as a benchmark for further experiments, not as proof that the simulator has learned a realistic spatial process. Reports are under data/derived/reports/central_italy_matched_q09996_episode_alignment/.
The first reduced-source (SOURCE_COUNT=64) screening run briefly met a causal potential-lead rule across three episodes, but this did not survive a nine-episode confirmation across 40 events. It is not a precursor feature. Its initial empty chronological test tail was traced to a fixed 84-hour modeling window after a 50-hour simulation; episode batches now derive the end timestamp from the number of simulated steps. With duration-aligned windows, the nine-episode reduced-source profile (SOURCE_COUNT=64, refill 470, removal interval 20, q=0.998/window=120) has 387 labeled rows, a 0.470284 positive rate, temporal drift 0.181728, and a weak simple temporal score of 0.518750 balanced accuracy. It is a valid synthetic target baseline, not evidence of a piezo precursor.
Pre-relaxation spatial contact, coherence, and stress-weighted criticality were then added as observable grid diagnostics. They also failed a nine-episode causal confirmation across 39 events. The conclusion is now stronger: within the present instantaneous-relaxation sandpile, available state readouts are associated with the avalanche at the same step but do not demonstrate a stable earlier analogue of a piezo precursor.
An opt-in delayed-failure extension then changed the simulation dynamics rather than merely its sensors. Near-critical cells accumulate bounded damage before relaxation, local damage lowers the failure threshold, and toppling resets damage. On nine damage-enabled episodes, pre-relaxation damage_total passes the causal 5--15 step lead rule: AUC 0.652315, standardized change difference 0.484616, and 6 of 9 episode effects positive. The associated target table has 387 labeled rows (197 positive, 190 negative) and temporal drift 0.245581, just inside the current gate. The simple temporal event-list head is still weak at 0.415810 balanced accuracy, so this validates the synthetic precursor mechanism, not useful forecasting or a model improvement.
A matched Transformer ablation on those same nine damage trajectories confirms the distinction. With seed 42, 8 epochs, and nine leave-one-episode-out folds, excluding damage_total, damage_max, and the active-cell count averages 0.599648 balanced accuracy; including them averages 0.586848. The causal feature therefore does not yet improve the current model architecture. It remains a synthetic diagnostic while parameter schedules or a short-horizon damage-specific head are evaluated.
A second matched experiment used 5-minute samples and a 15-minute future-event target, matching the confirmed damage lead. Its 5,373 labels contain 135 positives. With seed 42, four epochs, and nine leave-one-episode-out folds, the generic piezo Transformer averages 0.530826 without damage channels and 0.529700 with them. Shortening the target alone therefore does not make the architecture exploit damage; the next experiment needs an imbalance-aware, dedicated damage head.
Dedicated damage-only Transformer branches were also tested on the same target and folds. A 12-minute lookback averages 0.500824; expanding to 24 minutes gives 0.504695. An explicit, class-weighted logistic damage head using only current/past pre-relaxation level, one-step rise, short/long means, contrast, and range averages 0.506188 over nine unseen episodes (5 of 9 folds meet both 0.40 recall floors). Extending its causal history from 30 to 60 minutes improves the mean to 0.528976, but only 4 of 9 folds meet both floors and it remains below the 0.530826 no-damage control. The causal damage effect is therefore not yet sufficiently stable between trajectories for short-horizon event prediction. Keep it as a synthetic diagnostic, and improve the delayed-failure regime before adding it to the default model.
A controlled persistence screen then held all delayed-failure settings fixed except damage decay. The faster-decaying short-memory profile (0.970) produced a nine-episode causal damage_total lead at 15--30 minutes across 41 events and had near-zero temporal label drift (-0.007686). However, its 60-minute engineered head regressed to 0.400773 mean balanced accuracy, with only 2 of 9 folds meeting both recall floors. A slower-decaying long-memory profile (0.995) did not pass its initial three-episode causal screen. Neither profile improves the existing 0.985-decay control; the apparent precursor lead remains insufficiently transferable for prediction.
A reset-fraction screen reinforced the need for larger causal confirmation. Both residual damage (reset=0.75) and rapid reset (0.98) appeared to support a 15--30 minute lead in three episodes. Residual damage was expanded to nine episodes: its target timing remained stable (train/test positive-rate delta 0.005106), but the lead disappeared across 43 events. It was therefore stopped before model comparison. Three-episode causal screens are too noisy for profile selection; do not use them as a promotion gate.
Fixed local damage readouts were then added around each simulated piezo receiver: local mean, maximum, active fraction, and variation inside the receiver footprint. A causal top-three-receiver aggregation on nine fresh control-dynamics episodes and 43 events did not support a lead for any field. This rules out simple spatial dilution of damage_total as the explanation for the weak model results. The next simulation change should be a two-stage damage mechanism in which accumulated microdamage must mature over time before it alters failure behaviour.
That two-stage mechanism was implemented and tested with nine fresh episodes. Microdamage had to remain above 0.50 for five steps before mature weakness accumulated; only mature weakness lowered the local failure threshold. The extracted target table remained well aligned (5,373 rows, 130 positives, train/test positive-rate delta -0.001174), but neither microdamage nor mature weakness passed the causal lead rule across 44 events. No model was trained. The mechanism therefore does not yet produce a transferable precursor under its predeclared profile.
The first synthetic-to-real spatial transfer trial now provides a stricter real-data checkpoint. It pretrained a small CPU model on three synthetic avalanche event catalogs, then fine-tuned chronologically on 80% of the 2024--2026 INGV catalog and held out the final 20%. Because Italy as a whole has an M2.5+ event in nearly every week, the target is a fixed 1.5-degree Italy cell containing an M2.5+ event in the following week. On the holdout, the model reached 0.662743 balanced accuracy and 0.266667 precision, below a training-only historical spatial-rate baseline (0.686013 balanced accuracy, 0.343915 precision). It therefore has no demonstrated useful predictive value, but the end-to-end transfer, chronological evaluation, and independently located actual-versus-predicted map are now reproducible. VLF and astronomy are explicit missing-modality inputs in this run because their validated historical overlap is not yet long enough.
Scope
The matched transfer experiment suite now compares a historical spatial-rate baseline, a real seismic-only randomly initialized MLP, and synthetic-pretrained real fine-tuning under the same chronological split. The expanded corpus contains 79,976 dense synthetic records from four long episodes and 190 weekly spatial samples; independent simulation episodes are offset by 21 synthetic days because they share a demonstration start timestamp. On the final real holdout, the rate baseline reaches 0.686013 balanced accuracy and 0.343915 precision, random-init seismic reaches 0.666349 and 0.269113, and synthetic transfer reaches 0.680722 and 0.307054. Four rolling-origin transfer folds average 0.668175 balanced accuracy, with a worst fold of 0.656225. A train-only sweep selects M2.5 with 1.0-degree cells; its final holdout reaches 0.710283 balanced accuracy but only 0.210784 precision. These results establish useful controls and sensitivity checks, not predictive skill.
This report summarizes the current statistical comparison between real Italy-scoped seismic/VLF data and signals derived from the avalanche simulation, plus the current model-interface and PyTorch smoke results.
The comparison is diagnostic only. It does not demonstrate earthquake prediction capability.
Inputs
Real seismic:
data/derived/ingv/events_italy_2026-06-01_2026-07-08.combined.normalized.csv- 176 normalized all-Italy INGV event rows
data/derived/ingv/events_central_italy_2026-06-01_2026-07-08.combined.normalized.csv- 22 normalized central-Italy INGV event rows
data/derived/ingv/events_italy_all_available.combined.normalized.csv- 4836 normalized all-Italy rows, from
2024-01-01T21:38:30.320000Zto2026-07-07T08:51:55.110000Z data/derived/ingv/events_central_italy_all_available.combined.normalized.csv- 594 normalized central-Italy rows, from
2024-01-03T07:43:40.720000Zto2026-07-07T08:51:55.110000Z
Real VLF:
- 279 Cumiana
last_E_VLFspectrogram image-feature rows in the current refresh data/derived/multimodal/cumiana_last_E_VLF.image_features.csv- image sequence manifest:
data/derived/models/cumiana_vlf_image_sequence/manifest.json
Synthetic:
- current model training rows:
data/derived/models/mountain_256x256_seeds40-42_20000.aligned_hourly_synthetic_windows.csv - 501 combined relabeled hourly synthetic rows from seeds
40,41, and42 - source simulation outputs include
*.piezo.csv,*.avalanche_signal.csv, and*.avalanche_events.csv
Real prospective model rows:
data/derived/models/all_italy.real_vlf_aligned_windows.csvdata/derived/models/central_italy.real_vlf_aligned_windows.csv- each region-level table has 279 rows, 277 labeled rows, and 2 future pending rows
- all-Italy labels are currently 277 positive / 0 negative
- central-Italy labels are currently 0 positive / 277 negative
- both are
insufficient_class_variation
The separate fixed-cell table is the current real-data baseline input:
data/derived/multimodal/all_italy.spatial_vlf_image_windows.labeled.csv- 5,301 rows across 19 cells and 279 VLF timestamps
- 812 positive, 4,451 negative, and 38 pending labels
- the 19 spatial rows per timestamp are repeated observations of one VLF window, not 19 independent time samples
Method
Seismic event lists were converted into hourly binned energy traces using magnitude-derived energy proxies.
VLF spectrograms were reduced to column-intensity traces after cropping the plotted signal region.
Synthetic piezo and direct avalanche signals were reduced to per-step summed sensor traces.
For each trace, the report computes time-domain shape metrics and frequency-domain PSD metrics. Pairwise normalized_distance uses dimensionless shape and PSD-ratio metrics; amplitude, sample count, and duration are reported as deltas but are not included in that distance.
In model reports, ablation means removing one feature group or modality, such as VLF or astronomy, and comparing performance with and without it.
Output files:
data/derived/sim/mountain_256x256_seed42_10000.signal_shape_series.csvdata/derived/sim/mountain_256x256_seed42_10000.signal_shape_pairs.csvdata/derived/sim/mountain_256x256_seed42_10000.central_italy_signal_shape_series.csvdata/derived/sim/mountain_256x256_seed42_10000.central_italy_signal_shape_pairs.csvdata/derived/sim/mountain_256x256_seed42_10000.all_italy_signal_shape_series.csvdata/derived/sim/mountain_256x256_seed42_10000.all_italy_signal_shape_pairs.csvdata/derived/sim/mountain_256x256_seed42_10000.piezo_vlf_comparison.csvdata/derived/sim/mountain_256x256_seed42_10000.piezo_sensor_scan.csvdata/derived/models/model_family_comparison.jsondata/derived/models/sequence_full_regime/sequence_full_model_run_summary.jsondata/derived/models/sequence_full_regime/post_burn_in_temporal_split_diagnostics.jsondata/derived/models/sequence_full_balanced/sequence_full_balanced_model_run_summary.jsondata/derived/models/sequence_full_balanced/regime_balanced_split.jsondata/derived/models/tiny_patch_transformer/tiny_patch_transformer_model_run_summary.jsondata/derived/models/real_synthetic_compact_comparison.jsondata/derived/models/deep_patch_transformer/deep_patch_transformer_synthetic.jsondata/derived/models/deep_patch_transformer/all_italy.real_finetune.jsondata/derived/models/deep_patch_transformer/central_italy.real_finetune.jsondata/derived/models/self_supervised/real_vlf_image_autoencoder.jsondata/derived/models/self_supervised/real_vlf_image_embeddings.csvdata/derived/models/self_supervised/real_vlf_vs_synthetic_piezo_embedding_domain.jsondata/derived/models/self_supervised/real_vlf_vs_synthetic_piezo_embeddings.csvdata/derived/models/trial_forecast/mag_gt2_weekly_trial_forecast.jsondata/derived/models/trial_forecast/mag_gt2_weekly_trial_events.csvdata/derived/models/learned_forecast/mag_gt2_weekly_learned_forecast.jsondata/derived/models/learned_forecast/mag_gt2_weekly_learned_events.csvdata/derived/models/forecast_comparison/trial_vs_learned_weekly_forecast.jsondata/derived/models/forecast_comparison/trial_vs_learned_weekly_forecast.csvdata/derived/models/synthetic_event_list_patch_transformer/h6_patch_transformer.jsondata/derived/models/synthetic_event_list_patch_transformer_sweep/summary.json
Current Transformer Tuning
The current supervised Transformer work is still synthetic-only because real VLF-aligned labels remain one-class. A new target adapter maps the richer synthetic event-list labels into the standard Transformer input contract, preserving an 80/20 time-ordered split within each synthetic episode. On the warmed nine-episode h6 event-list target table, a short four-configuration CPU sweep found the best row at lookback 12, patch 3, dropout 0.1: the piezo/VLF-only ablation reached calibrated balanced accuracy 0.608629, while the full direct-avalanche plus piezo/VLF plus summary ablation reached 0.463892. This suggests the VLF-like precursor channel is the most useful synthetic signal in this setup, but the result is not yet stable enough for forecast promotion.
The current transfer baseline also has a selectable multiscale feature path. It adds causal 1--90 day local and neighbouring-cell event counts, maximum magnitudes, log energy, and recency to the compact seismic features. On the matched three-episode central-Italy trial, the multiscale MLP transfer fell to balanced accuracy 0.907098 and default precision 0.144231, versus 0.919624 and 0.163043 for the compact baseline. Its rate-matched precision was 0.208333, so it is useful as a diagnostic but is not an improvement.
A first domain-robust feature experiment adds FEATURE_MODE=relative, expressing local counts, magnitude, and energy against the preceding Italy-wide baseline. On the same matched corpus, its holdout balanced accuracy was 0.907098; rolling-origin mean balanced accuracy was 0.907935, with worst fold 0.895754. This was below the multiscale rolling mean 0.912232 and did not improve the compact transfer baseline. The implementation is retained for controlled comparisons, not promoted as the default.
A bounded CPU Transformer sweep over lookback {8,12}, patch size {2,3}, and dropout 0.1 was run on the synthetic h6 event-list interface. The longer five-seed, 36-epoch run is stored at data/derived/models/central_italy_transformer_sweep_long/summary.json. Across 20 runs, the piezo/VLF-only branch averaged calibrated balanced accuracy 0.5620, compared with 0.4770 for direct avalanche-only, 0.5472 for direct plus piezo/VLF, and 0.5340 for the full branch. The strongest configuration was lookback 8, patch 3, with piezo/VLF-only mean 0.5912 across seeds. The preference for piezo/VLF survives the longer run, but seed ranges remain wide and the split is still synthetic and fixed; this is not real-data evidence.
The next episode-held-out trial uses each of the nine synthetic episodes as a test group in turn, with 352 training rows and 44 held-out rows per fold. The piezo/VLF-only Transformer averaged calibrated balanced accuracy 0.5508, ranging from 0.3788 to 0.7009. The same model used auxiliary count and log magnitude-energy heads; their mean held-out MAEs were 0.4234 and 2.7124. This is a useful end-to-end multi-task check, but the wide fold range shows that regime generalization remains the main limitation.
An identical occurrence-only control averaged 0.5533, ranging from 0.4091 to 0.6830. The auxiliary heads therefore did not improve occurrence performance in this trial; retain them as optional targets for representation diagnostics, not as default loss terms.
A first domain-randomized batch used 12 complete episodes across three explicit simulator regimes: baseline, slow-fill, and fast-localized loading. The combined six-hour target table had 516 labeled rows and a 34.5% positive rate. Leave-one-episode-out piezo/VLF-only evaluation averaged calibrated balanced accuracy 0.5096; profile means were 0.5042, 0.4881, and 0.5364 respectively. This is close to chance and below the non-randomized episode result, so the model is not yet learning regime-invariant features. The batch is nevertheless valuable as a strict stress test.
An opt-in per-window sequence normalization control was then run on the same 12 folds. It reduced mean calibrated balanced accuracy to 0.4909, with profile means 0.4798, 0.4909, and 0.5021. This rules out simple lookback-level scaling as the immediate fix; the next work should examine target construction, event extraction, and physically meaningful regime-relative features before more model tuning.
* data/derived/models/missing_modality/missing_modality_seed42_summary.json
* data/derived/models/sequence_modality_diagnostic.json
* data/derived/models/all_italy.ingv_backfill_seismic_windows.temporal_holdout.json
* data/derived/models/central_italy.ingv_backfill_seismic_windows.temporal_holdout.json
Current Status
Real-data status:
- INGV prospective refresh now rebuilds stable current catalogs and labels only horizons fully covered by the fetched catalog; historical backfill currently runs through
2026-07-07T00:00:00Z. - Cumiana VLF image capture and image-feature extraction are working.
- Real VLF-aligned model tables are scaffolded for all-Italy and central Italy.
- Real PyTorch training should not start yet, because each real table has only one target class. The real deep patch Transformer wrapper records this blocker instead of training.
- Current region-level real VLF-aligned label counts are 277 positive / 0 negative for all-Italy and 0 positive / 277 negative for central Italy; each scope has 2 future pending rows.
- Self-supervised real VLF pretraining is available and is now the default model-development path until supervised labels have both classes.
- Historical seismic-only backfill for
2024-01-01to2026-07-07produced 130 weekly training windows per scope. - Backfilled all-Italy seismic windows are ready but heavily positive-skewed: train 95 positive / 9 negative, test 25 positive / 1 negative.
- Backfilled central-Italy seismic windows are more balanced: train 13 positive / 91 negative, test 6 positive / 20 negative.
- Current seismic-only temporal smoke scores are weak: all-Italy calibrated balanced accuracy
0.740000on a one-negative test set; central Italy calibrated balanced accuracy0.441667.
Synthetic-model status:
- The current relabeled synthetic model table has 501 rows, with 67 positives and 434 negatives.
- CPU PyTorch tabular and GRU sequence models are implemented and compared.
- Current corrected-label temporal smoke scores remain weak: tabular PyTorch calibrated balanced accuracy
0.507576; sequence GRU calibrated balanced accuracy0.500000. - Current corrected-label seed holdouts are stronger but synthetic-only: tabular PyTorch ranges from
0.732602to0.784169, while sequence GRU ranges from0.720690to0.821105. - Older sequence sweeps, matched comparisons, repeated training-seed runs, and tiny patch Transformer checks were produced before the target relabeling and should be rerun before model selection.
- Corrected-label temporal split diagnostics still show a label/regime shift:
gt0train positive rate0.080000, test positive rate0.346535; largest drift features are direct-avalanche active-topple and summary-topple aggregates. - A post-burn-in regime-balanced explicit split has matched train/test class rates and gives
sequence_fullcalibrated balanced accuracy0.650000; use this as an engineering diagnostic, not as forecasting evidence. - The selected deeper patch Transformer pretrain now writes
deep_patch_transformer_synthetic.pt; its latest synthetic calibrated scores are0.737879for piezo/VLF-only and0.583333for full sequence. - The first self-supervised real VLF autoencoder smoke used 247 Cumiana VLF rows and 224 windows. Test masked reconstruction MSE was
0.835488, better than the zero baseline1.074356. - A shared patch-Transformer self-supervision evaluation now compares five initialization regimes over seeds
7,42, and99. Synthetic pretraining raises mean held-out full-model balanced accuracy from random initialization0.451754to0.488678; balanced joint real/synthetic pretraining reaches0.487352, and real-only pretraining reaches0.470726. All remain below the0.60synthetic gate. - Synthetic masked reconstruction beats zero and last-patch baselines for all pretrained seeds. Real VLF reconstruction beats zero but not last-patch persistence, reflecting the small 186-window real training set and strong short-term persistence.
- Sequential synthetic-then-real pretraining has the strongest mean frozen probe (
0.519176) but weak fine-tuning (0.452264), so the next transfer check should use joint rehearsal or freeze the shared encoder during the real-only stage. - Positive/negative recall remains seed-sensitive; most runs fail to keep both above
0.40. The initialization gains are therefore useful engineering evidence but not a promotion result. - Transformer components now use stable name-derived initialization. Adding an unused modality adapter cannot alter shared parameters or advance the caller's PyTorch random stream, making architecture comparisons genuinely matched.
- In the controlled transfer rerun, random-init piezo/VLF-only is strongest at mean balanced accuracy
0.619033over three seeds, with a0.577723to0.663709range and both recalls above0.40in every seed. Synthetic pretraining falls to0.534272; current self-supervision does not improve this synthetic downstream target. - Controlled late gated fusion also does not improve the model. Random anchored full and direct-only variants average
0.609241and0.606385, below the0.619033anchor. Disabling direct avalanche raises the direct-only mean to0.611281, so that branch has not demonstrated incremental value. - A stricter leave-one-episode-out run holds out each of nine simulation episodes across three model seeds. Mean balanced accuracy drops to
0.578712, ranges from0.275641to0.758730, and only 14 of 27 folds keep both recalls above0.40. This fails the synthetic stability gate and shows that the promising within-episode score does not yet generalize reliably across simulation episodes. - Averaging the three predeclared model seeds raises unseen-episode mean balanced accuracy to
0.632634, but only 6 of 9 episodes retain both recall floors. Training-only threshold calibration outperforms a fixed0.5threshold (0.601361) without fixing the three unstable episodes, so the remaining failure is mainly episode generalization rather than threshold choice. - Episode diagnostics show that piezo feature effects change sign across trajectories. Shorter h3 labels (
0.580991ensemble mean), longer 60-minute h6 context (0.512103), and simple spatial sensor aggregates (0.544096) all regress. The h6/12-minute mean-only representation remains best, but zero of four variants passes the combined mean and recall-stability gate. - A second label-free real VLF anomaly layer now scores descriptor reconstruction and embedding novelty by window. The latest smoke artifact covers the current Cumiana sequence through
2026-07-16, with demo probability0.928922; this is not trained on earthquake labels and is not a validated forecast. - The tuned shape-profile synthetic-to-real embedding-domain diagnostic encoded 59,931 synthetic piezo/VLF windows through a descriptor autoencoder trained on real VLF windows. Synthetic centroid distance was
1.291640and synthetic-to-real nearest mean distance was1.846295. - The same diagnostic is still only a baseline, but it is stronger than the previous full-descriptor version: held-out real masked reconstruction MSE is
0.895188versus a zero baseline of0.960585; synthetic masked reconstruction is4.642818versus a zero baseline of4.721987. - A real-like synthetic inlier subset now marks the closest 25% of synthetic descriptor windows. It keeps 14,983 synthetic windows and reduces synthetic-to-real nearest mean distance to
1.162097, with scale mean absolute delta0.057490. - A synthetic-inlier transfer diagnostic now trains the masked descriptor autoencoder only on those 14,983 synthetic windows and evaluates on held-out real VLF descriptors. Held-out real masked reconstruction MSE is
0.688280versus a zero baseline of0.759011, but the transfer embedding centroid distance remains high at4.281796. - A mixed-domain alignment diagnostic now trains on real VLF plus 14,983 locally selected synthetic piezo/VLF windows with a CORAL embedding-alignment penalty. Held-out real masked reconstruction improves to
0.294475versus a zero baseline of0.588513, and held-out embedding centroid distance improves to1.033580. - Mixed-domain controls remain important: centroid-inlier selection scored
1.011474, random synthetic selection1.142438, and capped full-synthetic selection1.617345on held-out centroid distance. This means alignment training is useful, but the local inlier criterion is not yet clearly superior to centroid selection. - A first end-to-end trial weekly event-list forecast now combines current INGV history, VLF context, astronomy captures, and synthetic avalanche event artifacts. The
2026-07-08T00:00:00Zrun for the following week produced 25 capped>M2coordinate rows, with an uncapped expected count proxy of33.411669; this is a contract smoke test and not a validated prediction. - A first synthetic-trained learned weekly event-list forecast now preserves the same CSV contract. It trained a small logistic scorer on 501 aligned synthetic rows with 67 positives, produced a latest synthetic score of
0.899322, and emitted 25 capped coordinate rows with an uncapped expected count proxy of32.709463. - The learned scorer is only a scaffold: its held-out synthetic balanced accuracy is
0.507576, with 35/101 positives in the test split and very poor negative recall (0.015152). - A forecast comparison report now makes success criteria explicit. Stage 1 event-contract output passes; Stage 2 synthetic utility fails because balanced accuracy is below
0.60and negative recall is below0.40. - The learned forecast currently changes count/probability metadata but not event placement; candidate-to-baseline nearest distance is
0.0 km, so a model-informed spatial adapter remains a necessary next step. - Synthetic event-list target generation now derives forecast-shaped targets directly from avalanche event CSVs: future event count, occurrence, max/mean magnitude, centroid, first-event time, and time-to-first-event. The current h6 target table has 483 labeled rows with 213 positives and a
0.440994positive rate, making it the healthiest current horizon. - A dependency-light synthetic event-list model now trains separate heads for occurrence, count, max magnitude, and centroid. On the h6 temporal split it still fails occurrence utility (
0.500000balanced accuracy,0.000000negative recall), but on the h6 balanced engineering split it reaches0.887566balanced accuracy, count MAE0.506783, max-magnitude MAE0.009317, and centroid median error145.585806 km. - Current event-list model artifacts are
data/derived/models/synthetic_event_list_model/h6_event_list_model.json,h6_event_list_predictions.csv,h6_balanced_event_list_model.json, andh6_balanced_event_list_predictions.csv. Treat the balanced result as evidence that the synthetic target shape is learnable, not as forecasting validation. - Drift-aware synthetic event-list validation now writes
data/derived/models/synthetic_event_list_drift/h6_drift.jsonandh6_drift_buckets.csv. The temporal split confirms the failure mode: train positive rate0.318653, test positive rate0.927835, positive-rate delta0.609182, warninglarge_positive_rate_shift. - The largest feature shifts in the failing temporal split are direct-avalanche active-topple counts, summary topple counts, summary max height, and piezo near-critical cell counts. This points to a simulation lifecycle/regime issue rather than a purely model-capacity issue.
- Episode annotations now write
mountain_256x256_seeds40-42_20000.synthetic_event_list_targets_h6.episodes.csvwith 21 episode blocks. Episode-balanced validation remains learnable with balanced accuracy0.878079, positive recall0.928571, negative recall0.827586, count MAE0.497667, and centroid median error162.751648 km. - The first stationarity episode batch,
mountain_256x256_seeds4000-4201_5000, generated six 5000-step episodes. It reduced overall h6 positive rate to0.188312, but made chronological drift worse: train positive rate0.056911, test positive rate0.709677, positive-rate delta0.652766. The run still accumulated mass strongly, ending near mean height111, so the loading/removal balance remained wrong. sim.shnow exposesDEPOSITION_PROBABILITYandWARMUP_STEPS.run-synthetic-episode-batch.shnow defaults to a warmed aggressive stationarity profile: target fill limitWIDTH * HEIGHT / 128, bottom-layer removal every25steps, deposition probability0.45, and3000unrecorded warm-up steps.- The aggressive stationarity probe,
mountain_256x256_seeds4000-4200_3000, generated three 3000-step episodes. It kept final mean height near3.2and fixed the h6 drift warning: overall positive rate0.318182, train positive rate0.304762, test positive rate0.370370, positive-rate delta0.065608, warningok. - The aggressive probe is too small for model conclusions. Its temporal occurrence balanced accuracy is
0.505882; episode-balanced split reaches0.628571, but count and centroid heads remain weak. Use it as evidence that the simulation drift can be controlled, then scale the profile to more episodes and tune event extraction density. - Structured initial fill is now implemented via
INITIAL_FILL_MODE=structured,INITIAL_FILL_MEAN_HEIGHT,INITIAL_FILL_VARIATION, andINITIAL_FILL_SMOOTH_PASSES. The first structured-fill probe,mountain_256x256_seeds6000-6200_3000, started near the loaded regime and ended near mean height3.2, but did not improve drift versus the no-prefill aggressive probe: overall positive rate0.310606, train positive rate0.247619, test positive rate0.555556, positive-rate delta0.307937, warninglarge_positive_rate_shift. - Unrecorded warm-up is more promising than initial fill. The warm-up probe,
mountain_256x256_seeds7000-7200_3000, used1000unrecorded steps before each recorded episode. It produced overall positive rate0.356061, train positive rate0.352381, test positive rate0.370370, positive-rate delta0.017989, warningok. Treat this as the current best small stationarity profile, but still too small for model-selection claims. - Scaling the
1000-step warm-up profile to nine episodes,mountain_256x256_seeds8000-8202_3000, did not preserve the small-probe drift result. It produced 396 labeled rows, overall positive rate0.277778, train positive rate0.218354, test positive rate0.512500, positive-rate delta0.294146, warninglarge_positive_rate_shift. Mean height stayed bounded, so the remaining failure is early-vs-late target drift rather than overflow. - A
3000-step warm-up probe,mountain_256x256_seeds9000-9200_3000, started closer to the late-run regime and restored h6 drift control: 132 labeled rows, overall positive rate0.257576, train positive rate0.247619, test positive rate0.296296, positive-rate delta0.048677, warningok. The temporal event-list smoke model reached balanced accuracy0.674342, positive recall0.875000, negative recall0.473684, and count MAE0.330187; still treat this as a small engineering check. - Scaling the
3000-step warm-up profile to nine episodes,mountain_256x256_seeds10000-10202_3000, preserved acceptable h6 drift: 396 labeled rows, overall positive rate0.325758, train positive rate0.287975, test positive rate0.475000, positive-rate delta0.187025, warningok. The temporal event-list model remained weak (0.468045balanced accuracy), but balanced and episode-balanced checks reached0.609579and0.590008, respectively. - Direct avalanche extraction density was tested on the same nine warmed episodes without rerunning simulation. The dense
_q099_w60_m10profile produced 10 events per episode and stayed drift-ok, but saturated h6 targets: positive rate0.828283, temporal balanced accuracy0.530333, and poor negative recall. The midpoint_q0995_w120_m5profile produced 5 events per episode and a healthier positive rate0.512626, but model checks were weaker than the sparse default: temporal0.462475, balanced0.493902, episode-balanced0.554605. - Interpretation: keep sparse event extraction as the default for now. It is not dense enough for final synthetic training, but the denser profiles did not improve model learnability; the next improvement should come from better event-list heads, loss weighting, or richer target construction rather than simply lowering the extraction threshold.
- The dependency-light event-list occurrence head now supports deterministic feature-bag ensembles.
train-synthetic-event-list-model.shdefaults to 8 ensemble members with 50% feature bags. On the scaled sparse nine-episode table this improved the temporal split from0.468045to0.498120, the balanced split from0.609579to0.616473, and the episode-balanced split from0.590008to0.607551. The gain is modest but consistent; temporal utility remains below the0.60synthetic gate. - Synthetic event-list targets are now richer: they include event-rate, log magnitude energy, early/middle/final horizon event counts, peak timing, event duration, and spatial spread in addition to occurrence, count, magnitude, centroid, and first-event timing. Matching regression heads now write
predicted_*fields and event-shape metrics into the model reports. - On the scaled sparse nine-episode balanced split, the richer shape heads give event-rate MAE
0.068936, early/middle/final count MAE0.180167/0.246829/0.194170, spatial-spread MAE20.317859 km, and time-to-peak MAE6769.366174 s. These are useful shape checks, but they do not by themselves solve chronological generalization. - An optional boosted-stump occurrence head was added for nonlinear dependency-light checks. It improved the balanced engineering split to
0.629173balanced accuracy but failed the chronological split at0.347744, so the feature-bag logistic ensemble remains the default. - A synthetic event-list probe harness now compares horizon, burn-in, feature-cap, ensemble-size, boosted-stump, and balanced-control variants, then writes compact JSON/CSV summaries. A constrained h3 smoke run passed end to end; at 20 epochs, the 16-member feature-bag temporal variant reached
0.557692balanced accuracy, while boosted stumps remained weaker at0.371154. Treat this as a harness check, not a model-selection result. - The full event-list probe wrote
data/derived/models/synthetic_event_list_probes/summary.csvwith 54 model/drift reports. Among drift-ok chronological runs, the best result was h6 with a 32-feature cap: balanced accuracy0.511278, positive recall0.736842, negative recall0.285714, count MAE0.520321, and centroid median error640.573894 km. The nominal best chronological score was h12 with 20% burn-in and 64 features at0.579268, but that target hadlarge_positive_rate_shift(0.296459) and should not be promoted. - Balanced controls remain much stronger than chronological checks. The best drift-ok balanced control was h12 boosted stumps with balanced accuracy
0.651543; h12 episode-balanced reached0.626879, h6 boosted balanced reached0.629173, and h6 episode-balanced reached0.607551. This confirms target learnability under controlled splits but not chronological utility. - Interpretation: simple horizon changes, burn-in trimming, tabular feature caps, larger feature-bag ensembles, and boosted stumps did not solve temporal generalization. The next modeling improvement should add explicit temporal context or sequence/event-process heads rather than more one-row tabular occurrence heads.
- Lagged-context h6 probes add previous-row synthetic feature history while excluding target and diagnostic fields. This is the first explicit temporal-context check. The 128-feature cap improved drift-ok chronological balanced accuracy to
0.545739; the 256-feature cap improved it further to0.587093, with positive recall0.578947and negative recall0.595238. The 64-feature cap regressed to0.439223, adding lag 12 scored0.535088, and all features regressed to0.530702. - Interpretation: lagged context is the first change that materially narrows the temporal gap, but it is still just below the
0.60gate. The next implementation should use an explicit sequence/event-process head with regularization rather than simply widening tabular lag features. - A regularized CPU PyTorch event-list sequence head now reads synthetic event-list target tables directly, builds grouped lookback sequences, excludes target/diagnostic fields, and uses class weighting, dropout, AdamW weight decay, and gradient clipping. On h6 drift-ok chronological validation, the default lookback-12 seed-42 run reaches calibrated balanced accuracy
0.609649, positive recall0.552632, and negative recall0.666667, passing the0.60synthetic gate for the first time on this target family. - Sequence-head controls show the result is promising but not yet stable. Lookback 8 seed-42 scored
0.603383; lookback 4 scored0.467419; hidden size 32 scored0.578321; lookback 12 with seed 7 scored0.528822; lookback 12 with seed 99 scored0.584586; lookback 12 with dropout0.30scored0.569549. Treat the gate pass as a directionally useful engineering result, not a promoted forecast adapter. - A first sequence-head stability sweep wrote
data/derived/models/synthetic_event_list_sequence_sweep/summary.csvacross three seeds, lookbacks8and12, and dropout0.1/0.2/0.3. The best mean configuration is now lookback12, dropout0.1: mean balanced accuracy0.600459, min0.576441, max0.645363, stddev0.031777, with one of three seeds passing the0.60gate. The default sequence-head script now uses this setting. - Interpretation: the sequence head is the first event-list model family with mean performance at the synthetic gate, but it is still seed-sensitive. The next improvement should reduce variance, for example by model ensembling, longer training controls, or stronger regularization/early stopping.
- Sequence-head probability ensembling is now implemented. The three-seed lookback-12/dropout-0.1 ensemble scored
0.591479, below the gate. Pairwise ensembles show selection sensitivity: seed42+99scored0.644110, seed7+42scored0.604637, and seed7+99scored0.555764. This means ensembling can improve a chosen pair, but naive all-seed averaging does not yet solve stability. - Validation-selected sequence heads were tested with a 20% chronological validation slice inside the training period. This underperformed the train-calibrated baseline: lookback-12/dropout-0.1 mean balanced accuracy fell to
0.539265with no seed passing the gate, and dropout0.2fell to0.519632. - Early-stopped sequence heads were tested on the same validation slice with patience
10and train-calibrated thresholds. This also underperformed and increased variance: dropout0.1mean balanced accuracy was0.494570, min0.335213, max0.615288; dropout0.2mean was0.471178. Current validation slices appear too small or distribution-shifted to guide early stopping safely. - Interpretation: do not use validation-selected thresholding or early stopping as the next default. Keep the full-epoch lookback-12/dropout-0.1 sequence head as the current occurrence candidate, and pursue variance reduction through larger episode batches, more stable validation design, or representation sharing with count/location heads.
- Interpretation: initial fill can skip cold-start height growth, but with immediate bottom-layer removal and the current event-extraction settings it does not by itself produce a more stationary event process. Prefer warmed episodes over structured initial fill unless a later delayed-removal experiment beats the warm-up profile.
- A short piezo/VLF transform sweep added deterministic high-pass, burst, near-threshold, release-mix, and sensor-gain variants. The best transformed variant,
gain_burst, improved short-run held-out embedding centroid distance to1.757251versus1.841903for the current signal, but worsened held-out masked reconstruction to0.318687versus0.281600. - Refreshed missing-modality seed-42 checks give
0.632445calibrated balanced accuracy for piezo/VLF-only and0.722257for direct-avalanche-only. - Refreshed sequence modality diagnostics still rank direct-avalanche-only highest on grouped synthetic checks (
0.8359calibrated balanced accuracy), so direct seismic-like and piezo/VLF-like channels should remain separate. - A short diversity smoke run generated extra 128x128, 1000-step seeds
43and44and refreshed aligned tensors with evaluations disabled; use larger runs before drawing model conclusions. - Aligned synthetic targets now use true future look-ahead semantics:
target_horizon_rows=Nlabels an input row from the sum of direct avalanche events in the nextNcomplete rows, not from the current row or a single offset row.
Shape Diagnostics
Real seismic vs synthetic seismic event traces:
- central-Italy seed
42sparse-event distance:1.4709 - all-Italy seed
42sparse-event distance:1.3992 - central-Italy real seismic is very sparse: nonzero ratio
0.0245, PSD slope-0.0380 - all-Italy real seismic is less sparse: nonzero ratio
0.1756, PSD slope-0.0272 - current synthetic seismic events remain too dense: nonzero ratio
0.4275, PSD slope0.0689 - raw synthetic avalanche signal is effectively continuous: nonzero ratio
0.9991
Real VLF image columns vs synthetic piezo/VLF signal:
- shape distance:
1.4620 - real VLF PSD slope:
-0.3506 - synthetic piezo PSD slope:
-0.2705 - real VLF lag-1 autocorrelation:
0.5381 - synthetic piezo lag-1 autocorrelation:
0.9658
VLF image-level comparison:
- nearest Cumiana image:
data/raw/vlf/cumiana/captures/2026-06-30/last_E_VLF_2026-06-30T23-15-00Z.jpg - nearest normalized distance after render tuning:
13.6492 - simulated intensity mean:
0.5987 - real mean intensity:
0.4687 - simulated high-intensity ratio:
0.1695 - real high-intensity ratio:
0.1712 - simulated vertical streak count:
115 - real mean vertical streak count:
117.6957
Piezo sensor scan:
- all 16 piezo sensors were compared individually against the current Cumiana VLF image-column trace
- default pre-tuning best sensor:
9 - tuned best sensor:
5 - default pre-tuning best sensor lag-1 autocorrelation:
0.8081 - tuned best sensor lag-1 autocorrelation:
0.6002 - current VLF reference lag-1 autocorrelation in the scan: about
0.5930 - summed piezo lag-1 autocorrelation remains much smoother at about
0.9658 - corrected default pre-tuning shape score:
0.9701 - tuned threshold/locality shape score:
0.6217 - tuned burst-run rate:
0.0299; current VLF reference rate differs by about0.0073 - experimental threshold-40 release alone improved lag, but adding receiver locality improved the overall corrected score
Piezo multi-seed validation:
- output CSV:
data/derived/sim/piezo_seed_validation_summary.csv - seed
40: best sensor10, shape score0.6504, lag-10.5816 - seed
41: best sensor0, shape score0.6206, lag-10.6267 - seed
42: best sensor5, shape score0.6165, lag-10.6002 - mean shape score over seeds
40-42:0.6292 - mean lag-1 autocorrelation over seeds
40-42:0.6028
Full-size multi-seed check:
compare-simulation-grid.shran for full-size256 x 256,10000step runs on seeds40,41, and42- metrics-only runs used
RUN_HEATMAPS=0,RUN_VIDEO=0, andRUN_AUDIO=0 - summary CSV:
data/derived/sim/full_size_seed_comparison.csv
| seed | events | seismic distance | event nonzero | event PSD slope | piezo/VLF distance | piezo lag-1 | nearest VLF image distance |
|---|---|---|---|---|---|---|---|
| 40 | 189 | 1.5775 | 0.8030 | 0.2524 | 1.4041 | 0.9581 | 13.2128 |
| 41 | 191 | 1.6198 | 0.8106 | -0.4042 | 1.4663 | 0.9666 | 13.4896 |
| 42 | 172 | 1.5258 | 0.7652 | -0.1678 | 1.4620 | 0.9658 | 13.6492 |
Direct avalanche event extraction tuning:
- tuning grid: quantiles
0.90,0.95,0.975,0.99; local-max windows5,15,30,60 - output CSVs:
data/derived/sim/mountain_256x256_seed*_10000.avalanche_event_tuning.csv - best quantile for seeds
40,41, and42:0.99 - best local-max window is seed-dependent: seed
40prefers5, seed41prefers60, seed42prefers30 - current default
0.95/15is denser than the tuned0.99candidates - longer
20000step runs for seeds40,41, and42also favour quantile0.99 - in the longer runs, window
30is best for seeds40and41, and second-best for seed42 - refined sparse tuning adds
max_eventsand a stableshape_scoreindependent of candidate-grid size - refined 10000-step seed
42central-Italy sparse profile:q=0.99, window480, max events3 - refined 20000-step seeds
40-42consistently prefer:q=0.999, window240, max events5 - refined 20000-step sparse profile event shape scores: seed
400.193665, seed410.119737, seed420.144586
Interpretation
Sparse local-peak extraction improved the direct synthetic seismic event trace. The extended 2024-2026 INGV comparison showed the previous synthetic seismic event list was still too dense, especially against central Italy. The refined sparse profile improves 10000-step seed 42 central-Italy distance from 1.4709 to 1.0218, with nonzero ratio 0.0259 versus real 0.0245. Keep it as a sparse seismic profile until downstream target balance is re-evaluated.
Across the full-size seed check, seed 42 is still the best current default on direct seismic event-shape distance and sparsity, but all tested seeds remain much denser than real seismic events. The next tuning target is therefore direct avalanche event extraction thresholds and burst spacing, not map projection.
The old extraction tuning pass supported quantile 0.99 and local-max window 30 as a usable model-training default. The refined sparse profile uses far fewer events (q=0.999, window 240, max events 5 on 20000-step runs) and better matches central-Italy sparsity, but it may make synthetic model targets too sparse if promoted without redesigning the target windows.
The target-window redesign is now implemented. On the refined sparse seed 40-42, 20000-step profile, positive labels scale with the look-ahead horizon as expected: horizon 1 has 3/501 positives, horizon 3 has 9/495, horizon 6 has 18/486, horizon 12 has 36/468, and horizon 24 has 72/432. Temporal splits still have zero positive test rows because the sparse synthetic events occur too early in each run, so the sparse profile should not become the default until event timing is less clustered or substantially longer synthetic runs are available.
The piezo/VLF image rendering is now much closer to Cumiana image statistics for brightness, high-intensity coverage, and vertical streaks. The underlying piezo time series is still too smooth in time, with very high autocorrelation, so later tuning should focus on signal dynamics rather than adding display artifacts.
The per-sensor scan shows that summing all piezo sensors over-smooths the VLF-like channel. Single sensors are more plausible. Raw burst-run counts were misleading because the image-derived VLF trace has many more samples than the simulation trace, so comparisons now use burst-run rate.
The current tuned default uses thresholded accumulated-charge release plus a local receiver footprint. It improves lag-1 autocorrelation and the corrected overall shape score without artificial spikes. The seed 40-42 validation is consistent enough to keep these as current simulation defaults, but the best receiver is seed-dependent, so model data preparation should preserve sensor identity and support sensor selection or pooling.
Piezo receiver locality is now separated from the direct avalanche signal range. Tuning the VLF-like piezo channel should not silently change the seismic-like avalanche channel.
After that separation and the target relabeling fix, seed 40-42 simulation CSVs, aligned synthetic model rows, tensors, temporal diagnostics, tabular PyTorch reports, sequence GRU reports, and the tabular-vs-sequence comparison were refreshed. The chronological holdout remains weak, while leave-one-seed-out checks remain stronger and should be treated as synthetic transfer diagnostics only.
Longer synthetic run check:
- regenerated seed
40,41, and42at20000steps under current defaults - sparse event counts: seed
40has130, seed41has129, seed42has135 - combined
gt0aligned table after future look-ahead relabeling:501rows,67positives,434negatives - combined
gt1aligned table after future look-ahead relabeling:501rows,2positives,499negatives gt0chronological best calibrated balanced accuracy:0.507576gt0leave-one-seed-out best calibrated balanced accuracy range:0.591536to0.826389gt1remains mostly a sparsity check
Chronological split diagnostics:
- diagnostic outputs:
data/derived/models/mountain_256x256_seeds40-42_20000.aligned_hourly_synthetic_windows.temporal_diagnostics.jsondata/derived/models/mountain_256x256_seeds40-42_20000.aligned_hourly_synthetic_windows.temporal_diagnostics.csvdata/derived/models/mountain_256x256_seeds40-42_20000.aligned_hourly_synthetic_windows_gt1.temporal_diagnostics.jsondata/derived/models/mountain_256x256_seeds40-42_20000.aligned_hourly_synthetic_windows_gt1.temporal_diagnostics.csvgt0train positive rate is0.080000; test positive rate is0.346535gt1train positive rate is0.000000; test positive rate is0.019802- later test windows have much higher avalanche/topple intensity:
synthetic_direct_avalanche_active_topple_cell_count_*andsynthetic_summary_topple_count_*drift upward by more than one training standard deviation - several direct avalanche and summary features still change target-correlation strength between train and test in the
gt0split
The longer run increases target support but does not improve chronological generalization. It should be used to stress-test the model interface and synthetic-transfer workflow, not as evidence of predictive value.
The weak 20000 chronological holdout appears to be mostly a non-stationary split problem. The simulation trajectory changes regime over time: later windows are taller and more avalanche-active, and the target rate changes substantially. Time-split evaluation is still the right conservative check, but the current synthetic generator needs burn-in handling, regime-balanced splitting, or longer/more varied runs before model metrics are meaningful.
The direct avalanche signal should remain separate from the piezo/VLF channel. Cross-modality distances can be useful sanity checks, but tuning should compare real seismic primarily with direct avalanche outputs, and real VLF primarily with piezo-derived outputs.
Model Scaffold Update
Current CPU PyTorch tabular and sequence reports can now be compared directly with compare-model-runs.sh. The sequence path also has a bounded sweep script, a missing-modality smoke script, and a real Cumiana VLF image sequence materializer so synthetic piezo/VLF and real VLF inputs keep the same model-facing shape.
Corrected-label smoke outputs:
- tabular-vs-sequence comparison:
data/derived/models/mountain_256x256_seeds40-42_20000.tabular_vs_sequence_model_comparison.json - best calibrated row in that comparison:
0.826389,seismic_vlf_unified, held-outseed41 - tabular PyTorch temporal row: calibrated balanced accuracy
0.507576 - tabular PyTorch seed holdouts: calibrated balanced accuracy
0.784169,0.762832, and0.732602 - sequence GRU temporal row: calibrated balanced accuracy
0.500000 - sequence GRU seed holdouts: calibrated balanced accuracy
0.720690,0.821105, and0.768339 - current real VLF image sequence manifest:
data/derived/models/cumiana_vlf_image_sequence/manifest.json, with279time steps and25channels - the older region-level VLF-aligned smoke artifact had
23all-Italy positives and0negatives, while central Italy had0positives and23negatives; the current refresh supersedes it with 279 rows and 277 mature labels per region, still without class variation
Pre-relabel sweep outputs that need rerunning before model selection:
- tiny sequence sweep smoke:
data/derived/models/sequence_sweep_smoke/sequence_sweep_comparison.json; best calibrated row0.709624,sequence_full, held-outseed41 - missing-modality smoke:
data/derived/models/missing_modality/missing_modality_seed42_summary.json; no-piezo direct avalanche scored higher than piezo-only in this short run - full sequence sweep:
data/derived/models/sequence_sweep/sequence_sweep_comparison.json,24reports; best calibrated row0.766942,sequence_direct_avalanche_only,lookback=60,hidden=24, held-outseed42 - combined family comparison:
data/derived/models/model_family_comparison.json,37rows; pre-relabel best calibrated row0.772558,sequence_piezo_vlf_only, held-outseed42 - sequence modality diagnostic:
data/derived/models/sequence_modality_diagnostic.json,112evaluation rows; best default sequence row uses20epochs and piezo/VLF-only, while best sweep row uses10epochs and direct avalanche-only - matched 20-epoch sequence comparison:
data/derived/models/sequence_sweep_20epoch/default_vs_matched_sequence_diagnostic.json,64evaluation rows; pre-relabel best rowsequence_piezo_vlf_only,lookback=60,hidden=24, held-outseed42, calibrated balanced accuracy0.772558 - repeated training-seed comparison:
data/derived/models/sequence_training_seed_repeat/sequence_training_seed_selection.json; pre-relabel best single rowsequence_piezo_vlf_onlyat0.772558, withsequence_fullbest on mean group score (0.741342) and worst-held-out-seed score (0.712754) - tiny patch Transformer scaffold:
data/derived/models/tiny_patch_transformer/tiny_patch_transformer_model_run_summary.json; pre-relabel best calibrated row0.637500,sequence_piezo_vlf_only, explicit regime-balanced split - current region-level real model-input scaffold:
data/derived/models/all_italy.real_vlf_aligned_windows.csvanddata/derived/models/central_italy.real_vlf_aligned_windows.csv; each has 279 rows and 277 labeled rows but still lacks class variation
Sequence diagnostic interpretation: do not change the default GRU lookback from 60 on current evidence. The corrected-label temporal sequence row remains at balanced accuracy 0.5; older sweep and missing-modality reports need rerunning before choosing between direct avalanche, piezo/VLF, and full sequence inputs.
Post-burn-in regime interpretation: sequence_full does not yet show robust performance once the first 20 percent of each synthetic seed is removed and holdouts are made by seed/regime block. This supports the earlier concern that the current synthetic generator and split design still contain regime effects that should be understood before larger model runs.
Compact model comparison:
- artifact:
data/derived/models/real_synthetic_compact_comparison.json - CSV view:
data/derived/models/real_synthetic_compact_comparison.csv - central-Italy historical seismic-only temporal baseline: calibrated balanced accuracy
0.441667 - corrected-label synthetic temporal sequence run: calibrated balanced accuracy
0.500000 - corrected-label best synthetic seed-holdout row:
seismic_vlf_unified, held-outseed41, calibrated balanced accuracy0.826389 - post-burn-in
sequence_fullregime holdouts: mean calibrated balanced accuracy0.508413 - post-burn-in
sequence_fullregime-balanced split: calibrated balanced accuracy0.650000 - pre-relabel tiny patch Transformer regime-balanced split:
sequence_piezo_vlf_only, calibrated balanced accuracy0.637500
Limitations
The real VLF data is image-derived, not raw waveform data.
The current historical seismic sample covers January 1, 2024 through July 7, 2026. It is large enough for smoke baselines and shape diagnostics, but still not enough for robust predictive claims.
Simulation step time is an assumed mapping. Frequency-domain comparisons are therefore shape diagnostics, not physical frequency validation.
The current direct avalanche signal input uses *.avalanche_signal.csv; model smoke runs currently use the 20000-step seed 40-42 synthetic tables.