Finding faults with AI: Using high-resolution digital topography data (LiDAR)

Boseong Kang, Kim D. Blisniuk, Genya Ishigaki, Sally F. McGill, & Michael E. Oskin

Submitted August 30, 2026, SCEC Contribution #15617, 2026 SCEC Annual Meeting Poster #TBD

This project summarizes initial efforts to fine-tune an Earth Foundation Model (EFM) and develop a multimodal model through model fusion for the specific task of recognizing active faults. Machine learning through imagery has been used successfully in expanding earthquake catalogs by improved monitoring of land surface changes. EFMs are beginning to be fine-tuned for use in mapping floods, burn scars and landslides globally. This work attempts to apply AI-assisted methods to identify fault traces. We first fine-tuned the EFM, Prithvi-EO 2.0 (300M), on four different composites of the same area (spring, summer, fall, and winter) of HLS optical imagery over the Parkfield section of the San Andreas Fault. However, only 71 of the 672 training patches contained a fault, against 148 million trainable parameters, and the model overfit: training loss fell from 1.50 to 0.17 while fault IoU reached only 0.08. We therefore turned to high-resolution digital elevation data (LiDAR, through OpenTopography) and trained a smaller segmentation model on the USGS Quaternary Fault and Fold Database. This model identifies fault traces near those defined in the USGS Quaternary Fault and Fold Database, with 69 percent of predictions within 10 m, but 32 percent of database traces had no prediction within 400 m.

We then built a pipeline using detailed high-resolution mapping by expert geologists (referred to as expert labels) at two contrasting sites: the western Garlock Fault, on an alluvial-fan terrain, and the San Andreas Fault near CSUSB, along a steep mountain front, using 0.5 m airborne LiDAR derivatives in 256 by 256 pixel patches spanning 128 m of ground. Because digitized detailed high resolution mapped fault traces are scarce, we pretrained a SegFormer encoder with a masked autoencoder objective on 26,000 unlabeled California terrain patches, then fine-tuned this on expert labels. On the San Andreas site this improved buffered F1 by 44 to 60 percent over ImageNet initialization. Expanding the corpus to 59,000 patches with basin and desert regions reduced buffered F1 by 17 percent. Removing lower-confidence approximate traces reduced recall of certain traces from 73 to 56 percent.

Future work will expand the expert labeled training set, fuse this elevation branch with the fine-tuned optical model, and ask expert geologists to review detections on the alluvial fans that lie along strike from the mapped traces.

Key Words
Active Faults, Machine Learning

Citation
Kang, B., Blisniuk, K. D., Ishigaki, G., McGill, S. F., & Oskin, M. E. (2026, 08). Finding faults with AI: Using high-resolution digital topography data (LiDAR). Poster Presentation at 2026 SCEC Annual Meeting.


Related Projects & Working Groups
Earthquake Geology