Researchers at SLAC National Accelerator Laboratory and Stanford University have developed a neural-network-based data compression method that can reduce the size of scientific datasets by 10 to 100 times — while preserving the fine details that researchers need to recover later. The technique addresses one of the most pressing bottlenecks in next-generation scientific research: the growing gap between the data that experiments produce and the capacity to store and analyze it.

The system works by learning to compress experimental data into a compact representation that preserves structural information at multiple scales. Users can later select a specific region and level of resolution to reconstruct the data they need, similar to how a zoomable image allows navigation from a full view down to individual pixels. The compression is lossy but controlled — the algorithm retains the information that matters for scientific analysis while discarding redundant or noise-dominated components.

The challenge the system addresses is already acute at SLAC's Linac Coherent Light Source (LCLS), which produces some of the world's brightest X-ray pulses for studying materials, chemistry, and biology at the atomic scale. The facility's next-generation upgrade, LCLS-II, will generate up to one million X-ray pulses per second — producing data volumes that could overwhelm conventional storage systems. The compression method could allow researchers to capture and retain more of this data, enabling experiments that would otherwise be impossible due to storage constraints.

For the broader scientific community, the approach demonstrates how AI can serve as infrastructure — not just analyzing results, but determining how information should be preserved so that experiments remain usable long after they are conducted.