Next-Generation Science Experiments Defeat Storage Limits with Smart AI

Next-generation science experiments will collect more data than ever – so much so that they’ll surpass the capabilities of current data storage and analysis methods. To help, researchers at the Department of Energy’s SLAC National Accelerator Laboratory developed a method to use artificial intelligence to compress large amounts of raw data without losing subtle details critical to scientific discovery. They published the work in Nature Machine Intelligence.

There is going to be such a flood of data that there's really no way to handle it in the way we’ve done before,” said Joshua Turner, a lead scientist at SLAC and the Stanford Institute for Materials and Energy Sciences, a joint institute between SLAC and Stanford, and principal investigator of this work. “There are many applications in science now where data storage and analysis speed are really important problems, and I think this method is a clever way to solve them.

The method could be useful for data coming from SLAC’s Linac Coherent Light Source (LCLS), an ultrafast X-ray free-electron laser that will eventually generate up to a million X-ray pulses per second to take “snapshots” of atoms and molecules. Nearly one terabyte of data per second requires novel types of processing needed for instruments that draw on the full capabilities of the LCLS, such as the X-ray photon fluctuation spectroscopy (XPFS) instrument, which will study the movements of particles in exotic topological and quantum materials

Saving the Science-Rich Speckles

Conventional data-compression methods can erase some of the fine details in measurements that correspond to valuable scientific information. For example, the tiny speckles in X-ray images of molecules can contain important information about how materials transform. 

Those speckles often reflect the underlying arrangement, disorder or dynamics of a material,” said Yuan Ni, research associate at SLAC and lead author of the work. “If we lose them, we would lose unique scientific insights, like how a material is structured and how that changes over time.” 

The AI-based method uses neural networks to compress data, reducing the overall file size and making it easier to store, move and manage, while allowing users to control what information is kept, like the finer details. “Depending on the underlying data and the desired quality/fidelity, we can typically achieve 10- to 100-fold reductions in file size,” Ni said.

The team tested the method on various types of data, including measurements of molecules and materials from several experimental techniques, solar magnetic field measurements, and photographs. They found the neural network was able to adapt to different types of data, learning what features matter for different measurements.

A New Tool to Compress, Remember and Retrieve Data

Unlike traditional compression, this AI-based method doesn’t compress an entire dataset equally. First, it uses a mathematical tool, known as wavelet analysis, to separate features of the data by scale. Then, the neural network compresses the different-scaled features separately, ensuring the finer features aren’t lost by generalized compression. The neural network learns a compact representation of those features to preserve them.

In addition to lowering the cost of storing and managing massive datasets, the method makes retrieving data easier. Sometimes, researchers want to revisit a tiny slice of compressed data. With this method, they can decompress only the data they want, saving time and cost. 

If you are using a conventional compressor, you would need to decompress the entire file, which could take you minutes, hours or days,” said Zhantao Chen, assistant professor at the University of Texas at Austin who worked on this method while a SLAC research associate. “This method can decompress only the region of interest rather than the entire dataset, so it’s much more efficient.

This new data-compression approach can operate alongside broader data-reduction techniques such as selecting only the events and features of interest, the researchers noted. “Rather than replacing existing compression methods,” Ni said, “our work provides an additional AI-based approach.”

Other contributors include the University of California, Davis, and Carnegie Mellon University. To train the neural networks, the researchers used Perlmutter, a computational resource of the National Energy Research Scientific Computing Center (NERSC), a US Department of Energy Office of Science User Facility located at Lawrence Berkeley National Laboratory. This work was supported by DOE Office of Science and the Laboratory Directed Research and Development program at SLAC National Accelerator Laboratory. LCLS is an Office of Science user facility. 

Tell Us What You Think

Do you have a review, update or anything you would like to add to this news story?

Leave your feedback
Your comment type
Submit

While we only use edited and approved content for Azthena answers, it may on occasions provide incorrect responses. Please confirm any data provided with the related suppliers or authors. We do not provide medical advice, if you search for medical information you must always consult a medical professional before acting on any information provided.

Your questions, but not your email details will be shared with OpenAI and retained for 30 days in accordance with their privacy principles.

Please do not ask questions that use sensitive or confidential information.

Read the full Terms & Conditions.