← Latest papers
🔭 astrophysics

Connecting the Dots: A Machine Learning Ready Dataset for Ionospheric Forecasting Models

As part of the 2025 NASA Heliolab, this paper introduces a curated, open-access, machine learning-ready dataset that integrates diverse solar, heliospheric, and ionospheric measurements to train and benchmark spatiotemporal models for improved operational forecasting of vertical Total Electron Content (TEC) under varying geomagnetic conditions.

Original authors: Linnea M. Wolniewicz, Halil S. Kelebek, Simone Mestici, Michael D. Vergalla, Giacomo Acciarini, Bala Poduval, Olga Verkhoglyadova, Madhulika Guhathakurta, Thomas E. Berger, Atılım Güneş Baydin, Frank
Published 2026-05-14
📖 4 min read☕ Coffee break read

Original authors: Linnea M. Wolniewicz, Halil S. Kelebek, Simone Mestici, Michael D. Vergalla, Giacomo Acciarini, Bala Poduval, Olga Verkhoglyadova, Madhulika Guhathakurta, Thomas E. Berger, Atılım Güneş Baydin, Frank Soboczenski

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine the Earth is surrounded by a giant, invisible, and constantly shifting "ocean" of charged particles called the ionosphere. This layer is crucial because it acts like a mirror for our radio waves, helping GPS satellites, airplanes, and communication systems find their way. However, this ocean is turbulent. It gets churned up by the Sun, causing "storms" that can scramble our signals and knock out power grids.

For a long time, predicting these storms has been like trying to forecast the weather with a broken thermometer and a few scattered rain gauges. Scientists have all sorts of data—measurements of solar wind, magnetic fields, and electron counts—but they are all in different formats, at different speeds, and often have holes in them. It's a messy pile of puzzle pieces that don't quite fit together.

The Big Idea: Building a "Lego Set" for AI
This paper introduces a new project that acts like a master puzzle-maker. The team has taken all those messy, scattered data sources and cleaned them up, aligned them, and organized them into a single, neat package ready for Machine Learning (AI) to use.

Think of it this way:

  • Before: Imagine trying to teach a robot to drive by giving it a stack of papers with some notes written in pencil, some in ink, some in a different language, and some pages ripped out. The robot would be confused.
  • Now: The authors have taken all those notes, translated them into the same language, filled in the missing pages with the best guesses, and bound them into a single, perfect textbook. Now, the robot (the AI) can actually learn to drive.

What's Inside the Box?
The dataset is a "smoothie" made from many different ingredients, all blended to the same rhythm:

  1. The Sun's Mood: Data from NASA's SDO satellite showing the Sun's extreme ultraviolet light (like a camera taking pictures of the Sun's "mood swings").
  2. The Solar Wind: Measurements of the particles and magnetic fields blowing from the Sun toward Earth (the "wind" that pushes the ionosphere).
  3. Earth's Reaction: Data on how Earth's magnetic field is shaking (geomagnetic indices like Kp and AE).
  4. The Ocean's Surface: Maps of the Total Electron Content (TEC), which show how "thick" or "thin" the ionosphere is in different places. Some of these maps are very detailed (dense), while others are based on fewer measurements (sparse), including data crowdsourced from Android phones!
  5. Geometry: Calculations of where the Sun and Moon are relative to Earth, just to know the angles of the light hitting our atmosphere.

Cleaning Up the Mess
One of the hardest parts of this work was dealing with "missing data." Some sensors stop working for a few minutes; others have gaps of days. The authors created a smart system to handle this:

  • If a gap is tiny, they "forward-fill" it, meaning they assume the value stays the same as the last known reading for a short time.
  • If a gap is too big, they skip it entirely so the AI doesn't learn from stale, outdated information.
  • They also created a "storm catalog" to label specific times when the ionosphere was calm versus when it was in a violent storm, ensuring the AI learns the difference between a sunny day and a hurricane.

The Results: Teaching the AI to Predict
Using this new, clean dataset, the team trained several AI models (including ones inspired by how weather forecasters work) to predict the state of the ionosphere.

  • They tested these models to see if they could predict the "thickness" of the ionosphere (TEC) up to 12 hours in advance.
  • The results were promising: The AI models were much better at predicting the future than the old "persistence" method (which just assumes the weather will stay exactly the same as it is right now).

Why This Matters
The paper doesn't claim to have solved space weather forever. Instead, it claims to have built the foundation. By providing this "machine learning-ready" dataset, they are handing the keys to the rest of the scientific community. Now, researchers around the world can stop spending months cleaning up data and start building better, smarter models to protect our satellites, GPS, and power grids from solar storms. It's the difference between giving a chef a pile of unwashed, unpeeled vegetables versus a prepped, chopped, and measured ingredient kit ready for cooking.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →