← Latest papers
🔭 astrophysics

Improving Generalization and Uncertainty Quantification of Photometric Redshift Models

This paper demonstrates that combining spectroscopic and photometric redshift data through composite training or transfer learning significantly improves the generalization, bias, scatter, and outlier rates of photometric redshift models for future surveys like LSST and Euclid, while also employing Bayesian neural networks and conformal prediction to enhance uncertainty quantification.

Original authors: Jonathan Soriano, Tuan Do, Srinath Saikrishnan, Vikram Seenivasan, Bernie Boscoe, Jack Singal, Evan Jones

Published 2026-01-27
📖 5 min read🧠 Deep dive

Original authors: Jonathan Soriano, Tuan Do, Srinath Saikrishnan, Vikram Seenivasan, Bernie Boscoe, Jack Singal, Evan Jones

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Picture: Guessing How Far Away Stars Are

Imagine you are looking at a crowd of people from very far away. You can see their colors (red shirts, blue shirts) and how bright they are, but you can't see their faces clearly. You want to guess how far away each person is standing.

In astronomy, this is called estimating photometric redshifts (or "photo-z"). Instead of people, astronomers are looking at galaxies. Instead of shirts, they look at the light coming from the galaxy through different colored filters (like red, green, and blue).

The problem is: It's hard to guess the distance just by looking at the color. To get the exact distance, you usually need a special, expensive tool (spectroscopy) that acts like a high-powered zoom lens. But that tool is slow and can only look at a few bright, easy-to-see galaxies. Most galaxies are too faint or too weird-looking for that tool.

So, astronomers use Machine Learning (AI) to learn the pattern: "If a galaxy looks like this, it is probably that far away."

The Problem: The AI is Biased

The paper explains that the AI models we usually build have a major flaw. They are trained using the "zoom lens" data (spectroscopy).

  • The Analogy: Imagine you are teaching a student to identify animals, but you only show them pictures of golden retrievers. The student learns that "all dogs are golden retrievers."
  • The Reality: When you show that student a picture of a Chihuahua or a Great Dane (which are common in the real universe but rare in the "zoom lens" data), the student gets confused and makes bad guesses.

In astronomy, the "zoom lens" data is full of bright, common galaxies but misses the faint, weird, or distant ones. If we train our AI only on this, it fails when we try to use it on the billions of faint galaxies in upcoming giant surveys like the LSST (a massive telescope project).

The Solution: Two New Teaching Strategies

The authors tried two new ways to teach the AI so it can handle all types of galaxies, not just the easy ones. They used two different "textbooks" (datasets):

  1. The "Gold Standard" Book (GalaxiesML): Very accurate distances, but only for bright, common galaxies.
  2. The "Rough Draft" Book (TransferZ): Less accurate distances, but it covers a huge variety of galaxies, including faint and weird ones.

They tested two strategies to combine these books:

Strategy 1: The "Mixed Classroom" (Composite Dataset)

Instead of teaching the AI with just one book, they mixed the pages of both books together into one giant textbook.

  • The Analogy: You put the student in a classroom with both the "perfect" golden retrievers and the "rough draft" Chihuahuas. The student learns to recognize the differences and similarities between all of them at once.
  • The Result: This worked very well. The AI trained on this mixed book became much better at guessing distances for all types of galaxies. It was less biased and made fewer wild errors than the AI trained on just the "Gold Standard."

Strategy 2: The "Apprentice" (Transfer Learning)

First, they taught the AI using the "Rough Draft" book (the one with many galaxy types). Then, they took that trained AI and gave it a "finishing course" using the "Gold Standard" book to polish its accuracy.

  • The Analogy: You teach an apprentice carpenter how to build a house using cheap, rough wood (so they learn the basics of structure). Then, you give them a final exam using expensive, perfect wood to refine their skills.
  • The Result: This worked great for the "perfect wood" (bright galaxies), but the apprentice forgot how to handle the "rough wood" (faint galaxies). It became very good at one thing but bad at the other.

The "Uncertainty" Problem: Knowing When You Don't Know

In science, it's not enough to just make a guess; you need to know how confident you are in that guess.

  • The Analogy: If a weather forecaster says "It will rain," you want to know if they are 99% sure or just guessing.
  • The Paper's Tool: The authors used a mathematical trick called Split Conformal Prediction. Think of this as a "safety net." Even if the AI isn't 100% sure, this tool draws a box around the answer that says, "We are 68% sure the real answer is inside this box."
  • The Finding: They found that while the "Mixed Classroom" AI was the best at guessing the number, the "Apprentice" AI (and others) were better at knowing when to say, "I'm not sure."

The Final Verdict

The paper concludes that if you want an AI that can handle the messy, diverse reality of the universe (like the upcoming LSST survey), mixing the data sources together (The Mixed Classroom) is the best approach.

  • The Winner: The AI trained on the Composite Dataset (mixing both books) was the most consistent. It didn't get confused by weird galaxies and met most of the strict requirements for future space surveys.
  • The Trade-off: While it was the best at guessing the distance, it still needs help with the "safety net" (uncertainty). The authors note that for the most complex, probabilistic models, the AI sometimes gets too confident when it shouldn't be.

In short: To build a better telescope AI, don't just feed it the "perfect" data. Feed it a little bit of the "messy" data too, so it learns to recognize the whole universe, not just the easy parts.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →