2608.10280v1
Tabular foundation models for the estimation of probabilistic quasar photometric redshifts in S-PLUS
First listed 2026-08-12 | Last updated 2026-08-10
Abstract
We assess whether tabular foundation models can be used as off-the-shelf probabilistic photometric-redshift estimators for quasars in the 12-band S-PLUS DR6 survey, where colour-redshift degeneracies produce multi-modal posteriors and spectroscopic training sets are shifted relative to the photometric population. TabPFN 2.5, RealTabPFN 2.5, and TabICL are benchmarked against eight task-specific baselines, including linear conditional Gaussians, FlexZBoost, mixture-density networks, normalising flows, random forests, and gradient-boosted trees, with training sets from 500 to 121,626 quasars, using both density and point-prediction metrics, together with importance-weighted scores that approximate deployment on the photometric target sample. TabPFN 2.5 is best or statistically tied for best on all metrics except the unweighted CDE loss, on which the normalising flow is statistically tied and attains the lowest mean value; its largest gains occur for small training sets and in difficult regimes (very bright and faint sources, high redshift), while retaining near-nominal calibration under covariate shift. Its main practical cost is inference: with frozen weights, large support and target catalogues require substantial GPU/accelerator memory, and full-catalogue deployment may need support-set subsampling or distillation. SHAP attributions identify WISE W1/W2 as the strongest individual predictors, with UV and optical bands offering non-negligible refinements. We conclude that TabPFN 2.5 is a strong default for probabilistic quasar photo-z estimation, particularly when training data are limited or when calibration under covariate shift is critical.
Short digest
This paper tests whether frozen tabular foundation models can deliver full probabilistic quasar photo-z posteriors for the 12-band S-PLUS DR6 survey, where colour-redshift degeneracies and spectroscopic-to-photometric covariate shift make single-value estimates inadequate. Across training sets spanning 500 to 121,626 quasars and comparisons with eight specialised baselines, TabPFN 2.5 is best or statistically tied for best on nearly every density and point-estimation metric, with especially strong gains for small training samples, bright or faint sources, and high-redshift quasars. Its posteriors remain close to nominally calibrated after importance-weighted evaluation intended to emulate the photometric target catalogue, making it a strong default for catalogue-scale quasar redshift inference; the practical limitation is the accelerator-memory cost of inference on large support and target sets. SHAP results further identify WISE W1/W2 as the most informative individual inputs, with UV and optical bands providing useful refinements.
Key figures to inspect
- Figure 1. This is the paper's central benchmark synthesis, showing how TabPFN 2.5 ranks against specialised density estimators and point-prediction methods across both ordinary and importance-weighted evaluation regimes. It makes the broad, metric-level basis for the claimed default-model status immediately visible, while flagging that the raw weighting regime has a very small effective sample size.
- Figure 2. The PIT P-P curves provide the most direct visual test of whether predictive redshift densities remain calibrated when evaluation is shifted from the spectroscopic sample toward the photometric target population. This figure is essential for the paper's deployment-oriented claim that TabPFN 2.5 retains near-nominal calibration under covariate shift.
- Figure 3. These targeted catastrophic-outlier examples show what aggregate scores can conceal: cases where TabPFN 2.5 resolves a high-redshift ambiguity, cases that defeat all methods, and cases where it fails while competing estimators succeed. The figure grounds the probabilistic photo-z comparison in concrete multimodal posteriors and preserves the paper's caveat that no method eliminates catastrophic failures.
Discussion
Log in to view the paper discussion, see votes, and leave your own feedback.