[Astronomy and Computing] Transit surveys tend to underestimate how often habitable zone (HZ) planets appear. The core reason is geometric: the chance that an Earth-like planet transits a sun-like star is very low — less than 0.5%.

Their long orbital periods (200–400 days) add a further layer of observational bias, making them hard to detect with Kepler’s sensitivity. In simple terms, surveys are far more likely to catch hot Jupiters close to their stars than small, Earth-like planets at habitable distances.

We analysed 4,510 transit planets from the NASA Exoplanet Archive (March 2026). The raw HZ occurrence rates — the unadjusted numbers straight from the archive — are 0.33% for F/G stars, 0.72% for K stars, and 4.03% for M stars. These numbers are heavily influenced by how surveys select targets, not by nature alone. To correct for this, we developed a three-part analysis.

The first part uses a Bayesian Beta-Binomial occurrence-rate model with externally fixed completeness assumptions. It finds corrected rates of about 0.27 for K stars and 0.41 for M stars, with 68% credible intervals that remain stable when we vary the prior or rescale the assumed completeness by ±50%.

This indicates that the qualitative K and M dwarf ranking is robust to these specific assumptions; it does not, however, capture the full systematic uncertainty in the completeness model itself (see Sections 3.3 and 8). For F/G stars, the estimate (around 0.10) is far less certain, because only 3 HZ planets from this group appear in the dataset, and we treat it as schematic rather than a measurement. The second part applies a Random Forest regression.

Rather than labelling planets as simply HZ or non-HZ, we assign each planet a continuous habitability proximity score. This avoids class-imbalance problems and allows stable cross-validation. The model achieves a cross-validated R2 = 0.835 ± 0.039, but this performance is largely driven by orbital-distance features that are also used to define the target, so it should be read as a candidate-ranking tool rather than a measure of predictive skill.

When the two orbital-distance features are removed, R2 collapses to negative values, although the ablated model still recovers 90.5% of HZ planets by stellar properties alone. We therefore present the Random Forest primarily as a ranking aid, not as evidence that stellar features alone determine habitability.

The third part uses logistic regression to estimate detection probabilities for individual planets. These per-planet weights shift the median occurrence rates by less than 8%, suggesting the type-averaged occurrence-rate model is an adequate approximation at the current sample size.

All three methods agree on the stellar-type ranking. We also note a marginal signal (p ≈ 0.04) that M-dwarfs may host more HZ planets than expected; given the very small M-dwarf sample (k = 10, of which TRAPPIST-1 contributes three correlated detections), we regard this only as a falsifiable hypothesis for future surveys rather than a detection. The entire analysis can be reproduced using the public API.

Astrobiology,

Explorers Club Fellow, ex-NASA Space Station Payload manager/space biologist, Away Teams, Journalist, Lapsed climber, Synaesthete, Na’Vi-Jedi-Freman-Buddhist-mix, ASL, Devon Island and Everest Base Camp...

Leave a comment

Your email address will not be published. Required fields are marked *