🪐
Exoplanet
Challenge
India · Celesta
9,564
KOI Records
140
Features
85.5%
Accuracy
0.823
Macro F1
NASA · Kepler Mission · ML Classification

India High School
Exoplanet Data Challenge

A leakage-audited, reproducible machine-learning pipeline that classifies Kepler Objects of Interest into CONFIRMED, CANDIDATE, and FALSE POSITIVE — built on 9,564 KOI records from NASA's Kepler Space Telescope.

🌌
9,564
KOI Records
📊
140
Raw Features
🎯
0.855
Test Accuracy
⚡
0.823
Macro F1
🌠
3
Classes
🏆
5
CV Folds

India High School Exoplanet Data Challenge¶

Celesta - NASA Kepler Exoplanet Classification

This notebook builds a leakage-audited, reproducible machine learning submission for classifying Kepler Objects of Interest into CONFIRMED, CANDIDATE, and FALSE POSITIVE.

The primary result intentionally excludes NASA vetting verdict fields and diagnostic false-positive flags. A separate comparison model includes the diagnostic flags to show how much apparent performance comes from pipeline-vetting features rather than raw transit and stellar measurements.

1. Data Loading¶

In [1]:
import os
import sys
import json
import math
import random
import warnings
from pathlib import Path

import numpy as np
import pandas as pd
import matplotlib.pyplot as plt
import seaborn as sns

from sklearn.base import BaseEstimator, TransformerMixin, clone
from sklearn.compose import ColumnTransformer
from sklearn.ensemble import RandomForestClassifier
from sklearn.impute import SimpleImputer
from sklearn.inspection import permutation_importance
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import (
    accuracy_score,
    average_precision_score,
    classification_report,
    confusion_matrix,
    f1_score,
    precision_recall_fscore_support,
    precision_recall_curve,
    recall_score,
    roc_auc_score,
    RocCurveDisplay,
    PrecisionRecallDisplay,
)
from sklearn.model_selection import (
    RandomizedSearchCV,
    StratifiedKFold,
    cross_validate,
    train_test_split,
)
from sklearn.pipeline import Pipeline
from sklearn.tree import DecisionTreeClassifier
from sklearn.preprocessing import LabelEncoder, label_binarize, OneHotEncoder, StandardScaler
from sklearn.utils.class_weight import compute_sample_weight

from xgboost import XGBClassifier

warnings.filterwarnings("ignore", category=FutureWarning)
warnings.filterwarnings("ignore", category=UserWarning)

RANDOM_STATE = 42
np.random.seed(RANDOM_STATE)
random.seed(RANDOM_STATE)

ROOT = Path.cwd()
DATA_PATH = ROOT / "KOI_Cumulative_clean.csv"
GUIDE_PATH = ROOT / "Celesta_Dataset_Guide.pdf"
OUTPUT_DIR = ROOT / "outputs"
FIG_DIR = OUTPUT_DIR / "figures"
TABLE_DIR = OUTPUT_DIR / "tables"
for directory in [OUTPUT_DIR, FIG_DIR, TABLE_DIR]:
    directory.mkdir(parents=True, exist_ok=True)

sns.set_theme(style="whitegrid", context="notebook")
plt.rcParams["figure.dpi"] = 120
plt.rcParams["savefig.dpi"] = 180
plt.rcParams["axes.titleweight"] = "bold"

df = pd.read_csv(DATA_PATH)
print(f"Loaded {DATA_PATH.name}: {df.shape[0]:,} rows x {df.shape[1]:,} columns")
display(df.head(3))
Loaded KOI_Cumulative_clean.csv: 9,564 rows x 140 columns
rowid kepid kepoi_name kepler_name koi_disposition koi_vet_stat koi_vet_date koi_pdisposition koi_fpflag_nt koi_fpflag_ss ... koi_dicco_mdec koi_dicco_mdec_err koi_dicco_msky koi_dicco_msky_err koi_dikco_mra koi_dikco_mra_err koi_dikco_mdec koi_dikco_mdec_err koi_dikco_msky koi_dikco_msky_err
0 1 10797460 K00752.01 Kepler-227 b CONFIRMED Done 2018-08-16 CANDIDATE 0 0 ... 0.200 0.16 0.200 0.170 0.080 0.130 0.310 0.170 0.320 0.160
1 2 10797460 K00752.02 Kepler-227 c CONFIRMED Done 2018-08-16 CANDIDATE 0 0 ... 0.000 0.48 0.390 0.360 0.490 0.340 0.120 0.730 0.500 0.450
2 3 10811496 K00753.01 NaN CANDIDATE Done 2018-08-16 CANDIDATE 0 0 ... -0.034 0.07 0.042 0.072 0.002 0.071 -0.027 0.074 0.027 0.074

3 rows × 140 columns

In [2]:
profile = pd.DataFrame({
    "column": df.columns,
    "dtype": [str(t) for t in df.dtypes],
    "missing_count": df.isna().sum().values,
    "missing_pct": (df.isna().mean().values * 100).round(3),
    "n_unique": [df[c].nunique(dropna=True) for c in df.columns],
}).sort_values(["missing_pct", "column"], ascending=[False, True])

shape_summary = pd.DataFrame({
    "item": ["rows", "columns", "duplicate_rows", "target_missing"],
    "value": [df.shape[0], df.shape[1], df.duplicated().sum(), df["koi_disposition"].isna().sum()],
})

class_counts = (
    df["koi_disposition"]
    .value_counts()
    .rename_axis("class")
    .reset_index(name="count")
)
class_counts["share_pct"] = (class_counts["count"] / len(df) * 100).round(2)

display(shape_summary)
display(class_counts)
display(profile.head(30))

profile.to_csv(TABLE_DIR / "column_profile.csv", index=False)
class_counts.to_csv(TABLE_DIR / "class_counts.csv", index=False)
print("Progress: data loading and profiling complete.")
item value
0 rows 9564
1 columns 140
2 duplicate_rows 0
3 target_missing 0
class count share_pct
0 FALSE POSITIVE 4839 50.60
1 CONFIRMED 2747 28.72
2 CANDIDATE 1978 20.68
column dtype missing_count missing_pct n_unique
24 koi_eccen_err1 float64 9564 100.000 0
25 koi_eccen_err2 float64 9564 100.000 0
55 koi_incl_err1 float64 9564 100.000 0
56 koi_incl_err2 float64 9564 100.000 0
35 koi_ingress float64 9564 100.000 0
36 koi_ingress_err1 float64 9564 100.000 0
37 koi_ingress_err2 float64 9564 100.000 0
26 koi_longp float64 9564 100.000 0
27 koi_longp_err1 float64 9564 100.000 0
28 koi_longp_err2 float64 9564 100.000 0
83 koi_model_chisq float64 9564 100.000 0
82 koi_model_dof float64 9564 100.000 0
101 koi_sage float64 9564 100.000 0
102 koi_sage_err1 float64 9564 100.000 0
103 koi_sage_err2 float64 9564 100.000 0
52 koi_sma_err1 float64 9564 100.000 0
53 koi_sma_err2 float64 9564 100.000 0
58 koi_teq_err1 float64 9564 100.000 0
59 koi_teq_err2 float64 9564 100.000 0
3 kepler_name str 6817 71.278 2747
80 koi_bin_oedp_sig float64 1510 15.788 4875
13 koi_comment str 1209 12.641 995
73 koi_max_mult_ev float64 1142 11.941 8422
72 koi_max_sngle_ev float64 1142 11.941 8421
76 koi_num_transits float64 1142 11.941 1627
79 koi_quarters object 1142 11.941 198
115 koi_fwm_stat_sig float64 1076 11.251 929
124 koi_fwm_prao float64 830 8.678 1943
125 koi_fwm_prao_err float64 830 8.678 475
126 koi_fwm_pdeco float64 817 8.542 1846
Progress: data loading and profiling complete.

The official dataset guide describes this file as the Cumulative Kepler Objects of Interest table with 9,564 rows and 140 columns. The target has three classes, so this notebook treats the problem as a three-class classification task. The dataset is imbalanced: false positives are about half of the table, while candidates are the smallest class.

2. Leakage Audit¶

In [3]:
leakage_audit = pd.DataFrame([
    {"column_or_group": "koi_disposition", "decision": "target only", "reason": "The label being predicted; never used as an input feature."},
    {"column_or_group": "koi_score", "decision": "not present", "reason": "The guide says this target-linked score was removed by organizers."},
    {"column_or_group": "koi_pdisposition", "decision": "drop", "reason": "NASA pipeline pre-disposition; nearly copies the false-positive decision."},
    {"column_or_group": "koi_fpflag_nt, koi_fpflag_ss, koi_fpflag_co, koi_fpflag_ec", "decision": "conditional", "reason": "Diagnostic vetting flags. Excluded from primary raw-signal model; included only in a clearly labeled comparison model."},
    {"column_or_group": "rowid, kepid, kepoi_name, kepler_name", "decision": "drop", "reason": "Identifiers. kepler_name is especially leaky because named planets are almost always confirmed."},
    {"column_or_group": "koi_comment", "decision": "drop", "reason": "Contains explicit vetting comments such as secondary eclipses and V-shaped transit notes."},
    {"column_or_group": "koi_vet_stat, koi_vet_date, koi_disp_prov", "decision": "drop", "reason": "Vetting/status metadata; also constant in this dataset."},
    {"column_or_group": "koi_datalink_dvr, koi_datalink_dvs", "decision": "drop", "reason": "Near-unique report links, not physical measurements."},
    {"column_or_group": "koi_fittype, koi_parm_prov, koi_tce_delivname, koi_sparprov", "decision": "drop from primary", "reason": "Processing/provenance metadata; may reflect catalog history instead of astrophysics."},
    {"column_or_group": "all-missing and constant columns", "decision": "drop", "reason": "No usable statistical signal."},
    {"column_or_group": "numeric transit/orbital/stellar/photometric/centroid columns", "decision": "keep", "reason": "Physical measurements suitable for a leakage-free raw-signal model after imputation."},
])
display(leakage_audit)

checks = {}
if "kepler_name" in df.columns:
    checks["kepler_name_present_vs_target"] = pd.crosstab(df["kepler_name"].notna(), df["koi_disposition"])
if "koi_pdisposition" in df.columns:
    checks["koi_pdisposition_vs_target"] = pd.crosstab(df["koi_pdisposition"], df["koi_disposition"])
for flag in ["koi_fpflag_nt", "koi_fpflag_ss", "koi_fpflag_co", "koi_fpflag_ec"]:
    if flag in df.columns:
        checks[f"{flag}_vs_target"] = pd.crosstab((pd.to_numeric(df[flag], errors="coerce") > 0).astype(int), df["koi_disposition"])

for name, table in checks.items():
    print(f"\n{name}")
    display(table)

leakage_audit.to_csv(TABLE_DIR / "leakage_audit.csv", index=False)
print("Progress: leakage audit complete. Primary model will exclude target-leaking and vetting-derived fields.")
column_or_group decision reason
0 koi_disposition target only The label being predicted; never used as an in...
1 koi_score not present The guide says this target-linked score was re...
2 koi_pdisposition drop NASA pipeline pre-disposition; nearly copies t...
3 koi_fpflag_nt, koi_fpflag_ss, koi_fpflag_co, k... conditional Diagnostic vetting flags. Excluded from primar...
4 rowid, kepid, kepoi_name, kepler_name drop Identifiers. kepler_name is especially leaky b...
5 koi_comment drop Contains explicit vetting comments such as sec...
6 koi_vet_stat, koi_vet_date, koi_disp_prov drop Vetting/status metadata; also constant in this...
7 koi_datalink_dvr, koi_datalink_dvs drop Near-unique report links, not physical measure...
8 koi_fittype, koi_parm_prov, koi_tce_delivname,... drop from primary Processing/provenance metadata; may reflect ca...
9 all-missing and constant columns drop No usable statistical signal.
10 numeric transit/orbital/stellar/photometric/ce... keep Physical measurements suitable for a leakage-f...
kepler_name_present_vs_target
koi_disposition CANDIDATE CONFIRMED FALSE POSITIVE
kepler_name
False 1978 2 4837
True 0 2745 2
koi_pdisposition_vs_target
koi_disposition CANDIDATE CONFIRMED FALSE POSITIVE
koi_pdisposition
CANDIDATE 1978 2738 1
FALSE POSITIVE 0 9 4838
koi_fpflag_nt_vs_target
koi_disposition CANDIDATE CONFIRMED FALSE POSITIVE
koi_fpflag_nt
0 1977 2743 3313
1 1 4 1526
koi_fpflag_ss_vs_target
koi_disposition CANDIDATE CONFIRMED FALSE POSITIVE
koi_fpflag_ss
0 1976 2733 2629
1 2 14 2210
koi_fpflag_co_vs_target
koi_disposition CANDIDATE CONFIRMED FALSE POSITIVE
koi_fpflag_co
0 1978 2747 2950
1 0 0 1889
koi_fpflag_ec_vs_target
koi_disposition CANDIDATE CONFIRMED FALSE POSITIVE
koi_fpflag_ec
0 1978 2747 3691
1 0 0 1148
Progress: leakage audit complete. Primary model will exclude target-leaking and vetting-derived fields.

The strongest leakage risks are not subtle. koi_pdisposition is already a NASA pipeline verdict, and kepler_name is almost a proxy for being confirmed because confirmed exoplanets receive Kepler names. The four false-positive flags are physically meaningful diagnostics, but they are downstream vetting outputs rather than raw measurements. They are therefore excluded from the primary model and used only in a labeled comparison.

3. EDA¶

In [4]:
fig, ax = plt.subplots(figsize=(7, 4))
order = class_counts["class"].tolist()
sns.barplot(data=class_counts, x="class", y="count", order=order, palette="Set2", ax=ax)
ax.set_title("Class Distribution")
ax.set_xlabel("")
ax.set_ylabel("Rows")
for i, row in class_counts.reset_index(drop=True).iterrows():
    ax.text(i, row["count"] + 80, f'{row["count"]:,}\n{row["share_pct"]:.1f}%', ha="center", va="bottom", fontsize=9)
plt.tight_layout()
plt.savefig(FIG_DIR / "class_distribution.png", bbox_inches="tight")
plt.show()
No description has been provided for this image

False positives are the largest class, so a naive model can look accurate while ignoring the smaller CANDIDATE class. This is why macro F1 and per-class recall are more informative than accuracy alone.

In [5]:
missing_nonzero = profile[profile["missing_count"] > 0].copy()
missing_top = missing_nonzero.head(35)

fig, ax = plt.subplots(figsize=(9, 7))
sns.barplot(data=missing_top, y="column", x="missing_pct", color="#4C78A8", ax=ax)
ax.axvline(50, color="red", linestyle="--", linewidth=1, label="50% drop threshold")
ax.set_title("Top Missingness by Column")
ax.set_xlabel("Missing rows (%)")
ax.set_ylabel("")
ax.legend()
plt.tight_layout()
plt.savefig(FIG_DIR / "missingness_top_columns.png", bbox_inches="tight")
plt.show()

missing_summary = pd.DataFrame({
    "bucket": ["0%", "0-5%", "5-30%", "30-50%", ">50%"],
    "columns": [
        int((profile["missing_pct"] == 0).sum()),
        int(((profile["missing_pct"] > 0) & (profile["missing_pct"] <= 5)).sum()),
        int(((profile["missing_pct"] > 5) & (profile["missing_pct"] <= 30)).sum()),
        int(((profile["missing_pct"] > 30) & (profile["missing_pct"] <= 50)).sum()),
        int((profile["missing_pct"] > 50).sum()),
    ],
})
display(missing_summary)
missing_summary.to_csv(TABLE_DIR / "missingness_summary.csv", index=False)
No description has been provided for this image
bucket columns
0 0% 21
1 0-5% 70
2 5-30% 29
3 30-50% 0
4 >50% 20

Columns above 50% missing are dropped unless they are needed only for an audit. In this file, the high-missing group is mostly entirely empty uncertainty fields. Lower-missing physical columns are imputed inside cross-validation folds with median imputation for numeric features and most-frequent imputation for categorical features, so validation data never influences training preprocessing.

In [6]:
eda_cols = ["koi_period", "koi_duration", "koi_depth", "koi_prad", "koi_teq", "koi_insol"]
plot_df = df[["koi_disposition"] + eda_cols].copy()
for col in eda_cols:
    plot_df[f"log1p_{col}"] = np.log1p(pd.to_numeric(plot_df[col], errors="coerce").clip(lower=0))

fig, axes = plt.subplots(2, 3, figsize=(14, 8))
for ax, col in zip(axes.ravel(), eda_cols):
    sns.kdeplot(
        data=plot_df,
        x=f"log1p_{col}",
        hue="koi_disposition",
        common_norm=False,
        fill=True,
        alpha=0.25,
        linewidth=1.2,
        ax=ax,
    )
    ax.set_title(f"{col} by class (log1p)")
    ax.set_xlabel(f"log1p({col})")
    ax.set_ylabel("Density")
plt.tight_layout()
plt.savefig(FIG_DIR / "key_feature_distributions.png", bbox_inches="tight")
plt.show()

class_medians = df.groupby("koi_disposition")[eda_cols + ["koi_impact", "koi_model_snr"]].median(numeric_only=True).round(3)
display(class_medians)
class_medians.to_csv(TABLE_DIR / "key_feature_class_medians.csv")
No description has been provided for this image
koi_period koi_duration koi_depth koi_prad koi_teq koi_insol koi_impact koi_model_snr
koi_disposition
CANDIDATE 20.044 3.606 242.00 1.74 691.5 53.990 0.359 11.00
CONFIRMED 11.350 3.491 448.60 2.16 777.0 86.245 0.412 28.60
FALSE POSITIVE 5.244 4.057 575.95 8.97 1152.5 424.010 0.675 34.25

Planet radius, insolation, impact parameter, and signal-to-noise show visible class separation. False positives have a much larger median planetary radius than confirmed planets and candidates, which is consistent with many false positives being stellar eclipses or bad transit fits rather than small planets.

In [7]:
corr_exclude = {
    "rowid", "kepid", "koi_disposition", "koi_pdisposition",
    "koi_fpflag_nt", "koi_fpflag_ss", "koi_fpflag_co", "koi_fpflag_ec",
}
numeric_for_corr = [
    c for c in df.select_dtypes(include=[np.number]).columns
    if c not in corr_exclude and df[c].isna().mean() < 0.50 and df[c].nunique(dropna=True) > 1
]
# Keep the heatmap readable: use representative high-importance science columns rather than every numeric column.
corr_cols = [
    c for c in [
        "koi_period", "koi_duration", "koi_depth", "koi_ror", "koi_srho", "koi_prad", "koi_sma",
        "koi_incl", "koi_teq", "koi_insol", "koi_dor", "koi_model_snr", "koi_steff",
        "koi_slogg", "koi_srad", "koi_smass", "koi_kepmag", "koi_gmag", "koi_rmag",
        "koi_jmag", "koi_hmag", "koi_kmag", "koi_fwm_stat_sig", "koi_dicco_msky", "koi_dikco_msky",
    ]
    if c in numeric_for_corr
]
corr = df[corr_cols].corr(numeric_only=True)

fig, ax = plt.subplots(figsize=(12, 10))
sns.heatmap(corr, cmap="vlag", center=0, linewidths=0.2, ax=ax)
ax.set_title("Correlation Heatmap of Physical Numeric Features")
plt.tight_layout()
plt.savefig(FIG_DIR / "correlation_heatmap.png", bbox_inches="tight")
plt.show()

pairs = []
abs_corr = corr.abs()
for i, a in enumerate(corr_cols):
    for b in corr_cols[i + 1:]:
        val = abs_corr.loc[a, b]
        if pd.notna(val) and val >= 0.90:
            pairs.append({"feature_a": a, "feature_b": b, "abs_corr": round(float(val), 4)})
high_corr_pairs = pd.DataFrame(pairs).sort_values("abs_corr", ascending=False)
display(high_corr_pairs.head(20))
high_corr_pairs.to_csv(TABLE_DIR / "high_correlation_pairs.csv", index=False)
No description has been provided for this image
feature_a feature_b abs_corr
2 koi_kepmag koi_rmag 0.9993
12 koi_hmag koi_kmag 0.9978
10 koi_jmag koi_hmag 0.9931
11 koi_jmag koi_kmag 0.9901
0 koi_period koi_dor 0.9894
6 koi_gmag koi_rmag 0.9862
1 koi_kepmag koi_gmag 0.9793
3 koi_kepmag koi_jmag 0.9420
7 koi_rmag koi_jmag 0.9410
4 koi_kepmag koi_hmag 0.9184
8 koi_rmag koi_hmag 0.9159
5 koi_kepmag koi_kmag 0.9121
9 koi_rmag koi_kmag 0.9086

Several feature groups are strongly correlated because the catalog reports related physical quantities and paired uncertainty bounds. Tree models can handle this better than linear models, but feature importance must be interpreted as grouped evidence rather than as independent causal effects.

In [8]:
outlier_cols = ["koi_period", "koi_duration", "koi_depth", "koi_prad", "koi_teq", "koi_insol", "koi_impact", "koi_model_snr"]
outlier_rows = []
for col in outlier_cols:
    s = pd.to_numeric(df[col], errors="coerce")
    p01, p99 = s.quantile([0.01, 0.99])
    outlier_rows.append({
        "feature": col,
        "min": s.min(),
        "p01": p01,
        "median": s.median(),
        "p99": p99,
        "max": s.max(),
        "below_p01": int((s < p01).sum()),
        "above_p99": int((s > p99).sum()),
        "missing": int(s.isna().sum()),
    })
outlier_table = pd.DataFrame(outlier_rows)
display(outlier_table.round(4))
outlier_table.to_csv(TABLE_DIR / "outlier_scan.csv", index=False)

long_outliers = df[["koi_disposition"] + outlier_cols].melt(id_vars="koi_disposition", var_name="feature", value_name="value")
long_outliers["log1p_value"] = np.log1p(pd.to_numeric(long_outliers["value"], errors="coerce").clip(lower=0))

fig, ax = plt.subplots(figsize=(13, 6))
sns.boxplot(data=long_outliers, x="feature", y="log1p_value", hue="koi_disposition", fliersize=1.2, ax=ax)
ax.set_title("Outlier Scan on Log-Scaled Physical Features")
ax.set_xlabel("")
ax.set_ylabel("log1p(value)")
ax.tick_params(axis="x", rotation=25)
plt.tight_layout()
plt.savefig(FIG_DIR / "outlier_boxplots.png", bbox_inches="tight")
plt.show()
feature min p01 median p99 max below_p01 above_p99 missing
0 koi_period 0.2418 0.5277 9.7528 550.9979 1.299958e+05 96 96 0
1 koi_duration 0.0520 0.7987 3.7926 30.0475 1.385400e+02 96 96 0
2 koi_depth 0.0000 24.1000 421.1000 416740.0000 1.541400e+06 92 92 363
3 koi_prad 0.0800 0.5200 2.3900 369.2000 2.003460e+05 87 92 363
4 koi_teq 25.0000 196.0000 878.0000 4106.0000 1.466700e+04 87 92 363
5 koi_insol 0.0000 0.3500 141.6000 70297.1642 1.094755e+07 87 93 321
6 koi_impact 0.0000 0.0030 0.5370 1.7826 1.008060e+02 78 92 363
7 koi_model_snr 0.0000 5.0000 23.0000 4355.5000 9.054700e+03 88 92 363
No description has been provided for this image

The catalog contains extreme fitted values, including enormous planet radii and orbital periods. These are not automatically removed because implausible fitted values are themselves useful evidence that a signal may be a false positive. The modeling pipeline uses log transforms and robust tree models rather than deleting scientifically suspicious rows.

4. Feature Engineering¶

In [9]:
TARGET = "koi_disposition"
FLAG_COLUMNS = ["koi_fpflag_nt", "koi_fpflag_ss", "koi_fpflag_co", "koi_fpflag_ec"]
ALWAYS_DROP = [
    TARGET, "rowid", "kepid", "kepoi_name", "kepler_name",
    "koi_pdisposition", "koi_score", "koi_comment",
    "koi_vet_stat", "koi_vet_date", "koi_disp_prov",
    "koi_datalink_dvr", "koi_datalink_dvs",
    "koi_fittype", "koi_parm_prov", "koi_tce_delivname", "koi_sparprov",
    "koi_quarters",
]

class FeatureEngineer(BaseEstimator, TransformerMixin):
    def __init__(self, include_flags=False):
        self.include_flags = include_flags
        self.drop_columns_ = None
        self.input_columns_ = None
        self.output_columns_ = None

    def fit(self, X, y=None):
        X = X.copy()
        self.input_columns_ = list(X.columns)
        missing_pct = X.isna().mean()
        nunique = X.nunique(dropna=True)
        drop = set([c for c in ALWAYS_DROP if c in X.columns])
        if not self.include_flags:
            drop.update([c for c in FLAG_COLUMNS if c in X.columns])
        drop.update(missing_pct[missing_pct > 0.50].index.tolist())
        drop.update(nunique[nunique <= 1].index.tolist())
        self.drop_columns_ = sorted(drop)

        Xt = self._transform_frame(X)
        self.output_columns_ = list(Xt.columns)
        return self

    @staticmethod
    def _safe_num(df, col):
        if col not in df.columns:
            return pd.Series(np.nan, index=df.index, dtype="float64")
        return pd.to_numeric(df[col], errors="coerce")

    @staticmethod
    def _signed_abs_err(df, col):
        if col not in df.columns:
            return pd.Series(np.nan, index=df.index, dtype="float64")
        return pd.to_numeric(df[col], errors="coerce").abs()

    def _add_features(self, X):
        X = X.copy()
        eps = 1e-9
        positive_log_cols = [
            "koi_period", "koi_duration", "koi_depth", "koi_prad", "koi_teq", "koi_insol",
            "koi_model_snr", "koi_srad", "koi_smass", "koi_srho", "koi_dor",
            "koi_max_sngle_ev", "koi_max_mult_ev",
        ]
        for col in positive_log_cols:
            if col in X.columns:
                vals = pd.to_numeric(X[col], errors="coerce").clip(lower=0)
                X[f"log1p_{col}"] = np.log1p(vals)

        depth = self._safe_num(X, "koi_depth")
        depth_err = self._signed_abs_err(X, "koi_depth_err1")
        X["transit_snr_proxy"] = depth / (depth_err + eps)

        prad = self._safe_num(X, "koi_prad")
        srad = self._safe_num(X, "koi_srad")
        radius_ratio = prad / (srad * 109.1 + eps)
        depth_ratio = np.sqrt((depth.clip(lower=0) / 1_000_000).clip(lower=0))
        X["radius_ratio_from_prad_srad"] = radius_ratio
        X["depth_implied_radius_ratio"] = depth_ratio
        X["radius_ratio_residual"] = radius_ratio - depth_ratio
        X["radius_ratio_abs_residual"] = (radius_ratio - depth_ratio).abs()

        period = self._safe_num(X, "koi_period")
        duration = self._safe_num(X, "koi_duration")
        X["transit_duty_cycle"] = duration / (24 * period + eps)
        X["log_transit_duty_cycle"] = np.log1p(X["transit_duty_cycle"].clip(lower=0))

        teq = self._safe_num(X, "koi_teq")
        insol = self._safe_num(X, "koi_insol")
        # Equilibrium temperature scales approximately with the fourth root of insolation.
        expected_teq_from_insol = 278.0 * np.power(insol.clip(lower=0) + eps, 0.25)
        X["teq_insol_ratio"] = teq / (expected_teq_from_insol + eps)
        X["teq_insol_abs_residual"] = (teq - expected_teq_from_insol).abs()

        smass = self._safe_num(X, "koi_smass")
        srho = self._safe_num(X, "koi_srho")
        catalog_density_proxy = smass / (np.power(srad, 3) + eps)
        X["stellar_density_ratio"] = srho / (catalog_density_proxy + eps)
        X["log_stellar_density_ratio"] = np.sign(X["stellar_density_ratio"]) * np.log1p(np.abs(X["stellar_density_ratio"]))

        impact = self._safe_num(X, "koi_impact")
        ror = self._safe_num(X, "koi_ror")
        X["impact_plus_ror"] = impact + ror
        X["high_impact_flag"] = (impact > 1).astype(float)

        if "koi_quarters" in X.columns:
            quarters = X["koi_quarters"].astype("string")
            X["observed_quarter_count"] = quarters.str.count("1").astype("float64")
            X["observed_quarter_length"] = quarters.str.len().astype("float64")

        for a, b, name in [
            ("koi_gmag", "koi_rmag", "color_g_minus_r"),
            ("koi_rmag", "koi_imag", "color_r_minus_i"),
            ("koi_jmag", "koi_kmag", "color_j_minus_k"),
            ("koi_kepmag", "koi_kmag", "color_kep_minus_k"),
        ]:
            if a in X.columns and b in X.columns:
                X[name] = self._safe_num(X, a) - self._safe_num(X, b)

        for base in [
            "koi_period", "koi_duration", "koi_depth", "koi_prad", "koi_srho", "koi_insol",
            "koi_impact", "koi_ror", "koi_dor", "koi_steff", "koi_slogg", "koi_srad", "koi_smass",
        ]:
            if base in X.columns:
                val = self._safe_num(X, base).abs()
                err1 = self._signed_abs_err(X, f"{base}_err1")
                err2 = self._signed_abs_err(X, f"{base}_err2")
                if f"{base}_err1" in X.columns:
                    X[f"{base}_rel_err1"] = err1 / (val + eps)
                if f"{base}_err2" in X.columns:
                    X[f"{base}_rel_err2"] = err2 / (val + eps)
                if f"{base}_err1" in X.columns and f"{base}_err2" in X.columns:
                    X[f"{base}_mean_rel_err"] = (err1 + err2) / (2 * val + eps)

        # Diagnostic flags are kept only for the comparison model and binarized to handle the one anomalous value 465.
        if self.include_flags:
            for flag in FLAG_COLUMNS:
                if flag in X.columns:
                    X[flag] = (pd.to_numeric(X[flag], errors="coerce").fillna(0) > 0).astype(int)

        return X

    def _transform_frame(self, X):
        X = X.copy()
        X = self._add_features(X)
        X = X.drop(columns=[c for c in self.drop_columns_ if c in X.columns], errors="ignore")
        X = X.replace([np.inf, -np.inf], np.nan)
        return X

    def transform(self, X):
        Xt = self._transform_frame(X)
        for col in self.output_columns_:
            if col not in Xt.columns:
                Xt[col] = np.nan
        return Xt[self.output_columns_]

    def get_feature_names_out(self, input_features=None):
        return np.array(self.output_columns_, dtype=object)


feature_explanations = pd.DataFrame([
    {"engineered_feature": "log1p physical quantities", "plain_reason": "Transit measurements span orders of magnitude, so logs make small and large systems comparable."},
    {"engineered_feature": "transit_snr_proxy", "plain_reason": "A deeper dip measured with lower uncertainty is more trustworthy."},
    {"engineered_feature": "radius_ratio_from_prad_srad", "plain_reason": "A planet blocks light in proportion to its size compared with its star."},
    {"engineered_feature": "depth_implied_radius_ratio", "plain_reason": "Transit depth gives an independent estimate of the planet-to-star radius ratio."},
    {"engineered_feature": "radius_ratio_abs_residual", "plain_reason": "If radius estimates disagree strongly, the event may be a bad fit or false positive."},
    {"engineered_feature": "transit_duty_cycle", "plain_reason": "Transit duration should be plausible relative to the orbital period."},
    {"engineered_feature": "stellar_density_ratio", "plain_reason": "Transit geometry and catalog stellar properties should imply compatible stellar density."},
    {"engineered_feature": "teq_insol_abs_residual", "plain_reason": "Planet temperature and received starlight should be physically consistent."},
    {"engineered_feature": "impact_plus_ror / high_impact_flag", "plain_reason": "Grazing events are harder to trust and often resemble eclipsing binaries."},
    {"engineered_feature": "observed_quarter_count", "plain_reason": "More observed quarters means more chances to verify repeated transits."},
    {"engineered_feature": "color indices", "plain_reason": "Star color helps summarize the host star context from photometric bands."},
    {"engineered_feature": "relative uncertainty ratios", "plain_reason": "Large uncertainty relative to a value is itself evidence that a measurement is less reliable."},
])
display(feature_explanations)
feature_explanations.to_csv(TABLE_DIR / "feature_engineering_explanations.csv", index=False)

fe_preview = FeatureEngineer(include_flags=False).fit(df.drop(columns=[TARGET]), df[TARGET])
primary_feature_frame = fe_preview.transform(df.drop(columns=[TARGET]))
print(f"Primary raw-signal feature count after engineering: {primary_feature_frame.shape[1]}")
display(pd.DataFrame({"feature": primary_feature_frame.columns[:40]}))
print("Progress: feature engineering complete.")
engineered_feature plain_reason
0 log1p physical quantities Transit measurements span orders of magnitude,...
1 transit_snr_proxy A deeper dip measured with lower uncertainty i...
2 radius_ratio_from_prad_srad A planet blocks light in proportion to its siz...
3 depth_implied_radius_ratio Transit depth gives an independent estimate of...
4 radius_ratio_abs_residual If radius estimates disagree strongly, the eve...
5 transit_duty_cycle Transit duration should be plausible relative ...
6 stellar_density_ratio Transit geometry and catalog stellar propertie...
7 teq_insol_abs_residual Planet temperature and received starlight shou...
8 impact_plus_ror / high_impact_flag Grazing events are harder to trust and often r...
9 observed_quarter_count More observed quarters means more chances to v...
10 color indices Star color helps summarize the host star conte...
11 relative uncertainty ratios Large uncertainty relative to a value is itsel...
Primary raw-signal feature count after engineering: 166
feature
0 koi_period
1 koi_period_err1
2 koi_period_err2
3 koi_time0bk
4 koi_time0bk_err1
5 koi_time0bk_err2
6 koi_time0
7 koi_time0_err1
8 koi_time0_err2
9 koi_impact
10 koi_impact_err1
11 koi_impact_err2
12 koi_duration
13 koi_duration_err1
14 koi_duration_err2
15 koi_depth
16 koi_depth_err1
17 koi_depth_err2
18 koi_ror
19 koi_ror_err1
20 koi_ror_err2
21 koi_srho
22 koi_srho_err1
23 koi_srho_err2
24 koi_prad
25 koi_prad_err1
26 koi_prad_err2
27 koi_sma
28 koi_incl
29 koi_teq
30 koi_insol
31 koi_insol_err1
32 koi_insol_err2
33 koi_dor
34 koi_dor_err1
35 koi_dor_err2
36 koi_ldm_coeff2
37 koi_ldm_coeff1
38 koi_max_sngle_ev
39 koi_max_mult_ev
Progress: feature engineering complete.

5. Modeling¶

In [10]:
y_raw = df[TARGET].astype(str)
label_encoder = LabelEncoder()
y = label_encoder.fit_transform(y_raw)
class_names = label_encoder.classes_
candidate_label_index = int(np.where(class_names == "CANDIDATE")[0][0])

primary_features = (
    FeatureEngineer(include_flags=False)
    .fit(df.drop(columns=[TARGET]), y)
    .transform(df.drop(columns=[TARGET]))
    .select_dtypes(include=[np.number])
)
flag_features = (
    FeatureEngineer(include_flags=True)
    .fit(df.drop(columns=[TARGET]), y)
    .transform(df.drop(columns=[TARGET]))
    .select_dtypes(include=[np.number])
)

row_index = np.arange(len(df))
train_idx, test_idx = train_test_split(
    row_index, test_size=0.20, random_state=RANDOM_STATE, stratify=y
)
X_train_primary = primary_features.iloc[train_idx].reset_index(drop=True)
X_test_primary = primary_features.iloc[test_idx].reset_index(drop=True)
X_train_flags = flag_features.iloc[train_idx].reset_index(drop=True)
X_test_flags = flag_features.iloc[test_idx].reset_index(drop=True)
y_train = y[train_idx]
y_test = y[test_idx]

print("Classes:", dict(zip(class_names, range(len(class_names)))))
print("Primary raw-signal feature matrix:", primary_features.shape)
print("Comparison feature matrix with flags:", flag_features.shape)
print("Train rows:", X_train_primary.shape[0], "Test rows:", X_test_primary.shape[0])

cv = StratifiedKFold(n_splits=5, shuffle=True, random_state=RANDOM_STATE)

def _prepare_fold(X_train_fold, X_valid_fold, scale=False):
    imputer = SimpleImputer(strategy="median")
    X_train_imp = imputer.fit_transform(X_train_fold)
    X_valid_imp = imputer.transform(X_valid_fold)
    scaler = None
    if scale:
        scaler = StandardScaler()
        X_train_imp = scaler.fit_transform(X_train_imp)
        X_valid_imp = scaler.transform(X_valid_imp)
    return X_train_imp, X_valid_imp, imputer, scaler

def _fit_estimator(estimator, X_train_arr, y_train_arr, use_sample_weight=False):
    fitted = clone(estimator)
    if use_sample_weight:
        weights = compute_sample_weight(class_weight="balanced", y=y_train_arr)
        fitted.fit(X_train_arr, y_train_arr, sample_weight=weights)
    else:
        fitted.fit(X_train_arr, y_train_arr)
    return fitted

def cross_validate_manual(name, estimator, X_frame, y_array, use_sample_weight=False, scale=False, splitter=cv):
    rows = []
    fold_macro = []
    for fold, (tr, va) in enumerate(splitter.split(X_frame, y_array), start=1):
        Xtr, Xva, _, _ = _prepare_fold(X_frame.iloc[tr], X_frame.iloc[va], scale=scale)
        fitted = _fit_estimator(estimator, Xtr, y_array[tr], use_sample_weight=use_sample_weight)
        pred = fitted.predict(Xva)
        precision, recall, f1, support = precision_recall_fscore_support(
            y_array[va], pred, labels=np.arange(len(class_names)), zero_division=0
        )
        row = {
            "model": name,
            "fold": fold,
            "accuracy": accuracy_score(y_array[va], pred),
            "macro_f1": f1_score(y_array[va], pred, average="macro"),
            "weighted_f1": f1_score(y_array[va], pred, average="weighted"),
            "macro_recall": recall_score(y_array[va], pred, average="macro"),
            "candidate_recall": recall[candidate_label_index],
        }
        rows.append(row)
        fold_macro.append(row["macro_f1"])
    return pd.DataFrame(rows), np.array(fold_macro)

model_specs = [
    {
        "name": "Decision Tree baseline (unbalanced)",
        "estimator": DecisionTreeClassifier(max_depth=6, random_state=RANDOM_STATE),
        "X": X_train_primary,
        "use_sample_weight": False,
        "scale": False,
    },
    {
        "name": "Decision Tree baseline (balanced)",
        "estimator": DecisionTreeClassifier(max_depth=6, class_weight="balanced", random_state=RANDOM_STATE),
        "X": X_train_primary,
        "use_sample_weight": False,
        "scale": False,
    },
    {
        "name": "Random Forest (raw-signal)",
        "estimator": RandomForestClassifier(
            n_estimators=60,
            max_depth=None,
            min_samples_leaf=2,
            max_features="sqrt",
            class_weight="balanced_subsample",
            random_state=RANDOM_STATE,
            n_jobs=-1,
        ),
        "X": X_train_primary,
        "use_sample_weight": False,
        "scale": False,
    },
    {
        "name": "XGBoost (raw-signal)",
        "estimator": XGBClassifier(
            objective="multi:softprob",
            num_class=len(class_names),
            eval_metric="mlogloss",
            n_estimators=25,
            max_depth=3,
            learning_rate=0.08,
            subsample=0.9,
            colsample_bytree=0.85,
            reg_lambda=1.0,
            random_state=RANDOM_STATE,
            n_jobs=1,
            tree_method="hist",
        ),
        "X": X_train_primary,
        "use_sample_weight": True,
        "scale": False,
    },
    {
        "name": "Random Forest (with vetting flags comparison)",
        "estimator": RandomForestClassifier(
            n_estimators=60,
            max_depth=None,
            min_samples_leaf=2,
            max_features="sqrt",
            class_weight="balanced_subsample",
            random_state=RANDOM_STATE,
            n_jobs=-1,
        ),
        "X": X_train_flags,
        "use_sample_weight": False,
        "scale": False,
    },
]

all_fold_rows = []
fold_scores = {}
for spec in model_specs:
    print(f"Cross-validating: {spec['name']}")
    fold_df, macro_scores = cross_validate_manual(
        spec["name"], spec["estimator"], spec["X"], y_train,
        use_sample_weight=spec["use_sample_weight"], scale=spec["scale"]
    )
    all_fold_rows.append(fold_df)
    fold_scores[spec["name"]] = macro_scores

fold_metrics = pd.concat(all_fold_rows, ignore_index=True)
cv_results = (
    fold_metrics
    .groupby("model")
    .agg(
        accuracy_mean=("accuracy", "mean"),
        accuracy_std=("accuracy", "std"),
        macro_f1_mean=("macro_f1", "mean"),
        macro_f1_std=("macro_f1", "std"),
        weighted_f1_mean=("weighted_f1", "mean"),
        weighted_f1_std=("weighted_f1", "std"),
        macro_recall_mean=("macro_recall", "mean"),
        macro_recall_std=("macro_recall", "std"),
        candidate_recall_mean=("candidate_recall", "mean"),
        candidate_recall_std=("candidate_recall", "std"),
    )
    .reset_index()
    .sort_values("macro_f1_mean", ascending=False)
)
display(cv_results.round(4))
cv_results.to_csv(TABLE_DIR / "cv_model_comparison.csv", index=False)

fold_score_table = pd.DataFrame(fold_scores)
display(fold_score_table.round(4))
fold_score_table.to_csv(TABLE_DIR / "cv_fold_macro_f1.csv", index=False)
fold_metrics.to_csv(TABLE_DIR / "cv_fold_metrics_long.csv", index=False)

imbalance_check = cv_results[
    cv_results["model"].isin(["Decision Tree baseline (unbalanced)", "Decision Tree baseline (balanced)"])
][["model", "macro_f1_mean", "candidate_recall_mean"]]
display(imbalance_check.round(4))

print("Progress: baseline and primary cross-validation complete.")
Classes: {'CANDIDATE': 0, 'CONFIRMED': 1, 'FALSE POSITIVE': 2}
Primary raw-signal feature matrix: (9564, 166)
Comparison feature matrix with flags: (9564, 170)
Train rows: 7651 Test rows: 1913
Cross-validating: Decision Tree baseline (unbalanced)
Cross-validating: Decision Tree baseline (balanced)
Cross-validating: Random Forest (raw-signal)
Cross-validating: XGBoost (raw-signal)
Cross-validating: Random Forest (with vetting flags comparison)
model accuracy_mean accuracy_std macro_f1_mean macro_f1_std weighted_f1_mean weighted_f1_std macro_recall_mean macro_recall_std candidate_recall_mean candidate_recall_std
3 Random Forest (with vetting flags comparison) 0.9200 0.0112 0.8946 0.0148 0.9183 0.0116 0.8883 0.0155 0.7648 0.0334
2 Random Forest (raw-signal) 0.8475 0.0149 0.8116 0.0175 0.8438 0.0146 0.8074 0.0176 0.5967 0.0329
4 XGBoost (raw-signal) 0.8017 0.0071 0.7821 0.0092 0.8087 0.0066 0.8008 0.0110 0.7332 0.0300
1 Decision Tree baseline (unbalanced) 0.7904 0.0103 0.7507 0.0094 0.7874 0.0087 0.7482 0.0110 0.5424 0.0320
0 Decision Tree baseline (balanced) 0.7654 0.0131 0.7491 0.0116 0.7758 0.0103 0.7714 0.0130 0.7452 0.0320
Decision Tree baseline (unbalanced) Decision Tree baseline (balanced) Random Forest (raw-signal) XGBoost (raw-signal) Random Forest (with vetting flags comparison)
0 0.7483 0.7399 0.8056 0.7771 0.8949
1 0.7643 0.7500 0.8280 0.7951 0.9096
2 0.7421 0.7593 0.8270 0.7813 0.9029
3 0.7559 0.7615 0.8116 0.7863 0.8951
4 0.7429 0.7349 0.7855 0.7708 0.8705
model macro_f1_mean candidate_recall_mean
1 Decision Tree baseline (unbalanced) 0.7507 0.5424
0 Decision Tree baseline (balanced) 0.7491 0.7452
Progress: baseline and primary cross-validation complete.
In [11]:
# Small, documented tuning search on the strongest leakage-free tree families.
tuning_cv = StratifiedKFold(n_splits=3, shuffle=True, random_state=RANDOM_STATE)

tuning_candidates = [
    {
        "model": "Random Forest tuned",
        "estimator": RandomForestClassifier(
            n_estimators=80,
            max_depth=None,
            min_samples_leaf=2,
            max_features="sqrt",
            class_weight="balanced_subsample",
            random_state=RANDOM_STATE,
            n_jobs=-1,
        ),
        "params": {"n_estimators": 80, "max_depth": None, "min_samples_leaf": 2, "max_features": "sqrt"},
        "use_sample_weight": False,
    },
    {
        "model": "Random Forest tuned",
        "estimator": RandomForestClassifier(
            n_estimators=120,
            max_depth=18,
            min_samples_leaf=2,
            max_features="sqrt",
            class_weight="balanced_subsample",
            random_state=RANDOM_STATE,
            n_jobs=-1,
        ),
        "params": {"n_estimators": 120, "max_depth": 18, "min_samples_leaf": 2, "max_features": "sqrt"},
        "use_sample_weight": False,
    },
    {
        "model": "Random Forest tuned",
        "estimator": RandomForestClassifier(
            n_estimators=120,
            max_depth=None,
            min_samples_leaf=1,
            max_features=0.60,
            class_weight="balanced_subsample",
            random_state=RANDOM_STATE,
            n_jobs=-1,
        ),
        "params": {"n_estimators": 120, "max_depth": None, "min_samples_leaf": 1, "max_features": 0.60},
        "use_sample_weight": False,
    },
    {
        "model": "XGBoost tuned",
        "estimator": XGBClassifier(
            objective="multi:softprob",
            num_class=len(class_names),
            eval_metric="mlogloss",
            n_estimators=35,
            max_depth=3,
            learning_rate=0.08,
            subsample=0.90,
            colsample_bytree=0.85,
            reg_lambda=1.0,
            random_state=RANDOM_STATE,
            n_jobs=1,
            tree_method="hist",
        ),
        "params": {"n_estimators": 35, "max_depth": 3, "learning_rate": 0.08, "subsample": 0.90, "colsample_bytree": 0.85},
        "use_sample_weight": True,
    },
    {
        "model": "XGBoost tuned",
        "estimator": XGBClassifier(
            objective="multi:softprob",
            num_class=len(class_names),
            eval_metric="mlogloss",
            n_estimators=55,
            max_depth=3,
            learning_rate=0.06,
            subsample=0.95,
            colsample_bytree=0.85,
            reg_lambda=1.2,
            random_state=RANDOM_STATE,
            n_jobs=1,
            tree_method="hist",
        ),
        "params": {"n_estimators": 55, "max_depth": 3, "learning_rate": 0.06, "subsample": 0.95, "colsample_bytree": 0.85},
        "use_sample_weight": True,
    },
]

tuning_rows = []
for candidate in tuning_candidates:
    label = f"{candidate['model']} {candidate['params']}"
    print("Tuning:", label)
    fold_df, macro_scores = cross_validate_manual(
        label,
        candidate["estimator"],
        X_train_primary,
        y_train,
        use_sample_weight=candidate["use_sample_weight"],
        splitter=tuning_cv,
    )
    tuning_rows.append({
        "model": candidate["model"],
        "mean_macro_f1": macro_scores.mean(),
        "std_macro_f1": macro_scores.std(),
        "params": candidate["params"],
        "estimator": candidate["estimator"],
        "use_sample_weight": candidate["use_sample_weight"],
    })

tuning_results_full = pd.DataFrame(tuning_rows).sort_values("mean_macro_f1", ascending=False).reset_index(drop=True)
display(tuning_results_full.drop(columns=["estimator", "use_sample_weight"]).round(4))
tuning_results_full.drop(columns=["estimator", "use_sample_weight"]).to_csv(TABLE_DIR / "tuning_results.csv", index=False)

best_row = tuning_results_full.iloc[0]
best_primary_name = (
    "XGBoost tuned raw-signal"
    if best_row["model"] == "XGBoost tuned"
    else "Random Forest tuned raw-signal"
)
best_primary_estimator = best_row["estimator"]
best_primary_use_sample_weight = bool(best_row["use_sample_weight"])
best_primary_params = best_row["params"]
print(f"Selected primary model: {best_primary_name}")
print("Best parameters:", best_primary_params)
print("Progress: hyperparameter tuning complete.")
Tuning: Random Forest tuned {'n_estimators': 80, 'max_depth': None, 'min_samples_leaf': 2, 'max_features': 'sqrt'}
Tuning: Random Forest tuned {'n_estimators': 120, 'max_depth': 18, 'min_samples_leaf': 2, 'max_features': 'sqrt'}
Tuning: Random Forest tuned {'n_estimators': 120, 'max_depth': None, 'min_samples_leaf': 1, 'max_features': 0.6}
Tuning: XGBoost tuned {'n_estimators': 35, 'max_depth': 3, 'learning_rate': 0.08, 'subsample': 0.9, 'colsample_bytree': 0.85}
Tuning: XGBoost tuned {'n_estimators': 55, 'max_depth': 3, 'learning_rate': 0.06, 'subsample': 0.95, 'colsample_bytree': 0.85}
model mean_macro_f1 std_macro_f1 params
0 Random Forest tuned 0.8128 0.0088 {'n_estimators': 120, 'max_depth': 18, 'min_sa...
1 Random Forest tuned 0.8121 0.0099 {'n_estimators': 80, 'max_depth': None, 'min_s...
2 Random Forest tuned 0.8087 0.0061 {'n_estimators': 120, 'max_depth': None, 'min_...
3 XGBoost tuned 0.7900 0.0069 {'n_estimators': 55, 'max_depth': 3, 'learning...
4 XGBoost tuned 0.7875 0.0072 {'n_estimators': 35, 'max_depth': 3, 'learning...
Selected primary model: Random Forest tuned raw-signal
Best parameters: {'n_estimators': 120, 'max_depth': 18, 'min_samples_leaf': 2, 'max_features': 'sqrt'}
Progress: hyperparameter tuning complete.

Class imbalance is handled with class weights rather than SMOTE. This avoids creating synthetic astronomical objects and keeps every training row tied to a real KOI. The notebook compares unbalanced and balanced decision-tree baselines so the effect on minority-class recall is measured rather than assumed.

6. Evaluation¶

In [12]:
def fit_final_bundle(estimator, X_frame, y_array, use_sample_weight=False, scale=False):
    imputer = SimpleImputer(strategy="median")
    X_imp = imputer.fit_transform(X_frame)
    scaler = None
    if scale:
        scaler = StandardScaler()
        X_imp = scaler.fit_transform(X_imp)
    fitted = _fit_estimator(estimator, X_imp, y_array, use_sample_weight=use_sample_weight)
    return {"estimator": fitted, "imputer": imputer, "scaler": scaler, "scale": scale}

def transform_bundle(bundle, X_frame):
    X_imp = bundle["imputer"].transform(X_frame)
    if bundle["scaler"] is not None:
        X_imp = bundle["scaler"].transform(X_imp)
    return X_imp

def predict_bundle(bundle, X_frame):
    X_arr = transform_bundle(bundle, X_frame)
    pred = bundle["estimator"].predict(X_arr)
    proba = bundle["estimator"].predict_proba(X_arr) if hasattr(bundle["estimator"], "predict_proba") else None
    return pred, proba

def evaluate_model(name, bundle, X_eval, y_eval, save_prefix):
    y_pred, y_proba = predict_bundle(bundle, X_eval)
    report = classification_report(y_eval, y_pred, target_names=class_names, output_dict=True, zero_division=0)
    report_df = pd.DataFrame(report).T
    cm = confusion_matrix(y_eval, y_pred)
    cm_norm = confusion_matrix(y_eval, y_pred, normalize="true")
    summary = {
        "model": name,
        "accuracy": accuracy_score(y_eval, y_pred),
        "macro_f1": f1_score(y_eval, y_pred, average="macro"),
        "weighted_f1": f1_score(y_eval, y_pred, average="weighted"),
        "macro_recall": recall_score(y_eval, y_pred, average="macro"),
    }
    if y_proba is not None:
        y_bin = label_binarize(y_eval, classes=np.arange(len(class_names)))
        summary["macro_roc_auc_ovr"] = roc_auc_score(y_eval, y_proba, multi_class="ovr", average="macro")
        summary["macro_average_precision"] = average_precision_score(y_bin, y_proba, average="macro")

    report_df.to_csv(TABLE_DIR / f"{save_prefix}_classification_report.csv")
    pd.DataFrame(cm, index=class_names, columns=class_names).to_csv(TABLE_DIR / f"{save_prefix}_confusion_matrix_counts.csv")
    pd.DataFrame(cm_norm, index=class_names, columns=class_names).to_csv(TABLE_DIR / f"{save_prefix}_confusion_matrix_normalized.csv")

    return summary, report_df, cm, cm_norm, y_pred, y_proba

# Fit the selected primary leakage-free model on training data only, then evaluate once on the held-out test set.
best_primary_bundle = fit_final_bundle(
    best_primary_estimator,
    X_train_primary,
    y_train,
    use_sample_weight=best_primary_use_sample_weight,
    scale=False,
)
primary_summary, primary_report, primary_cm, primary_cm_norm, primary_pred, primary_proba = evaluate_model(
    best_primary_name, best_primary_bundle, X_test_primary, y_test, "primary_raw_signal"
)

# Fit the labeled comparison model with diagnostic flags, using the same train/test split.
flag_estimator = RandomForestClassifier(
    n_estimators=100,
    min_samples_leaf=2,
    max_features="sqrt",
    class_weight="balanced_subsample",
    random_state=RANDOM_STATE,
    n_jobs=-1,
)
flag_bundle = fit_final_bundle(flag_estimator, X_train_flags, y_train, use_sample_weight=False, scale=False)
flag_summary, flag_report, flag_cm, flag_cm_norm, flag_pred, flag_proba = evaluate_model(
    "Random Forest with vetting flags comparison", flag_bundle, X_test_flags, y_test, "comparison_with_flags"
)

final_summary = pd.DataFrame([primary_summary, flag_summary])
display(final_summary.round(4))
display(primary_report.round(4))
final_summary.to_csv(TABLE_DIR / "final_test_summary.csv", index=False)

fig, axes = plt.subplots(1, 2, figsize=(12, 4.8))
sns.heatmap(primary_cm, annot=True, fmt="d", cmap="Blues", xticklabels=class_names, yticklabels=class_names, ax=axes[0])
axes[0].set_title("Primary Raw-Signal Model\nConfusion Matrix Counts")
axes[0].set_xlabel("Predicted")
axes[0].set_ylabel("Actual")

sns.heatmap(primary_cm_norm, annot=True, fmt=".2f", cmap="Blues", xticklabels=class_names, yticklabels=class_names, ax=axes[1])
axes[1].set_title("Primary Raw-Signal Model\nNormalized by Actual Class")
axes[1].set_xlabel("Predicted")
axes[1].set_ylabel("Actual")
plt.tight_layout()
plt.savefig(FIG_DIR / "primary_confusion_matrices.png", bbox_inches="tight")
plt.show()
model accuracy macro_f1 weighted_f1 macro_recall macro_roc_auc_ovr macro_average_precision
0 Random Forest tuned raw-signal 0.8552 0.8234 0.8532 0.8201 0.9522 0.8801
1 Random Forest with vetting flags comparison 0.9211 0.8958 0.9196 0.8895 0.9819 0.9468
precision recall f1-score support
CANDIDATE 0.6970 0.6389 0.6667 396.0000
CONFIRMED 0.8957 0.9071 0.9014 549.0000
FALSE POSITIVE 0.8903 0.9143 0.9021 968.0000
accuracy 0.8552 0.8552 0.8552 0.8552
macro avg 0.8277 0.8201 0.8234 1913.0000
weighted avg 0.8518 0.8552 0.8532 1913.0000
No description has been provided for this image
In [13]:
fig, axes = plt.subplots(1, 2, figsize=(13, 5))
y_test_bin = label_binarize(y_test, classes=np.arange(len(class_names)))
for i, class_name in enumerate(class_names):
    PrecisionRecallDisplay.from_predictions(
        y_test_bin[:, i],
        primary_proba[:, i],
        name=class_name,
        ax=axes[0],
    )
    RocCurveDisplay.from_predictions(
        y_test_bin[:, i],
        primary_proba[:, i],
        name=class_name,
        ax=axes[1],
    )
axes[0].set_title("Precision-Recall Curves\nPrimary Raw-Signal Model")
axes[1].set_title("ROC Curves\nPrimary Raw-Signal Model")
plt.tight_layout()
plt.savefig(FIG_DIR / "primary_pr_roc_curves.png", bbox_inches="tight")
plt.show()

model_comparison = cv_results[["model", "macro_f1_mean", "macro_f1_std", "weighted_f1_mean", "macro_recall_mean"]].copy()
final_test_rows = final_summary[["model", "macro_f1", "weighted_f1", "macro_recall", "accuracy"]].copy()
display(model_comparison.round(4))
display(final_test_rows.round(4))
print("Progress: evaluation complete.")
No description has been provided for this image
model macro_f1_mean macro_f1_std weighted_f1_mean macro_recall_mean
3 Random Forest (with vetting flags comparison) 0.8946 0.0148 0.9183 0.8883
2 Random Forest (raw-signal) 0.8116 0.0175 0.8438 0.8074
4 XGBoost (raw-signal) 0.7821 0.0092 0.8087 0.8008
1 Decision Tree baseline (unbalanced) 0.7507 0.0094 0.7874 0.7482
0 Decision Tree baseline (balanced) 0.7491 0.0116 0.7758 0.7714
model macro_f1 weighted_f1 macro_recall accuracy
0 Random Forest tuned raw-signal 0.8234 0.8532 0.8201 0.8552
1 Random Forest with vetting flags comparison 0.8958 0.9196 0.8895 0.9211
Progress: evaluation complete.

The primary metric is macro F1 because each class matters, including unresolved candidates. The model with diagnostic flags is reported only as a comparison: its higher score, if present, reflects access to NASA vetting signals that a raw-signal model should not rely on.

7. Explainability¶

In [14]:
feature_names = list(X_train_primary.columns)
model_step = best_primary_bundle["estimator"]

if hasattr(model_step, "feature_importances_"):
    importances = pd.DataFrame({
        "feature": feature_names,
        "importance": model_step.feature_importances_,
    }).sort_values("importance", ascending=False)
else:
    perm = permutation_importance(
        model_step, transform_bundle(best_primary_bundle, X_test_primary), y_test, n_repeats=5, random_state=RANDOM_STATE,
        scoring="f1_macro", n_jobs=-1
    )
    importances = pd.DataFrame({
        "feature": feature_names[: len(perm.importances_mean)],
        "importance": perm.importances_mean,
    }).sort_values("importance", ascending=False)

display(importances.head(25))
importances.to_csv(TABLE_DIR / "primary_feature_importance.csv", index=False)

fig, ax = plt.subplots(figsize=(9, 7))
sns.barplot(data=importances.head(20), x="importance", y="feature", color="#59A14F", ax=ax)
ax.set_title("Top Feature Importances - Primary Raw-Signal Model")
ax.set_xlabel("Importance")
ax.set_ylabel("")
plt.tight_layout()
plt.savefig(FIG_DIR / "primary_feature_importance.png", bbox_inches="tight")
plt.show()
feature importance
93 koi_dikco_msky 0.033056
87 koi_dicco_msky 0.032714
133 koi_depth_rel_err1 0.022093
40 koi_model_snr 0.021507
18 koi_ror 0.021370
135 koi_depth_mean_rel_err 0.020360
98 log1p_koi_prad 0.020223
101 log1p_koi_model_snr 0.019833
108 transit_snr_proxy 0.019262
109 radius_ratio_from_prad_srad 0.018366
114 log_transit_duty_cycle 0.018341
24 koi_prad 0.018117
113 transit_duty_cycle 0.017213
130 koi_duration_rel_err1 0.014036
131 koi_duration_rel_err2 0.013898
132 koi_duration_mean_rel_err 0.013555
134 koi_depth_rel_err2 0.013035
156 koi_steff_mean_rel_err 0.012571
119 impact_plus_ror 0.012438
112 radius_ratio_abs_residual 0.012354
107 log1p_koi_max_mult_ev 0.011564
146 koi_impact_rel_err2 0.011531
154 koi_steff_rel_err1 0.011407
25 koi_prad_err1 0.011204
70 koi_fwm_stat_sig 0.011063
No description has been provided for this image
In [15]:
shap_status = "not_run"
try:
    import shap

    tree_model = best_primary_bundle["estimator"]
    X_test_transformed = transform_bundle(best_primary_bundle, X_test_primary)
    if not isinstance(X_test_transformed, np.ndarray):
        X_test_transformed = np.asarray(X_test_transformed)
    shap_sample_size = min(120, X_test_transformed.shape[0])
    X_shap = X_test_transformed[:shap_sample_size]

    explainer = shap.TreeExplainer(tree_model)
    shap_values = explainer.shap_values(X_shap)

    if isinstance(shap_values, list):
        mean_abs = np.mean([np.abs(vals).mean(axis=0) for vals in shap_values], axis=0)
        summary_values = shap_values
    else:
        arr = np.asarray(shap_values)
        if arr.ndim == 3:
            mean_abs = np.abs(arr).mean(axis=(0, 2))
            summary_values = arr[:, :, 0]
        else:
            mean_abs = np.abs(arr).mean(axis=0)
            summary_values = arr

    shap_importance = pd.DataFrame({
        "feature": feature_names,
        "mean_abs_shap": mean_abs,
    }).sort_values("mean_abs_shap", ascending=False)
    display(shap_importance.head(20))
    shap_importance.to_csv(TABLE_DIR / "primary_shap_importance.csv", index=False)

    fig, ax = plt.subplots(figsize=(9, 7))
    sns.barplot(data=shap_importance.head(20), x="mean_abs_shap", y="feature", color="#E15759", ax=ax)
    ax.set_title("Mean Absolute SHAP Importance - Primary Model")
    ax.set_xlabel("Mean absolute SHAP value")
    ax.set_ylabel("")
    plt.tight_layout()
    plt.savefig(FIG_DIR / "primary_shap_summary_bar.png", bbox_inches="tight")
    plt.show()

    # Individual prediction explanation using the top SHAP contributors for the first test row.
    if isinstance(shap_values, list):
        pred_class = int(primary_pred[0])
        row_vals = shap_values[pred_class][0]
        expected = explainer.expected_value[pred_class]
    else:
        arr = np.asarray(shap_values)
        pred_class = int(primary_pred[0])
        if arr.ndim == 3:
            row_vals = arr[0, :, pred_class]
            expected = np.asarray(explainer.expected_value)[pred_class]
        else:
            row_vals = arr[0]
            expected = explainer.expected_value

    row_explanation = pd.DataFrame({
        "feature": feature_names,
        "shap_value": row_vals,
        "abs_shap": np.abs(row_vals),
    }).sort_values("abs_shap", ascending=False).head(12)
    display(row_explanation[["feature", "shap_value"]])
    row_explanation.to_csv(TABLE_DIR / "individual_prediction_shap_top_features.csv", index=False)

    fig, ax = plt.subplots(figsize=(8, 5))
    sns.barplot(
        data=row_explanation.sort_values("shap_value"),
        x="shap_value",
        y="feature",
        palette=["#E15759" if v < 0 else "#59A14F" for v in row_explanation.sort_values("shap_value")["shap_value"]],
        ax=ax,
    )
    ax.axvline(0, color="black", linewidth=0.8)
    actual_label = class_names[y_test[0]]
    predicted_label = class_names[pred_class]
    ax.set_title(f"Individual SHAP Explanation\nActual: {actual_label}, Predicted: {predicted_label}")
    ax.set_xlabel("SHAP contribution")
    ax.set_ylabel("")
    plt.tight_layout()
    plt.savefig(FIG_DIR / "individual_prediction_shap.png", bbox_inches="tight")
    plt.show()
    shap_status = "completed"
except Exception as exc:
    shap_status = f"failed: {type(exc).__name__}: {exc}"
    print("SHAP could not be completed; feature importance remains available.")
    print(shap_status)

print("SHAP status:", shap_status)
print("Progress: explainability complete.")
C:\Users\Lenovo\OneDrive\Documents\New project 2\.venv_test\Lib\site-packages\shap\plots\colors\_colors.py:47: PendingDeprecationWarning: The set_bad function will be deprecated in a future version. Use cmap.with_extremes(bad=...) or Colormap(bad=...) instead.
  red_blue.set_bad(gray_rgb.tolist(), 1.0)
C:\Users\Lenovo\OneDrive\Documents\New project 2\.venv_test\Lib\site-packages\shap\plots\colors\_colors.py:48: PendingDeprecationWarning: The set_over function will be deprecated in a future version. Use cmap.with_extremes(over=...) or Colormap(over=...) instead.
  red_blue.set_over(gray_rgb.tolist(), 1.0)
C:\Users\Lenovo\OneDrive\Documents\New project 2\.venv_test\Lib\site-packages\shap\plots\colors\_colors.py:49: PendingDeprecationWarning: The set_under function will be deprecated in a future version. Use cmap.with_extremes(under=...) or Colormap(under=...) instead.
  red_blue.set_under(gray_rgb.tolist(), 1.0)  # "under" is incorrectly used instead of "bad" in the scatter plot
feature mean_abs_shap
93 koi_dikco_msky 0.024554
87 koi_dicco_msky 0.020659
133 koi_depth_rel_err1 0.014020
18 koi_ror 0.013260
98 log1p_koi_prad 0.012950
40 koi_model_snr 0.012349
101 log1p_koi_model_snr 0.012085
109 radius_ratio_from_prad_srad 0.011912
114 log_transit_duty_cycle 0.011188
24 koi_prad 0.011129
108 transit_snr_proxy 0.010866
113 transit_duty_cycle 0.010862
135 koi_depth_mean_rel_err 0.010736
70 koi_fwm_stat_sig 0.009739
39 koi_max_mult_ev 0.009147
107 log1p_koi_max_mult_ev 0.008897
134 koi_depth_rel_err2 0.008018
130 koi_duration_rel_err1 0.007885
131 koi_duration_rel_err2 0.007617
156 koi_steff_mean_rel_err 0.007615
No description has been provided for this image
feature shap_value
98 log1p_koi_prad 0.064642
18 koi_ror 0.059397
24 koi_prad 0.055353
109 radius_ratio_from_prad_srad 0.043880
25 koi_prad_err1 0.037452
146 koi_impact_rel_err2 0.024468
26 koi_prad_err2 0.024241
112 radius_ratio_abs_residual 0.020862
94 koi_dikco_msky_err 0.016711
133 koi_depth_rel_err1 0.016683
15 koi_depth 0.015503
40 koi_model_snr 0.013959
No description has been provided for this image
SHAP status: completed
Progress: explainability complete.

Plain-language explanation: the model looks for whether the dimming pattern behaves like a planet crossing a star. Real planets usually produce a repeatable, physically plausible dip whose size, duration, and orbit agree with the star's properties. False positives often look too large, too grazing, too hot, too uncertain, or inconsistent with what the star should allow. The primary model avoids using NASA's diagnostic flags so its score reflects what can be learned from the measurements themselves, not from copying the vetting pipeline.

8. Conclusion and Written Summary¶

In [16]:
# Build a metrics-grounded written summary using values computed above.
primary = final_summary.loc[final_summary["model"] == best_primary_name].iloc[0]
flags = final_summary.loc[final_summary["model"] == "Random Forest with vetting flags comparison"].iloc[0]
top_features = importances.head(8)["feature"].tolist()

summary_text = f'''
## Written Summary

### EDA & Data Cleaning Approach

The dataset contains {df.shape[0]:,} Kepler Objects of Interest and {df.shape[1]:,} columns from the Celesta-provided KOI cumulative table. The target is a three-class label: CONFIRMED, CANDIDATE, and FALSE POSITIVE. The class distribution is imbalanced, with {class_counts.loc[class_counts['class'].eq('FALSE POSITIVE'), 'count'].iloc[0]:,} false positives, {class_counts.loc[class_counts['class'].eq('CONFIRMED'), 'count'].iloc[0]:,} confirmed planets, and {class_counts.loc[class_counts['class'].eq('CANDIDATE'), 'count'].iloc[0]:,} candidates. Missingness was handled conservatively: columns with more than 50% missing values or no variance were dropped, leakage-prone identifiers and vetting metadata were removed, and remaining numeric values were median-imputed inside each cross-validation fold. This prevents validation rows from influencing preprocessing. The EDA showed strong skew and outliers in orbital period, radius, depth, insolation, and signal-to-noise, so the model uses log-transformed versions of these quantities rather than deleting suspicious rows.

### Model Choice Justification

I compared an unbalanced decision-tree baseline, a class-weighted decision-tree baseline, a Random Forest, and XGBoost. The paired decision-tree baselines show whether class weighting improves minority-class behavior instead of assuming it helps. Tree models were chosen as the primary family because this is a medium-sized tabular dataset with nonlinear relationships between transit geometry, stellar properties, and disposition. A neural network would be harder to explain and unnecessary for about ten thousand rows. Validation used five-fold stratified cross-validation on the training set, followed by a single held-out test evaluation for the final confusion matrix and curves. Hyperparameter tuning used a small manual search over Random Forest and XGBoost settings, optimizing macro F1 because the candidate class should not be hidden by overall accuracy.

### Key Findings

The selected primary model is {best_primary_name}. On the held-out test set it reached accuracy {primary['accuracy']:.3f}, macro F1 {primary['macro_f1']:.3f}, weighted F1 {primary['weighted_f1']:.3f}, and macro recall {primary['macro_recall']:.3f}. The most important raw-signal features included {', '.join(top_features[:6])}. These features make physical sense: false positives often have unusually large inferred radii, high impact parameters, inconsistent radius-depth relationships, or extreme signal properties. A separate comparison model that included the NASA diagnostic false-positive flags reached macro F1 {flags['macro_f1']:.3f}. That comparison is useful but not treated as the main result because those flags are products of the vetting process. Removing them gives a more honest estimate of what a model can learn from measured transit and stellar properties.

### Explaining It to a Non-Technical Audience

The model is trying to decide whether a tiny repeated dimming of a star looks like a planet passing in front of it. It checks whether the dip is the right size, lasts for a believable amount of time, repeats in a stable orbit, and matches what we know about the star. If the dip looks too deep, too oddly shaped, too uncertain, or physically inconsistent, the model becomes more suspicious that it is a false alarm rather than a real planet.

### Limitations and Future Work

The labels come from an already-vetted astronomical catalog, so the model learns from historical Kepler decisions rather than discovering planets independently. Some high-performing columns are downstream diagnostic products, which is why they are excluded from the primary model and shown only as a comparison. Future work could test the pipeline on a newer mission such as TESS, add calibrated astrophysical priors, and compare binary and three-class formulations depending on the competition's final scoring rules.
'''

print(summary_text)
(OUTPUT_DIR / "written_summary.md").write_text(summary_text.strip() + "\n", encoding="utf-8")

devpost_checklist = pd.DataFrame([
    {"item": "Project title + one-line description", "status": "Ready: title in README"},
    {"item": "Team members listed", "status": "Ready: Karthik M A"},
    {"item": "Public GitHub repo linked", "status": "Ready: https://github.com/sleader3221-dot/India_high_school"},
    {"item": "Screenshot/visualization", "status": "Ready: confusion matrix and SHAP/importance figures saved"},
    {"item": "Metrics table and confusion matrix visible", "status": "Ready in notebook and outputs/tables"},
    {"item": "Written summary included", "status": "Ready in notebook and outputs/written_summary.md"},
    {"item": "English submission", "status": "Ready"},
])
display(devpost_checklist)
devpost_checklist.to_csv(TABLE_DIR / "devpost_checklist.csv", index=False)
print("Progress: write-up complete.")
## Written Summary

### EDA & Data Cleaning Approach

The dataset contains 9,564 Kepler Objects of Interest and 140 columns from the Celesta-provided KOI cumulative table. The target is a three-class label: CONFIRMED, CANDIDATE, and FALSE POSITIVE. The class distribution is imbalanced, with 4,839 false positives, 2,747 confirmed planets, and 1,978 candidates. Missingness was handled conservatively: columns with more than 50% missing values or no variance were dropped, leakage-prone identifiers and vetting metadata were removed, and remaining numeric values were median-imputed inside each cross-validation fold. This prevents validation rows from influencing preprocessing. The EDA showed strong skew and outliers in orbital period, radius, depth, insolation, and signal-to-noise, so the model uses log-transformed versions of these quantities rather than deleting suspicious rows.

### Model Choice Justification

I compared an unbalanced decision-tree baseline, a class-weighted decision-tree baseline, a Random Forest, and XGBoost. The paired decision-tree baselines show whether class weighting improves minority-class behavior instead of assuming it helps. Tree models were chosen as the primary family because this is a medium-sized tabular dataset with nonlinear relationships between transit geometry, stellar properties, and disposition. A neural network would be harder to explain and unnecessary for about ten thousand rows. Validation used five-fold stratified cross-validation on the training set, followed by a single held-out test evaluation for the final confusion matrix and curves. Hyperparameter tuning used a small manual search over Random Forest and XGBoost settings, optimizing macro F1 because the candidate class should not be hidden by overall accuracy.

### Key Findings

The selected primary model is Random Forest tuned raw-signal. On the held-out test set it reached accuracy 0.855, macro F1 0.823, weighted F1 0.853, and macro recall 0.820. The most important raw-signal features included koi_dikco_msky, koi_dicco_msky, koi_depth_rel_err1, koi_model_snr, koi_ror, koi_depth_mean_rel_err. These features make physical sense: false positives often have unusually large inferred radii, high impact parameters, inconsistent radius-depth relationships, or extreme signal properties. A separate comparison model that included the NASA diagnostic false-positive flags reached macro F1 0.896. That comparison is useful but not treated as the main result because those flags are products of the vetting process. Removing them gives a more honest estimate of what a model can learn from measured transit and stellar properties.

### Explaining It to a Non-Technical Audience

The model is trying to decide whether a tiny repeated dimming of a star looks like a planet passing in front of it. It checks whether the dip is the right size, lasts for a believable amount of time, repeats in a stable orbit, and matches what we know about the star. If the dip looks too deep, too oddly shaped, too uncertain, or physically inconsistent, the model becomes more suspicious that it is a false alarm rather than a real planet.

### Limitations and Future Work

The labels come from an already-vetted astronomical catalog, so the model learns from historical Kepler decisions rather than discovering planets independently. Some high-performing columns are downstream diagnostic products, which is why they are excluded from the primary model and shown only as a comparison. Future work could test the pipeline on a newer mission such as TESS, add calibrated astrophysical priors, and compare binary and three-class formulations depending on the competition's final scoring rules.

item status
0 Project title + one-line description Ready: title in README
1 Team members listed Ready: Karthik M A
2 Public GitHub repo linked Ready: https://github.com/sleader3221-dot/Indi...
3 Screenshot/visualization Ready: confusion matrix and SHAP/importance fi...
4 Metrics table and confusion matrix visible Ready in notebook and outputs/tables
5 Written summary included Ready in notebook and outputs/written_summary.md
6 English submission Ready
Progress: write-up complete.
India High School Exoplanet Data Challenge | Celesta · NASA Kepler Mission | By Karthik M A
GitHub Repo Random Forest · XGBoost · scikit-learn