India High School Exoplanet Data Challenge¶
Celesta - NASA Kepler Exoplanet Classification
This notebook builds a leakage-audited, reproducible machine learning submission for classifying Kepler Objects of Interest into CONFIRMED, CANDIDATE, and FALSE POSITIVE.
The primary result intentionally excludes NASA vetting verdict fields and diagnostic false-positive flags. A separate comparison model includes the diagnostic flags to show how much apparent performance comes from pipeline-vetting features rather than raw transit and stellar measurements.
1. Data Loading¶
import os
import sys
import json
import math
import random
import warnings
from pathlib import Path
import numpy as np
import pandas as pd
import matplotlib.pyplot as plt
import seaborn as sns
from sklearn.base import BaseEstimator, TransformerMixin, clone
from sklearn.compose import ColumnTransformer
from sklearn.ensemble import RandomForestClassifier
from sklearn.impute import SimpleImputer
from sklearn.inspection import permutation_importance
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import (
accuracy_score,
average_precision_score,
classification_report,
confusion_matrix,
f1_score,
precision_recall_fscore_support,
precision_recall_curve,
recall_score,
roc_auc_score,
RocCurveDisplay,
PrecisionRecallDisplay,
)
from sklearn.model_selection import (
RandomizedSearchCV,
StratifiedKFold,
cross_validate,
train_test_split,
)
from sklearn.pipeline import Pipeline
from sklearn.tree import DecisionTreeClassifier
from sklearn.preprocessing import LabelEncoder, label_binarize, OneHotEncoder, StandardScaler
from sklearn.utils.class_weight import compute_sample_weight
from xgboost import XGBClassifier
warnings.filterwarnings("ignore", category=FutureWarning)
warnings.filterwarnings("ignore", category=UserWarning)
RANDOM_STATE = 42
np.random.seed(RANDOM_STATE)
random.seed(RANDOM_STATE)
ROOT = Path.cwd()
DATA_PATH = ROOT / "KOI_Cumulative_clean.csv"
GUIDE_PATH = ROOT / "Celesta_Dataset_Guide.pdf"
OUTPUT_DIR = ROOT / "outputs"
FIG_DIR = OUTPUT_DIR / "figures"
TABLE_DIR = OUTPUT_DIR / "tables"
for directory in [OUTPUT_DIR, FIG_DIR, TABLE_DIR]:
directory.mkdir(parents=True, exist_ok=True)
sns.set_theme(style="whitegrid", context="notebook")
plt.rcParams["figure.dpi"] = 120
plt.rcParams["savefig.dpi"] = 180
plt.rcParams["axes.titleweight"] = "bold"
df = pd.read_csv(DATA_PATH)
print(f"Loaded {DATA_PATH.name}: {df.shape[0]:,} rows x {df.shape[1]:,} columns")
display(df.head(3))
Loaded KOI_Cumulative_clean.csv: 9,564 rows x 140 columns
| rowid | kepid | kepoi_name | kepler_name | koi_disposition | koi_vet_stat | koi_vet_date | koi_pdisposition | koi_fpflag_nt | koi_fpflag_ss | ... | koi_dicco_mdec | koi_dicco_mdec_err | koi_dicco_msky | koi_dicco_msky_err | koi_dikco_mra | koi_dikco_mra_err | koi_dikco_mdec | koi_dikco_mdec_err | koi_dikco_msky | koi_dikco_msky_err | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 0 | 1 | 10797460 | K00752.01 | Kepler-227 b | CONFIRMED | Done | 2018-08-16 | CANDIDATE | 0 | 0 | ... | 0.200 | 0.16 | 0.200 | 0.170 | 0.080 | 0.130 | 0.310 | 0.170 | 0.320 | 0.160 |
| 1 | 2 | 10797460 | K00752.02 | Kepler-227 c | CONFIRMED | Done | 2018-08-16 | CANDIDATE | 0 | 0 | ... | 0.000 | 0.48 | 0.390 | 0.360 | 0.490 | 0.340 | 0.120 | 0.730 | 0.500 | 0.450 |
| 2 | 3 | 10811496 | K00753.01 | NaN | CANDIDATE | Done | 2018-08-16 | CANDIDATE | 0 | 0 | ... | -0.034 | 0.07 | 0.042 | 0.072 | 0.002 | 0.071 | -0.027 | 0.074 | 0.027 | 0.074 |
3 rows × 140 columns
profile = pd.DataFrame({
"column": df.columns,
"dtype": [str(t) for t in df.dtypes],
"missing_count": df.isna().sum().values,
"missing_pct": (df.isna().mean().values * 100).round(3),
"n_unique": [df[c].nunique(dropna=True) for c in df.columns],
}).sort_values(["missing_pct", "column"], ascending=[False, True])
shape_summary = pd.DataFrame({
"item": ["rows", "columns", "duplicate_rows", "target_missing"],
"value": [df.shape[0], df.shape[1], df.duplicated().sum(), df["koi_disposition"].isna().sum()],
})
class_counts = (
df["koi_disposition"]
.value_counts()
.rename_axis("class")
.reset_index(name="count")
)
class_counts["share_pct"] = (class_counts["count"] / len(df) * 100).round(2)
display(shape_summary)
display(class_counts)
display(profile.head(30))
profile.to_csv(TABLE_DIR / "column_profile.csv", index=False)
class_counts.to_csv(TABLE_DIR / "class_counts.csv", index=False)
print("Progress: data loading and profiling complete.")
| item | value | |
|---|---|---|
| 0 | rows | 9564 |
| 1 | columns | 140 |
| 2 | duplicate_rows | 0 |
| 3 | target_missing | 0 |
| class | count | share_pct | |
|---|---|---|---|
| 0 | FALSE POSITIVE | 4839 | 50.60 |
| 1 | CONFIRMED | 2747 | 28.72 |
| 2 | CANDIDATE | 1978 | 20.68 |
| column | dtype | missing_count | missing_pct | n_unique | |
|---|---|---|---|---|---|
| 24 | koi_eccen_err1 | float64 | 9564 | 100.000 | 0 |
| 25 | koi_eccen_err2 | float64 | 9564 | 100.000 | 0 |
| 55 | koi_incl_err1 | float64 | 9564 | 100.000 | 0 |
| 56 | koi_incl_err2 | float64 | 9564 | 100.000 | 0 |
| 35 | koi_ingress | float64 | 9564 | 100.000 | 0 |
| 36 | koi_ingress_err1 | float64 | 9564 | 100.000 | 0 |
| 37 | koi_ingress_err2 | float64 | 9564 | 100.000 | 0 |
| 26 | koi_longp | float64 | 9564 | 100.000 | 0 |
| 27 | koi_longp_err1 | float64 | 9564 | 100.000 | 0 |
| 28 | koi_longp_err2 | float64 | 9564 | 100.000 | 0 |
| 83 | koi_model_chisq | float64 | 9564 | 100.000 | 0 |
| 82 | koi_model_dof | float64 | 9564 | 100.000 | 0 |
| 101 | koi_sage | float64 | 9564 | 100.000 | 0 |
| 102 | koi_sage_err1 | float64 | 9564 | 100.000 | 0 |
| 103 | koi_sage_err2 | float64 | 9564 | 100.000 | 0 |
| 52 | koi_sma_err1 | float64 | 9564 | 100.000 | 0 |
| 53 | koi_sma_err2 | float64 | 9564 | 100.000 | 0 |
| 58 | koi_teq_err1 | float64 | 9564 | 100.000 | 0 |
| 59 | koi_teq_err2 | float64 | 9564 | 100.000 | 0 |
| 3 | kepler_name | str | 6817 | 71.278 | 2747 |
| 80 | koi_bin_oedp_sig | float64 | 1510 | 15.788 | 4875 |
| 13 | koi_comment | str | 1209 | 12.641 | 995 |
| 73 | koi_max_mult_ev | float64 | 1142 | 11.941 | 8422 |
| 72 | koi_max_sngle_ev | float64 | 1142 | 11.941 | 8421 |
| 76 | koi_num_transits | float64 | 1142 | 11.941 | 1627 |
| 79 | koi_quarters | object | 1142 | 11.941 | 198 |
| 115 | koi_fwm_stat_sig | float64 | 1076 | 11.251 | 929 |
| 124 | koi_fwm_prao | float64 | 830 | 8.678 | 1943 |
| 125 | koi_fwm_prao_err | float64 | 830 | 8.678 | 475 |
| 126 | koi_fwm_pdeco | float64 | 817 | 8.542 | 1846 |
Progress: data loading and profiling complete.
The official dataset guide describes this file as the Cumulative Kepler Objects of Interest table with 9,564 rows and 140 columns. The target has three classes, so this notebook treats the problem as a three-class classification task. The dataset is imbalanced: false positives are about half of the table, while candidates are the smallest class.
2. Leakage Audit¶
leakage_audit = pd.DataFrame([
{"column_or_group": "koi_disposition", "decision": "target only", "reason": "The label being predicted; never used as an input feature."},
{"column_or_group": "koi_score", "decision": "not present", "reason": "The guide says this target-linked score was removed by organizers."},
{"column_or_group": "koi_pdisposition", "decision": "drop", "reason": "NASA pipeline pre-disposition; nearly copies the false-positive decision."},
{"column_or_group": "koi_fpflag_nt, koi_fpflag_ss, koi_fpflag_co, koi_fpflag_ec", "decision": "conditional", "reason": "Diagnostic vetting flags. Excluded from primary raw-signal model; included only in a clearly labeled comparison model."},
{"column_or_group": "rowid, kepid, kepoi_name, kepler_name", "decision": "drop", "reason": "Identifiers. kepler_name is especially leaky because named planets are almost always confirmed."},
{"column_or_group": "koi_comment", "decision": "drop", "reason": "Contains explicit vetting comments such as secondary eclipses and V-shaped transit notes."},
{"column_or_group": "koi_vet_stat, koi_vet_date, koi_disp_prov", "decision": "drop", "reason": "Vetting/status metadata; also constant in this dataset."},
{"column_or_group": "koi_datalink_dvr, koi_datalink_dvs", "decision": "drop", "reason": "Near-unique report links, not physical measurements."},
{"column_or_group": "koi_fittype, koi_parm_prov, koi_tce_delivname, koi_sparprov", "decision": "drop from primary", "reason": "Processing/provenance metadata; may reflect catalog history instead of astrophysics."},
{"column_or_group": "all-missing and constant columns", "decision": "drop", "reason": "No usable statistical signal."},
{"column_or_group": "numeric transit/orbital/stellar/photometric/centroid columns", "decision": "keep", "reason": "Physical measurements suitable for a leakage-free raw-signal model after imputation."},
])
display(leakage_audit)
checks = {}
if "kepler_name" in df.columns:
checks["kepler_name_present_vs_target"] = pd.crosstab(df["kepler_name"].notna(), df["koi_disposition"])
if "koi_pdisposition" in df.columns:
checks["koi_pdisposition_vs_target"] = pd.crosstab(df["koi_pdisposition"], df["koi_disposition"])
for flag in ["koi_fpflag_nt", "koi_fpflag_ss", "koi_fpflag_co", "koi_fpflag_ec"]:
if flag in df.columns:
checks[f"{flag}_vs_target"] = pd.crosstab((pd.to_numeric(df[flag], errors="coerce") > 0).astype(int), df["koi_disposition"])
for name, table in checks.items():
print(f"\n{name}")
display(table)
leakage_audit.to_csv(TABLE_DIR / "leakage_audit.csv", index=False)
print("Progress: leakage audit complete. Primary model will exclude target-leaking and vetting-derived fields.")
| column_or_group | decision | reason | |
|---|---|---|---|
| 0 | koi_disposition | target only | The label being predicted; never used as an in... |
| 1 | koi_score | not present | The guide says this target-linked score was re... |
| 2 | koi_pdisposition | drop | NASA pipeline pre-disposition; nearly copies t... |
| 3 | koi_fpflag_nt, koi_fpflag_ss, koi_fpflag_co, k... | conditional | Diagnostic vetting flags. Excluded from primar... |
| 4 | rowid, kepid, kepoi_name, kepler_name | drop | Identifiers. kepler_name is especially leaky b... |
| 5 | koi_comment | drop | Contains explicit vetting comments such as sec... |
| 6 | koi_vet_stat, koi_vet_date, koi_disp_prov | drop | Vetting/status metadata; also constant in this... |
| 7 | koi_datalink_dvr, koi_datalink_dvs | drop | Near-unique report links, not physical measure... |
| 8 | koi_fittype, koi_parm_prov, koi_tce_delivname,... | drop from primary | Processing/provenance metadata; may reflect ca... |
| 9 | all-missing and constant columns | drop | No usable statistical signal. |
| 10 | numeric transit/orbital/stellar/photometric/ce... | keep | Physical measurements suitable for a leakage-f... |
kepler_name_present_vs_target
| koi_disposition | CANDIDATE | CONFIRMED | FALSE POSITIVE |
|---|---|---|---|
| kepler_name | |||
| False | 1978 | 2 | 4837 |
| True | 0 | 2745 | 2 |
koi_pdisposition_vs_target
| koi_disposition | CANDIDATE | CONFIRMED | FALSE POSITIVE |
|---|---|---|---|
| koi_pdisposition | |||
| CANDIDATE | 1978 | 2738 | 1 |
| FALSE POSITIVE | 0 | 9 | 4838 |
koi_fpflag_nt_vs_target
| koi_disposition | CANDIDATE | CONFIRMED | FALSE POSITIVE |
|---|---|---|---|
| koi_fpflag_nt | |||
| 0 | 1977 | 2743 | 3313 |
| 1 | 1 | 4 | 1526 |
koi_fpflag_ss_vs_target
| koi_disposition | CANDIDATE | CONFIRMED | FALSE POSITIVE |
|---|---|---|---|
| koi_fpflag_ss | |||
| 0 | 1976 | 2733 | 2629 |
| 1 | 2 | 14 | 2210 |
koi_fpflag_co_vs_target
| koi_disposition | CANDIDATE | CONFIRMED | FALSE POSITIVE |
|---|---|---|---|
| koi_fpflag_co | |||
| 0 | 1978 | 2747 | 2950 |
| 1 | 0 | 0 | 1889 |
koi_fpflag_ec_vs_target
| koi_disposition | CANDIDATE | CONFIRMED | FALSE POSITIVE |
|---|---|---|---|
| koi_fpflag_ec | |||
| 0 | 1978 | 2747 | 3691 |
| 1 | 0 | 0 | 1148 |
Progress: leakage audit complete. Primary model will exclude target-leaking and vetting-derived fields.
The strongest leakage risks are not subtle. koi_pdisposition is already a NASA pipeline verdict, and kepler_name is almost a proxy for being confirmed because confirmed exoplanets receive Kepler names. The four false-positive flags are physically meaningful diagnostics, but they are downstream vetting outputs rather than raw measurements. They are therefore excluded from the primary model and used only in a labeled comparison.
3. EDA¶
fig, ax = plt.subplots(figsize=(7, 4))
order = class_counts["class"].tolist()
sns.barplot(data=class_counts, x="class", y="count", order=order, palette="Set2", ax=ax)
ax.set_title("Class Distribution")
ax.set_xlabel("")
ax.set_ylabel("Rows")
for i, row in class_counts.reset_index(drop=True).iterrows():
ax.text(i, row["count"] + 80, f'{row["count"]:,}\n{row["share_pct"]:.1f}%', ha="center", va="bottom", fontsize=9)
plt.tight_layout()
plt.savefig(FIG_DIR / "class_distribution.png", bbox_inches="tight")
plt.show()
False positives are the largest class, so a naive model can look accurate while ignoring the smaller CANDIDATE class. This is why macro F1 and per-class recall are more informative than accuracy alone.
missing_nonzero = profile[profile["missing_count"] > 0].copy()
missing_top = missing_nonzero.head(35)
fig, ax = plt.subplots(figsize=(9, 7))
sns.barplot(data=missing_top, y="column", x="missing_pct", color="#4C78A8", ax=ax)
ax.axvline(50, color="red", linestyle="--", linewidth=1, label="50% drop threshold")
ax.set_title("Top Missingness by Column")
ax.set_xlabel("Missing rows (%)")
ax.set_ylabel("")
ax.legend()
plt.tight_layout()
plt.savefig(FIG_DIR / "missingness_top_columns.png", bbox_inches="tight")
plt.show()
missing_summary = pd.DataFrame({
"bucket": ["0%", "0-5%", "5-30%", "30-50%", ">50%"],
"columns": [
int((profile["missing_pct"] == 0).sum()),
int(((profile["missing_pct"] > 0) & (profile["missing_pct"] <= 5)).sum()),
int(((profile["missing_pct"] > 5) & (profile["missing_pct"] <= 30)).sum()),
int(((profile["missing_pct"] > 30) & (profile["missing_pct"] <= 50)).sum()),
int((profile["missing_pct"] > 50).sum()),
],
})
display(missing_summary)
missing_summary.to_csv(TABLE_DIR / "missingness_summary.csv", index=False)
| bucket | columns | |
|---|---|---|
| 0 | 0% | 21 |
| 1 | 0-5% | 70 |
| 2 | 5-30% | 29 |
| 3 | 30-50% | 0 |
| 4 | >50% | 20 |
Columns above 50% missing are dropped unless they are needed only for an audit. In this file, the high-missing group is mostly entirely empty uncertainty fields. Lower-missing physical columns are imputed inside cross-validation folds with median imputation for numeric features and most-frequent imputation for categorical features, so validation data never influences training preprocessing.
eda_cols = ["koi_period", "koi_duration", "koi_depth", "koi_prad", "koi_teq", "koi_insol"]
plot_df = df[["koi_disposition"] + eda_cols].copy()
for col in eda_cols:
plot_df[f"log1p_{col}"] = np.log1p(pd.to_numeric(plot_df[col], errors="coerce").clip(lower=0))
fig, axes = plt.subplots(2, 3, figsize=(14, 8))
for ax, col in zip(axes.ravel(), eda_cols):
sns.kdeplot(
data=plot_df,
x=f"log1p_{col}",
hue="koi_disposition",
common_norm=False,
fill=True,
alpha=0.25,
linewidth=1.2,
ax=ax,
)
ax.set_title(f"{col} by class (log1p)")
ax.set_xlabel(f"log1p({col})")
ax.set_ylabel("Density")
plt.tight_layout()
plt.savefig(FIG_DIR / "key_feature_distributions.png", bbox_inches="tight")
plt.show()
class_medians = df.groupby("koi_disposition")[eda_cols + ["koi_impact", "koi_model_snr"]].median(numeric_only=True).round(3)
display(class_medians)
class_medians.to_csv(TABLE_DIR / "key_feature_class_medians.csv")
| koi_period | koi_duration | koi_depth | koi_prad | koi_teq | koi_insol | koi_impact | koi_model_snr | |
|---|---|---|---|---|---|---|---|---|
| koi_disposition | ||||||||
| CANDIDATE | 20.044 | 3.606 | 242.00 | 1.74 | 691.5 | 53.990 | 0.359 | 11.00 |
| CONFIRMED | 11.350 | 3.491 | 448.60 | 2.16 | 777.0 | 86.245 | 0.412 | 28.60 |
| FALSE POSITIVE | 5.244 | 4.057 | 575.95 | 8.97 | 1152.5 | 424.010 | 0.675 | 34.25 |
Planet radius, insolation, impact parameter, and signal-to-noise show visible class separation. False positives have a much larger median planetary radius than confirmed planets and candidates, which is consistent with many false positives being stellar eclipses or bad transit fits rather than small planets.
corr_exclude = {
"rowid", "kepid", "koi_disposition", "koi_pdisposition",
"koi_fpflag_nt", "koi_fpflag_ss", "koi_fpflag_co", "koi_fpflag_ec",
}
numeric_for_corr = [
c for c in df.select_dtypes(include=[np.number]).columns
if c not in corr_exclude and df[c].isna().mean() < 0.50 and df[c].nunique(dropna=True) > 1
]
# Keep the heatmap readable: use representative high-importance science columns rather than every numeric column.
corr_cols = [
c for c in [
"koi_period", "koi_duration", "koi_depth", "koi_ror", "koi_srho", "koi_prad", "koi_sma",
"koi_incl", "koi_teq", "koi_insol", "koi_dor", "koi_model_snr", "koi_steff",
"koi_slogg", "koi_srad", "koi_smass", "koi_kepmag", "koi_gmag", "koi_rmag",
"koi_jmag", "koi_hmag", "koi_kmag", "koi_fwm_stat_sig", "koi_dicco_msky", "koi_dikco_msky",
]
if c in numeric_for_corr
]
corr = df[corr_cols].corr(numeric_only=True)
fig, ax = plt.subplots(figsize=(12, 10))
sns.heatmap(corr, cmap="vlag", center=0, linewidths=0.2, ax=ax)
ax.set_title("Correlation Heatmap of Physical Numeric Features")
plt.tight_layout()
plt.savefig(FIG_DIR / "correlation_heatmap.png", bbox_inches="tight")
plt.show()
pairs = []
abs_corr = corr.abs()
for i, a in enumerate(corr_cols):
for b in corr_cols[i + 1:]:
val = abs_corr.loc[a, b]
if pd.notna(val) and val >= 0.90:
pairs.append({"feature_a": a, "feature_b": b, "abs_corr": round(float(val), 4)})
high_corr_pairs = pd.DataFrame(pairs).sort_values("abs_corr", ascending=False)
display(high_corr_pairs.head(20))
high_corr_pairs.to_csv(TABLE_DIR / "high_correlation_pairs.csv", index=False)
| feature_a | feature_b | abs_corr | |
|---|---|---|---|
| 2 | koi_kepmag | koi_rmag | 0.9993 |
| 12 | koi_hmag | koi_kmag | 0.9978 |
| 10 | koi_jmag | koi_hmag | 0.9931 |
| 11 | koi_jmag | koi_kmag | 0.9901 |
| 0 | koi_period | koi_dor | 0.9894 |
| 6 | koi_gmag | koi_rmag | 0.9862 |
| 1 | koi_kepmag | koi_gmag | 0.9793 |
| 3 | koi_kepmag | koi_jmag | 0.9420 |
| 7 | koi_rmag | koi_jmag | 0.9410 |
| 4 | koi_kepmag | koi_hmag | 0.9184 |
| 8 | koi_rmag | koi_hmag | 0.9159 |
| 5 | koi_kepmag | koi_kmag | 0.9121 |
| 9 | koi_rmag | koi_kmag | 0.9086 |
Several feature groups are strongly correlated because the catalog reports related physical quantities and paired uncertainty bounds. Tree models can handle this better than linear models, but feature importance must be interpreted as grouped evidence rather than as independent causal effects.
outlier_cols = ["koi_period", "koi_duration", "koi_depth", "koi_prad", "koi_teq", "koi_insol", "koi_impact", "koi_model_snr"]
outlier_rows = []
for col in outlier_cols:
s = pd.to_numeric(df[col], errors="coerce")
p01, p99 = s.quantile([0.01, 0.99])
outlier_rows.append({
"feature": col,
"min": s.min(),
"p01": p01,
"median": s.median(),
"p99": p99,
"max": s.max(),
"below_p01": int((s < p01).sum()),
"above_p99": int((s > p99).sum()),
"missing": int(s.isna().sum()),
})
outlier_table = pd.DataFrame(outlier_rows)
display(outlier_table.round(4))
outlier_table.to_csv(TABLE_DIR / "outlier_scan.csv", index=False)
long_outliers = df[["koi_disposition"] + outlier_cols].melt(id_vars="koi_disposition", var_name="feature", value_name="value")
long_outliers["log1p_value"] = np.log1p(pd.to_numeric(long_outliers["value"], errors="coerce").clip(lower=0))
fig, ax = plt.subplots(figsize=(13, 6))
sns.boxplot(data=long_outliers, x="feature", y="log1p_value", hue="koi_disposition", fliersize=1.2, ax=ax)
ax.set_title("Outlier Scan on Log-Scaled Physical Features")
ax.set_xlabel("")
ax.set_ylabel("log1p(value)")
ax.tick_params(axis="x", rotation=25)
plt.tight_layout()
plt.savefig(FIG_DIR / "outlier_boxplots.png", bbox_inches="tight")
plt.show()
| feature | min | p01 | median | p99 | max | below_p01 | above_p99 | missing | |
|---|---|---|---|---|---|---|---|---|---|
| 0 | koi_period | 0.2418 | 0.5277 | 9.7528 | 550.9979 | 1.299958e+05 | 96 | 96 | 0 |
| 1 | koi_duration | 0.0520 | 0.7987 | 3.7926 | 30.0475 | 1.385400e+02 | 96 | 96 | 0 |
| 2 | koi_depth | 0.0000 | 24.1000 | 421.1000 | 416740.0000 | 1.541400e+06 | 92 | 92 | 363 |
| 3 | koi_prad | 0.0800 | 0.5200 | 2.3900 | 369.2000 | 2.003460e+05 | 87 | 92 | 363 |
| 4 | koi_teq | 25.0000 | 196.0000 | 878.0000 | 4106.0000 | 1.466700e+04 | 87 | 92 | 363 |
| 5 | koi_insol | 0.0000 | 0.3500 | 141.6000 | 70297.1642 | 1.094755e+07 | 87 | 93 | 321 |
| 6 | koi_impact | 0.0000 | 0.0030 | 0.5370 | 1.7826 | 1.008060e+02 | 78 | 92 | 363 |
| 7 | koi_model_snr | 0.0000 | 5.0000 | 23.0000 | 4355.5000 | 9.054700e+03 | 88 | 92 | 363 |
The catalog contains extreme fitted values, including enormous planet radii and orbital periods. These are not automatically removed because implausible fitted values are themselves useful evidence that a signal may be a false positive. The modeling pipeline uses log transforms and robust tree models rather than deleting scientifically suspicious rows.
4. Feature Engineering¶
TARGET = "koi_disposition"
FLAG_COLUMNS = ["koi_fpflag_nt", "koi_fpflag_ss", "koi_fpflag_co", "koi_fpflag_ec"]
ALWAYS_DROP = [
TARGET, "rowid", "kepid", "kepoi_name", "kepler_name",
"koi_pdisposition", "koi_score", "koi_comment",
"koi_vet_stat", "koi_vet_date", "koi_disp_prov",
"koi_datalink_dvr", "koi_datalink_dvs",
"koi_fittype", "koi_parm_prov", "koi_tce_delivname", "koi_sparprov",
"koi_quarters",
]
class FeatureEngineer(BaseEstimator, TransformerMixin):
def __init__(self, include_flags=False):
self.include_flags = include_flags
self.drop_columns_ = None
self.input_columns_ = None
self.output_columns_ = None
def fit(self, X, y=None):
X = X.copy()
self.input_columns_ = list(X.columns)
missing_pct = X.isna().mean()
nunique = X.nunique(dropna=True)
drop = set([c for c in ALWAYS_DROP if c in X.columns])
if not self.include_flags:
drop.update([c for c in FLAG_COLUMNS if c in X.columns])
drop.update(missing_pct[missing_pct > 0.50].index.tolist())
drop.update(nunique[nunique <= 1].index.tolist())
self.drop_columns_ = sorted(drop)
Xt = self._transform_frame(X)
self.output_columns_ = list(Xt.columns)
return self
@staticmethod
def _safe_num(df, col):
if col not in df.columns:
return pd.Series(np.nan, index=df.index, dtype="float64")
return pd.to_numeric(df[col], errors="coerce")
@staticmethod
def _signed_abs_err(df, col):
if col not in df.columns:
return pd.Series(np.nan, index=df.index, dtype="float64")
return pd.to_numeric(df[col], errors="coerce").abs()
def _add_features(self, X):
X = X.copy()
eps = 1e-9
positive_log_cols = [
"koi_period", "koi_duration", "koi_depth", "koi_prad", "koi_teq", "koi_insol",
"koi_model_snr", "koi_srad", "koi_smass", "koi_srho", "koi_dor",
"koi_max_sngle_ev", "koi_max_mult_ev",
]
for col in positive_log_cols:
if col in X.columns:
vals = pd.to_numeric(X[col], errors="coerce").clip(lower=0)
X[f"log1p_{col}"] = np.log1p(vals)
depth = self._safe_num(X, "koi_depth")
depth_err = self._signed_abs_err(X, "koi_depth_err1")
X["transit_snr_proxy"] = depth / (depth_err + eps)
prad = self._safe_num(X, "koi_prad")
srad = self._safe_num(X, "koi_srad")
radius_ratio = prad / (srad * 109.1 + eps)
depth_ratio = np.sqrt((depth.clip(lower=0) / 1_000_000).clip(lower=0))
X["radius_ratio_from_prad_srad"] = radius_ratio
X["depth_implied_radius_ratio"] = depth_ratio
X["radius_ratio_residual"] = radius_ratio - depth_ratio
X["radius_ratio_abs_residual"] = (radius_ratio - depth_ratio).abs()
period = self._safe_num(X, "koi_period")
duration = self._safe_num(X, "koi_duration")
X["transit_duty_cycle"] = duration / (24 * period + eps)
X["log_transit_duty_cycle"] = np.log1p(X["transit_duty_cycle"].clip(lower=0))
teq = self._safe_num(X, "koi_teq")
insol = self._safe_num(X, "koi_insol")
# Equilibrium temperature scales approximately with the fourth root of insolation.
expected_teq_from_insol = 278.0 * np.power(insol.clip(lower=0) + eps, 0.25)
X["teq_insol_ratio"] = teq / (expected_teq_from_insol + eps)
X["teq_insol_abs_residual"] = (teq - expected_teq_from_insol).abs()
smass = self._safe_num(X, "koi_smass")
srho = self._safe_num(X, "koi_srho")
catalog_density_proxy = smass / (np.power(srad, 3) + eps)
X["stellar_density_ratio"] = srho / (catalog_density_proxy + eps)
X["log_stellar_density_ratio"] = np.sign(X["stellar_density_ratio"]) * np.log1p(np.abs(X["stellar_density_ratio"]))
impact = self._safe_num(X, "koi_impact")
ror = self._safe_num(X, "koi_ror")
X["impact_plus_ror"] = impact + ror
X["high_impact_flag"] = (impact > 1).astype(float)
if "koi_quarters" in X.columns:
quarters = X["koi_quarters"].astype("string")
X["observed_quarter_count"] = quarters.str.count("1").astype("float64")
X["observed_quarter_length"] = quarters.str.len().astype("float64")
for a, b, name in [
("koi_gmag", "koi_rmag", "color_g_minus_r"),
("koi_rmag", "koi_imag", "color_r_minus_i"),
("koi_jmag", "koi_kmag", "color_j_minus_k"),
("koi_kepmag", "koi_kmag", "color_kep_minus_k"),
]:
if a in X.columns and b in X.columns:
X[name] = self._safe_num(X, a) - self._safe_num(X, b)
for base in [
"koi_period", "koi_duration", "koi_depth", "koi_prad", "koi_srho", "koi_insol",
"koi_impact", "koi_ror", "koi_dor", "koi_steff", "koi_slogg", "koi_srad", "koi_smass",
]:
if base in X.columns:
val = self._safe_num(X, base).abs()
err1 = self._signed_abs_err(X, f"{base}_err1")
err2 = self._signed_abs_err(X, f"{base}_err2")
if f"{base}_err1" in X.columns:
X[f"{base}_rel_err1"] = err1 / (val + eps)
if f"{base}_err2" in X.columns:
X[f"{base}_rel_err2"] = err2 / (val + eps)
if f"{base}_err1" in X.columns and f"{base}_err2" in X.columns:
X[f"{base}_mean_rel_err"] = (err1 + err2) / (2 * val + eps)
# Diagnostic flags are kept only for the comparison model and binarized to handle the one anomalous value 465.
if self.include_flags:
for flag in FLAG_COLUMNS:
if flag in X.columns:
X[flag] = (pd.to_numeric(X[flag], errors="coerce").fillna(0) > 0).astype(int)
return X
def _transform_frame(self, X):
X = X.copy()
X = self._add_features(X)
X = X.drop(columns=[c for c in self.drop_columns_ if c in X.columns], errors="ignore")
X = X.replace([np.inf, -np.inf], np.nan)
return X
def transform(self, X):
Xt = self._transform_frame(X)
for col in self.output_columns_:
if col not in Xt.columns:
Xt[col] = np.nan
return Xt[self.output_columns_]
def get_feature_names_out(self, input_features=None):
return np.array(self.output_columns_, dtype=object)
feature_explanations = pd.DataFrame([
{"engineered_feature": "log1p physical quantities", "plain_reason": "Transit measurements span orders of magnitude, so logs make small and large systems comparable."},
{"engineered_feature": "transit_snr_proxy", "plain_reason": "A deeper dip measured with lower uncertainty is more trustworthy."},
{"engineered_feature": "radius_ratio_from_prad_srad", "plain_reason": "A planet blocks light in proportion to its size compared with its star."},
{"engineered_feature": "depth_implied_radius_ratio", "plain_reason": "Transit depth gives an independent estimate of the planet-to-star radius ratio."},
{"engineered_feature": "radius_ratio_abs_residual", "plain_reason": "If radius estimates disagree strongly, the event may be a bad fit or false positive."},
{"engineered_feature": "transit_duty_cycle", "plain_reason": "Transit duration should be plausible relative to the orbital period."},
{"engineered_feature": "stellar_density_ratio", "plain_reason": "Transit geometry and catalog stellar properties should imply compatible stellar density."},
{"engineered_feature": "teq_insol_abs_residual", "plain_reason": "Planet temperature and received starlight should be physically consistent."},
{"engineered_feature": "impact_plus_ror / high_impact_flag", "plain_reason": "Grazing events are harder to trust and often resemble eclipsing binaries."},
{"engineered_feature": "observed_quarter_count", "plain_reason": "More observed quarters means more chances to verify repeated transits."},
{"engineered_feature": "color indices", "plain_reason": "Star color helps summarize the host star context from photometric bands."},
{"engineered_feature": "relative uncertainty ratios", "plain_reason": "Large uncertainty relative to a value is itself evidence that a measurement is less reliable."},
])
display(feature_explanations)
feature_explanations.to_csv(TABLE_DIR / "feature_engineering_explanations.csv", index=False)
fe_preview = FeatureEngineer(include_flags=False).fit(df.drop(columns=[TARGET]), df[TARGET])
primary_feature_frame = fe_preview.transform(df.drop(columns=[TARGET]))
print(f"Primary raw-signal feature count after engineering: {primary_feature_frame.shape[1]}")
display(pd.DataFrame({"feature": primary_feature_frame.columns[:40]}))
print("Progress: feature engineering complete.")
| engineered_feature | plain_reason | |
|---|---|---|
| 0 | log1p physical quantities | Transit measurements span orders of magnitude,... |
| 1 | transit_snr_proxy | A deeper dip measured with lower uncertainty i... |
| 2 | radius_ratio_from_prad_srad | A planet blocks light in proportion to its siz... |
| 3 | depth_implied_radius_ratio | Transit depth gives an independent estimate of... |
| 4 | radius_ratio_abs_residual | If radius estimates disagree strongly, the eve... |
| 5 | transit_duty_cycle | Transit duration should be plausible relative ... |
| 6 | stellar_density_ratio | Transit geometry and catalog stellar propertie... |
| 7 | teq_insol_abs_residual | Planet temperature and received starlight shou... |
| 8 | impact_plus_ror / high_impact_flag | Grazing events are harder to trust and often r... |
| 9 | observed_quarter_count | More observed quarters means more chances to v... |
| 10 | color indices | Star color helps summarize the host star conte... |
| 11 | relative uncertainty ratios | Large uncertainty relative to a value is itsel... |
Primary raw-signal feature count after engineering: 166
| feature | |
|---|---|
| 0 | koi_period |
| 1 | koi_period_err1 |
| 2 | koi_period_err2 |
| 3 | koi_time0bk |
| 4 | koi_time0bk_err1 |
| 5 | koi_time0bk_err2 |
| 6 | koi_time0 |
| 7 | koi_time0_err1 |
| 8 | koi_time0_err2 |
| 9 | koi_impact |
| 10 | koi_impact_err1 |
| 11 | koi_impact_err2 |
| 12 | koi_duration |
| 13 | koi_duration_err1 |
| 14 | koi_duration_err2 |
| 15 | koi_depth |
| 16 | koi_depth_err1 |
| 17 | koi_depth_err2 |
| 18 | koi_ror |
| 19 | koi_ror_err1 |
| 20 | koi_ror_err2 |
| 21 | koi_srho |
| 22 | koi_srho_err1 |
| 23 | koi_srho_err2 |
| 24 | koi_prad |
| 25 | koi_prad_err1 |
| 26 | koi_prad_err2 |
| 27 | koi_sma |
| 28 | koi_incl |
| 29 | koi_teq |
| 30 | koi_insol |
| 31 | koi_insol_err1 |
| 32 | koi_insol_err2 |
| 33 | koi_dor |
| 34 | koi_dor_err1 |
| 35 | koi_dor_err2 |
| 36 | koi_ldm_coeff2 |
| 37 | koi_ldm_coeff1 |
| 38 | koi_max_sngle_ev |
| 39 | koi_max_mult_ev |
Progress: feature engineering complete.
5. Modeling¶
y_raw = df[TARGET].astype(str)
label_encoder = LabelEncoder()
y = label_encoder.fit_transform(y_raw)
class_names = label_encoder.classes_
candidate_label_index = int(np.where(class_names == "CANDIDATE")[0][0])
primary_features = (
FeatureEngineer(include_flags=False)
.fit(df.drop(columns=[TARGET]), y)
.transform(df.drop(columns=[TARGET]))
.select_dtypes(include=[np.number])
)
flag_features = (
FeatureEngineer(include_flags=True)
.fit(df.drop(columns=[TARGET]), y)
.transform(df.drop(columns=[TARGET]))
.select_dtypes(include=[np.number])
)
row_index = np.arange(len(df))
train_idx, test_idx = train_test_split(
row_index, test_size=0.20, random_state=RANDOM_STATE, stratify=y
)
X_train_primary = primary_features.iloc[train_idx].reset_index(drop=True)
X_test_primary = primary_features.iloc[test_idx].reset_index(drop=True)
X_train_flags = flag_features.iloc[train_idx].reset_index(drop=True)
X_test_flags = flag_features.iloc[test_idx].reset_index(drop=True)
y_train = y[train_idx]
y_test = y[test_idx]
print("Classes:", dict(zip(class_names, range(len(class_names)))))
print("Primary raw-signal feature matrix:", primary_features.shape)
print("Comparison feature matrix with flags:", flag_features.shape)
print("Train rows:", X_train_primary.shape[0], "Test rows:", X_test_primary.shape[0])
cv = StratifiedKFold(n_splits=5, shuffle=True, random_state=RANDOM_STATE)
def _prepare_fold(X_train_fold, X_valid_fold, scale=False):
imputer = SimpleImputer(strategy="median")
X_train_imp = imputer.fit_transform(X_train_fold)
X_valid_imp = imputer.transform(X_valid_fold)
scaler = None
if scale:
scaler = StandardScaler()
X_train_imp = scaler.fit_transform(X_train_imp)
X_valid_imp = scaler.transform(X_valid_imp)
return X_train_imp, X_valid_imp, imputer, scaler
def _fit_estimator(estimator, X_train_arr, y_train_arr, use_sample_weight=False):
fitted = clone(estimator)
if use_sample_weight:
weights = compute_sample_weight(class_weight="balanced", y=y_train_arr)
fitted.fit(X_train_arr, y_train_arr, sample_weight=weights)
else:
fitted.fit(X_train_arr, y_train_arr)
return fitted
def cross_validate_manual(name, estimator, X_frame, y_array, use_sample_weight=False, scale=False, splitter=cv):
rows = []
fold_macro = []
for fold, (tr, va) in enumerate(splitter.split(X_frame, y_array), start=1):
Xtr, Xva, _, _ = _prepare_fold(X_frame.iloc[tr], X_frame.iloc[va], scale=scale)
fitted = _fit_estimator(estimator, Xtr, y_array[tr], use_sample_weight=use_sample_weight)
pred = fitted.predict(Xva)
precision, recall, f1, support = precision_recall_fscore_support(
y_array[va], pred, labels=np.arange(len(class_names)), zero_division=0
)
row = {
"model": name,
"fold": fold,
"accuracy": accuracy_score(y_array[va], pred),
"macro_f1": f1_score(y_array[va], pred, average="macro"),
"weighted_f1": f1_score(y_array[va], pred, average="weighted"),
"macro_recall": recall_score(y_array[va], pred, average="macro"),
"candidate_recall": recall[candidate_label_index],
}
rows.append(row)
fold_macro.append(row["macro_f1"])
return pd.DataFrame(rows), np.array(fold_macro)
model_specs = [
{
"name": "Decision Tree baseline (unbalanced)",
"estimator": DecisionTreeClassifier(max_depth=6, random_state=RANDOM_STATE),
"X": X_train_primary,
"use_sample_weight": False,
"scale": False,
},
{
"name": "Decision Tree baseline (balanced)",
"estimator": DecisionTreeClassifier(max_depth=6, class_weight="balanced", random_state=RANDOM_STATE),
"X": X_train_primary,
"use_sample_weight": False,
"scale": False,
},
{
"name": "Random Forest (raw-signal)",
"estimator": RandomForestClassifier(
n_estimators=60,
max_depth=None,
min_samples_leaf=2,
max_features="sqrt",
class_weight="balanced_subsample",
random_state=RANDOM_STATE,
n_jobs=-1,
),
"X": X_train_primary,
"use_sample_weight": False,
"scale": False,
},
{
"name": "XGBoost (raw-signal)",
"estimator": XGBClassifier(
objective="multi:softprob",
num_class=len(class_names),
eval_metric="mlogloss",
n_estimators=25,
max_depth=3,
learning_rate=0.08,
subsample=0.9,
colsample_bytree=0.85,
reg_lambda=1.0,
random_state=RANDOM_STATE,
n_jobs=1,
tree_method="hist",
),
"X": X_train_primary,
"use_sample_weight": True,
"scale": False,
},
{
"name": "Random Forest (with vetting flags comparison)",
"estimator": RandomForestClassifier(
n_estimators=60,
max_depth=None,
min_samples_leaf=2,
max_features="sqrt",
class_weight="balanced_subsample",
random_state=RANDOM_STATE,
n_jobs=-1,
),
"X": X_train_flags,
"use_sample_weight": False,
"scale": False,
},
]
all_fold_rows = []
fold_scores = {}
for spec in model_specs:
print(f"Cross-validating: {spec['name']}")
fold_df, macro_scores = cross_validate_manual(
spec["name"], spec["estimator"], spec["X"], y_train,
use_sample_weight=spec["use_sample_weight"], scale=spec["scale"]
)
all_fold_rows.append(fold_df)
fold_scores[spec["name"]] = macro_scores
fold_metrics = pd.concat(all_fold_rows, ignore_index=True)
cv_results = (
fold_metrics
.groupby("model")
.agg(
accuracy_mean=("accuracy", "mean"),
accuracy_std=("accuracy", "std"),
macro_f1_mean=("macro_f1", "mean"),
macro_f1_std=("macro_f1", "std"),
weighted_f1_mean=("weighted_f1", "mean"),
weighted_f1_std=("weighted_f1", "std"),
macro_recall_mean=("macro_recall", "mean"),
macro_recall_std=("macro_recall", "std"),
candidate_recall_mean=("candidate_recall", "mean"),
candidate_recall_std=("candidate_recall", "std"),
)
.reset_index()
.sort_values("macro_f1_mean", ascending=False)
)
display(cv_results.round(4))
cv_results.to_csv(TABLE_DIR / "cv_model_comparison.csv", index=False)
fold_score_table = pd.DataFrame(fold_scores)
display(fold_score_table.round(4))
fold_score_table.to_csv(TABLE_DIR / "cv_fold_macro_f1.csv", index=False)
fold_metrics.to_csv(TABLE_DIR / "cv_fold_metrics_long.csv", index=False)
imbalance_check = cv_results[
cv_results["model"].isin(["Decision Tree baseline (unbalanced)", "Decision Tree baseline (balanced)"])
][["model", "macro_f1_mean", "candidate_recall_mean"]]
display(imbalance_check.round(4))
print("Progress: baseline and primary cross-validation complete.")
Classes: {'CANDIDATE': 0, 'CONFIRMED': 1, 'FALSE POSITIVE': 2}
Primary raw-signal feature matrix: (9564, 166)
Comparison feature matrix with flags: (9564, 170)
Train rows: 7651 Test rows: 1913
Cross-validating: Decision Tree baseline (unbalanced)
Cross-validating: Decision Tree baseline (balanced)
Cross-validating: Random Forest (raw-signal)
Cross-validating: XGBoost (raw-signal)
Cross-validating: Random Forest (with vetting flags comparison)
| model | accuracy_mean | accuracy_std | macro_f1_mean | macro_f1_std | weighted_f1_mean | weighted_f1_std | macro_recall_mean | macro_recall_std | candidate_recall_mean | candidate_recall_std | |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 3 | Random Forest (with vetting flags comparison) | 0.9200 | 0.0112 | 0.8946 | 0.0148 | 0.9183 | 0.0116 | 0.8883 | 0.0155 | 0.7648 | 0.0334 |
| 2 | Random Forest (raw-signal) | 0.8475 | 0.0149 | 0.8116 | 0.0175 | 0.8438 | 0.0146 | 0.8074 | 0.0176 | 0.5967 | 0.0329 |
| 4 | XGBoost (raw-signal) | 0.8017 | 0.0071 | 0.7821 | 0.0092 | 0.8087 | 0.0066 | 0.8008 | 0.0110 | 0.7332 | 0.0300 |
| 1 | Decision Tree baseline (unbalanced) | 0.7904 | 0.0103 | 0.7507 | 0.0094 | 0.7874 | 0.0087 | 0.7482 | 0.0110 | 0.5424 | 0.0320 |
| 0 | Decision Tree baseline (balanced) | 0.7654 | 0.0131 | 0.7491 | 0.0116 | 0.7758 | 0.0103 | 0.7714 | 0.0130 | 0.7452 | 0.0320 |
| Decision Tree baseline (unbalanced) | Decision Tree baseline (balanced) | Random Forest (raw-signal) | XGBoost (raw-signal) | Random Forest (with vetting flags comparison) | |
|---|---|---|---|---|---|
| 0 | 0.7483 | 0.7399 | 0.8056 | 0.7771 | 0.8949 |
| 1 | 0.7643 | 0.7500 | 0.8280 | 0.7951 | 0.9096 |
| 2 | 0.7421 | 0.7593 | 0.8270 | 0.7813 | 0.9029 |
| 3 | 0.7559 | 0.7615 | 0.8116 | 0.7863 | 0.8951 |
| 4 | 0.7429 | 0.7349 | 0.7855 | 0.7708 | 0.8705 |
| model | macro_f1_mean | candidate_recall_mean | |
|---|---|---|---|
| 1 | Decision Tree baseline (unbalanced) | 0.7507 | 0.5424 |
| 0 | Decision Tree baseline (balanced) | 0.7491 | 0.7452 |
Progress: baseline and primary cross-validation complete.
# Small, documented tuning search on the strongest leakage-free tree families.
tuning_cv = StratifiedKFold(n_splits=3, shuffle=True, random_state=RANDOM_STATE)
tuning_candidates = [
{
"model": "Random Forest tuned",
"estimator": RandomForestClassifier(
n_estimators=80,
max_depth=None,
min_samples_leaf=2,
max_features="sqrt",
class_weight="balanced_subsample",
random_state=RANDOM_STATE,
n_jobs=-1,
),
"params": {"n_estimators": 80, "max_depth": None, "min_samples_leaf": 2, "max_features": "sqrt"},
"use_sample_weight": False,
},
{
"model": "Random Forest tuned",
"estimator": RandomForestClassifier(
n_estimators=120,
max_depth=18,
min_samples_leaf=2,
max_features="sqrt",
class_weight="balanced_subsample",
random_state=RANDOM_STATE,
n_jobs=-1,
),
"params": {"n_estimators": 120, "max_depth": 18, "min_samples_leaf": 2, "max_features": "sqrt"},
"use_sample_weight": False,
},
{
"model": "Random Forest tuned",
"estimator": RandomForestClassifier(
n_estimators=120,
max_depth=None,
min_samples_leaf=1,
max_features=0.60,
class_weight="balanced_subsample",
random_state=RANDOM_STATE,
n_jobs=-1,
),
"params": {"n_estimators": 120, "max_depth": None, "min_samples_leaf": 1, "max_features": 0.60},
"use_sample_weight": False,
},
{
"model": "XGBoost tuned",
"estimator": XGBClassifier(
objective="multi:softprob",
num_class=len(class_names),
eval_metric="mlogloss",
n_estimators=35,
max_depth=3,
learning_rate=0.08,
subsample=0.90,
colsample_bytree=0.85,
reg_lambda=1.0,
random_state=RANDOM_STATE,
n_jobs=1,
tree_method="hist",
),
"params": {"n_estimators": 35, "max_depth": 3, "learning_rate": 0.08, "subsample": 0.90, "colsample_bytree": 0.85},
"use_sample_weight": True,
},
{
"model": "XGBoost tuned",
"estimator": XGBClassifier(
objective="multi:softprob",
num_class=len(class_names),
eval_metric="mlogloss",
n_estimators=55,
max_depth=3,
learning_rate=0.06,
subsample=0.95,
colsample_bytree=0.85,
reg_lambda=1.2,
random_state=RANDOM_STATE,
n_jobs=1,
tree_method="hist",
),
"params": {"n_estimators": 55, "max_depth": 3, "learning_rate": 0.06, "subsample": 0.95, "colsample_bytree": 0.85},
"use_sample_weight": True,
},
]
tuning_rows = []
for candidate in tuning_candidates:
label = f"{candidate['model']} {candidate['params']}"
print("Tuning:", label)
fold_df, macro_scores = cross_validate_manual(
label,
candidate["estimator"],
X_train_primary,
y_train,
use_sample_weight=candidate["use_sample_weight"],
splitter=tuning_cv,
)
tuning_rows.append({
"model": candidate["model"],
"mean_macro_f1": macro_scores.mean(),
"std_macro_f1": macro_scores.std(),
"params": candidate["params"],
"estimator": candidate["estimator"],
"use_sample_weight": candidate["use_sample_weight"],
})
tuning_results_full = pd.DataFrame(tuning_rows).sort_values("mean_macro_f1", ascending=False).reset_index(drop=True)
display(tuning_results_full.drop(columns=["estimator", "use_sample_weight"]).round(4))
tuning_results_full.drop(columns=["estimator", "use_sample_weight"]).to_csv(TABLE_DIR / "tuning_results.csv", index=False)
best_row = tuning_results_full.iloc[0]
best_primary_name = (
"XGBoost tuned raw-signal"
if best_row["model"] == "XGBoost tuned"
else "Random Forest tuned raw-signal"
)
best_primary_estimator = best_row["estimator"]
best_primary_use_sample_weight = bool(best_row["use_sample_weight"])
best_primary_params = best_row["params"]
print(f"Selected primary model: {best_primary_name}")
print("Best parameters:", best_primary_params)
print("Progress: hyperparameter tuning complete.")
Tuning: Random Forest tuned {'n_estimators': 80, 'max_depth': None, 'min_samples_leaf': 2, 'max_features': 'sqrt'}
Tuning: Random Forest tuned {'n_estimators': 120, 'max_depth': 18, 'min_samples_leaf': 2, 'max_features': 'sqrt'}
Tuning: Random Forest tuned {'n_estimators': 120, 'max_depth': None, 'min_samples_leaf': 1, 'max_features': 0.6}
Tuning: XGBoost tuned {'n_estimators': 35, 'max_depth': 3, 'learning_rate': 0.08, 'subsample': 0.9, 'colsample_bytree': 0.85}
Tuning: XGBoost tuned {'n_estimators': 55, 'max_depth': 3, 'learning_rate': 0.06, 'subsample': 0.95, 'colsample_bytree': 0.85}
| model | mean_macro_f1 | std_macro_f1 | params | |
|---|---|---|---|---|
| 0 | Random Forest tuned | 0.8128 | 0.0088 | {'n_estimators': 120, 'max_depth': 18, 'min_sa... |
| 1 | Random Forest tuned | 0.8121 | 0.0099 | {'n_estimators': 80, 'max_depth': None, 'min_s... |
| 2 | Random Forest tuned | 0.8087 | 0.0061 | {'n_estimators': 120, 'max_depth': None, 'min_... |
| 3 | XGBoost tuned | 0.7900 | 0.0069 | {'n_estimators': 55, 'max_depth': 3, 'learning... |
| 4 | XGBoost tuned | 0.7875 | 0.0072 | {'n_estimators': 35, 'max_depth': 3, 'learning... |
Selected primary model: Random Forest tuned raw-signal
Best parameters: {'n_estimators': 120, 'max_depth': 18, 'min_samples_leaf': 2, 'max_features': 'sqrt'}
Progress: hyperparameter tuning complete.
Class imbalance is handled with class weights rather than SMOTE. This avoids creating synthetic astronomical objects and keeps every training row tied to a real KOI. The notebook compares unbalanced and balanced decision-tree baselines so the effect on minority-class recall is measured rather than assumed.
6. Evaluation¶
def fit_final_bundle(estimator, X_frame, y_array, use_sample_weight=False, scale=False):
imputer = SimpleImputer(strategy="median")
X_imp = imputer.fit_transform(X_frame)
scaler = None
if scale:
scaler = StandardScaler()
X_imp = scaler.fit_transform(X_imp)
fitted = _fit_estimator(estimator, X_imp, y_array, use_sample_weight=use_sample_weight)
return {"estimator": fitted, "imputer": imputer, "scaler": scaler, "scale": scale}
def transform_bundle(bundle, X_frame):
X_imp = bundle["imputer"].transform(X_frame)
if bundle["scaler"] is not None:
X_imp = bundle["scaler"].transform(X_imp)
return X_imp
def predict_bundle(bundle, X_frame):
X_arr = transform_bundle(bundle, X_frame)
pred = bundle["estimator"].predict(X_arr)
proba = bundle["estimator"].predict_proba(X_arr) if hasattr(bundle["estimator"], "predict_proba") else None
return pred, proba
def evaluate_model(name, bundle, X_eval, y_eval, save_prefix):
y_pred, y_proba = predict_bundle(bundle, X_eval)
report = classification_report(y_eval, y_pred, target_names=class_names, output_dict=True, zero_division=0)
report_df = pd.DataFrame(report).T
cm = confusion_matrix(y_eval, y_pred)
cm_norm = confusion_matrix(y_eval, y_pred, normalize="true")
summary = {
"model": name,
"accuracy": accuracy_score(y_eval, y_pred),
"macro_f1": f1_score(y_eval, y_pred, average="macro"),
"weighted_f1": f1_score(y_eval, y_pred, average="weighted"),
"macro_recall": recall_score(y_eval, y_pred, average="macro"),
}
if y_proba is not None:
y_bin = label_binarize(y_eval, classes=np.arange(len(class_names)))
summary["macro_roc_auc_ovr"] = roc_auc_score(y_eval, y_proba, multi_class="ovr", average="macro")
summary["macro_average_precision"] = average_precision_score(y_bin, y_proba, average="macro")
report_df.to_csv(TABLE_DIR / f"{save_prefix}_classification_report.csv")
pd.DataFrame(cm, index=class_names, columns=class_names).to_csv(TABLE_DIR / f"{save_prefix}_confusion_matrix_counts.csv")
pd.DataFrame(cm_norm, index=class_names, columns=class_names).to_csv(TABLE_DIR / f"{save_prefix}_confusion_matrix_normalized.csv")
return summary, report_df, cm, cm_norm, y_pred, y_proba
# Fit the selected primary leakage-free model on training data only, then evaluate once on the held-out test set.
best_primary_bundle = fit_final_bundle(
best_primary_estimator,
X_train_primary,
y_train,
use_sample_weight=best_primary_use_sample_weight,
scale=False,
)
primary_summary, primary_report, primary_cm, primary_cm_norm, primary_pred, primary_proba = evaluate_model(
best_primary_name, best_primary_bundle, X_test_primary, y_test, "primary_raw_signal"
)
# Fit the labeled comparison model with diagnostic flags, using the same train/test split.
flag_estimator = RandomForestClassifier(
n_estimators=100,
min_samples_leaf=2,
max_features="sqrt",
class_weight="balanced_subsample",
random_state=RANDOM_STATE,
n_jobs=-1,
)
flag_bundle = fit_final_bundle(flag_estimator, X_train_flags, y_train, use_sample_weight=False, scale=False)
flag_summary, flag_report, flag_cm, flag_cm_norm, flag_pred, flag_proba = evaluate_model(
"Random Forest with vetting flags comparison", flag_bundle, X_test_flags, y_test, "comparison_with_flags"
)
final_summary = pd.DataFrame([primary_summary, flag_summary])
display(final_summary.round(4))
display(primary_report.round(4))
final_summary.to_csv(TABLE_DIR / "final_test_summary.csv", index=False)
fig, axes = plt.subplots(1, 2, figsize=(12, 4.8))
sns.heatmap(primary_cm, annot=True, fmt="d", cmap="Blues", xticklabels=class_names, yticklabels=class_names, ax=axes[0])
axes[0].set_title("Primary Raw-Signal Model\nConfusion Matrix Counts")
axes[0].set_xlabel("Predicted")
axes[0].set_ylabel("Actual")
sns.heatmap(primary_cm_norm, annot=True, fmt=".2f", cmap="Blues", xticklabels=class_names, yticklabels=class_names, ax=axes[1])
axes[1].set_title("Primary Raw-Signal Model\nNormalized by Actual Class")
axes[1].set_xlabel("Predicted")
axes[1].set_ylabel("Actual")
plt.tight_layout()
plt.savefig(FIG_DIR / "primary_confusion_matrices.png", bbox_inches="tight")
plt.show()
| model | accuracy | macro_f1 | weighted_f1 | macro_recall | macro_roc_auc_ovr | macro_average_precision | |
|---|---|---|---|---|---|---|---|
| 0 | Random Forest tuned raw-signal | 0.8552 | 0.8234 | 0.8532 | 0.8201 | 0.9522 | 0.8801 |
| 1 | Random Forest with vetting flags comparison | 0.9211 | 0.8958 | 0.9196 | 0.8895 | 0.9819 | 0.9468 |
| precision | recall | f1-score | support | |
|---|---|---|---|---|
| CANDIDATE | 0.6970 | 0.6389 | 0.6667 | 396.0000 |
| CONFIRMED | 0.8957 | 0.9071 | 0.9014 | 549.0000 |
| FALSE POSITIVE | 0.8903 | 0.9143 | 0.9021 | 968.0000 |
| accuracy | 0.8552 | 0.8552 | 0.8552 | 0.8552 |
| macro avg | 0.8277 | 0.8201 | 0.8234 | 1913.0000 |
| weighted avg | 0.8518 | 0.8552 | 0.8532 | 1913.0000 |
fig, axes = plt.subplots(1, 2, figsize=(13, 5))
y_test_bin = label_binarize(y_test, classes=np.arange(len(class_names)))
for i, class_name in enumerate(class_names):
PrecisionRecallDisplay.from_predictions(
y_test_bin[:, i],
primary_proba[:, i],
name=class_name,
ax=axes[0],
)
RocCurveDisplay.from_predictions(
y_test_bin[:, i],
primary_proba[:, i],
name=class_name,
ax=axes[1],
)
axes[0].set_title("Precision-Recall Curves\nPrimary Raw-Signal Model")
axes[1].set_title("ROC Curves\nPrimary Raw-Signal Model")
plt.tight_layout()
plt.savefig(FIG_DIR / "primary_pr_roc_curves.png", bbox_inches="tight")
plt.show()
model_comparison = cv_results[["model", "macro_f1_mean", "macro_f1_std", "weighted_f1_mean", "macro_recall_mean"]].copy()
final_test_rows = final_summary[["model", "macro_f1", "weighted_f1", "macro_recall", "accuracy"]].copy()
display(model_comparison.round(4))
display(final_test_rows.round(4))
print("Progress: evaluation complete.")
| model | macro_f1_mean | macro_f1_std | weighted_f1_mean | macro_recall_mean | |
|---|---|---|---|---|---|
| 3 | Random Forest (with vetting flags comparison) | 0.8946 | 0.0148 | 0.9183 | 0.8883 |
| 2 | Random Forest (raw-signal) | 0.8116 | 0.0175 | 0.8438 | 0.8074 |
| 4 | XGBoost (raw-signal) | 0.7821 | 0.0092 | 0.8087 | 0.8008 |
| 1 | Decision Tree baseline (unbalanced) | 0.7507 | 0.0094 | 0.7874 | 0.7482 |
| 0 | Decision Tree baseline (balanced) | 0.7491 | 0.0116 | 0.7758 | 0.7714 |
| model | macro_f1 | weighted_f1 | macro_recall | accuracy | |
|---|---|---|---|---|---|
| 0 | Random Forest tuned raw-signal | 0.8234 | 0.8532 | 0.8201 | 0.8552 |
| 1 | Random Forest with vetting flags comparison | 0.8958 | 0.9196 | 0.8895 | 0.9211 |
Progress: evaluation complete.
The primary metric is macro F1 because each class matters, including unresolved candidates. The model with diagnostic flags is reported only as a comparison: its higher score, if present, reflects access to NASA vetting signals that a raw-signal model should not rely on.
7. Explainability¶
feature_names = list(X_train_primary.columns)
model_step = best_primary_bundle["estimator"]
if hasattr(model_step, "feature_importances_"):
importances = pd.DataFrame({
"feature": feature_names,
"importance": model_step.feature_importances_,
}).sort_values("importance", ascending=False)
else:
perm = permutation_importance(
model_step, transform_bundle(best_primary_bundle, X_test_primary), y_test, n_repeats=5, random_state=RANDOM_STATE,
scoring="f1_macro", n_jobs=-1
)
importances = pd.DataFrame({
"feature": feature_names[: len(perm.importances_mean)],
"importance": perm.importances_mean,
}).sort_values("importance", ascending=False)
display(importances.head(25))
importances.to_csv(TABLE_DIR / "primary_feature_importance.csv", index=False)
fig, ax = plt.subplots(figsize=(9, 7))
sns.barplot(data=importances.head(20), x="importance", y="feature", color="#59A14F", ax=ax)
ax.set_title("Top Feature Importances - Primary Raw-Signal Model")
ax.set_xlabel("Importance")
ax.set_ylabel("")
plt.tight_layout()
plt.savefig(FIG_DIR / "primary_feature_importance.png", bbox_inches="tight")
plt.show()
| feature | importance | |
|---|---|---|
| 93 | koi_dikco_msky | 0.033056 |
| 87 | koi_dicco_msky | 0.032714 |
| 133 | koi_depth_rel_err1 | 0.022093 |
| 40 | koi_model_snr | 0.021507 |
| 18 | koi_ror | 0.021370 |
| 135 | koi_depth_mean_rel_err | 0.020360 |
| 98 | log1p_koi_prad | 0.020223 |
| 101 | log1p_koi_model_snr | 0.019833 |
| 108 | transit_snr_proxy | 0.019262 |
| 109 | radius_ratio_from_prad_srad | 0.018366 |
| 114 | log_transit_duty_cycle | 0.018341 |
| 24 | koi_prad | 0.018117 |
| 113 | transit_duty_cycle | 0.017213 |
| 130 | koi_duration_rel_err1 | 0.014036 |
| 131 | koi_duration_rel_err2 | 0.013898 |
| 132 | koi_duration_mean_rel_err | 0.013555 |
| 134 | koi_depth_rel_err2 | 0.013035 |
| 156 | koi_steff_mean_rel_err | 0.012571 |
| 119 | impact_plus_ror | 0.012438 |
| 112 | radius_ratio_abs_residual | 0.012354 |
| 107 | log1p_koi_max_mult_ev | 0.011564 |
| 146 | koi_impact_rel_err2 | 0.011531 |
| 154 | koi_steff_rel_err1 | 0.011407 |
| 25 | koi_prad_err1 | 0.011204 |
| 70 | koi_fwm_stat_sig | 0.011063 |
shap_status = "not_run"
try:
import shap
tree_model = best_primary_bundle["estimator"]
X_test_transformed = transform_bundle(best_primary_bundle, X_test_primary)
if not isinstance(X_test_transformed, np.ndarray):
X_test_transformed = np.asarray(X_test_transformed)
shap_sample_size = min(120, X_test_transformed.shape[0])
X_shap = X_test_transformed[:shap_sample_size]
explainer = shap.TreeExplainer(tree_model)
shap_values = explainer.shap_values(X_shap)
if isinstance(shap_values, list):
mean_abs = np.mean([np.abs(vals).mean(axis=0) for vals in shap_values], axis=0)
summary_values = shap_values
else:
arr = np.asarray(shap_values)
if arr.ndim == 3:
mean_abs = np.abs(arr).mean(axis=(0, 2))
summary_values = arr[:, :, 0]
else:
mean_abs = np.abs(arr).mean(axis=0)
summary_values = arr
shap_importance = pd.DataFrame({
"feature": feature_names,
"mean_abs_shap": mean_abs,
}).sort_values("mean_abs_shap", ascending=False)
display(shap_importance.head(20))
shap_importance.to_csv(TABLE_DIR / "primary_shap_importance.csv", index=False)
fig, ax = plt.subplots(figsize=(9, 7))
sns.barplot(data=shap_importance.head(20), x="mean_abs_shap", y="feature", color="#E15759", ax=ax)
ax.set_title("Mean Absolute SHAP Importance - Primary Model")
ax.set_xlabel("Mean absolute SHAP value")
ax.set_ylabel("")
plt.tight_layout()
plt.savefig(FIG_DIR / "primary_shap_summary_bar.png", bbox_inches="tight")
plt.show()
# Individual prediction explanation using the top SHAP contributors for the first test row.
if isinstance(shap_values, list):
pred_class = int(primary_pred[0])
row_vals = shap_values[pred_class][0]
expected = explainer.expected_value[pred_class]
else:
arr = np.asarray(shap_values)
pred_class = int(primary_pred[0])
if arr.ndim == 3:
row_vals = arr[0, :, pred_class]
expected = np.asarray(explainer.expected_value)[pred_class]
else:
row_vals = arr[0]
expected = explainer.expected_value
row_explanation = pd.DataFrame({
"feature": feature_names,
"shap_value": row_vals,
"abs_shap": np.abs(row_vals),
}).sort_values("abs_shap", ascending=False).head(12)
display(row_explanation[["feature", "shap_value"]])
row_explanation.to_csv(TABLE_DIR / "individual_prediction_shap_top_features.csv", index=False)
fig, ax = plt.subplots(figsize=(8, 5))
sns.barplot(
data=row_explanation.sort_values("shap_value"),
x="shap_value",
y="feature",
palette=["#E15759" if v < 0 else "#59A14F" for v in row_explanation.sort_values("shap_value")["shap_value"]],
ax=ax,
)
ax.axvline(0, color="black", linewidth=0.8)
actual_label = class_names[y_test[0]]
predicted_label = class_names[pred_class]
ax.set_title(f"Individual SHAP Explanation\nActual: {actual_label}, Predicted: {predicted_label}")
ax.set_xlabel("SHAP contribution")
ax.set_ylabel("")
plt.tight_layout()
plt.savefig(FIG_DIR / "individual_prediction_shap.png", bbox_inches="tight")
plt.show()
shap_status = "completed"
except Exception as exc:
shap_status = f"failed: {type(exc).__name__}: {exc}"
print("SHAP could not be completed; feature importance remains available.")
print(shap_status)
print("SHAP status:", shap_status)
print("Progress: explainability complete.")
C:\Users\Lenovo\OneDrive\Documents\New project 2\.venv_test\Lib\site-packages\shap\plots\colors\_colors.py:47: PendingDeprecationWarning: The set_bad function will be deprecated in a future version. Use cmap.with_extremes(bad=...) or Colormap(bad=...) instead. red_blue.set_bad(gray_rgb.tolist(), 1.0) C:\Users\Lenovo\OneDrive\Documents\New project 2\.venv_test\Lib\site-packages\shap\plots\colors\_colors.py:48: PendingDeprecationWarning: The set_over function will be deprecated in a future version. Use cmap.with_extremes(over=...) or Colormap(over=...) instead. red_blue.set_over(gray_rgb.tolist(), 1.0) C:\Users\Lenovo\OneDrive\Documents\New project 2\.venv_test\Lib\site-packages\shap\plots\colors\_colors.py:49: PendingDeprecationWarning: The set_under function will be deprecated in a future version. Use cmap.with_extremes(under=...) or Colormap(under=...) instead. red_blue.set_under(gray_rgb.tolist(), 1.0) # "under" is incorrectly used instead of "bad" in the scatter plot
| feature | mean_abs_shap | |
|---|---|---|
| 93 | koi_dikco_msky | 0.024554 |
| 87 | koi_dicco_msky | 0.020659 |
| 133 | koi_depth_rel_err1 | 0.014020 |
| 18 | koi_ror | 0.013260 |
| 98 | log1p_koi_prad | 0.012950 |
| 40 | koi_model_snr | 0.012349 |
| 101 | log1p_koi_model_snr | 0.012085 |
| 109 | radius_ratio_from_prad_srad | 0.011912 |
| 114 | log_transit_duty_cycle | 0.011188 |
| 24 | koi_prad | 0.011129 |
| 108 | transit_snr_proxy | 0.010866 |
| 113 | transit_duty_cycle | 0.010862 |
| 135 | koi_depth_mean_rel_err | 0.010736 |
| 70 | koi_fwm_stat_sig | 0.009739 |
| 39 | koi_max_mult_ev | 0.009147 |
| 107 | log1p_koi_max_mult_ev | 0.008897 |
| 134 | koi_depth_rel_err2 | 0.008018 |
| 130 | koi_duration_rel_err1 | 0.007885 |
| 131 | koi_duration_rel_err2 | 0.007617 |
| 156 | koi_steff_mean_rel_err | 0.007615 |
| feature | shap_value | |
|---|---|---|
| 98 | log1p_koi_prad | 0.064642 |
| 18 | koi_ror | 0.059397 |
| 24 | koi_prad | 0.055353 |
| 109 | radius_ratio_from_prad_srad | 0.043880 |
| 25 | koi_prad_err1 | 0.037452 |
| 146 | koi_impact_rel_err2 | 0.024468 |
| 26 | koi_prad_err2 | 0.024241 |
| 112 | radius_ratio_abs_residual | 0.020862 |
| 94 | koi_dikco_msky_err | 0.016711 |
| 133 | koi_depth_rel_err1 | 0.016683 |
| 15 | koi_depth | 0.015503 |
| 40 | koi_model_snr | 0.013959 |
SHAP status: completed Progress: explainability complete.
Plain-language explanation: the model looks for whether the dimming pattern behaves like a planet crossing a star. Real planets usually produce a repeatable, physically plausible dip whose size, duration, and orbit agree with the star's properties. False positives often look too large, too grazing, too hot, too uncertain, or inconsistent with what the star should allow. The primary model avoids using NASA's diagnostic flags so its score reflects what can be learned from the measurements themselves, not from copying the vetting pipeline.
8. Conclusion and Written Summary¶
# Build a metrics-grounded written summary using values computed above.
primary = final_summary.loc[final_summary["model"] == best_primary_name].iloc[0]
flags = final_summary.loc[final_summary["model"] == "Random Forest with vetting flags comparison"].iloc[0]
top_features = importances.head(8)["feature"].tolist()
summary_text = f'''
## Written Summary
### EDA & Data Cleaning Approach
The dataset contains {df.shape[0]:,} Kepler Objects of Interest and {df.shape[1]:,} columns from the Celesta-provided KOI cumulative table. The target is a three-class label: CONFIRMED, CANDIDATE, and FALSE POSITIVE. The class distribution is imbalanced, with {class_counts.loc[class_counts['class'].eq('FALSE POSITIVE'), 'count'].iloc[0]:,} false positives, {class_counts.loc[class_counts['class'].eq('CONFIRMED'), 'count'].iloc[0]:,} confirmed planets, and {class_counts.loc[class_counts['class'].eq('CANDIDATE'), 'count'].iloc[0]:,} candidates. Missingness was handled conservatively: columns with more than 50% missing values or no variance were dropped, leakage-prone identifiers and vetting metadata were removed, and remaining numeric values were median-imputed inside each cross-validation fold. This prevents validation rows from influencing preprocessing. The EDA showed strong skew and outliers in orbital period, radius, depth, insolation, and signal-to-noise, so the model uses log-transformed versions of these quantities rather than deleting suspicious rows.
### Model Choice Justification
I compared an unbalanced decision-tree baseline, a class-weighted decision-tree baseline, a Random Forest, and XGBoost. The paired decision-tree baselines show whether class weighting improves minority-class behavior instead of assuming it helps. Tree models were chosen as the primary family because this is a medium-sized tabular dataset with nonlinear relationships between transit geometry, stellar properties, and disposition. A neural network would be harder to explain and unnecessary for about ten thousand rows. Validation used five-fold stratified cross-validation on the training set, followed by a single held-out test evaluation for the final confusion matrix and curves. Hyperparameter tuning used a small manual search over Random Forest and XGBoost settings, optimizing macro F1 because the candidate class should not be hidden by overall accuracy.
### Key Findings
The selected primary model is {best_primary_name}. On the held-out test set it reached accuracy {primary['accuracy']:.3f}, macro F1 {primary['macro_f1']:.3f}, weighted F1 {primary['weighted_f1']:.3f}, and macro recall {primary['macro_recall']:.3f}. The most important raw-signal features included {', '.join(top_features[:6])}. These features make physical sense: false positives often have unusually large inferred radii, high impact parameters, inconsistent radius-depth relationships, or extreme signal properties. A separate comparison model that included the NASA diagnostic false-positive flags reached macro F1 {flags['macro_f1']:.3f}. That comparison is useful but not treated as the main result because those flags are products of the vetting process. Removing them gives a more honest estimate of what a model can learn from measured transit and stellar properties.
### Explaining It to a Non-Technical Audience
The model is trying to decide whether a tiny repeated dimming of a star looks like a planet passing in front of it. It checks whether the dip is the right size, lasts for a believable amount of time, repeats in a stable orbit, and matches what we know about the star. If the dip looks too deep, too oddly shaped, too uncertain, or physically inconsistent, the model becomes more suspicious that it is a false alarm rather than a real planet.
### Limitations and Future Work
The labels come from an already-vetted astronomical catalog, so the model learns from historical Kepler decisions rather than discovering planets independently. Some high-performing columns are downstream diagnostic products, which is why they are excluded from the primary model and shown only as a comparison. Future work could test the pipeline on a newer mission such as TESS, add calibrated astrophysical priors, and compare binary and three-class formulations depending on the competition's final scoring rules.
'''
print(summary_text)
(OUTPUT_DIR / "written_summary.md").write_text(summary_text.strip() + "\n", encoding="utf-8")
devpost_checklist = pd.DataFrame([
{"item": "Project title + one-line description", "status": "Ready: title in README"},
{"item": "Team members listed", "status": "Ready: Karthik M A"},
{"item": "Public GitHub repo linked", "status": "Ready: https://github.com/sleader3221-dot/India_high_school"},
{"item": "Screenshot/visualization", "status": "Ready: confusion matrix and SHAP/importance figures saved"},
{"item": "Metrics table and confusion matrix visible", "status": "Ready in notebook and outputs/tables"},
{"item": "Written summary included", "status": "Ready in notebook and outputs/written_summary.md"},
{"item": "English submission", "status": "Ready"},
])
display(devpost_checklist)
devpost_checklist.to_csv(TABLE_DIR / "devpost_checklist.csv", index=False)
print("Progress: write-up complete.")
## Written Summary ### EDA & Data Cleaning Approach The dataset contains 9,564 Kepler Objects of Interest and 140 columns from the Celesta-provided KOI cumulative table. The target is a three-class label: CONFIRMED, CANDIDATE, and FALSE POSITIVE. The class distribution is imbalanced, with 4,839 false positives, 2,747 confirmed planets, and 1,978 candidates. Missingness was handled conservatively: columns with more than 50% missing values or no variance were dropped, leakage-prone identifiers and vetting metadata were removed, and remaining numeric values were median-imputed inside each cross-validation fold. This prevents validation rows from influencing preprocessing. The EDA showed strong skew and outliers in orbital period, radius, depth, insolation, and signal-to-noise, so the model uses log-transformed versions of these quantities rather than deleting suspicious rows. ### Model Choice Justification I compared an unbalanced decision-tree baseline, a class-weighted decision-tree baseline, a Random Forest, and XGBoost. The paired decision-tree baselines show whether class weighting improves minority-class behavior instead of assuming it helps. Tree models were chosen as the primary family because this is a medium-sized tabular dataset with nonlinear relationships between transit geometry, stellar properties, and disposition. A neural network would be harder to explain and unnecessary for about ten thousand rows. Validation used five-fold stratified cross-validation on the training set, followed by a single held-out test evaluation for the final confusion matrix and curves. Hyperparameter tuning used a small manual search over Random Forest and XGBoost settings, optimizing macro F1 because the candidate class should not be hidden by overall accuracy. ### Key Findings The selected primary model is Random Forest tuned raw-signal. On the held-out test set it reached accuracy 0.855, macro F1 0.823, weighted F1 0.853, and macro recall 0.820. The most important raw-signal features included koi_dikco_msky, koi_dicco_msky, koi_depth_rel_err1, koi_model_snr, koi_ror, koi_depth_mean_rel_err. These features make physical sense: false positives often have unusually large inferred radii, high impact parameters, inconsistent radius-depth relationships, or extreme signal properties. A separate comparison model that included the NASA diagnostic false-positive flags reached macro F1 0.896. That comparison is useful but not treated as the main result because those flags are products of the vetting process. Removing them gives a more honest estimate of what a model can learn from measured transit and stellar properties. ### Explaining It to a Non-Technical Audience The model is trying to decide whether a tiny repeated dimming of a star looks like a planet passing in front of it. It checks whether the dip is the right size, lasts for a believable amount of time, repeats in a stable orbit, and matches what we know about the star. If the dip looks too deep, too oddly shaped, too uncertain, or physically inconsistent, the model becomes more suspicious that it is a false alarm rather than a real planet. ### Limitations and Future Work The labels come from an already-vetted astronomical catalog, so the model learns from historical Kepler decisions rather than discovering planets independently. Some high-performing columns are downstream diagnostic products, which is why they are excluded from the primary model and shown only as a comparison. Future work could test the pipeline on a newer mission such as TESS, add calibrated astrophysical priors, and compare binary and three-class formulations depending on the competition's final scoring rules.
| item | status | |
|---|---|---|
| 0 | Project title + one-line description | Ready: title in README |
| 1 | Team members listed | Ready: Karthik M A |
| 2 | Public GitHub repo linked | Ready: https://github.com/sleader3221-dot/Indi... |
| 3 | Screenshot/visualization | Ready: confusion matrix and SHAP/importance fi... |
| 4 | Metrics table and confusion matrix visible | Ready in notebook and outputs/tables |
| 5 | Written summary included | Ready in notebook and outputs/written_summary.md |
| 6 | English submission | Ready |
Progress: write-up complete.