Beyond Random Splits: Class-Conditioned Site Exposure and Attribution Change in Phishing URL Evaluation
DOI:
https://doi.org/10.54692/ijeci.2026.1002/278Keywords:
Phishing URL detection, registrable-site-disjoint evaluation, unseen-site generalization, Public Suffix List, machine learning evaluation.Abstract
Random URL-level evaluation can place test URLs from registrable sites already represented in training, making its generalization condition different from evaluation on previously unseen sites. Prior phishing research has studied site-disjoint, temporal, cross-dataset, robustness, and explanation settings, but the controlled relationship among class-conditioned site exposure, performance retention, and attribution change across random and site-disjoint evaluation remains limited. We evaluated a provenance-controlled URL-only corpus of 124,154 URLs (52,589 benign and 71,565 phishing) under random URL-level Regime A and zero-overlap private-PSL registrable-site-disjoint Regime C, using three compact representation/model families, five frozen seeds per primary regime, site-cluster percentile bootstrap intervals, and within-model attribution comparisons. In A42, 99.7% of benign and 64.6% of phishing test URLs belonged to registrable sites represented in training; Regime C imposed 0.0% private-site overlap by construction. The A42–C42 MCC gaps were 0.095 for M1 (95% interval [-0.039, 0.235]), 0.200 for M2 [0.102, 0.315], and 0.151 for M3 [0.082, 0.247]; the M2 and M3 intervals excluded zero, whereas M1’s included zero. M3’s character n-gram TF-IDF logistic-regression representation retained the strongest absolute unseen-site MCC, although M2 and M3 had larger A-to-C decreases than M1. Attribution rankings remained broadly consistent within models while selected feature/token magnitudes and ranks changed across regimes; these summaries are descriptive and non-causal. The findings support reporting class-conditioned site exposure and using registrable-site-disjoint evaluation alongside conventional random splits when the intended target is performance on previously unseen sites.