%0 Journal Article %T Larsson S. Random Data Splits Create Artificial Confidence in Virtual Screening, Molecular Property Prediction, and Drug–Target Modeling %A Anders Nilsson %A Emma Lindholm %A James Cooper %A Sven Larsson %J Pharmacophore %@ 2229-5402 %D 2024 %V 15 %N 6 %R 10.51847/0p8QlYktAq %P 87-97 %X Machine-learning models are increasingly used to prioritize compounds, estimate molecular properties, and infer drug–target relationships, yet their reported performance often depends more strongly on evaluation design than headline metrics reveal. Random data splitting remains common because it is simple, statistically efficient, and compatible with standardized model comparison. In chemically structured pharmaceutical datasets, however, random allocation frequently places closely related molecules, shared scaffolds, homologous targets, replicated assays, or records from the same experimental source on both sides of the training–test boundary. The resulting test set may therefore measure interpolation within familiar evidence rather than performance on the novel compounds, targets, or experimental conditions implied by a prospective deployment claim. This Original Validation-Standards Article develops a conceptual account of how chemical similarity, scaffold leakage, and hidden dependence generate artificial confidence across virtual screening, molecular-property prediction, and drug–target modeling. It proposes Deployment-Aligned Validation Credibility as a structured construct linking the intended pharmaceutical use to a dependence audit, a shift-matched split strategy, task-appropriate metrics, uncertainty assessment, external corroboration, and transparent reporting. The analysis distinguishes random interpolation from scaffold transfer, chemical-cluster transfer, temporal prediction, entity-disjoint evaluation, source-independent testing, and prospective experimental confirmation. It also emphasizes that no single performance measure can establish pharmaceutical usefulness, mechanistic validity, experimental reproducibility, or therapeutic value. The proposed construct is methodological rather than empirically validated and remains conditional on dataset provenance, assay comparability, similarity definitions, and the practical decision being supported. Its purpose is to make validation claims narrower, more interpretable, and more consistent with the distribution shifts that molecular models will encounter in actual pharmaceutical research. %U https://pharmacophorejournal.com/article/larsson-s-random-data-splits-create-artificial-confidence-in-virtual-screening-molecular-property-pvf35autdppolxc