<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD with MathML3 v1.3 20210610//EN" "JATS-archivearticle1-3-mathml3.dtd"><article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"
  dtd-version="1.3" xml:lang="en" article-type="research-article">
  <?DTDIdentifier.IdentifierValue -//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.2 20190208//EN?>
  <?DTDIdentifier.IdentifierType public?>
  <?SourceDTD.DTDName JATS-journalpublishing1.dtd?>
  <?SourceDTD.Version 1.2?>
  <?ConverterInfo.XSLTName jats2jats3.xsl?>
  <?ConverterInfo.Version 1?>
  <?properties open_access?>
  <front>
    <journal-meta>
      <journal-id journal-id-type="iso-abbrev">Pharmacophore</journal-id>
      <journal-id journal-id-type="publisher-id">pharmacophorejournal.com</journal-id>
      <journal-id journal-id-type="publisher-id">Pharmacophore</journal-id>
      <journal-title-group>
        <journal-title>Pharmacophore</journal-title>
      </journal-title-group>
      <issn pub-type="epub">2229-5402</issn>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="publisher-id">pharmacophorejournal.com-6927</article-id>
      <article-id pub-id-type="doi">10.51847/0p8QlYktAq</article-id>
      <article-categories>
        <subj-group subj-group-type="heading">
          <subject>Original research</subject>
        </subj-group>
      </article-categories>
      <title-group>
        <article-title>Larsson S. Random Data Splits Create Artificial Confidence in Virtual Screening, Molecular Property Prediction, and Drug–Target Modeling</article-title>
      </title-group>
                    <contrib-group>
                      <contrib contrib-type="author">
              <name>
                <surname>Nilsson</surname>
                <given-names>Anders</given-names>
              </name>
                              <xref rid="aff1" ref-type="aff">1</xref>
                                                            <xref rid="cor1" ref-type="corresp" />
                          </contrib>
                      <contrib contrib-type="author">
              <name>
                <surname>Lindholm</surname>
                <given-names>Emma</given-names>
              </name>
                              <xref rid="aff2" ref-type="aff">2</xref>
                                        </contrib>
                      <contrib contrib-type="author">
              <name>
                <surname>Cooper</surname>
                <given-names>James</given-names>
              </name>
                              <xref rid="aff3" ref-type="aff">3</xref>
                                        </contrib>
                      <contrib contrib-type="author">
              <name>
                <surname>Larsson</surname>
                <given-names>Sven</given-names>
              </name>
                              <xref rid="aff4" ref-type="aff">4</xref>
                                        </contrib>
                  </contrib-group>
                  <aff id="aff1">
            <label>1</label>Department of Robust Model Validation and Data Splitting, Faculty of Pharmacy, Swedish University of Agricultural Sciences, Uppsala, Sweden.
          </aff>
                  <aff id="aff2">
            <label>2</label>Department of Artificial Confidence and Leakage Detection, Faculty of Pharmaceutical Sciences, University of Copenhagen, Copenhagen, Denmark.
          </aff>
                  <aff id="aff3">
            <label>3</label>Department of Virtual Screening Reliability, Faculty of Pharmacy, University of Leeds, Leeds, United Kingdom.
          </aff>
                  <aff id="aff4">
            <label>4</label>Department of Drug–Target Modeling and Evaluation, Faculty of Pharmacy, University of Helsinki, Helsinki, Finland.
          </aff>
                          <author-notes>
            <corresp id="cor1">
              <bold>Address for correspondence:</bold> Prof. Wael Abu Dayyih, Department of
              Pharmaceutical Chemistry, Faculty of Pharmacy, Mutah University, Al-Karak 61710, Jordan.
                              E-mail: <email xlink:href="anders.nilsson@slu.se">anders.nilsson@slu.se</email>
                          </corresp>
          </author-notes>
                    <pub-date pub-type="epub">
        <day>28</day>
        <month>12</month>
        <year>2024</year>
      </pub-date>
      <volume>15</volume>
      <issue>6</issue>
      <fpage>87</fpage>
      <lpage>97</lpage>
      <permissions>
        <copyright-statement>
          Copyright: &#x000a9; 2026 Pharmacophore
        </copyright-statement>
        <copyright-year>2026</copyright-year>
        <license>
          <ali:license_ref xmlns:ali="http://www.niso.org/schemas/ali/1.0/"
            specific-use="textmining" content-type="ccbyncsalicense">
            https://creativecommons.org/licenses/by-nc-sa/4.0/</ali:license_ref>
          <license-p>This is an open access journal, and articles are distributed under the terms of
            the Creative Commons Attribution-NonCommercial-ShareAlike 4.0 License, which allows
            others to remix, tweak, and build upon the work non-commercially, as long as appropriate
            credit is given and the new creations are licensed under the identical terms.</license-p>
        </license>
      </permissions>
      <abstract>
        <title>A<sc>BSTRACT</sc></title>
        <p>Machine-learning models are increasingly used to prioritize compounds, estimate molecular properties, and infer drug–target relationships, yet their reported performance often depends more strongly on evaluation design than headline metrics reveal. Random data splitting remains common because it is simple, statistically efficient, and compatible with standardized model comparison. In chemically structured pharmaceutical datasets, however, random allocation frequently places closely related molecules, shared scaffolds, homologous targets, replicated assays, or records from the same experimental source on both sides of the training–test boundary. The resulting test set may therefore measure interpolation within familiar evidence rather than performance on the novel compounds, targets, or experimental conditions implied by a prospective deployment claim. This Original Validation-Standards Article develops a conceptual account of how chemical similarity, scaffold leakage, and hidden dependence generate artificial confidence across virtual screening, molecular-property prediction, and drug–target modeling. It proposes Deployment-Aligned Validation Credibility as a structured construct linking the intended pharmaceutical use to a dependence audit, a shift-matched split strategy, task-appropriate metrics, uncertainty assessment, external corroboration, and transparent reporting. The analysis distinguishes random interpolation from scaffold transfer, chemical-cluster transfer, temporal prediction, entity-disjoint evaluation, source-independent testing, and prospective experimental confirmation. It also emphasizes that no single performance measure can establish pharmaceutical usefulness, mechanistic validity, experimental reproducibility, or therapeutic value. The proposed construct is methodological rather than empirically validated and remains conditional on dataset provenance, assay comparability, similarity definitions, and the practical decision being supported. Its purpose is to make validation claims narrower, more interpretable, and more consistent with the distribution shifts that molecular models will encounter in actual pharmaceutical research.</p>
      </abstract>
      <kwd-group>
                <kwd>Data splitting</kwd>
                <kwd>Scaffold leakage</kwd>
                <kwd>Virtual screening</kwd>
                <kwd>Molecular-property prediction</kwd>
                <kwd>Drug–target interaction</kwd>
                <kwd>Distribution shift</kwd>
              </kwd-group>
    </article-meta>
  </front>
</article>