<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD with MathML3 v1.3 20210610//EN" "JATS-archivearticle1-3-mathml3.dtd"><article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"
  dtd-version="1.3" xml:lang="en" article-type="research-article">
  <?DTDIdentifier.IdentifierValue -//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.2 20190208//EN?>
  <?DTDIdentifier.IdentifierType public?>
  <?SourceDTD.DTDName JATS-journalpublishing1.dtd?>
  <?SourceDTD.Version 1.2?>
  <?ConverterInfo.XSLTName jats2jats3.xsl?>
  <?ConverterInfo.Version 1?>
  <?properties open_access?>
  <front>
    <journal-meta>
      <journal-id journal-id-type="iso-abbrev">Pharmacophore</journal-id>
      <journal-id journal-id-type="publisher-id">pharmacophorejournal.com</journal-id>
      <journal-id journal-id-type="publisher-id">Pharmacophore</journal-id>
      <journal-title-group>
        <journal-title>Pharmacophore</journal-title>
      </journal-title-group>
      <issn pub-type="epub">2229-5402</issn>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="publisher-id">pharmacophorejournal.com-6928</article-id>
      <article-id pub-id-type="doi">10.51847/rbFe0o1wZP</article-id>
      <article-categories>
        <subj-group subj-group-type="heading">
          <subject>Original research</subject>
        </subj-group>
      </article-categories>
      <title-group>
        <article-title>A Foundation Model Should Not Flatten Pharmaceutical Evidence across Molecules, Sequences, Omics, Images, and Scientific Text</article-title>
      </title-group>
                    <contrib-group>
                      <contrib contrib-type="author">
              <name>
                <surname>Khoury</surname>
                <given-names>Omar</given-names>
              </name>
                              <xref rid="aff1" ref-type="aff">1</xref>
                                                            <xref rid="cor1" ref-type="corresp" />
                          </contrib>
                      <contrib contrib-type="author">
              <name>
                <surname>Ross</surname>
                <given-names>Elizabeth</given-names>
              </name>
                              <xref rid="aff2" ref-type="aff">2</xref>
                                        </contrib>
                      <contrib contrib-type="author">
              <name>
                <surname>Lefevre</surname>
                <given-names>Jean-Marc</given-names>
              </name>
                              <xref rid="aff3" ref-type="aff">3</xref>
                                        </contrib>
                      <contrib contrib-type="author">
              <name>
                <surname>Al-Hammadi</surname>
                <given-names>Aisha</given-names>
              </name>
                              <xref rid="aff4" ref-type="aff">4</xref>
                                        </contrib>
                  </contrib-group>
                  <aff id="aff1">
            <label>1</label>Department of Multimodal Pharmaceutical Foundation Models, Faculty of Pharmacy, American University of Beirut, Beirut, Lebanon.
          </aff>
                  <aff id="aff2">
            <label>2</label>Department of Evidence Preservation and Modality Integrity, College of Pharmacy, University of Arizona, Tucson, United States.
          </aff>
                  <aff id="aff3">
            <label>3</label>Department of Non-Flattened Molecular and Omics Representation, Faculty of Sciences, University of Montpellier, Montpellier, France.
          </aff>
                  <aff id="aff4">
            <label>4</label>Department of Scientific Text and Image Integration, College of Medicine and Health Sciences, UAE University, Al Ain, United Arab Emirates.
          </aff>
                          <author-notes>
            <corresp id="cor1">
              <bold>Address for correspondence:</bold> Prof. Wael Abu Dayyih, Department of
              Pharmaceutical Chemistry, Faculty of Pharmacy, Mutah University, Al-Karak 61710, Jordan.
                              E-mail: <email xlink:href="omar.khoury@aub.edu">omar.khoury@aub.edu</email>
                          </corresp>
          </author-notes>
                    <pub-date pub-type="epub">
        <day>28</day>
        <month>12</month>
        <year>2024</year>
      </pub-date>
      <volume>15</volume>
      <issue>6</issue>
      <fpage>108</fpage>
      <lpage>117</lpage>
      <permissions>
        <copyright-statement>
          Copyright: &#x000a9; 2026 Pharmacophore
        </copyright-statement>
        <copyright-year>2026</copyright-year>
        <license>
          <ali:license_ref xmlns:ali="http://www.niso.org/schemas/ali/1.0/"
            specific-use="textmining" content-type="ccbyncsalicense">
            https://creativecommons.org/licenses/by-nc-sa/4.0/</ali:license_ref>
          <license-p>This is an open access journal, and articles are distributed under the terms of
            the Creative Commons Attribution-NonCommercial-ShareAlike 4.0 License, which allows
            others to remix, tweak, and build upon the work non-commercially, as long as appropriate
            credit is given and the new creations are licensed under the identical terms.</license-p>
        </license>
      </permissions>
      <abstract>
        <title>A<sc>BSTRACT</sc></title>
        <p>Pharmaceutical foundation models are increasingly expected to learn across chemical structures, biological sequences, molecular profiles, experimental images, and scientific language. Although joint representation learning may increase task coverage, naive multimodal fusion can erase the differences that determine what each source actually measures, predicts, or claims. Molecular structures encode chemical identity through representation-dependent abstractions; sequences reflect evolutionary and functional regularities; omics measurements describe context-sensitive biological states; images capture assay-conditioned phenotypes; and scientific text contains reported interpretations shaped by publication, terminology, and evidentiary quality. This article develops a proposed evidence-preserving multimodal pharmaceutical foundation-model architecture for organizing these non-equivalent evidence classes without reducing them to anonymous latent vectors. The architecture combines modality-specific evidence contracts, specialized encoders, typed alignment interfaces, conditional fusion, missing-modality management, disagreement preservation, uncertainty calibration, abstention, and bounded scientific-use controls. Its central principle is that fusion should be authorized by the scientific task, relation type, provenance, biological context, modality availability, uncertainty state, and intended decision rather than applied automatically whenever multiple inputs are available. Evaluation is correspondingly separated into within-modality representation validity, cross-modal alignment, transfer, grounding, calibration, disagreement handling, missingness robustness, and prospective scientific usefulness. The proposed architecture remains conceptual: it does not establish empirical superiority, mechanistic validity, clinical utility, regulatory acceptability, or deployment readiness. Its contribution is a methodological basis for designing pharmaceutical foundation models that preserve evidentiary distinctions, expose unsupported inferential transitions, and support more disciplined evaluation of when multimodal learning may contribute to drug discovery and development.</p>
      </abstract>
      <kwd-group>
                <kwd>Multimodal foundation models</kwd>
                <kwd>Pharmaceutical data science</kwd>
                <kwd>Evidence preservation</kwd>
                <kwd>Molecular representation</kwd>
                <kwd>Multi-omics integration</kwd>
                <kwd>Conditional fusion</kwd>
              </kwd-group>
    </article-meta>
  </front>
</article>