<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD with MathML3 v1.3 20210610//EN" "JATS-archivearticle1-3-mathml3.dtd"><article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"
  dtd-version="1.3" xml:lang="en" article-type="research-article">
  <?DTDIdentifier.IdentifierValue -//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.2 20190208//EN?>
  <?DTDIdentifier.IdentifierType public?>
  <?SourceDTD.DTDName JATS-journalpublishing1.dtd?>
  <?SourceDTD.Version 1.2?>
  <?ConverterInfo.XSLTName jats2jats3.xsl?>
  <?ConverterInfo.Version 1?>
  <?properties open_access?>
  <front>
    <journal-meta>
      <journal-id journal-id-type="iso-abbrev">Pharmacophore</journal-id>
      <journal-id journal-id-type="publisher-id">pharmacophorejournal.com</journal-id>
      <journal-id journal-id-type="publisher-id">Pharmacophore</journal-id>
      <journal-title-group>
        <journal-title>Pharmacophore</journal-title>
      </journal-title-group>
      <issn pub-type="epub">2229-5402</issn>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="publisher-id">pharmacophorejournal.com-6943</article-id>
      <article-id pub-id-type="doi">10.51847/bWOObpCFdP</article-id>
      <article-categories>
        <subj-group subj-group-type="heading">
          <subject>Original research</subject>
        </subj-group>
      </article-categories>
      <title-group>
        <article-title>One Molecular Language for Chemical Structures, Biological Sequences, Three-Dimensional Conformations, Spectra, Images, and Scientific Text</article-title>
      </title-group>
                    <contrib-group>
                      <contrib contrib-type="author">
              <name>
                <surname>Fedorova</surname>
                <given-names>Elena</given-names>
              </name>
                              <xref rid="aff1" ref-type="aff">1</xref>
                                                            <xref rid="cor1" ref-type="corresp" />
                          </contrib>
                      <contrib contrib-type="author">
              <name>
                <surname>Volkov</surname>
                <given-names>Alexey</given-names>
              </name>
                              <xref rid="aff1" ref-type="aff">1</xref>
                                        </contrib>
                      <contrib contrib-type="author">
              <name>
                <surname>Morozova</surname>
                <given-names>Irina</given-names>
              </name>
                              <xref rid="aff2" ref-type="aff">2</xref>
                                        </contrib>
                      <contrib contrib-type="author">
              <name>
                <surname>Kuznetsov</surname>
                <given-names>Dmitry</given-names>
              </name>
                              <xref rid="aff3" ref-type="aff">3</xref>
                                        </contrib>
                  </contrib-group>
                  <aff id="aff1">
            <label>1</label>Department of Unified Molecular Language, Faculty of Pharmacy, Novosibirsk State University, Novosibirsk, Russia.
          </aff>
                  <aff id="aff2">
            <label>2</label>Department of Chemical Structures and Biological Sequences, Faculty of Pharmaceutical Sciences, Ural Federal University, Yekaterinburg, Russia.
          </aff>
                  <aff id="aff3">
            <label>3</label>Department of Conformations, Spectra, Images, and Text Integration, Faculty of Pharmacy, Kazan Federal University, Kazan, Russia.
          </aff>
                          <author-notes>
            <corresp id="cor1">
              <bold>Address for correspondence:</bold> Prof. Wael Abu Dayyih, Department of
              Pharmaceutical Chemistry, Faculty of Pharmacy, Mutah University, Al-Karak 61710, Jordan.
                              E-mail: <email xlink:href="elena.fedorova@nsu.ru">elena.fedorova@nsu.ru</email>
                          </corresp>
          </author-notes>
                    <pub-date pub-type="epub">
        <day>28</day>
        <month>12</month>
        <year>2025</year>
      </pub-date>
      <volume>16</volume>
      <issue>6</issue>
      <fpage>89</fpage>
      <lpage>98</lpage>
      <permissions>
        <copyright-statement>
          Copyright: &#x000a9; 2026 Pharmacophore
        </copyright-statement>
        <copyright-year>2026</copyright-year>
        <license>
          <ali:license_ref xmlns:ali="http://www.niso.org/schemas/ali/1.0/"
            specific-use="textmining" content-type="ccbyncsalicense">
            https://creativecommons.org/licenses/by-nc-sa/4.0/</ali:license_ref>
          <license-p>This is an open access journal, and articles are distributed under the terms of
            the Creative Commons Attribution-NonCommercial-ShareAlike 4.0 License, which allows
            others to remix, tweak, and build upon the work non-commercially, as long as appropriate
            credit is given and the new creations are licensed under the identical terms.</license-p>
        </license>
      </permissions>
      <abstract>
        <title>A<sc>BSTRACT</sc></title>
        <p>Pharmaceutical knowledge is distributed across chemical graphs and strings, biological sequences, three-dimensional conformations, analytical spectra, cellular images, and scientific text. Although contemporary representation-learning systems can encode several of these sources, their outputs commonly remain divided by incompatible units, acquisition processes, invariances, uncertainty structures, and validation conventions. This fragmentation limits cross-modal retrieval, weakens transfer between molecular and biological contexts, and can encourage the unsupported treatment of correlated representations as interchangeable scientific evidence. This article develops a proposed multimodal molecular-language architecture that distinguishes semantic content that may be shared across modalities from information that must remain modality-specific. The construct combines versioned identity anchors, modality-native representations, a bounded shared semantic layer, private residual channels, provenance records, uncertainty descriptions, contradiction handling, and task-specific evidence gates. Its central premise is that molecular identity, substructure, perturbation, phenotype, function, and experimental context may provide alignment anchors, but alignment does not erase the distinct meanings of coordinates, spectral peaks, image features, sequence patterns, or textual claims. The architecture therefore treats missing modalities, inferred proxies, contextual disagreement, and unresolved contradiction as explicit representational states rather than hidden preprocessing problems. Evaluation is framed as a multidimensional obligation covering retrieval, generation, transfer, native-modality fidelity, uncertainty, contradiction sensitivity, and scientific grounding. The proposal remains conceptual and does not establish empirical superiority, mechanistic validity, therapeutic utility, regulatory acceptability, or deployment readiness. Its principal implication is that progress toward a reusable molecular language requires not merely larger models or unified embeddings, but disciplined preservation of modality-specific evidence and transparent boundaries between learned correspondence and pharmaceutical meaning.</p>
      </abstract>
      <kwd-group>
                <kwd>Multimodal molecular representation</kwd>
                <kwd>Chemical language models</kwd>
                <kwd>Protein language models</kwd>
                <kwd>Three-dimensional molecular learning</kwd>
                <kwd>Molecular spectra</kwd>
                <kwd>Phenotypic imaging</kwd>
              </kwd-group>
    </article-meta>
  </front>
</article>