TY - JOUR T1 - Large Language Models as Pharmaceutical Knowledge Workers: Capabilities, Failure Modes, and Evidence Requirements across Scientific Tasks A1 - Luis Benitez A1 - Andrea Gomez A1 - Jazmin Franco A1 - Roberto Cardozo JF - Pharmacophore JO - Pharmacophore SN - 2229-5402 Y1 - 2025 VL - 16 IS - 4 DO - 10.51847/IwEKFH04Ct SP - 102 EP - 111 N2 - Large language models are increasingly positioned as pharmaceutical knowledge workers capable of searching literature, extracting evidence, summarizing documents, organizing knowledge, proposing hypotheses, assisting molecular design, generating code, and coordinating scientific tools. Yet the evidentiary meaning of these capabilities remains unclear. Strong benchmark performance, fluent explanations, or successful demonstrations in constrained environments do not necessarily establish reliable scientific reasoning, experimental value, workflow benefit, or readiness to influence pharmaceutical decisions. This critical state-of-the-field review evaluates large language models according to the scientific tasks they perform, the evidence units used to assess them, and the consequences of their errors. The review develops a proposed task taxonomy spanning evidence processing, knowledge organization, hypothesis generation, molecular assistance, computational support, workflow orchestration, and decision support. It also distinguishes output validity, source grounding, task-valid performance, expert-comparator evidence, prospective workflow validation, and decision-consequence evidence as separate levels of support that should not be treated as interchangeable. Across the literature, the strongest evidence concerns bounded text-processing and tool-mediated tasks, whereas evidence for autonomous scientific judgment, experimentally confirmed molecular contribution, and routine pharmaceutical deployment remains limited. Recurring limitations include hallucinated content, citation failure, data leakage, shortcut learning, hidden substitution of easier tasks for intended scientific constructs, inadequate source attribution, and under-specified human oversight. The central contribution is a role-based evidentiary interpretation of pharmaceutical language-model use in which validation requirements increase with autonomy, irreversibility, and decision consequence. The proposed synthesis is conceptual rather than empirically validated and should be used to structure evaluation, not to certify systems. Reliable adoption will require source-auditable outputs, task-specific comparators, prospective workflow studies, explicit escalation rules, configuration traceability, and governance that treats the model as one component of a broader scientific system. UR - https://pharmacophorejournal.com/article/large-language-models-as-pharmaceutical-knowledge-workers-capabilities-failure-modes-and-evidence-wxytktxavubqyka ER -