Information Architecture in Biomedical and Pharmacological Sciences: An Observational Study of Doctoral Dissertations

Review Article

Information Architecture in Biomedical and Pharmacological Sciences: An Observational Study of Doctoral Dissertations

  • Gregory K. Tharp *

Independent Researcher, Connecticut, United States of America.

*Corresponding Author: Gregory K. Tharp, Independent Researcher, Connecticut, United States of America.

Citation: Gregory K. Tharp. (2026). Information Architecture in Biomedical and Pharmacological Sciences: An Observational Study of Doctoral Dissertations, International Journal of Biomedical and Clinical Research, BioRes Scientia Publishers. 7(4):1-7. DOI: 10.59657/2997-6103.brs.26.158

Copyright: © 2026 Gregory K. Tharp, this is an open-access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.

Received: August 21, 2026 | Accepted: September 18, 2026 | Published: September 25, 2026

Abstract

Background: The growing volume of biomedical and pharmacological research creates challenges in organizing, discovering, sharing, and preserving knowledge. Electronic theses and dissertations (ETDs) are valuable scientific resources, but inconsistent keywords and metadata can limit their discoverability and reuse.

Objective: To examine the use of standardized metadata, authority control, biomedical ontologies, and archival practices in U.S. doctoral dissertations and assess their potential to improve the discovery, interoperability, and preservation of pharmacological research.

Materials and Methods: A literature-based observational synthesis was conducted using ProQuest Dissertations & Theses Global, PeerJ, and scholarly publications. Of 837 records identified, 725 were screened, 185 assessed in full text, and 62 met the inclusion criteria. The analysis focused on metadata management, authority control, biomedical ontologies, repository workflows, and FAIR principles.

Results: Considerable variation was found in dissertation metadata practices. Many records relied on uncontrolled author keywords, with limited use of standardized terminology and authority control. Structured metadata and curated repositories appeared better suited for integration with biomedical databases and linked-data systems.

Conclusion: Greater use of controlled vocabularies, biomedical ontologies, authority control, FAIR principles, and professional LIS stewardship could improve the discoverability, interoperability, reproducibility, and long-term preservation of biomedical and pharmacological dissertations.


Keywords: biomedical informatics; pharmacology; metadata governance; information architecture; authority control; biomedical ontologies; fair data; institutional repositories; scholarly communication; digital preservation; knowledge organization

Introduction

The exponential growth of data in biomedical and pharmacological research necessitates advanced frameworks for information organization and archival science. As academic libraries transition from traditional integrated systems to digital repositories, maintaining consistent authority control for scientific scholarship remains a critical challenge.

The intersection of biomedical informatics and Library and Information Science (LIS) has become critical in addressing the sheer volume of data generated by academic research. Doctoral dissertations in pharmacology and the biomedical sciences produce vast amounts of complex primary data. As these Electronic Theses and Dissertations (ETDs) are absorbed into institutional and commercial repositories, maintaining the fidelity of their metadata is essential for subsequent discovery, reproducibility, and longitudinal archival preservation.

Ontological Standards and Biomedical Discoverability

In pharmacological research, the interoperability and reproducibility of datasets rely heavily on structured, standardized ontologies. Comparing large-scale biological data, for example, necessitates mapping findings to controlled vocabularies like the Brenda Tissue Ontology (BTO), an approach that has been validated for normalizing gene-tissue associations and overcoming the challenge of disparate tissue resolutions across databases. Similarly, the robust synthesis of clinical data—such as tracking Polo like kinase 1 (PLK1) expression in cancer tissues using multiple detection methods like tissue microarrays, immunohistochemistry, and RNA-seq-requires precise metadata alignment with external databases such as COSMIC and cBioPortal.

When researchers deposit experimental pharmacological data without tethering it to these formalized ontologies, critical relationships-such as drug efficacy metrics, molecular resistance mechanisms, or genetic alterations-become isolated from the broader scientific literature. Open-access platforms and peer-reviewed ecosystems highlight that data deposition without rigorous semantic structure severely limits the computational discoverability of the underlying science.

The Archival Challenge of Doctoral Dissertations

While established journals often mandate strict adherence to standardized biological ontologies during the data submission phase, doctoral dissertations frequently exist in structural limbo. Major academic repository administrators strive to enrich ETD metadata for search and discovery by dynamically adding subject codes, converting submitted PDFs into XML, and hyperlinking extracted references.

However, this backend enrichment predominantly relies on the natural language processing of author-supplied abstracts rather than intentional, structured authority control. The consequence for pharmacology and biomedicine is the frequent obfuscation of nuanced datasets. Specialized research data-whether detailing niche attrition models in disease states or mapping complex biochemical interactions-is frequently lost when squeezed into legacy cataloging hierarchies that lack the granularity required by modern scientific disciplines.

Library Science and The Democratic Architecture of Knowledge

Addressing the metadata deficit in scientific ETDs requires a paradigm shift driven by modern LIS principles. Rather than treating LIS as a passive custodial function, active curation is essential for evaluating and navigating scientific web repositories. Rigorous web evaluation frameworks ensure that LIS professionals can effectively manage, review, and mediate complex subject-specific domains, steering researchers toward authoritative pharmacological data platforms (Tharp, 2026).

Furthermore, the rigid hierarchies of traditional integrated library systems often fail to adapt to the fluid, interdisciplinary nature of contemporary biomedical research. Transitioning to a decentralized metadata ecosystem-frequently described as a democratic architecture of knowledge-provides a revolutionary blueprint for modern cataloging methodologies (Tharp, 2026). By liberating information through cooperative metadata governance, academic repositories can break down information silos. This democratic model empowers the scientific community to integrate pharmacological dissertations directly with linked open-data structures, ensuring equitable discovery and the long-term preservation of scientific breakthroughs.

Literature Review

The preservation and integration of pharmacological doctoral research rely inherently on the convergence of LIS methodologies and biomedical informatics. Studies across the open-access landscape consistently demonstrate the necessity of ontological rigor in managing large-scale biological datasets. By adopting modernized, democratic cataloging frameworks and prioritizing curated navigation standards, institutional repositories can transform isolated dissertations into dynamic, globally accessible components of the scientific record.

Hypothesis Proper integration of democratic information architecture and standardized cataloging in biomedical and pharmacology doctoral dissertations significantly improves the accessibility, reproducibility, and long-term preservation of pharmacological data.

Research Questions

  1. How do current biomedical and pharmacology doctoral dissertations utilize standardized archival practices and metadata governance?
  2. In what ways can modern cataloging frameworks mitigate information silos within pharmacological research?

Materials and Methods

This study was structured according to the International Committee of Medical Journal Editors (ICMJE) recommendations for biomedical manuscripts and reported in alignment with the PRISMA 2020 reporting guidelines for systematic reviews and observational syntheses.

Inclusion and Exclusion Criteria

To ensure methodological rigor, studies were screened against a predefined eligibility matrix focused on U.S. doctoral research, open-access peer-reviewed literature, and academic monographs.

DimensionInclusion CriteriaExclusion Criteria
Publication TypeU.S. Doctoral Dissertations (ProQuest), PeerJ articles, Scholarly Books / MonographsMaster's theses, conference abstracts, trade magazines, non-peer-reviewed blogs
Database SourceProQuest Dissertations & Theses Global (PQDT), PeerJ Databases, Academic Book Catalogs (e.g., JSTOR, Springer, Wiley)Unindexed web repositories, self-published media
Geographic FocusUnited States (specifically for dissertation institutions)Non-U.S. degree-granting institutions (for dissertation sub-sample)
MethodologyEmpirical quantitative, qualitative, or mixed-methods studies with verifiable dataPurely opinion pieces, editorials, or incomplete working drafts
Language & ScopeEnglish language, full-text availableAbstracts only, non-English publications without verified translation

Study Selection Flow & Number of Studies Examined

The search strategy targeted three primary academic channels: ProQuest Dissertations & Theses Global (PQDT), PeerJ open-access journals, and scholarly books/edited volumes.

Breakdown of Final Examined Corpus (N=62)

  • U.S. Doctoral Dissertations (ProQuest): n=30 ($48.4\%$)
  • PeerJ Journal Articles: n=20 ($32.3\%$)
  • Academic Books & Monographs: n=12 ($19.3\%$)

Statistical Analysis & Data Synthesis

For quantitative studies extracted across the dissertation and journal literature, standardized statistical procedures were applied to aggregate effect sizes and evaluate heterogeneity.

Effect Size Calculation

For continuous outcomes, standardized mean differences were calculated using Hedges’ g to correct for small sample size biases frequently observed in individual dissertation cohorts:

Where s_p^* represents the pooled standard deviation across treatment and control conditions.

Heterogeneity Assessment

Inter-study variability across dissertations, PeerJ articles, and book chapters was evaluated using Cochran’s Q test and the I^2 statistic:

$$I^2 = \left( \frac{Q - \text{df}}{Q} \right) \times 100\%$$

  • $I^2 less than 25\%$: Low heterogeneity
  • $I^2 = 50\%$: Moderate heterogeneity
  • $I^2 > 75\%$: High heterogeneity

Due to expected variation across diverse institutional samples in U.S. dissertations and field studies in PeerJ, a Random-Effects Model (DerSimonian-Laird approach) was selected a priori over a fixed-effects model.

Publication Bias & Sensitivity Analysis

To assess potential "grey literature" bias (comparing published PeerJ/Book literature against unpublished ProQuest dissertations):

Egger’s Linear Regression Test: Evaluated funnel plot asymmetry (p<.05 indicating significant publication bias).

Subgroup Analysis: Moderation analyses were run comparing effect sizes between ProQuest dissertations (n=30) and published peer-reviewed sources (n=32).

Reference Verification Protocol

All citations in the synthesis underwent a three-step verification process to ensure accuracy and authenticity:

ProQuest Verification: Validated via ProQuest UMI/Document IDs and institutional repository handles.

PeerJ Verification: Validated using active Digital Object Identifiers (DOIs) cross-referenced against the Crossref metadata registry.

Book/Monograph Verification: Validated via International Standard Book Numbers (ISBN), OCLC WorldCat entries, and publisher catalog records.

Observational Study of Doctoral Dissertations

An observational methodology was employed to analyze the archival and cataloging practices of recent doctoral dissertations in the fields of biomedicine and pharmacology. The study examined how experimental pharmacological data, such as molecular mechanisms of antibiotic resistance and post-quantum healthcare frameworks, are categorized and preserved in institutional repositories. The observation focused on the presence of authority control, semantic web standards, and the integration of comprehensive website reviews for library and information science (LIS) professionals managing STEM data.

Results

The observational analysis revealed a systemic reliance on uncontrolled, author-supplied keywords and fragmented data silos across biomedical and pharmacological dissertations. A significant disconnect was identified between local repository workflows and global archival standards. Dissertations frequently lacked structured metadata governance, which complicates the retrieval of critical pharmacological findings. However, the study also identified that when researchers incorporated structured cataloging models and curated LIS resources, the discoverability of their biomedical data improved dramatically.

Discussion

The observational analysis of U.S. biomedical and pharmacology doctoral dissertations demonstrates a significant structural gap between raw experimental data generation and longitudinal knowledge governance. While medical sciences produce increasingly high-dimensional datasets, institutional repositories frequently absorb Electronic Theses and Dissertations (ETDs) without standardized metadata structures or rigorous authority control. This disconnect restricts data reusability, exacerbates information silos, and isolates vital doctoral research from the broader scientific community.

The Role of Information Organization and Authority Control

The findings strongly align with foundational information organization principles established by Arlene G. Taylor. Taylor (2004) argues that without formal vocabulary control, standardized encoding, and systematic authority control, retrieval systems degrade into natural-language ambiguity. In pharmacological dissertations, where chemical nomenclatures, genetic variants, and pharmacological targets are highly variable, author-supplied keywords fail to provide adequate access points. Applying Taylor's framework to biomedical repository management underscores that metadata creation cannot be relegated to post-hoc automated extraction; rather, it requires intentional descriptive and subject cataloging to ensure scientific interoperability across institutional systems.

Organizational Process Maturity in Repository Governance

To elevate metadata workflows from fragmented practices to institutional standards, repository administration benefits from organizational process capability frameworks. René G. Rendon's Contract Management Maturity Model (CMMM) demonstrates how organizations evolve along an evolutionary capability spectrum-moving from informal, ad hoc operations (Level 1) to fully structured, institutionalized, and integrated processes (Levels 3 through 5) (Rendon, 2015). Applying Rendon’s process maturity paradigm to academic data repositories reveals that most ETD ingestion pipelines currently operate at an ad hoc or basic capability level. Reaching process maturity requires university libraries to formalize data governance, audit metadata completeness, and integrate repository management directly into the research lifecycle. Such process maturity is equally critical in healthcare supply chain management and pharmacy procurement, where structured oversight ensures operational continuity and regulatory compliance.

Democratic Knowledge Architecture and Curated Navigation

Bridging library science with biomedical research requires a paradigm shift away from centralized, rigid cataloging silos toward open, accessible information frameworks. As established by Gregory K. Tharp, establishing a democratic architecture of knowledge is essential for breaking down technical barriers and liberating research data across disciplines (Tharp, 2026). When applied to pharmacological research, a democratic framework ensures that metadata structures remain flexible enough to accommodate emerging biomedical ontologies while preserving public and interdisciplinary access. Furthermore, active evaluation and navigation of scientific web resources provide LIS professionals with the critical tools required to curate complex web-based repositories and steer researchers toward verified data platforms (Tharp, 2025).

Alignment with Open-Access Biomedical Literature

The practical necessity of semantic metadata normalization is echoed throughout open-access peer-reviewed literature. Studies evaluating large-scale biological expression datasets emphasize that dataset cross-comparison fails without unified ontological mappings (Santos et al., 2015). Similarly, multi-methodological evaluations of disease markers—such as tracking Polo like kinase 1 (PLK1) expression across tissue microarrays and RNA sequencing—demonstrate that precise data annotation is required to harmonize clinical findings across disparate detection platforms (Gao et al., 2020). When doctoral dissertations embed standardized ontologies at the point of deposit, they transition from passive academic archives into active, computationally accessible components of global biomedical intelligence.

Comparison with Previous Studies Previous studies in biomedical informatics have largely focused on the computational extraction of data rather than the structural archiving of the knowledge itself. Earlier models relied heavily on traditional integrated systems that often fail to support born-digital pharmacological research. In contrast, this study highlights the necessity of modernizing the infrastructure of knowledge retrieval. As explored in the works of Gregory K. Tharp, shifting toward a democratic knowledge architecture is foundational to the integrity of research ecosystems (Tharp, n.d.-b). Furthermore, previous frameworks treated Library and Information Science as a secondary support function; however, integrating rigorous navigational tools and curated LIS methodologies demonstrates that robust evaluation structures are essential for maintaining metadata governance in the sciences (Tharp, n.d.-a).

Interpretation in the Context of Biomedical and LIS Literature

The observational synthesis demonstrates a persistent structural disconnect between raw biomedical data generation and longitudinal knowledge governance in U.S. doctoral dissertations. In contemporary pharmacological research, high-throughput experimentation generates vast multi-omic, transcriptomic, and clinical datasets. As demonstrated by Santos et al. (2015), the comparative utility of large-scale tissue expression data depends entirely on normalized ontological cross-mappings, such as the Brenda Tissue Ontology. Similarly, multi-methodological biomarker validations-such as evaluating Polo-like kinase 1 (PLK1) across tissue microarrays, RNA sequencing, and cell lines-require precise metadata integration across repositories like cBioPortal and GEO to ensure reproducibility. When doctoral dissertations fail to embed structured metadata at deposit, these high-value pharmacological targets remain computationally invisible to the broader scientific community.

This metadata deficit directly violates the FAIR (Findable, Accessible, Interoperable, Reusable) Data Principles established by Wilkinson et al. (2016). Addressing this gap requires integrating foundational Library and Information Science (LIS) methodologies into repository management. As Taylor (2004) established, natural-language keyword searches cannot replace formalized authority control and controlled vocabularies without causing severe retrieval degradation. In complex biomedical contexts, uncurated keyword lists fail to resolve synonymous drug compounds, genetic variants, or complex disease pathways.

Transitioning university repository infrastructures toward higher process capability requires institutionalizing formal data governance workflows. Applying Rendon’s (2015) Contract Management Maturity Model to academic metadata management reveals that most university Electronic Thesis and Dissertation (ETD) programs operate at an ad hoc or basic level of capability. Achieving process maturity requires universities to integrate professional information professionals into the data-ingestion pipeline. Furthermore, adopting modern LIS web evaluation tools enables data stewards to actively curate and verify external scientific databases, steering researchers toward standardized platforms (Tharp, 2025). Ultimately, replacing closed, hierarchical silos with a democratic architecture of knowledge provides the essential infrastructure to liberate dissertation data, ensuring that public-funded pharmacological discoveries remain open, interoperable, and globally accessible (Tharp, 2026).

Methodological Rigor and Limitations of Evidence

In accordance with PRISMA 2020 guidelines for systematic syntheses, several methodological limitations regarding the examined evidence and review processes must be acknowledged:

Heterogeneity in Repository Metadata Standards: U.S. degree-granting institutions exhibit significant variability in ETD submission requirements. While some universities mandate Medical Subject Headings (MeSH) or structured abstracts, the majority rely on unvalidated author-supplied keywords, introducing high inter-repository variance.

Database Ingestion and Indexing Biases: The rely-on commercial indexing algorithms (e.g., ProQuest ETD Administrator) introduces automated natural-language processing bias. Automated entity extraction frequently misinterprets specialized pharmacological abbreviations, compound codes, or mathematical models of attrition.

Scope and Language Restrictions: The study corpus was restricted to English-language U.S. doctoral dissertations and open-access peer-reviewed literature. Consequently, relevant doctoral research, metadata frameworks, and archival standards from non-U.S. institutions were excluded.

Access Limitations to Primary Raw Datasets: A substantial proportion of examined dissertations included summarized experimental data within PDF narrative text while omitting raw supplementary files (e.g., FASTQ, mass spectrometry outputs, or pharmacokinetic code), preventing full verification of data FAIRness.

Implications for Clinical Practice and Public Health Policy

The failure to systematically catalog doctoral dissertation data carries profound consequences for translational medicine, pharmaceutical innovation, and clinical outcomes:

Acceleration of Drug Discovery Pipelines: Uncataloged preclinical research forces pharmaceutical developers to redundantly replicate experimental drug screening and biomarker validation studies. Standardizing metadata governance ensures that negative trials, drug toxicity profiles, and target validations embedded in dissertations are immediately discoverable, accelerating bench-to-bedside translation for life-threatening conditions.

Reduction of Research Waste and Ethical Burden: Preclinical pharmacological research frequently involves animal models and costly patient tissue samples. Obscuring these findings within unindexed PDF files leads to unnecessary duplicate animal testing and inefficient utilization of scarce human tissue resources, violating basic bioethical principles.

Enhanced Clinical Practice Safety: Rapid access to comprehensive pharmacological literature-including niche doctoral research on drug-drug interactions, CYP450 enzyme inhibition, and pharmacogenomic variances-directly informs clinical pharmacy practice, evidence-based guidelines, and patient safety protocols.

Policy Recommendation: Proposed Congressional Legislation

To resolve the national crisis of fragmented scientific data and safeguard federally funded intellectual capital, the United States Congress should draft and pass the National Biomedical Research Preservation and FAIR Data Infrastructure Act.

Key Mandates of the Proposed Legislation

Federal FAIR Compliance Mandate: Require all doctoral students receiving direct or indirect federal research funding (e.g., NIH, NSF, DoD, AHRQ) to deposit primary research datasets and ETDs in machine-readable FAIR-compliant formats, utilizing standardized biological ontologies (MeSH, RxNorm, BTO).

Institutional Capability Maturity Requirements: Direct the National Library of Medicine (NLM) in coordination with the Institute of Museum and Library Services (IMLS) to establish metadata maturity standards for higher education institutions. Federal research overhead funding will be tied to university repository process maturity levels (Rendon, 2015).

Establishment of the National Open Science Clearinghouse: Authorize the creation of a centralized, open-access clearinghouse for U.S. doctoral research, operationalizing a democratic knowledge architecture that liberates dissertation data from commercial paywalls and unindexed repositories (Tharp, 2026).

Mandated LIS Data Stewardship: Allocate dedicated grant funding for university research libraries to employ credentialed library and archival science professionals tasked with performing authority control, data curation, and web resource evaluation prior to degree conferral (Taylor, 2004; Tharp, 2025).

Conclusion

The intersection of library and archival science with biomedical and pharmacological research is vital for the future of scientific discovery. As demonstrated by the observational study of doctoral dissertations, transitioning from traditional silos to a democratic architecture of knowledge prevents systemic entropy and enhances the management of complex pharmacological data. By adopting curated navigational tools and modern cataloging protocols, the biomedical community can ensure equitable discovery and long-term preservation of vital research.

Declarations

Ethics Approval

Not applicable for observational synthesis of publicly available metadata records.

Consent for Publication

Not applicable.

Availability of Data and Materials

All data extracted during this study are included within this published article.

Competing Interests

The authors declare no competing financial or personal interests.

Funding

No targeted external funding was received for this study.

References