Close Menu
NCIJ Network NCIJ Network
    What's Hot

    Brinks Home Discloses Data Breach as Hackers Leak Files

    August 3, 2026

    Ten mystery investors are using 2,380 BTC to completely hijack a Nasdaq company and gut its leadership

    August 3, 2026

    EIG’s MidOcean Energy lines up new investment as NYK spreads its LNG wings

    August 3, 2026
    Facebook X (Twitter) Instagram
    Trending
    • Brinks Home Discloses Data Breach as Hackers Leak Files
    • Ten mystery investors are using 2,380 BTC to completely hijack a Nasdaq company and gut its leadership
    • EIG’s MidOcean Energy lines up new investment as NYK spreads its LNG wings
    • Rejected Wisconsin data center proposal had guaranteed tax revenue, housing
    • Washington’s Badger Mountain Solar Project Canceled by Developer — ProPublica
    • Did Taylor Farms or its CEO donate to Trump? We followed the money
    • Red Cross visits Aung San Suu Kyi after Myanmar military pushed to prove she is alive | Myanmar
    • Richard Tice investigated by parliamentary standards watchdog
    • About
      • Our Team
      • Editorial Policy
      • Editorial Independence
      • International Support
    • Trust & Standards
      • AI Usage Policy
      • Conflict of Interest Policy
      • Corrections Policy
      • Ethics Policy
      • Fact-Checking Policy
      • Source Protection
    • Get Involved
      • Guide for Sources
      • Support Independent Journalism
    • Legal
      • Cookie Policy
      • Privacy Policy
      • Terms of Use
    Facebook X (Twitter) Instagram
    NCIJ Network NCIJ Network
    Monday, August 3
    • Home
    • World
    • Ai
    • Business
    • Politics
    • Health
    • Crypto
    • Science
    • Technology
    • Cybersecurity
    • Defense & Security
    • Economy
    • Energy
    • Europe
    • More
      • Fact Check
      • Investigations
      • Opinion & Analysis
      • Environment
    NCIJ Network NCIJ Network
    Home»Artificial Intelligence

    Why biological data matters more in AI drug discovery

    NCIJ NETWNCIJ NETWORKBy NCIJ NETWNCIJ NETWORKAugust 3, 2026 Artificial Intelligence No Comments5 Mins Read
    Share
    Facebook Twitter LinkedIn Pinterest Email

    GSK has entered into a research collaboration with British biotechnology company Relation Therapeutics worth up to $110 million, expanding the companies’ existing work in AI-assisted drug discovery.

    Under the agreement, Relation will generate large-scale datasets measuring how human cells respond to genetic changes and drug interventions. The data will be used to train AI models designed to identify potential drug targets, including models within Relation’s MORGAN platform.

    The agreement places biological data generation alongside AI model development. Relation’s research approach links computational analysis with experiments that generate new information on human cells.

    The collaboration builds on earlier agreements between GSK and Relation focused on fibrotic diseases and osteoarthritis. Those projects involved observational studies designed to create two functional disease datasets for analysis using Relation’s Lab-in-the-Loop platform.

    The earlier work combined human genetics, single-cell multi-omics generated from human tissue, functional assays, and machine learning to identify and validate potential disease targets.

    How Relation generates biological data

    Relation describes its Lab-in-the-Loop approach as a combination of laboratory experimentation and computational analysis. Its work includes tissue profiling, single-cell and spatial transcriptomics, sequencing, and target validation, while machine learning is used for target identification, prioritisation, validation, and experimental design.

    The company also conducts perturbation experiments that measure how genetic changes affect cellular characteristics associated with disease. Those results can then be analysed alongside genetic and patient-derived biological data.

    Public repositories remain an important source of training material for biological foundation models, although combining information produced across different studies can introduce technical challenges.

    A 2025 review in Experimental & Molecular Medicine noted that repositories including CZ CELLxGENE, the Human Cell Atlas, and NCBI Gene Expression Omnibus give researchers access to large volumes of single-cell data. CZ CELLxGENE alone provides access to more than 100 million standardised cells, according to the review.

    Sampling methods, sequencing protocols, experimental procedures, and processing pipelines can differ between studies. Single-cell data can also contain technical noise and other artefacts, requiring careful dataset selection, filtering, composition balancing, and quality control during foundation-model training.

    Dataset overlap presents another issue. The review noted that the same or similar cells can appear across multiple public resources, potentially giving them disproportionate influence during training and creating data-leakage risks when training and test datasets overlap.

    The review found that assembling a high-quality, non-redundant dataset is as important as model architecture when building robust single-cell foundation models.

    Bigger biological datasets do not guarantee better models

    Research published in Nature Methods in June this year examined how the size and diversity of pretraining data affected single-cell foundation models using a corpus of 22.2 million cells. Researchers trained 400 models and evaluated them across 6,400 experiments.

    The study found that current single-cell foundation models tended to reach performance plateaus after training on only a fraction of the available corpus. Unlike large language models, the systems assessed did not display clear data-scaling laws in which continually increasing training data consistently produced better results.

    The researchers found that model capacity, dataset size, and computational resources need to be balanced rather than simply increased together. The study did not establish that smaller or proprietary datasets are inherently better, but it found that adding more biological training data did not consistently lead to further performance gains.

    A separate study published in Genome Biology in 2025 assessed two single-cell foundation models, Geneformer and scGPT, across several zero-shot evaluation tasks. The models did not consistently outperform simpler approaches, while the researchers also identified challenges involving batch effects and cautioned against assuming that larger pretrained models automatically produce better biological representations.

    Pharma companies pursue specialised datasets

    Relation has already applied its data-generation approach to Osteomics, which it describes as a proprietary functional single-cell bone atlas. The project uses patient-derived samples and combines single-cell and spatial omics with imaging, genomics, proteomics, and clinical phenotype data.

    According to the company, Osteomics is being used to investigate disease biology, therapeutic targets, biomarkers, and patient subgroups in osteoporosis. Hospitals and research partners in the UK and Australia are involved in the observational study.

    Research published in Nature Genetics last month also examined the cellular and genetic determinants of skeletal disease using single-cell analysis, genetic data, and functional validation. Several Relation researchers were among the study’s authors.

    A 2025 Nature Biotechnology analysis of AI-focused biopharma deals identified specialised dataset providers as one of several trends emerging from recent partnerships. Other trends included larger upfront payments, new therapeutic modalities, and greater participation from larger biotechnology companies.

    The analysis said high-quality, disease-specific datasets are becoming an important input for causal and generative machine-learning models. It cited GSK’s separate agreement with Ochre Bio, worth $37.5 million for data licensing involving human liver single-cell and perfused-organ data.

    Another example involved AstraZeneca and Pathos AI entering a $200 million agreement with Tempus in 2025. Under the arrangement, Pathos was to develop oncology foundation models using de-identified clinical, genomic, and imaging data covering more than 150,000 patients.

    Access to sufficient high-quality data remains a constraint in AI drug discovery. A Nature research highlight on federated learning in pharmaceutical research identified limited access to suitable training data as a major bottleneck for AI applications, while noting that companies can also face restrictions on sharing proprietary information.

    AI-biopharma agreements therefore vary in how companies obtain data and computational capabilities. Some centre on access to AI platforms, while others cover joint development, data licensing, or the creation of new biological datasets.

    The GSK–Relation agreement includes both data generation and model development. Relation will produce human cellular datasets as part of the collaboration and use them to train AI models for identifying potential drug targets.

    (Photo by CDC)

    See also: How AI is shortening drug discovery timelines in China

    Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information.

    AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.

    biological data discovery drug matters
    NCIJ NETWNCIJ NETWORK
    • Website

    Keep Reading

    Brinks Home Discloses Data Breach as Hackers Leak Files

    Rejected Wisconsin data center proposal had guaranteed tax revenue, housing

    Alibaba Qwen Releases Qwen3.8-Max: A 2.4 Trillion Parameter MoE Model and the Most Capable One in the Qwen Family to Date

    Cogent AI Team Releases VR-1: A Frontier Cyber Reasoning Model That Composes and Verifies Enterprise Attack Paths

    Onton Releases Ontology 1: A Neurosymbolic Search Model That is 2.7x More Accurate than the World’s Best E-commerce Search Engines

    A Tutorial on GeoAI: Designing Footprint Extraction from NAIP Imagery Using U-Net, Grounding DINO, SAM, and Mask R-CNN

    Add A Comment
    Leave A Reply Cancel Reply

    Editors Picks

    Brinks Home Discloses Data Breach as Hackers Leak Files

    August 3, 2026

    Ten mystery investors are using 2,380 BTC to completely hijack a Nasdaq company and gut its leadership

    August 3, 2026

    EIG’s MidOcean Energy lines up new investment as NYK spreads its LNG wings

    August 3, 2026

    Rejected Wisconsin data center proposal had guaranteed tax revenue, housing

    August 3, 2026
    Latest Posts

    Oil has harmed the nature and people of the Niger Delta; human rights may save it

    July 23, 2026

    Drought announcement looms for parts of Wales over river levels

    July 23, 2026

    Swiss Bank BancaStato Launches Bitcoin Trading Through Sygnum And Avaloq

    July 23, 2026

    Subscribe to News

    Get the latest sports news from NewsSite about world, sports and politics.

    NCIJ Network is an independent digital news platform delivering trusted investigative journalism, European and global news, in-depth analysis, and fact-based reporting with accuracy, transparency, and integrity.

    Facebook X (Twitter) Instagram Pinterest YouTube

    Brinks Home Discloses Data Breach as Hackers Leak Files

    August 3, 2026

    Ten mystery investors are using 2,380 BTC to completely hijack a Nasdaq company and gut its leadership

    August 3, 2026

    EIG’s MidOcean Energy lines up new investment as NYK spreads its LNG wings

    August 3, 2026

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    Type above and press Enter to search. Press Esc to cancel.