Assessing the effects of environmental exposures on population health requires reliable, comprehensive, and standardized toxicological data to support regulatory decision-making. However, existing databases remain highly fragmented, with narrow and uneven coverage across species and organs, incomplete dose–response relationships, and inconsistent validation chains, all of which limit their utility for risk prediction. At the same time, regulatory science and technological development are rapidly shifting toward artificial intelligence-based computational models, highlighting the urgent need to establish high-quality toxicology databases as the foundation for next-generation methodologies. To address these challenges, a team led by Bin Wang at Peking University and Mingliang Fang at Fudan University systematically examined the major barriers to toxicological data curation, harmonization, and open access, and proposed corresponding strategies. These include establishing confidence-tiering frameworks to make effective use of heterogeneous datasets, applying emerging high-throughput platforms to generate interaction data efficiently and accurately, and promoting research community-based data sharing.
The authors further explained how AI-enabled approaches could improve the organization, interoperability, and usability of toxicological data infrastructure, thereby supporting AI-driven environmental toxicology and mechanism-based human health risk assessment. Specific strategies include integrating knowledge networks to develop mechanism-guided AI models; adopting transfer-learning frameworks that connect large-scale pretraining with fine-tuning on small datasets; and combining knowledge graph augmentation with prompt learning to systematically predict exposure-biology-disease interactions. Together, these efforts aim to transform fragmented resources into a systematic, interpretable, and predictive evidence framework. The authors conclude that building high-quality toxicology databases will accelerate the transition of toxicology toward an AI-driven paradigm and provide a robust foundation for more reliable risk assessment and stronger protection of global environmental health.。

Graphical Abstract
Article type: Perspective
Lead institution: Peking University
First author: Bin Wang
Corresponding author: Mingliang Fang (Fudan University)
Journal: Environmental Science & Technology
Article information: Wang B, Wu T, Shou Y, Ma Y, Ren M, Gago-Ferrero P, Schlenk D, Fang M*. Building the Foundations of AI-Driven Toxicology: How to Use Fragmented Data for Mechanism-Based Human Health Risk Assessment. Environmental Science & Technology. 2026, https://doi.org/10.1021/acs.est.5c17092
全文链接:https://pubs.acs.org/doi/10.1021/acs.est.5c17092

Figure 1. Title Page
Human environmental exposure typically involves complex mixtures of multiple chemicals whose mechanisms of action remain poorly understood and may involve unintended or off-target effects. Traditional toxicological approaches are therefore poorly suited to high-throughput, mechanism-based, and systematic health risk assessment. In recent years, regulatory demands, new approach methodologies, and advances in artificial intelligence have jointly driven environmental toxicology away from its traditional reliance on animal testing and conventional quantitative structure-activity relationship models toward high-throughput screening, omics technologies, organ-on-a-chip systems, knowledge graphs, and mechanism-driven computational models. However, existing toxicological data remain dispersed across institutions and platforms, with uneven coverage of chemicals, species, organs, toxicity endpoints, and dose–response relationships. Inconsistencies in data formats, identifier systems, experimental conditions, and metadata standards further hinder cross-database integration and model generalization. To provide a systematic overview of the data landscape in this field, Table 1 summarizes representative toxicology databases and analytical platforms, including ECOTOX, CompTox, PubChem, ChEMBL, ToxCast, CTD, the OECD QSAR Toolbox, and ExposomeX. It compares their host institutions, data sources, coverage, and core functions, thereby establishing a foundation for examining the limitations of current databases and proposing AI-driven strategies for data integration.
Table 1. Overview of Current Representative Toxicology Databases and Analytical Platforms (partial snapshot)

Current toxicology databases are diverse and can broadly be classified into primary data resources and integrated analytical platforms, covering fields such as ecotoxicology, chemical and biological activity, drug–target interactions, endocrine disruption, and the exposome. ECOTOX, NORMAN, and MarineTox Predictor support multispecies ecological risk assessment, whereas PubChem, ChEMBL, and CompTox provide large-scale information on chemical structures, biological activities, and toxicity endpoints. ExposomeX further connects exposure–biology–disease relationships. However, existing data remain distributed across heterogeneous experimental systems and databases, with substantial differences in species and organ coverage, toxicity endpoints, exposure conditions, and metadata completeness. Inconsistent chemical identifiers, data formats, terminologies, and experimental reporting standards impede cross-platform integration and interoperability. Research-generated data capture substantial diversity but are often highly heterogeneous, whereas regulatory data are generally more standardized but may be constrained by limited accessibility and transparency. Together, these limitations contribute to the key challenges summarized in Figure 1, including fragmented evidence, imbalanced data coverage, and insufficient experimental consistency.

Figure 1 | Three major challenges in building high-quality toxicology databases. Toxicological evidence is highly heterogeneous and often lacks a complete validation chain, while data coverage across different targets and toxicity endpoints is markedly imbalanced. Existing data are concentrated in a limited number of model species and organs, constraining cross-species extrapolation and mechanistic inference. In addition, experimental testing is costly, highly endpoint-specific, and subject to interlaboratory variability, further reducing the reliability of data integration, reuse, and risk prediction.
To address the major bottlenecks of fragmented toxicological data, uneven data quality, and insufficient mechanistic information, the study proposes a systematic framework for AI-driven health risk assessment (Figure 2). First, rather than simply excluding all low-quality data, the framework stratifies datasets according to the degree of experimental standardization, reproducibility, and metadata completeness, and establishes a confidence assessment system based on dimensions such as completeness, consistency, and task relevance. High-confidence data can be used directly for model training, medium-confidence data can be incorporated through weighting, and low-confidence data can be used for robustness testing or hypothesis generation. This approach expands the coverage of chemicals and toxicity endpoints while maintaining analytical reliability. Second, the study emphasizes the use of emerging high-throughput technologies, including protein–ligand screening, DNA-encoded chemical libraries, affinity-selection mass spectrometry, transcriptomics, metabolomics, and single-cell multi-omics, to generate standardized biological data at scale under harmonized experimental conditions. These technologies can help identify chemical targets, molecular initiating events, key pathways, and potential adverse outcomes, thereby providing a more comprehensive data foundation for mechanistic research and toxicity prediction. Finally, knowledge graphs and large language model agents can further integrate chemical structures, biological responses, omics information, and evidence from the literature to support chemical–target prediction, omics-to-pathway mapping, mixture-effect analysis, and risk prioritization. This framework is expected to transform dispersed toxicological data into traceable, interpretable, and predictive mechanistic evidence, advancing environmental health risk assessment from experience-based judgment toward a new stage driven by knowledge and enabled by AI.

Figure 2 | Building high-quality toxicology databases to advance AI-driven human health risk assessment. This goal can be pursued from three complementary directions. First, toxicological data should be stratified according to the degree of data characterization, and a confidence evaluation system should be established based on dimensions such as completeness, consistency, and relevance. Second, emerging high-throughput technologies, including protein–ligand screening, transcriptomics, and metabolomics, can be used to generate biological data at scale and expand mechanistic coverage. Finally, knowledge graphs and large language model agents can integrate chemical information, biological responses, and knowledge from the literature to support AI-based toxicological analysis and risk assessment.
Figure 3 further illustrates how sparse toxicological data can be transformed through AI into health risk assessment tools with both generalizability and mechanistic interpretability. Early models mainly relied on transductive approaches such as matrix factorization, network diffusion, and node2vec, which depended on the complete graph structure observed during training and therefore struggled to handle unseen chemicals, unknown targets, and new toxicity endpoints. With the development of graph neural networks, models have gradually shifted toward inductive learning, enabling prediction for new entities through message passing and local subgraph aggregation. Building on this, self-supervised pretraining and few-shot fine-tuning allow knowledge transfer from large-scale chemical–biological networks, while graph prompt learning further injects task-specific information to reduce the mismatch between pretrained knowledge and downstream toxicological tasks.
At the application level, large language models can automatically extract and structure toxicological evidence from databases, tables, reports, and the literature. Multimodal learning can further integrate chemical structures, knowledge graphs, and omics features into a unified biological representation space. By constructing exposome knowledge graphs, multidimensional individual exposure profiles, biomarkers, and health outcomes can be linked, thereby supporting risk identification under complex environmental conditions. At the same time, models can trace key mechanisms along the chain of “chemical–molecular target–pathway–disease endpoint,” shifting risk prediction from black-box outputs toward interpretable biological inference.
The study also points out that foundation models may drive toxicology from one-to-one prediction toward many-to-many systems modeling, while enabling zero-shot or few-shot prediction for data-scarce chemicals. In the future, the deep integration of environmental exposure data, ecotoxicological evidence, omics information, and biological knowledge networks—combined with external validation, benchmark testing, and uncertainty assessment to ensure model reliability—will be crucial for AI to genuinely serve environmental health risk assessment and regulatory decision-making.

Figure 3 | Paradigm evolution and integrative framework for transforming sparse toxicological data into generalizable, mechanism-driven health risk assessment. AI models have evolved from transductive approaches, such as matrix factorization and graph embedding, which depend on complete graph structures, toward inductive graph neural networks capable of predicting previously unseen chemicals and toxicity endpoints. The integration of pretraining, fine-tuning, and graph prompt learning further improves model adaptability to few-shot tasks. Meanwhile, large language models facilitate the extraction of evidence from databases, tables, and the literature, while multimodal learning integrates chemical structures, knowledge graphs, and omics information. Exposome knowledge graphs can then support individualized risk assessment, and the tracing of “chemical–pathway–disease endpoint” chains enhances biological interpretability and mechanistic inference.
Summary and Outlook
To meet the need for reliable population health risk assessment, future toxicology database development must move beyond simple data accumulation toward a systematic enterprise centered on high-quality, shareable, interpretable, and machine-actionable data. Emerging high-throughput technologies, tiered data-use strategies, community-based data sharing, knowledge-guided AI, and transfer learning should be further integrated to improve the harmonization of heterogeneous data and the generalizability of predictive models. At the same time, greater emphasis must be placed on standardization, traceability, external validation, and uncertainty assessment to ensure that risk predictions can meaningfully inform regulatory decision-making. Building high-quality toxicology databases is not only essential for methodological advancement but also represents a shared global scientific responsibility and public health mission.
Author Introduction
First Author:

Bin Wang is a tenured Associate Professor and Researcher at the Institute of Reproductive and Child Health, Peking University, and an Adjunct Professor at the College of Urban and Environmental Sciences, Peking University.
His research focuses primarily on environmental health, exposomics, big data, and artificial intelligence. He led the development of the ExposomeX platform (www.exposomex.cn), which supports causal-inference research on exposure–biology–disease relationships. To date, he has served as principal investigator for four Young Scientists Fund and General Program grants from the National Natural Science Foundation of China and has contributed as a key investigator to three National Key Research and Development Program projects. As first or corresponding author, he has published more than 70 papers in internationally recognized journals, including Environmental Health Perspectives, Environmental Science & Technology, The Lancet Regional Health – Western Pacific, and The Innovation. He has an H-index of 54 and has received more than 6,500 citations. He serves as an Associate Editor of Environmental Science & Technology, with editorial responsibilities covering artificial intelligence, big data, and environmental health. He developed and teaches the Peking University undergraduate and graduate curriculum-innovation course “Exposomics”, which was recognized as one of Peking University’s Top Ten Outstanding Undergraduate Teaching Cases. He also serves as Chair of the “Environment and Population Health” group of the China Cohort Consortium and Deputy Secretary-General of the Specialized Committee on Environment and Reproductive Health of the Chinese Environmental Mutagen Society. His honors include the Second Prize for Science and Technology from the Beijing Preventive Medicine Association and the national title of “Outstanding Individual in China’s Science and Technology System for Combating COVID-19”.
Corresponding Author

Mingliang Fang is currently working as a full professor at Fudan University, Shanghai, China.
Dr. Fang received his Ph.D. degree in Environmental Chemistry and Toxicology from Duke University, in the United States in 2015. His Ph.D. work mainly focused on the toxicity of flame retardants, as well as the application of effect-directed analysis in identifying causal compounds in mixtures. Then, he moved to The Scripps Research Institute (TSRI) as a Postdoctoral Research Associate in the field of Metabolomics. Currently, his main research is on the development of human exposome methods and the use of omics techniques such as chemoproteomics and metabolomics to evaluate the toxicity of pollutants or their mixtures. Dr. Fang has published more than 150 peer-reviewed papers in high-tier journals, including Nature Nanotechnology, Nature Water, Nature Chemical Biology, PNAS, Science Advances, Environmental Health Perspectives, ES&T, and Analytical Chemistry as a leading author. He also works as the Associate Editor of Environmental Pollution (Elsevier) and is an Editorial Board member for ES&T, ES&T Letters. He is also the recipient of the 2023 Chemical Research in Toxicology Young Investigator Award.