INSIGHT researchers from Hanyang University have developed a multimodal large language model workflow for automatically extracting structured nanotoxicity data from scientific articles. The paper entitled ‘Automated Large-Scale Extraction of Nanotoxicity Data from the Literature Using Multimodal Large Language Models’ focuses on oxide nanomaterials and integrates information from full text, figures, tables, captions, legends, and Materials and Methods sections.

The extracted material identities, physicochemical properties, exposure conditions, biological contexts, and cytotoxicity results are subsequently used to develop AutoML-based nanotoxicity prediction models. The study also examines how data quality, material composition, and model scope affect predictive performance and applicability-domain coverage.
Data-driven nanosafety assessment requires large and well-curated datasets, but manual literature curation is slow and difficult to scale. Nanotoxicity data are particularly challenging because relevant information is often distributed across different parts of an article. The authors’ previous text-based workflow could extract material and contextual information, but quantitative variables such as exposure dose, exposure time, and cell viability frequently had to be recovered manually from figures. The motivation was therefore to develop a multimodal workflow that could reduce these remaining manual steps and generate modeling-ready datasets at a larger scale.
The study moves beyond text-only extraction by jointly analyzing article text and visual content. It automatically extracts quantitative assay results and physicochemical values from figures and tables while preserving information about their source, including page number, caption, panel, axes, and supporting evidence.
Four multimodal LLMs—Claude Sonnet 4.6, Gemini 3.1 Pro, GPT-5.4, and Grok 4.20—were benchmarked within the same workflow for accuracy, runtime, and efficiency. The paper also extends beyond information extraction by using the resulting data for AutoML modeling. It compares broad multi-material, quality-filtered, material-balanced, and single-material prediction strategies and evaluates both predictive performance and applicability domain.
GPT-5.4 provided the best balance between extraction accuracy and efficiency. It achieved a mean extraction F1 score of 96.1% and an assay-result F1 score of approximately 86.5%, with the latter representing one of the most difficult figure-derived endpoints.
The optimized workflow was applied to an initial pool of 821 articles. After screening, 465 relevant in vitro oxide-nanomaterial articles were used to construct Park-OxideNM2026, containing 16,786 structured records.
The physicochemical-quality-filtered dataset achieved the best overall predictive performance, with a mean F1 score of 0.824 compared with 0.792 for the full dataset. This indicates that improved descriptor completeness can compensate for reduced data quantity. In contrast, balancing record numbers across the four most represented materials through undersampling reduced performance. Single-material models improved target-specific predictions for ZnO, TiO₂, Fe₃O₄, and SiO₂. However, these improvements occurred within substantially narrower applicability domains. The results therefore show a trade-off between target-specific performance and broad descriptor-space coverage. Exposure dose was the most influential predictor, followed by variables related to exposure duration, cell context, assay type, and intrinsic material properties.
This paper is primarily relevant to the data and model layers of SSbD framework for INSIGHT. It provides a scalable method for linking material properties with exposure conditions, biological context, and toxicity outcomes. The retention of source-level evidence also supports the traceability and transparency required for auditable assessment workflows.
The comparison of broad and material-specific models is relevant to INSIGHT’s model-selection process. It shows that the most accurate model is not always the most broadly applicable one and that model choice should consider data quality, intended use, and applicability domain. The extracted records could also provide input data for future mechanistic or IOP-based pipelines. However, this study does not itself construct an AOP or IOP and does not assess environmental, social, economic, or full life-cycle impacts. Its main contribution to INSIGHT is therefore a scalable and provenance-aware foundation for the human-health hazard and predictive-modeling components of the wider SSbD framework.
Follow this link to read the full paper.



