Artificial Intelligence and Machine Learning for Drug Solubility Enhancement: Predictive Modelling, Formulation Optimization, and Translational Challenges: AI and ML for drug solubility enhancement
Keywords:
Artificial intelligence; Machine learning; Solubility, Drug delivery systems; Formulation design;, Pharmaceutical preparationsAbstract
Background: Poor aqueous solubility remains a major constraint in drug development because it can
limit dissolution, oral absorption, dose feasibility and formulation robustness. Artificial intelligence
(AI) and machine learning (ML) are increasingly used to predict solubility-related properties and guide
enabling-formulation design. Objective: This review critically evaluates AI/ML approaches for drug-
solubility enhancement, with particular emphasis on the relationship between predicted endpoints and
experimentally meaningful pharmaceutical outcomes. Methods: Peer-reviewed literature published
from January 2010 to 16 July 2026 was identified through structured searches of PubMed/MEDLINE
and major scholarly publisher databases, supplemented by backward and forward citation tracking.
Evidence concerning aqueous-solubility prediction, amorphous solid dispersions, cocrystals and salts,
nanocrystals and nanoparticulate systems, cyclodextrin complexes, solvent and excipient selection,
generative design, high-throughput screening and automated experimentation was synthesised
critically. Key findings: Evidence is strongest for molecular-solubility prediction, amorphous solid-
dispersion development, cocrystal screening and nanoformulation optimization. AI/ML can narrow
experimental search spaces, integrate molecular, material and process variables, and support adaptive
experiment selection. However, most studies remain retrospective and optimise surrogate endpoints
such as logS, miscibility, cocrystal formation, particle size or complexation affinity rather than
prospectively verified biorelevant dissolution or systemic exposure. Heterogeneous datasets,
inconsistent endpoint definitions, information leakage, limited external validation, uncertain
applicability domains, incomplete uncertainty reporting and restricted interpretability constrain
translation. Conclusion: AI is most defensible as an uncertainty-aware decision-support layer
integrated with mechanistic modelling, Quality by Design and prospective experimentation.
Regulatory credibility will require predefined contexts of use, transparent data provenance,
independent validation, human oversight and controlled lifecycle governance.



















