Machine Learning Assisted Molecular Property Prediction for Sustainable Chemical Design

Authors

  • Dhalvinder Raj Author

Keywords:

Molecular property prediction; Machine learning; Graph neural networks; Sustainable chemistry; ChemBERTa; Ensemble learning; Cheminformatics

Abstract

The challenge of predicting the molecule properties is one of the cornerstones of sustainable chemical design and drug discovery. The conventional approaches to computation, including density functional theory (DFT) and molecular dynamics (MD) are precise but too costly to be done at large scale. The paper introduces a machine learning (ML) aided system to predict molecular properties in high-throughput with the use of Graph Convolutional Networks (GCN), Deep Neural Networks (DNN), Random Forest (RF), and a fine-tuned ChemBERTa Transformer in an ensemble architecture. The encoding of molecules is done through Morgan fingerprints as well as RDKit descriptors and graph-based representations, and dimensionality reduction is done through Principal Component Analysis (PCA) and LASSO-based feature selection. On the benchmark ChEMBL and MoleculeNet logP, aqueous solubility, toxicity, and bioactivity prediction datasets, the proposed ensemble model attains an R 2 of 0.93, RMSE of 0.53, and MAE of 0.39. The results of the experiment show that the framework is as much as 15.3% more accurate at predictions than state-of-the-art baselines. The targeted system has a direct implication as it fosters the principles of green chemistry by facilitating the rapid in-silico screening of viable sustainable molecular candidates, thus minimizing the wet-lab experiments that are very expensive. There are open-source code and data.

Downloads

Published

2026-08-19