Interpretable machine learning for predicting membrane performance of organic solvent nanofiltration membranes using an open-access database
The discovery and optimization of solvent-resistant nanofiltration (SRNF) membranes remain limited by the complexity of polymeric material systems and the scarcity of harmonized experimental data. Here, we demon-strate how the open membrane database - a preexisting, diverse Findable, Accessible, Interoperable and Reusable (FAIR), open-access membrane database - can be transformed into a robust platform for data-driven materials understanding. Using 5600+ curated SRNF filtration experiments comprising up to 154 descriptors, we construct interpretable machine-learning models that predict two key crucial performance metrics—solvent permeance and molecular-weight cut-off (MWCO). After systematic data cleaning and feature engineering, ensemble tree- based models outperform linear and distance-based methods, achieving test-set R2 values of 0.76 for per-meance and 0.70 for MWCO. Model explainability via permutation importance and SHapley Additive exPlana-tions (SHAP) analysis reveals that selective-layer chemistry, deposition method, and nanocomposite components dominate membrane performance, whereas solution conditions influence permeance but not MWCO. External validation of four previously unreported membranes and three literature experiments not in the database con-firms the approach's predictive capability, while systematic overestimation of some data points suggests publi-cation bias in the underlying literature. Our results provide mechanistic insight into structure–property relationships in SRNF membranes and establish a possible modeling pipeline for leveraging open experimental data to accelerate the rational design of complex material systems.