digilib@itb.ac.id +62 812 2508 8800

Surfactant flooding is a recognized method within the realm of Chemical Enhanced Oil Recovery (CEOR); however, its implementation in the field is constrained by the susceptibility of surfactant efficacy to the unique brine chemistry of specific reservoirs. Traditional formulation methodologies predominantly focus on optimal salinity as the key design parameter, frequently neglecting the unique influences of specific ions (Na?, Ca²?, Mg²?, K?, Ba²?, SO?²?, HCO??, CO?²?, and Cl?) on interfacial tension (IFT), microemulsion phase behavior, and solubilization efficacy. This research establishes a framework utilizing machine learning techniques to forecast and analyze surfactant performance, drawing on experimental data collected from four oil reservoirs in Indonesia, which include three light oil fields and one heavy oil field. The workflow encompasses several key components, including data preprocessing, the engineering of features related to brine–surfactant interaction descriptors, and the application of regression and classification modeling techniques such as Random Forest, Support Vector Regression (SVR), and XGBoost. Additionally, it incorporates SHAP-based interpretability methods and utilizes Monte Carlo simulation for the purpose of formulation screening. Results indicate that the composition of brine ions contains predictive information regarding interfacial tension (IFT), the type of microemulsion, and the oil/water solubilization ratios (OSR/WSR). However, isolating the effects of individual ions proves challenging due to significant multicollinearity among the brine variables. Among the ions examined, carbonate (CO?²?) and potassium (K?) exhibited the strongest correlations with the type of microemulsion. Additionally, the characteristics of engineered surfactant-brine interactions had a more substantial impact on overall model predictions. Random Forest demonstrated superior performance in the classification of IFT, WSR, and microemulsions, whereas Support Vector Regression (SVR) exhibited optimal results for OSR. Monte Carlo screening of 10,000 virtual formulations revealed that 7.21% were identified as predicted Winsor Type III systems. Among the five hypotheses examined, three received partial support, one was supported within the confines of the study's domain, and the hypothesis regarding cross-reservoir transferability could not be validated due to significant confounding factors associated with reservoir-related variables. The integrated SHAP–Monte Carlo framework has been converted into practical formulation guidelines, providing a swift, data-driven decision-support tool aimed at diminishing dependence on laboratory trial-and-error screening in the design of surfactant enhanced oil recovery (EOR).