Abstract
Alkali-activated lunar regolith is considered one of the most promising lunar regolith-based construction materials for large-scale lunar construction. This study presents a comprehensive framework integrating machine learning (ML) algorithms to investigate the influence of various features on the mechanical strength of alkali-activated lunar regolith simulant (AALRS), aiming to achieve the strength prediction and optimization design of AALRS. The properties of lunar regolith simulant, mixture proportions, preparation parameters, environmental conditions, and enhancement methods were employed as input features for ML modeling. The compressive and flexural strength predictive models were constructed using eight ML algorithms and evaluated through statistical indicators. Among the models, Extreme Gradient Boosting demonstrated the best performance, yielding an R2 of 0.8684, a root mean square error of 6.2007 MPa, and a mean absolute error of 4.0874 MPa on the testing dataset. Using the best prediction models, four design strategies with three objectives, namely, compressive strength, flexural strength, and transport payload (TP), were optimized and evaluated using the non-dominated sorting genetic algorithm II and the technique for order preference by similarity to ideal solution methods. The proposed prediction and optimization framework for the mechanical performance of AALRS, which integrates extreme environmental effects and TP-driven design, provides a robust data-driven approach for the design, prediction, and optimization of AALRS, advancing the development of extraterrestrial construction materials and technologies.
Keywords
1. Introduction
The Moon, Earth’s nearest celestial body, has transitioned from a target of initial exploration to a strategic point for sustained human presence, resource development, and scientific research. Leading space agencies such as the National Aeronautics and Space Administration, the China National Space Administration, and the European Space Agency are actively pursuing ambitious lunar programs, indicating their determination not only to return humans to the Moon but also to establish permanent habitats there[1,2]. The key to achieving this long-term vision lies in overcoming the long distances and economic barriers associated with transporting construction materials from Earth[3]. Consequently, the principle of in-situ resource utilization (ISRU), based on the development and utilization of lunar resources, has become a widely accepted consensus[4-6]. Lunar regolith, which predominantly covers the surface of the Moon, is a layer of fragmented material resulting from meteorite impacts and solar wind bombardment[7]. Its composition is dominated by various minerals that are primarily metallic oxides, which show great potential for alkali activation to produce construction materials[8-10]. Specifically, alkali-activated lunar regolith simulant (AALRS) exhibits adequate mechanical strength, excellent durability, lower energy requirements for production, high ISRU efficiency, and a relatively simple preparation process[8]. These compelling features make AALRS a promising candidate for constructing large-scale lunar infrastructure and have garnered considerable research interest. However, realizing the full potential of AALRS necessitates a deep understanding of its mechanical performance under extreme environments, including dramatic thermal cycles, high vacuum, and pervasive radiation, which is fundamental to constructing durable lunar engineering structures.
Currently, extensive experimental work has been conducted to understand how various factors influence the properties of AALRS. Researchers have systematically investigated a wide range of parameters. These include the chemical and mineralogical compositions of different lunar regolith simulants (LRS), the specific types and concentrations of alkaline activators, curing conditions such as temperature and pressure, various additives, and activation methods[11-18]. Geng et al.[9] reported that the strength of AALRS increases with increasing environmental temperature. They also found that sodium hydroxide exhibits a superior alkali activation capability compared to sodium silicate. Xue et al.[19] found that the silicate modulus (SM) and dosage significantly impact both the thermodynamic equilibrium of the pore solution and the resulting reaction products, thereby affecting the compressive strength (CS) of the geopolymer. Wu et al.[18] investigated the mechanical strength of AALRS using various LRS with different particle sizes and compositions. These findings underscored the critical importance of appropriately matching alkaline activators and external environmental conditions with the specific reactivity of the LRS. This involves ensuring sufficient alkali dissolution and the addition of silicate ions to form adequate aluminosilicate oligomers, along with utilizing elevated temperatures to accelerate the reorganization of the aluminosilicate network. Shao et al.[17] analyzed the effects of lunar highlands simulant (LHS-1) and lunar mare simulant on the CS of AALRS. The results showed that the LHS-1 samples consistently demonstrated better performance attributed to favorable Si/Al and Ca/Si ratios. Although researchers have gained some qualitative understanding of how different factors affect the mechanical strength of AALRS, many conclusions remain controversial. The inherent variations in composition and particle characteristics among different LRS, along with the complex chemical reactions in alkali activation, lead to a reaction system driven by numerous interacting variables. These variables frequently exhibit highly non-linear and interconnected effects on the material properties, hindering accurate predictions. Consequently, traditional empirical models developed from limited experimental datasets often show poor predictive accuracy and limited generalization ability, leading to uncertainties regarding the most effective strategies for optimal material design of AALRS. Establishing AALRS strength prediction models that possess both high precision and strong generalization is indispensable. This is vital for guiding material design strategies, streamlining the optimization of manufacturing parameters, and effectively controlling the resultant properties of AALRS, thereby accelerating the development cycle and reducing reliance on empirical testing.
To overcome the shortcomings of traditional empirical models in predicting the mechanical strength of AALRS, machine learning (ML) provides a powerful data-driven solution. As an interdisciplinary field at the intersection of computer science and statistics, ML establishes “black-box” correlations between feature variables and target properties through data analysis, without necessitating an understanding of the underlying mechanistic interactions[20-25]. This approach excels in addressing nonlinear problems and transforms conventional mechanism-driven models into data-driven models. Sun et al.[26] extracted control factors from the perspective of the reaction mechanism and collected 871 data points to build a CS predictive model of slag and fly ash-based alkali-activated concrete using random forests (RFs), regression trees, and an artificial neural network. Ding et al.[27] chose three key components, including Al2O3, CaO, and SiO2, as inputs to develop a strength predictive model. The Shapley additive explanations (SHAP) analysis was performed to reveal the intricate relationships among key elements.
Although ML has been widely applied to predict the performance of traditional alkali-activated construction materials, its use in predicting the mechanical properties of AALRS has not been reported. Neser[28] pioneeringly reviewed extraterrestrial construction materials and highlighted the integration of ML into material-process-product-suitability frameworks to predict the properties and responses of AALRS, thereby enhancing the understanding of extraterrestrial materials. The development of ML prediction models would streamline material design, optimize manufacturing parameters, and reduce the reliance on extensive experimental testing. These advancements are essential for the application of ISRU and sustainable lunar construction.
Given the above discussions and analyses, this study integrated strength prediction modeling and multi-objective optimization to investigate the mechanical strength of AALRS under extreme environments. Specifically, 17 input features were selected based on the properties of LRS, mixture proportions, preparation parameters, extreme environmental conditions, and performance enhancement methods. Eight representative algorithms were used to fit 518 CS data points and 301 flexural strength (FS) data points. Through evaluation using statistical indicators, robust compressive and flexural strength predictive models for AALRS were developed based on the extreme gradient boosting (XGBoost) model, which demonstrated the best performance. The SHAP analysis was employed to analyze the feature importance of input variables. Considering three key output objectives inspired by practical lunar construction, including CS, flexural strength, and transport payload (TP), a multi-objective optimization framework was proposed using non-dominated sorting genetic algorithm-II (NSGA-II) and the technique for order preference by similarity to ideal solution (TOPSIS). In addition, graphical user interfaces (GUI) for strength prediction and performance optimization were developed. This study advances the development of AALRS for lunar construction by providing a comprehensive data-driven framework, which facilitates the design, prediction, and optimization of AALRS materials under extraterrestrial conditions.
2. Data Preparation and Dataset Construction
CS and FS are critical indicators for evaluating the performance and structural integrity of AALRS-based materials. To develop reliable and robust strength prediction models, this study conducted a systematic review of existing literature on AALRS and experimental results[9-16,18,19,29-53]. The literature identification process followed a structured search strategy. Relevant publications were retrieved from major academic databases, including Web of Science, Scopus, and Google Scholar, using keywords such as “alkali-activated lunar regolith”, “lunar regolith geopolymer”, “lunar regolith binder”, “compressive strength”, and “flexural strength”. The search covered publications up to 2025. Inclusion criteria required that: (i) the study involved alkali-activated or geopolymer-based lunar regolith simulant materials; (ii) quantitative data on CS and flexural strength were explicitly reported; and (iii) key mixture design parameters and preparation conditions were clearly documented. Studies were excluded if data were incomplete, duplicated, or could not be reliably extracted. Following this screening process, a CS dataset containing 518 data points and an FS dataset containing 301 data points were ultimately established. It should be noted that the datasets were derived from studies using LRS, which cannot fully reproduce certain characteristics of actual lunar regolith, such as agglutinate content and space weathering effects. These differences may influence particle reactivity and bonding behavior, and thus affect the direct applicability of the model. Future work may require calibration using higher-fidelity simulants or real lunar samples. For the CS dataset, 17 feature variables were selected as input parameters, and these variables cover key aspects relevant to the preparation process and properties of AALRS, including the properties of LRS, mixture proportions, preparation parameters, extreme environmental conditions, and performance enhancement methods. In contrast, only 10 feature variables were chosen for the FS dataset due to its smaller number of data points. Table 1 presents detailed information and corresponding abbreviations for all the input parameters. The first 10 variables, from median particle diameter (MPD) to vacuum degree (VD), represent the common features available across all collected studies. Among them, IT and TT denote the initial and termination temperatures during curing, respectively, describing the temperature range applied in the curing process. VD is used to characterize the low-pressure curing environment. These three variables jointly represent the simulated extreme environmental conditions relevant to lunar construction. The variables after VD correspond to performance enhancement methods for AALRS, including fiber reinforcement, nanomaterial addition, reactive additive incorporation, and thermal activation of lunar regolith. Since these enhancement strategies were not reported in all literature sources, the corresponding missing values were uniformly assigned to zero.
| Input parameters | Abbreviation | Input parameters | Abbreviation |
| Median particle diameter | MPD | Si/Al of LRS | Si/Al |
| Water-to-binder | W/B | Alkali content | AC |
| Silicate modulus | SM | Curing age | CA |
| Specimen volume | CV/FV | Initial temperature during Curing | IT |
| Termination temperature during Curing | TT | Vacuum degree | VD |
| Basalt fiber content | BF | Ratio of length-to-diameter of fiber | RLD |
| Carbon nanotubes | CNT | Reactive Al2O3 content | Al2O3 |
| Ca(OH)2 content | Ca(OH)2 | Heat activation temperature | HAT |
| Heat activation time | HATT |
The workflow of practical lunar construction is shown in Figure 1. Two key considerations are highlighted: first, cost-effective payload transport, with a reported cost of $1.2 million per kilogram of payload delivered to the lunar surface[6]; second, the mechanical performance of AALRS, specifically its CS and FS. Therefore, this study designated CS, FS, and TP as output objectives. The TP represents the mass of raw materials required to be transported from Earth to the Moon. The TP rate and the ISRU rate are in a complementary relationship, where TP + ISRU = 1. This relationship quantifies the balance between Earth dependency and lunar autonomy. Regarding water resources, although research identifies potential sources such as water ice at the lunar South Pole and hydroxyl groups in lunar regolith, mature and viable technologies for in-situ extraction remain unavailable[54,55]. Therefore, within the scope of this study, water is treated as a resource requiring transport from Earth, and its associated mass is incorporated into the calculation of the TP.
Figure 1. The workflow and key considerations of lunar construction. AALRS: alkali-activated lunar regolith simulant.
To gain profound insights into the inherent characteristics of the collected data and the intricate interrelationships among variables, a comprehensive statistical analysis was performed. Key statistical parameters were calculated for all the variables, including minimum, maximum, mean, standard deviation (SD), kurtosis, skewness, and coefficient of variation (CV), as summarized in Table 2 and Table 3. These parameters collectively provide a comprehensive overview of the value range coverage, distribution patterns, and dispersion characteristics of data across all variables. This detailed statistical profiling confirms that the datasets adequately capture the complexity and variability of the system under investigation, encompassing diverse material properties, processing parameters, environmental conditions, and AALRS-related performance metrics. Such representativeness ensures that the datasets are suitable for developing a reliable and robust predictive model for AALRS.
| Feature parameters | unit | Minimum | Maximum | Mean | Standard deviation | Kurtosis | Skewness | Coefficient of variation |
| MPD | μm | 5.760 | 247.450 | 53.120 | 37.040 | 3.700 | 1.570 | 0.700 |
| Si/Al | / | 0.980 | 5.610 | 2.460 | 0.410 | 15.510 | 0.840 | 0.170 |
| W/B | / | 0.040 | 0.490 | 0.270 | 0.090 | 0.890 | -0.840 | 0.330 |
| AC | / | 0 | 0.150 | 0.060 | 0.030 | 0.050 | -0.180 | 0.440 |
| SM | / | 0 | 3.340 | 0.730 | 0.750 | 0.220 | 0.840 | 1.030 |
| CA | day | 0.188 | 108.000 | 10.140 | 12.410 | 20.820 | 3.340 | 1.220 |
| CV | cm3 | 0 | 353.390 | 61.450 | 77.700 | 4.950 | 2.130 | 1.260 |
| IT | ℃ | -196.000 | 120.000 | 59.990 | 43.140 | 12.910 | -2.550 | 0.720 |
| TT | ℃ | -196.000 | 600.000 | 52.440 | 66.640 | 28.600 | 1.990 | 1.270 |
| VD | Pa | 0.200 | 101,325 | 88,913 | 33,241 | 3.340 | -2.310 | 0.370 |
| BF | % | 0 | 4.1250 | 0.030 | 0.260 | 166.760 | 12.340 | 9.230 |
| RLD | / | 0 | 750.000 | 10.510 | 72.190 | 84.980 | 8.840 | 6.870 |
| CNT | % | 0 | 0.500 | 0.010 | 0.060 | 54.930 | 7.200 | 6.220 |
| Al2O3 | % | 0 | 10.000 | 0.210 | 1.010 | 42.590 | 6.030 | 4.820 |
| Ca(OH)2 | % | 0 | 8.000 | 0.080 | 0.680 | 97.230 | 9.630 | 8.730 |
| HAT | ℃ | 0 | 1,300.000 | 56.180 | 254.690 | 16.990 | 4.340 | 4.530 |
| HATT | min | 0 | 60.000 | 2.160 | 10.250 | 23.220 | 4.890 | 4.740 |
| CS | MPa | 0 | 81.500 | 17.990 | 16.970 | 1.630 | 1.350 | 0.940 |
MPD: median particle diameter; W/B: water-to-binder; AC: alkali content; SM: silicate modulus; CA: curing age; CV: coefficient of variation; IT: initial temperature during curing; TT: termination temperature during curing; VD: vacuum degree; BF: basalt fiber content; RLD: ratio of length-to-diameter of fiber; CNT: carbon nanotubes; HAT: heat activation temperature; HATT: heat activation time; CS: compressive strength.
| Feature parameters | unit | Minimum | Maximum | Mean | Standard deviation | Kurtosis | Skewness | Coefficient of variation |
| MPD | μm | 5.760 | 214.000 | 32.930 | 17.190 | 57.200 | 6.040 | 0.520 |
| Si/Al | / | 2.190 | 2.990 | 2.600 | 0.190 | 0.190 | -1.080 | 0.070 |
| W/B | / | 0.098 | 0.300 | 0.290 | 0.040 | 15.860 | -4.030 | 0.140 |
| AC | % | 0.011 | 0.104 | 0.070 | 0.020 | 3.180 | -1.540 | 0.260 |
| SM | / | 0 | 3.000 | 0.780 | 0.740 | -0.220 | 0.720 | 0.960 |
| CA | day | 0.188 | 108.000 | 20.830 | 28.210 | 1.410 | 1.600 | 1.350 |
| FV | cm3 | 4.000 | 256.000 | 98.840 | 88.390 | -0.460 | 1.070 | 0.890 |
| IT | ℃ | -196.000 | 120.000 | 58.460 | 49.790 | 11.730 | -2.630 | 0.850 |
| TT | ℃ | -80.730 | 120.000 | 64.490 | 35.140 | 1.340 | -0.510 | 0.540 |
| VD | Pa | 10.000 | 101,325 | 86,860 | 35,491 | 2.220 | -2.050 | 0.410 |
| FS | MPa | 0.030 | 11.410 | 3.370 | 2.360 | 0.010 | 0.560 | 0.700 |
MPD: median particle diameter; W/B: water-to-binder; AC: alkali content; SM: silicate modulus; CA: curing age; FV: specimen volume (flexural); IT: initial temperature during curing; TT: termination temperature during curing; VD: vacuum degree; FS: flexural strength.
To ensure the independence of variables, the Pearson correlation coefficient (PCC) was employed to assess the linear correlation between variables, aiming to detect potential multicollinearity issues. The PCC is a widely utilized statistical metric that quantifies the strength and direction of the linear relationship between two variables. It is calculated as the covariance of two variables divided by the product of their SDs, with values ranging from -1 to +1. A coefficient of +1 indicates a perfect positive linear correlation, -1 signifies a perfect negative linear correlation, and 0 implies the absence of any linear correlation[56].
Figure 2 and Figure 3 present the frequency distributions and PCC matrix heatmaps of variables in the CS and FS datasets, respectively. In these figures, the diagonal elements show the frequency histograms of each variable; the left side of the off-diagonal elements displays the scatter distribution characteristics between variables; and the right side of the off-diagonal elements is labeled with the corresponding PCC value. Histograms can intuitively reflect the central tendency of the variable data, and scatter plots can reveal the correlation trends between variables. PCC heatmaps visually delineate the linear interrelationships among the input variables and between these input variables and the output variables. As depicted in Figure 2, the PCC values between CS and each input variable range from -0.26 to 0.26. This low range of PCC values suggests that the linear relationship between CS and any single input variable is weak, implying that non-linear relationships are likely dominant or that multiple variables interact to influence CS. Moreover, since the absolute PCC values between nearly all input variables are less than 0.7, this confirms that these input variables provide distinct information without redundancy or multicollinearity issues. Figure 3 shows that the PCC range between FS and the input variables is from -0.11 to 0.63, indicating a higher degree of linear correlation between the input variables and FS compared to that observed with CS.
Figure 2. Frequency distributions and Pearson Correlation Coefficient matrix heatmaps of variables in the CS dataset. CS: compressive strength.
Figure 3. Frequency distributions and Pearson Correlation Coefficient matrix heatmaps of variables in the FS dataset. FS: flexural strength.
The results clearly demonstrate that simple linear models are insufficient to capture the intricate interdependencies between these variables. Thus, traditional explicit equations are unsuitable for accurately quantifying the impact of inputs or for reliable prediction. Given the intrinsic non-linearity and complex interactions between input variables and output variables, employing data-driven ML algorithms is essential. ML methods are uniquely capable of learning the non-linear mappings between inputs and outputs, enabling the development of predictive models crucial for evaluating the material performance of AALRS.
Figure 4 outlines the proposed predictive models and multi-objective optimization framework. During the model building process, the dataset was randomly partitioned using the scikit-learn library in Python. Seventy percent (70%) of the data was allocated for model training (training set), while the remaining thirty percent (30%) was used for performance evaluation (testing set). This procedure aims to minimize potential biases and random effects associated with a single data split. To eliminate differences in the scales of variable features and address potential feature weight imbalance, thereby enhancing model robustness and convergence speed, all input and output variables in the dataset were scaled to the range between 0 and 1 through data normalization, as described in Eq. (1).
Figure 4. Framework of the proposed predictive models and multi-objective optimization. SVR: support vector regression; GBDT: gradient boosting decision tree; RF: random forest; RMSE: root mean square error; RRMSE: relative root mean square error; MAE: mean absolute error; NSGA-II: non-dominated sorting genetic algorithm-II.
where Y and Ynorm represent the data values before and after normalization, respectively, and Ymin and Ymax represent the minimum and maximum values in the dataset, respectively.
3. Methodology
3.1 ML models
This study employed eight representative ML algorithms, including support vector regression (SVR), adaptive boosting (AdaBoost), automated machine learning (AutoML), categorical boosting (CatBoost), gradient boosting decision tree (GBDT), light gradient boosting machine (LightGBM), RF, and XGBoost. This selection encompasses traditional single learning algorithms, various ensemble learning algorithms, and the AutoML framework. Considering that individual ML models can be sensitive to data distribution and potentially overfit specific training datasets, their ability to accurately predict performance under diverse conditions may be limited. Ensemble learning models mitigate this issue by combining multiple base learners and integrating their predictions. Such approaches typically reduce errors inherent in single models, decrease the likelihood of overfitting, and enhance the overall generalization ability of the models[57-59]. Bagging methods such as the RF model exemplify a parallel ensemble technique[60]. These methods involve training multiple base models in parallel, typically on different subsets of the data. The results from these models are combined using averaging or voting to produce the final prediction. Figure 5 illustrates the working principle of this approach. Boosting represents a sequential ensemble approach, exemplified in this work by AdaBoost, GBDT, CatBoost, LightGBM, and XGBoost[61]. This approach trains multiple weak learners iteratively, with each subsequent learner placing greater emphasis on instances mispredicted by the previous ones. The final prediction is derived from a weighted aggregation of the outputs from these individual weak learners, as illustrated in Figure 6.
3.1.1 SVR
SVR, an extension of Support Vector Machines, employs structural risk minimization to construct a regression function that predicts continuous values, ignoring errors within a tolerance threshold (ε) and penalizing only errors exceeding this threshold[60]. It maps data from the original input space into a high-dimensional feature space using kernel functions. This transformation allows SVR to model non-linear relationships in the original input space by finding a linear regression function in the high-dimensional feature space. The linear kernel, polynomial kernel, and radial basis function kernel are frequently utilized kernel functions. This mechanism enhances robustness by allowing the model to disregard minor noise and focus on larger deviations. The working principle of SVR is illustrated in Figure 7.
3.1.2 AdaBoost model
AdaBoost minimizes overall model error by integrating multiple weak learners into a robust classifier, applicable to a wide range of prediction tasks. Its core mechanism involves iteratively training weak learners on weighted training data: initially, all observations are assigned equal weights; after each iteration, the weights of observations with large errors are increased, while those with smaller errors are reduced. This allows subsequent learners to focus on observations that are difficult to predict accurately. The predictions of weak learners are aggregated through a weighted sum, with the requirement that each weak learner performs better than random guessing to ensure model effectiveness.
3.1.3 AutoML
AutoML is an efficient ML framework that automates data preprocessing, feature engineering, model selection, and hyperparameter optimization, significantly enhancing modeling efficiency while maintaining high prediction accuracy. Compared to traditional ML approaches, AutoML leverages innovative techniques such as Bayesian optimization and genetic algorithms to automatically identify optimal model configurations[62]. Its core strength lies in reducing development barriers and time costs.
3.1.4 GBDT model
GBDT is a robust ML framework that iteratively builds decision trees, combining them additively to optimize a loss function via gradient descent, thereby enhancing prediction accuracy. It fits the residuals of the previous tree in each iteration, gradually approximating the target function, with a learning rate to balance overfitting and performance. GBDT excels in classification and regression tasks, offering interpretability through feature importance.
3.1.5 CatBoost model
CatBoost is an efficient ML framework built on the GBDT model. It particularly excels at handling categorical features. The framework aims to enhance training speed and efficiency while maintaining high prediction accuracy. A key innovation within the model is the ordered target statistics technique, which automates the encoding of categorical variables. This automation thus reduces the data preprocessing workload and effectively mitigates target leakage and overfitting, which are issues often encountered with traditional encoding approaches. Furthermore, CatBoost incorporates mechanisms to automatically handle missing values and frequently employs symmetric oblivious decision trees.
3.1.6 LightGBM model
LightGBM is a highly efficient gradient boosting framework based on the GBDT model, significantly enhancing training speed and prediction accuracy through optimized data structures and algorithms. It incorporates several innovative techniques including histogram-based algorithms for accelerated split point discovery, Gradient-based One-Side Sampling to focus on informative data instances, and Exclusive Feature Bundling for effective handling of sparse features. Additionally, its adoption of a leaf-wise tree growth strategy typically yields superior results in fewer iterations compared to traditional level-wise growth methods.
3.1.7 RF model
RF is a widely used ensemble learning method that constructs a multitude of decision trees during training. It employs bagging, with each individual tree trained on a distinct bootstrap sample of the original dataset. Furthermore, during the tree building process, only a random subset of features is considered when determining the best split at each node. The final prediction is then derived by aggregating the outputs from all trees, typically using averaging for regression tasks or majority voting for classification problems.
3.1.8 XGBoost model
XGBoost is an efficient ML framework based on the GBDT model, enhancing prediction accuracy and computational efficiency by optimizing its loss function and incorporating regularization terms[63]. The core mechanisms include additive tree building, loss function minimization utilizing second-order gradient information, capabilities for handling sparse data, and feature importance evaluation. As an ensemble method, it sequentially adds weak decision tree learners. To minimize overall model errors, each new tree is trained to correct the errors made by the preceding ensemble. This correction is achieved specifically by modeling the negative gradients of the loss function, and the tree construction process itself utilizes second-order derivative information. Each new iteration corrects errors made by the preceding ones, and the final prediction is the sum of all iteration outputs.
where
where n is the number of samples;
XGBoost integrates and reorganizes the Taylor expansion of the objective function into polynomials related to the predicted residuals, yielding the model complexity formula as follows:
where λ is a regularization parameter; γ is the penalty coefficient of the model; gi and hi are the first and second derivatives of the loss function, respectively, and T is the number of leaf nodes.
3.2 Model performance evaluation
The performance of the ML models was evaluated using several statistical indicators[64]. The coefficient of determination (R2) quantified the explained variance, with values closer to 1 reflecting a better fit to the actual data. The magnitude of prediction error was quantified using mean absolute error (MAE), root mean square error (RMSE), and relative root mean square error (RRMSE). For these indicators, smaller values (closer to zero) signify higher predictive accuracy, as they reflect smaller deviations between predicted and observed values. The performance metric ρ (normalized RMSE ratio) was also considered; values approaching 0 indicate a smaller deviation between predicted and observed values, thus signifying higher model precision. In summary, a superior model should exhibit an R2 value approaching 1, coupled with low values for MAE, RMSE, and RRMSE. The calculation formulas for the relevant statistical indicators are shown in Eqs. (8)-(12):
where n represents the total number of test samples; Xa represents the actual mechanical strength values; Xp represents the predicted mechanical strength values; and
3.3 SHAP analysis
Despite their widely recognized predictive performance, ML models are often considered black boxes owing to their complexity and lack of transparency. To enhance the interpretability of model behavior, this study employed the SHAP analysis method. SHAP quantifies the contribution of individual input variables to the model’s prediction outcomes, consequently fostering a deeper understanding of the model’s decision-making process. The importance of the input variables relative to the output can be assessed using Eq. (13). In essence, features whose value changes lead to more significant fluctuations in the model’s prediction outcomes are considered to have greater importance.
where p represents the number of features; xj represents the j-th feature variable; φj(f) denotes the SHAP value of the j-th feature, quantifying its contribution to the model’s prediction; S represents a subset of input features excluding xj; f(S) denotes the expected model output when only the features in S are used for prediction; and f(S∪{Xj}) - f(S) reflects the change in the predicted value after adding xj to the feature set.
3.4 NSGA-II
This study employed the NSGA-II algorithm to identify Pareto optimal solutions. Compared with its predecessors, NSGA-II represents a significant advancement by reducing computational complexity while incorporating elitism and a crowding distance metric during the iterative process[65]. The algorithm begins with the random generation of an initial population, followed by non-dominated sorting to establish Pareto ranks. Subsequently, a new population is generated through genetic operators including selection, crossover, and mutation. The algorithm then merges the parent and offspring populations, applying non-dominated sorting and crowding distance calculations to select individuals for the next generation. This maintains both convergence toward the Pareto front and solution diversity. The iterative process continues as the algorithm repeatedly generates, merges, and selectively retains populations, progressively approaching the true Pareto optimal front. This evolutionary approach enables effective exploration of complex solution spaces in multi-objective optimization while balancing multiple competing objectives, without requiring prior articulation of preferences.
3.5 TOPSIS analysis for decision making in multi-objective optimization
After the Pareto optimal front is obtained, a preferred optimal solution needs to be selected from it. TOPSIS analysis is a simple, flexible, logical, and efficient method for identifying the optimal solution in multi-attribute decision-making problems[21]. Its core mechanism involves calculating the distances of each alternative from the positive ideal solution, comprising the best values for all criteria, and the negative ideal solution, comprising the worst values, ranking alternatives based on a weighted closeness index to select the one closest to the positive ideal and farthest from the negative ideal. The formulas for calculating the positive ideal distance, negative ideal distance, and relative closeness are presented in Eqs. (14)-(16). In this study, to select the preferred AALRS solution from the Pareto optimal set, TOPSIS considers three performance indicators as criteria: CS, FS, and TP. Different weights are assigned to these three indicators (CS, FS, TP) under four strategies to determine the overall optimal solution, as provided in Table 4. The weights satisfy the normalization condition f(CS) + f(FS) + f(TP) = 1. This ensures that the weighted contributions of the three criteria collectively form a comprehensive evaluation framework, thereby enabling a systematic comparison of solutions in the Pareto optimal set.
| f(CS) | f(FS) | f(TP) | |
| Strategy 1 | 0.10 | 0.10 | 0.80 |
| Strategy 2 | 0.33 | 0.33 | 0.34 |
| Strategy 3 | 0.10 | 0.80 | 0.10 |
| Strategy 4 | 0.80 | 0.10 | 0.10 |
CS: compressive strength; FS: flexural strength; TP: transport payload.
where
4. Model Performance Evaluation of CS
4.1 Model validation
Figure 8 presents a comparison of the actual and predicted CS for eight ML models, with Figure 8a representing the training set and Figure 8b representing the testing set. The red and blue deviation lines represent the ±20% and ±10% error ranges, respectively. Results demonstrate that CatBoost, LightGBM, XGBoost, and GBDT models exhibit high prediction accuracy, characterized by high R2 values and data points predominantly within the ±10% error range. This indicates their potential for accurately predicting CS, which is a critical factor for ensuring the safety and stability of lunar base structures. Specifically, the XGBoost model achieves an R2 of 0.9951 on the training set and maintains an R2 of 0.8724 on the testing set, reflecting strong generalization ability. Conversely, RF and SVR models show inferior predictive performance, with lower R2 values and a large number of data points falling outside the ±10% or ±20% error ranges.
Figure 8. Comparison of actual and predicted compressive strength. (a) Training set; (b) Testing set. ML: machine learning; LightGBM: light gradient boosting machine; RF: random forest; SVR: support vector regression; GBDT: gradient boosting decision tree.
Figure 9 shows the error fluctuations between the actual values and predicted values based on the eight ML models. The blue filled region represents the training set, while the red filled region represents the testing set. Errors, defined as the differences between the predicted and actual values, are plotted as histograms, which provide insights into model robustness. An error distribution tightly centered around zero with low variance generally indicates high reliability and unbiased predictive performance. Among these models, CatBoost, GBDT, and XGBoost models exhibit minimal error fluctuations, reflecting more consistent predictive performance. Figure 10 displays the error values of predicted CS for the eight models. Errors in the training set are lower than those in the testing set, with most training-set errors concentrated between -15 and 15 MPa. However, the magnitude of these error fluctuations is amplified in the testing set. Notably, the RF and SVR models show larger error fluctuations, which suggests limitations in their ability to handle data noise and capture non-linear relationships. This is consistent with the lower R2 values reported earlier.
Figure 9. Error fluctuations between the actual values and predicted values based on the eight machine learning models. (a) AdaBoost model; (b) AutoML; (c) CatBoost model; (d) GBDT model; (e) LightGBM model; (f) RF model; (g) SVR model; (h) XGBoost model. ML: machine learning; GBDT: gradient boosting decision tree; LightGBM: light gradient boosting machine; RF: random forest; SVR: support vector regression.
Figure 10. The error values of the predicted compressive strength for eight models. (a) Training set; (b) Testing set. ML: machine learning; LightGBM: light gradient boosting machine; RF: random forest; SVR: support vector regression; GBDT: gradient boosting decision tree.
Table 5 presents the evaluation indicators for the eight ML algorithms on the training, testing, and overall datasets, with Figure 11 providing a radar chart visualization of these indicators. The results indicate that most ensemble boosting models achieve strong predictive performance and generalization ability, with all the models achieving R2 values above 0.83. XGBoost and CatBoost demonstrate excellent performance on the testing set, with R2 values of 0.8684 and 0.8638, respectively. Notably, these two models exhibit significantly higher R2 values on the training dataset than on the testing dataset, indicating a degree of overfitting, where the models learn the training data extremely well but generalize slightly less effectively to the testing data. The ρ metric, a composite indicator combining RRMSE and R2, reflects the model’s fitting capacity, with lower values indicating better performance. The maximum ρ value across all models is 0.2500 for the SVR model, which exceeds the 0.2 threshold recommended by Gandomi et al.[66], suggesting the SVR model is unsuitable for providing reliable predictions. The ρ values of the XGBoost and CatBoost models are 0.1659 and 0.1690, respectively. This indicates a high consistency between the two models’ predicted results and the actual results. Furthermore, the XGBoost, CatBoost, and GBDT models also display relatively low values for RMSE, RRMSE, and MAE.
Figure 11. Radar chart of evaluation indicators based on the compressive strength dataset. (a) AdaBoost model; (b) AutoML; (c) CatBoost model; (d) GBDT model; (e) LightGBM model; (f) RF model; (g) SVR model; (h) XGBoost model. ML: machine learning; GBDT: gradient boosting decision tree; LightGBM: light gradient boosting machine; RF: random forest; SVR: support vector regression; MAE: mean absolute error; RMSE: root mean square error; RRMSE: relative root mean square error.
| Developed model | Dataset | R2 | RMSE (MPa) | RRMSE | MAE (MPa) | ρ |
| AdaBoost | Training | 0.9710 | 2.8712 | 0.1650 | 2.4908 | 0.0831 |
| Testing | 0.8370 | 6.9018 | 0.3567 | 4.9615 | 0.1863 | |
| Overall | 0.9300 | 4.4858 | 0.2494 | 3.2358 | 0.1270 | |
| AutoML | Training | 0.9587 | 3.4241 | 0.1968 | 2.2977 | 0.0994 |
| Testing | 0.8340 | 6.9649 | 0.3600 | 4.8686 | 0.1882 | |
| Overall | 0.9206 | 4.7767 | 0.2656 | 3.0730 | 0.1355 | |
| CatBoost | Training | 0.9999 | 0.1918 | 0.0110 | 0.1412 | 0.0055 |
| Testing | 0.8638 | 6.3074 | 0.3260 | 4.1670 | 0.1690 | |
| Overall | 0.9582 | 3.4673 | 0.1928 | 1.3552 | 0.0974 | |
| GBDT | Training | 0.9962 | 1.0341 | 0.0594 | 0.2776 | 0.0297 |
| Testing | 0.8686 | 6.1953 | 0.3202 | 4.1815 | 0.1657 | |
| Overall | 0.9571 | 3.5101 | 0.1951 | 1.4549 | 0.0986 | |
| LightGBM | Training | 0.9892 | 1.7550 | 0.1009 | 1.2042 | 0.0506 |
| Testing | 0.8461 | 6.7045 | 0.3465 | 4.4191 | 0.1805 | |
| Overall | 0.9453 | 3.9631 | 0.2203 | 2.1737 | 0.1117 | |
| RF | Training | 0.9420 | 4.0592 | 0.2333 | 2.7933 | 0.1184 |
| Testing | 0.8032 | 7.5819 | 0.3919 | 5.2869 | 0.2067 | |
| Overall | 0.8996 | 5.3706 | 0.2986 | 3.5453 | 0.1532 | |
| SVR | Training | 0.9618 | 3.2947 | 0.1893 | 1.4800 | 0.0956 |
| Testing | 0.7253 | 8.9580 | 0.4630 | 5.6388 | 0.2500 | |
| Overall | 0.8894 | 5.6374 | 0.3134 | 2.7341 | 0.1613 | |
| XGBoost | Training | 0.9950 | 1.1883 | 0.0683 | 0.8363 | 0.0342 |
| Testing | 0.8724 | 6.2007 | 0.3205 | 4.0874 | 0.1659 | |
| Overall | 0.9562 | 3.5469 | 0.1972 | 1.8167 | 0.0997 |
RMSE: root mean square error; RRMSE: relative root mean square error; MAE: mean absolute error; ML: machine learning; GBDT: gradient boosting decision tree;LightGBM: light gradient boosting machine; RF: random forest; SVR: support vector regression.
Model performance was further assessed using a Taylor diagram based on the testing set, as depicted in Figure 12. This diagram simultaneously evaluates the agreement between actual and predicted data using three indicators: SD, RMSE, and Pearson correlation coefficient (R). This enables the selection of the model with the strongest overall predictive capability. The observation point, positioned at the bottom right, represents the SD of the actual dataset with a correlation coefficient of 1. The radial distance from a model point to the origin (0,0) indicates the SD of the predictions. The angle between the model point and the x-axis represents the correlation coefficient between the predicted and actual values. The direct distance from the model point to the observation point corresponds to the RMSE. The optimal model point is the one nearest to the observation point, indicating a low RMSE, high R, and an SD closely aligned with the observations. As observed, the XGBoost model demonstrates a high correlation coefficient of 0.9318, a SD of 15.9495 MPa (most similar to that of the actual observations), and a low RMSE of 6.2007 MPa. Although the GBDT model achieves a marginally higher R value and a slightly lower RMSE than the XGBoost model, the SD value of the XGBoost model aligns more closely with that of the actual observations. Consequently, a comprehensive evaluation confirms that the XGBoost model outperforms all the other predictive models.
Figure 12. Taylor diagram of the performance metrics based on the compressive strength dataset. ML: machine learning; LightGBM: light gradient boosting machine; RF: random forest; SVR: support vector regression; GBDT: gradient boosting decision tree.
4.2 Interpretability analysis
An in-depth understanding of how ML models predict the CS of AALRS based on input features including properties of LRS, mixture proportions, environmental conditions, and enhancement methods is vital for gaining profound insight into the strength development mechanisms of AALRS under extreme environments. However, the inherent “black box” nature of many powerful ML algorithms can impede in-depth exploration of these mechanisms. To address this crucial transparency gap, SHAP analysis (a powerful, theoretically grounded tool designed to elucidate predictions from complex ML algorithms) was employed in this study. Specifically, it was applied to interpret the predictions of the best-performing XGBoost model and quantify the impact of each input feature on these predictions. This process significantly enhances the model’s transparency regarding its prediction logic and internal decision-making, providing intuitive insights and enabling reliable interpretation.
The SHAP beeswarm plot visualizes the distribution of the SHAP values for each input feature, as shown in Figure 13. The X-axis represents the SHAP value of each data point for a given input feature, indicating the contribution of feature values to the model’s prediction relative to a baseline. Positive SHAP values indicate that the feature value increases the predicted CS, while negative values indicate that the feature value decreases the predicted CS; larger absolute values signify a stronger positive or negative influence. The color of the points represents the original value of the input feature, with red representing high feature values and blue representing low feature values. For MPD, most red points have negative SHAP values, whereas blue points predominantly have positive SHAP values. This indicates a negative correlation between MPD and predicted CS. This correlation likely arises because a reduction in MPD increases the reactive surface area of LRS, promoting more sufficient reactions and thus leading to higher CS. Similarly, Si/Al ratio, W/B, and VD exhibit negative correlations with predicted CS. Regarding the Si/Al ratio, the Al-O bond is more easily broken than the Si-O bond; thus, higher Al content (lower Si/Al ratio) facilitates the dissolution of aluminosilicates in alkaline solutions, which contributes to stronger CS development. The strength enhancement associated with lower VD may be attributed to silicate clusters formed during dehydration. These clusters bind LRS particles to enhance the CS of AALRS. Moreover, lower VD may also reduce weathering effects, which induce CS loss. Conversely, IT, alkali content (AC), curing age (CA), TT, and SM show positive correlations with CS, with IT and AC exerting particularly strong positive impacts. Additionally, feature variables related to enhancement methods, specifically heat activation temperature (HAT), Al2O3, heat activation time (HATT), and Ca(OH)2, are all positively correlated with CS.
Figure 13. Shapley value contribution and normalization importance ranking. MPD: median particle diameter of lunar regolith simulant; IT: initial temperature during curing; AC: alkali content of alkali-activated lunar regolith simulant; CA: curing age; TT: termination temperature during curing; CV: coefficient of variation; SM: silicate modulus; VD: vacuum degree; HAT: heat activation temperature; HATT: heat activation time; BF: basalt fiber content; CNF: carbon nanofiber; RLD: ratio of length-to-diameter of fiber.
The Y-axis ranks all input features from top to bottom based on their mean absolute SHAP value, representing the average magnitude of each feature’s impact on the model’s predictive outcomes. The results reveal that MPD, IT, AC, CA, Si/Al, and TT are the key features determining predictions. In practical lunar construction engineering applications, these features should be prioritized for monitoring and optimization to ensure reliable CS of AALRS. Additionally, the points for MPD, IT, and AC are widely spread horizontally, indicating that the direction and magnitude of their impact on predicted CS vary considerably across different samples. For example, high AC strongly promotes CS in some samples but have a milder effect in others.
5. Model Performance Evaluation of Flexural Strength
5.1 Model validation
The precise prediction of flexural strength is crucial for designing reliable lunar structures such as roads and platforms that withstand bending loads. Figure 14 presents a comparison between the actual and predicted FS values for the eight models. Consistent with the findings for CS, ensemble models based on gradient boosting, namely AdaBoost, CatBoost, XGBoost, and GBDT, generally exhibit superior predictive capabilities. This is demonstrated by their high R2 values and a significant proportion of data points falling within the ±10% error range, which is particularly evident in the training set results. Compared with the CS prediction results, the predictions for FS exhibit notably greater scatter, especially in the testing set. This increased dispersion is highlighted by a larger proportion of test data points falling outside the ±10% error range, and even beyond the ±20% error range.
Figure 14. Comparison of actual and predicted flexural strength. (a) Training set; (b) Testing set. ML: machine learning; LightGBM: light gradient boosting machine; RF: random forest; SVR: support vector regression; GBDT: gradient boosting decision tree.
Figure 15 illustrates the error fluctuations between the actual and predicted flexural strength values for the eight ML models. These error histograms provide insights into model robustness, where a distribution more concentrated around zero indicates higher reliability. Similar to the observations for CS, AdaBoost, CatBoost, XGBoost, and GBDT models exhibit minimal fluctuations in FS predictions, again demonstrating their superior robustness in capturing the properties of the material. Figure 16 displays the prediction error values of FS for the eight models. Most prediction errors were concentrated between -4 and 4 MPa. The RF and SVR models exhibit larger errors and lower reliability due to sensitivity to data noise. The error fluctuations for FS are largely similar to those for CS.
Figure 15. Error fluctuations between the actual and predicted values based on the eight machine learning models. (a) AdaBoost model; (b) AutoML; (c) CatBoost model; (d) GBDT model; (e) LightGBM model; (f) RF model; (g) SVR model; (h) XGBoost model. ML: machine learning; GBDT: gradient boosting decision tree; LightGBM: light gradient boosting machine; RF: random forest; SVR: support vector regression.
Figure 16. The error values of the predicted flexural strength for eight models. (a) Training set; (b) Testing set. ML: machine learning; LightGBM: light gradient boosting machine; RF: random forest; SVR: support vector regression; GBDT: gradient boosting decision tree.
Table 6 presents the performance indicators for the eight ML algorithms evaluated on the training, testing, and overall datasets for flexural strength, with Figure 17 providing a radar chart visualization of these indicators. The XGBoost model demonstrates the best performance among all the algorithms. In the testing set, it achieves R2 = 0.8507, RMSE = 0.8382 MPa, RRMSE = 0.2269, MAE = 0.5725 MPa, and ρ = 0.1180. The AdaBoost model also performs well on the testing set, showing competitive predictive capability. It achieves R2 = 0.8414, RMSE = 0.8807 MPa, RRMSE = 0.2384, MAE = 0.5624 MPa, and ρ = 0.1243. The AutoML exhibits the weakest generalization ability. It achieves R2 = 0.6865, RMSE = 1.2498 MPa, RRMSE = 0.3384, MAE = 0.7582 MPa, and ρ = 0.1850. These results confirm that the ensemble boosting models, particularly XGBoost, exhibit the best generalization ability for predicting FS among the predictive models, which is consistent with the findings for CS.
Figure 17. Radar chart of evaluation indicators based on the flexural strength dataset. (a) AdaBoost model; (b) AutoML; (c) CatBoost model; (d) GBDT model; (e) LightGBM model; (f) RF model; (g) SVR model; (h) XGBoost model. ML: machine learning; GBDT: gradient boosting decision tree; LightGBM: light gradient boosting machine; RF: random forest; SVR: support vector regression; MAE: mean absolute error; RMSE: root mean square error; RRMSE: relative root mean square error.
| Developed model | Dataset | R2 | RMSE (MPa) | RRMSE | MAE (MPa) | ρ |
| AdaBoost | Training | 0.9932 | 0.2078 | 0.0642 | 0.1449 | 0.0321 |
| Testing | 0.8414 | 0.8807 | 0.2384 | 0.5624 | 0.1243 | |
| Overall | 0.9524 | 0.5144 | 0.1525 | 0.2711 | 0.0771 | |
| AutoML | Training | 0.9999 | 0.0034 | 0.0011 | 0.0003 | 0.0005 |
| Testing | 0.6865 | 1.2498 | 0.3384 | 0.7582 | 0.1850 | |
| Overall | 0.9161 | 0.6872 | 0.2037 | 0.2294 | 0.1040 | |
| CatBoost | Training | 0.9537 | 0.5486 | 0.1696 | 0.3240 | 0.0858 |
| Testing | 0.8416 | 0.8645 | 0.2340 | 0.5361 | 0.1220 | |
| Overall | 0.9237 | 0.6603 | 0.1957 | 0.3881 | 0.0998 | |
| GBDT | Training | 0.9759 | 0.3784 | 0.1170 | 0.2018 | 0.0588 |
| Testing | 0.8451 | 0.8595 | 0.2320 | 0.5217 | 0.1212 | |
| Overall | 0.9418 | 0.5686 | 0.1685 | 0.2985 | 0.0855 | |
| LightGBM | Training | 0.9151 | 0.7113 | 0.2199 | 0.4488 | 0.1124 |
| Testing | 0.8005 | 0.9775 | 0.2646 | 0.6582 | 0.1396 | |
| Overall | 0.8846 | 0.8011 | 0.2375 | 0.5121 | 0.1223 | |
| RF | Training | 0.8944 | 0.7992 | 0.2471 | 0.4385 | 0.1270 |
| Testing | 0.7704 | 1.0494 | 0.2840 | 0.6352 | 0.1512 | |
| Overall | 0.8611 | 0.8823 | 0.2615 | 0.4979 | 0.1356 | |
| SVR | Training | 0.9003 | 0.7093 | 0.2057 | 0.3415 | 0.1055 |
| Testing | 0.8458 | 0.8567 | 0.2319 | 0.5058 | 0.1208 | |
| Overall | 0.8487 | 0.9235 | 0.2737 | 0.3775 | 0.1425 | |
| XGBoost | Training | 0.9596 | 0.5083 | 0.1571 | 0.3223 | 0.0793 |
| Testing | 0.8507 | 0.8382 | 0.2269 | 0.5725 | 0.1180 | |
| Overall | 0.9524 | 0.5144 | 0.1525 | 0.2711 | 0.0771 |
RMSE: root mean square error; RRMSE: relative root mean square error; MAE: mean absolute error; ML: machine learning; GBDT: gradient boosting decision tree; LightGBM: light gradient boosting machine; RF: random forest; SVR: support vector regression.
Model performance for predicting flexural strength was assessed using a Taylor diagram based on the testing set, as depicted in Figure 18. Consistent with the findings for CS, the Taylor diagram for FS clearly identifies XGBoost as the best-performing model, exhibiting the highest R value of 0.9223, the lowest RMSE of 0.8382 MPa, and a SD of 1.9927 MPa. Moreover, the SDs of AdaBoost and GBDT models (2.1616 MPa and 2.0952 MPa, respectively) are remarkably close to that of the observation (SD = 2.1690 MPa), indicating a strong ability to replicate the variability of the actual FS data. Despite some models showing specific advantages in matching the SD, a comprehensive evaluation that simultaneously considers R, SD, and RMSE emphasizes that the XGBoost model provides the most accurate and reliable overall predictions for FS.
Figure 18. Taylor diagram of the performance metrics based on the flexural strength dataset. ML: machine learning; LightGBM: light gradient boosting machine; RF: random forest; SVR: support vector regression; GBDT: gradient boosting decision tree.
5.2 Interpretability analysis
Figure 19 shows the SHAP beeswarm plot, which displays the relative importance of input features for the XGBoost model’s predictions of AALRS flexural strength. The analysis reveals that the correlations between most input features and FS are largely similar to those observed for CS. Specifically, IT, SM, TT, CA, AC, and specimen volume (flexural) (FV) exhibit positive correlations with FS, as indicated by the general trend of higher feature values corresponding to positive SHAP values. Conversely, Si/Al ratio, VD, and W/B demonstrate negative correlations with FS. Notably, high values of MPD do not yield consistently positive or negative contributions to FS. The distribution of their corresponding data points on both the positive and negative sides of the SHAP value axis suggests that their contribution to FS could be governed by interaction effects with other features. Furthermore, IT and SM, the most impactful features, show significant positive correlations with FS and the widest dispersion of SHAP values. This emphasizes their substantial overall influence. Consequently, these features should be prioritized during the design and optimization of AALRS materials.
Figure 19. Shapley value contribution and normalization importance ranking. IT: initial temperature during curing; TT: termination temperature during curing; CA: curing age; VD: vacuum degree; AC: alkali content of alkali-activated lunar regolith simulant; MPD: median particle diameter of lunar regolith simulant; SM: silicate modulus; FV: specimen volume (flexural).
6. Performance Optimization Design of AALRS
This section details the multi-objective optimization for AALRS. Initially, the optimal ML algorithm selected based on the evaluation indicators was employed to establish performance-based objective functions for CS and FS. To incorporate resource efficiency, the TP was designated as a third objective function, which is inversely proportional to the ISRU rate. Constrained by the allowable ranges of input parameters, the NSGA-II algorithm was then applied to perform multi-objective optimization, generating a set of Pareto optimal solutions. These solutions were subsequently subjected to comprehensive evaluation and ranking using the TOPSIS methodology. This approach enabled the determination of the optimal parameters, including properties of LRS, mixture proportions, preparation parameters, and environmental conditions, that offer the best compromise between mechanical performance and ISRU efficiency. The overall optimization workflow is depicted in Figure 4.
6.1 Objective functions
The objective functions aim to maximize the mechanical strength (including CS and FS) of AALRS while minimizing its TP. Based on previous evaluations, the XGBoost model was chosen to formulate objective functions for mechanical performance in the multi-objective optimization design, owing to its high predictive accuracy and robustness. These functions are represented in Eqs. (17)-(18). TP was adopted as the direct evaluation criterion for ISRU efficiency, quantifying the mass of water, alkali, and extra silicon required for AALRS preparation that must be transported from Earth. The system boundary for this objective is defined as the total mass of constituents (not sourced in-situ) required for the AALRS preparation, with the calculation formula shown in Eq. (19). Notably, increasing the proportion of processed LRS in preparation is anticipated to significantly reduce the overall TP, which is a critical factor for the economic viability and sustainability of lunar construction.
where x1,x2,x3,x4,x5,x6,x7,x8,x9,x10,x11,x12,x13,x14,x15,x16,x17 are the MPD, Si/Al, W/B, AC, SM, CA, CV/FV, IT, TT, VD, basalt fiber content (BF), ratio of length-to-diameter of fiber (RLD), carbon nanotubes, Al2O3, Ca(OH)2, HAT, and HATT, respectively. The parameter VD was treated as a discrete variable, with permissible values of 10 Pa, 100 Pa, and 101,325 Pa.
6.2 Constraint boundary
The establishment of well-defined constraints is a critical prerequisite for achieving meaningful and practically relevant outcomes in multi-objective optimization. Therefore, prior to initiating the mixture proportion optimization design for AALRS, comprehensive constraint boundaries were established. These constraints are fundamentally derived from experimental programs in existing studies and practical condition limitations, defining the permissible ranges, interdependencies, and specific conditions for the decision variables, ensuring that the optimized solutions are both physically meaningful and experimentally feasible. The constraints are detailed as follows:
(1) Each decision variable xi is subjected to the constraints of the following value intervals.
(2) This study aims to optimize AALRS performance with discrete choices for certain parameters. For instance, VD is treated as a discrete decision variable, and its values are restricted to specific levels observed or targeted in experiments. Therefore, the value of discrete variable x10 is as follows.
6.3 Multi-objective optimization
In this study, the NSGA-II algorithm was configured with a population size of 500, a maximum number of iterations of 500, a judgment threshold for stagnation in objective optimization of 1 × 10-6, and a maximum upper limit for the evolutionary stagnation counter set at 100. The performance of the NSGA-II algorithm under four distinct objective weighting strategies was evaluated using spacing trace plots during iterative generations, as presented in Figure 20. The spacing indicator evaluates the uniformity of solution distribution along the Pareto front. A low spacing value suggests that solutions are evenly distributed along the Pareto front, with no large gaps or excessive clustering in specific regions. The four weighting strategies are defined as follows: The first strategy (CS = 0.1, FS = 0.1, TP = 0.8) heavily emphasizes TP optimization, which corresponds to maximizing ISRU efficiency. The second strategy (CS = 0.33, FS = 0.33, TP = 0.34) applies relatively balanced weights across all three objectives. The third and fourth strategies (CS = 0.1, FS = 0.8, TP = 0.1 and CS = 0.8, FS = 0.1, TP = 0.1) primarily focus on optimizing FS and CS, respectively. For all strategies, the spacing values exhibit a general decreasing trend and eventually stabilize. This further confirms the efficiency of the algorithm in generating a well-distributed set of solutions.
Figure 20. Spacing trace plots. CS: compressive strength; FS: flexural strength; TP: transport payload.
Figure 21 shows the multi-objective optimization results of the AALRS design, where green dots denote actual data, blue dots represent the Pareto front of non-dominated solutions derived from the NSGA-II algorithm, and red stars highlight the top 50 optimal designs selected using TOPSIS. Four distinct weighting strategies were set, and each strategy leads to a different emphasis in the optimization. When the strategy prioritizes minimizing TP (CS = 0.1, FS = 0.1, TP = 0.8), the Pareto front and top 50 solutions cluster toward lower TP values and mechanical strength, which demonstrates that significant sacrifices in CS and FS are required to achieve maximal ISRU efficiency. A balanced weighting strategy (CS = 0.33, FS = 0.33, TP = 0.34) yields a more broadly distributed top 50 selection, indicating that solutions offering various trade-offs among CS, FS, and TP are all considered viable. When the focus is on maximizing CS (CS = 0.8, FS = 0.1, TP = 0.1), these optimal solution sets gravitate towards regions with higher CS. This strategy underscores a distinct trend of significant increase in CS, though this often necessitates higher TP values or a less optimized FS. When the focus is on maximizing FS (CS = 0.1, FS = 0.8, TP = 0.1), this strategy prioritizes FS. The results are similar to those of the CS-maximizing strategy, but with the optimization emphasis shifted to FS. These visualizations demonstrate the intrinsic trade-offs in AALRS design. The increase in the mechanical strength of AALRS often conflicts with the practical demand of minimizing the amount of materials transported for lunar base construction. This analysis is vital for decision-making, as it enables the selection of a strategy that best aligns with specific priorities and resource constraints, thereby optimizing the material performance of AALRS for practical lunar construction applications. Additionally, the clear delineation of TOPSIS solutions within each strategy effectively highlights the preferred solutions from the broader Pareto front.
Figure 21. Optimization results of AALRS. AALRS: alkali-activated lunar regolith simulant; CS: compressive strength; FS: flexural strength; TP: transport payload.
6.4 GUI design for AALRS
The complexities of database compilation, model training, and testing present challenges to the practical application of ML. Therefore, the development and utilization of a GUI is critically important for designing AALRS under extreme environments. These interfaces bridge the gap between complex predictive models and user-friendly applications, allowing researchers to leverage powerful analytical tools without needing deep expertise in the underlying computational complexity. The GUI aims to achieve strength prediction and performance optimization of AALRS based on the multi-objective functions (Eqs. (17)-(19)) and variable constraint ranges (Eqs. (20)-(25)), as shown in Figure 22, Figure 23, Figure 24, Figure 25 and Figure 26.
Figure 22. GUI for strength prediction of AALRS based on the XGBoost model. GUI: graphical user interface; AALRS: alkali-activated lunar regolith simulant; MPD: median particle diameter of lunar regolith simulant; AC: alkali content of alkali-activated lunar regolith simulant; IT: initial temperature during curing; BF: basalt fiber content; HATT: heat activation time; SM: silicate modulus; TT: termination temperature during curing; RLD: ratio of length-to-diameter of fiber; CV: coefficient of variation; CS: compressive strength; FS: flexural strength; TP: transport payload.
Figure 23. GUI for performance optimization design of AALRS (CS = 0.10, FS = 0.10, TP = 0.80). GUI: graphical user interface; AALRS: alkali-activated lunar regolith simulant; CS: compressive strength; FS: flexural strength; TP: transport payload.
Figure 24. GUI for performance optimization design of AALRS (CS = 0.33, FS = 0.33, TP = 0.34). GUI: graphical user interface; AALRS: alkali-activated lunar regolith simulant; CS: compressive strength; FS: flexural strength; TP: transport payload.
Figure 25. GUI for performance optimization design of AALRS (CS = 0.10, FS = 0.80, TP = 0.10). GUI: graphical user interface; AALRS: alkali-activated lunar regolith simulant; CS: compressive strength; FS: flexural strength; TP: transport payload.
Figure 26. GUI for performance optimization design of AALRS (CS = 0.80, FS = 0.10, TP = 0.10). GUI: graphical user interface; AALRS: alkali-activated lunar regolith simulant; CS: compressive strength; FS: flexural strength; TP: transport payload.
Figure 22 details the GUI for strength prediction. The interface is divided into an “Input Panel” on the left, an “Output Panel” on the right, and a “Control Panel” situated at the bottom. In the input panel, users can define input parameters, including MPD, Si/Al, W/B, AC, SM, CA, IT, TT, VD, BF, RLD, carbon nanofiber, Al2O3, Ca(OH)2, HAT, HATT, CV, and FV. Activating the “Predict” button located in the control panel initiates the calculation, with the output panel subsequently displaying three output values, including CS, FS, and TP. Experimental data from the study[45] were randomly selected and input into the GUI for validation. The validation results confirm only a slight deviation between the GUI-predicted value of 12.34 MPa and the actual CS value of 13.20 MPa.
Figure 23, Figure 24, Figure 25 and Figure 26 provide a comprehensive overview of the GUI designed for the performance optimization of AALRS, showing various interactive panels. These panels include “Objective Weights”, “Fixed Parameters”, “Optimization Results”, “Output Objectives”, “TOPSIS Indicators”, “Optimization Visualization”, and “Control Panel”. In the “Objective Weights” panel, users prioritize different scenarios and strategies by assigning specific weights to objectives such as CS, FS, and TP. Additionally, the “Fixed Parameters” panel allows users to establish boundary conditions and constraints for optimization. Upon activation of the “Optimization” button, the XGBoost algorithm calculates the optimal input variables. The corresponding predicted values for each output objective are shown in the “Output Objectives” panel. In addition, the “Normalized Pareto Front” is presented in the “Optimization Visualization” panel. The Pareto optimal solution set exhibits performance within reasonable limits, demonstrating that the established multi-objective optimization model is effective. Table 7 summarizes the optimal mix proportions and output results of AALRS under different strategies. For instance, under the strategy prioritizing maximizing ISRU (CS = 0.1, FS = 0.1, TP = 0.8), the TP achieves a minimum value of 19.09%, while AALRS maintains adequate mechanical strength. This exemplifies a viable pathway towards minimizing TP and associated costs. In contrast, when prioritizing CS (CS = 0.8, FS = 0.1, TP = 0.1), the AALRS attains a maximum CS of 57.83 MPa. Notably, the AC reaches its optimal upper limit of 6.98%, underscoring the significant role of alkali in enhancing CS. Conversely, strategies targeting low TP minimize the demand for alkali and water, as these components directly contribute to the increase in TP. The GUI effectively translates complex optimization processes into a practical tool, enabling data-driven design, rapid scenario exploration, and systematic balancing of conflicting goals.
| Parameters | Strategy 1 | Strategy 2 | Strategy 3 | Strategy 4 |
| CS = 0.10, FS = 0.10, TP = 0.80 | CS = 0.33, FS = 0.33, TP = 0.34 | CS = 0.10, FS = 0.80, TP = 0.10 | CS = 0.80, FS = 0.10, TP = 0.10 | |
| MPD (μm) | 15.76 | 15.04 | 36.39 | 15.08 |
| Si/Al (--) | 2.60 | 2.60 | 2.35 | 2.60 |
| W/B (--) | 0.05 | 0.05 | 0.10 | 0.05 |
| AC (%) | 6.20 | 6.90 | 6.00 | 6.98 |
| SM (--) | 2.00 | 2.00 | 2.00 | 2.00 |
| IT (℃) | 80.09 | 80.15 | 115.95 | 81.24 |
| TT (℃) | 105.16 | 105.18 | 105.00 | 80.11 |
| VD (Pa) | 101,325.00 | 101,325.00 | 101,325.00 | 101,325.00 |
| f(CS) (MPa) | 52.96 | 54.99 | 36.87 | 57.83 |
| f(FS) (MPa) | 6.00 | 6.02 | 9.16 | 4.81 |
| f(TP) (%) | 19.09 | 20.45 | 21.75 | 20.60 |
AALRS: alkali-activated lunar regolith simulant; MPD: median particle diameter of lunar regolith simulant; W/B: water-to-binder; AC: alkali content of alkali-activated lunar regolith simulant; SM: silicate modulus; IT: initial temperature during curing; TT: termination temperature during curing; VD: vacuum degree.
7. Conclusion and Future Works
This study established an integrated ML prediction and multi-objective optimization framework for the design of AALRS, with the objectives of maximizing mechanical performance and minimizing Earth-to-Moon TP. The main conclusions are as follows.
(1) Among the eight evaluated ML algorithms encompassing different categories, XGBoost demonstrated superior performance in predicting the compressive and flexural strengths of AALRS. For CS prediction, XGBoost achieved an R2 of 0.8684, an RMSE of 6.2007 MPa, and a MAE of 4.0874 MPa on the testing dataset, indicating strong generalization ability. Similarly, for flexural strength prediction, XGBoost exhibited the best performance with an R2 of 0.8507, an RMSE of 0.8382 MPa, and a MAE of 0.5725 MPa on the testing dataset. The superiority of XGBoost is attributed to its gradient boosting architecture with built-in regularization, which effectively balances model complexity and generalization capacity under a moderate-sized dataset.
(2) SHAP interpretability analysis revealed that the dominant factors governing CS and FS follow distinct patterns. For CS, MPD, IT, AC, CA, Si/Al, and TT were identified as the key influencing factors. MPD, Si/Al, W/B, and VD exhibited negative correlations with CS, while IT, AC, CA, TT, and SM showed positive correlations. For flexural strength, IT and SM were the most influential factors with significant positive correlations. The SHAP analysis provided valuable insights into the underlying mechanisms governing AALRS strength development.
(3) The NSGA-II multi-objective optimization coupled with TOPSIS decision-making demonstrated the feasibility of achieving favorable trade-offs between mechanical strength and ISRU efficiency. Under the strategy of prioritizing ISRU maximization (CS = 0.1, FS = 0.1, TP = 0.8), an optimal mix design achieved a TP of only 19.09% while maintaining adequate mechanical strength (CS of 52.96 MPa and FS of 6.00 MPa). When prioritizing CS (CS = 0.8, FS = 0.1, TP = 0.1), the optimized AALRS achieved a maximum CS of 57.83 MPa with a slightly higher TP of 20.60%. The developed GUI tools for strength prediction and performance optimization convert the complex computational processes into practical and visual tools for data-driven AALRS design.
Future research should advance in three directions. First, the optimization framework should be expanded by integrating SHAP-identified key factors with lunar-specific functional indicators, such as thermal conductivity, radiation shielding capacity, and long-term durability under thermal cycling, to address both mechanical and service performance requirements. Second, low-energy preparation routes via solar-driven alkali activation should be explored to enhance ISRU efficiency and minimize Earth-sourced payload dependency. Third, the GUI-based design tool should be coupled with lunar structural analysis modules, evolving from material-level prediction into an integrated material-structure design platform for engineering application.
Authors contribution
Yao Y: Conceptualization, methodology, data curation, formal analysis, writing-original draft.
Zhang W: Data curation, investigation, validation.
Liu H: Software, formal analysis, visualization.
Zhu C: Supervision, project administration, funding acquisition, writing-review & editing.
Li X: Investigation, resources.
Chen X: Validation, data curation.
Wang Y: Visualization, formal analysis.
Liu C: Supervision, writing-review & editing.
Conflicts of interest
Chao Liu and Huawei Liu are Editorial Board Members of Journal of Building Design and Environment. The other authors declare no conflicts of interest.
Ethical approval
Not applicable.
Consent to participate
Not applicable.
Availability of data and materials
Data supporting the findings of this study are available from the corresponding author upon reasonable request.
Funding
This work was supported by the National Natural Science Foundation of China (Grant No. 52508304), the Shaanxi Provincial Natural Science Basic Research Program (Grant No. 2025JC-YBMS-550), and Shaanxi University Youth Innovation Team Project (Grant No. 2023–2026).
Copyright
© The Author(s) 2026.
References
-
1. Bao C, Zhang D, Wang Q, Cui Y, Feng P. Lunar in situ large-scale construction: Quantitative evaluation of regolith solidification techniques. Engineering. 2024;39:204-221.[DOI]
-
2. Zhou C, Gao Y, Zhou Y, She W, Shi Y, Ding L, et al. Properties and characteristics of regolith-based materials for extraterrestrial construction. Engineering. 2024;37:159-181.[DOI]
-
3. Azami M, Kazemi Z, Moazen S, Dubé M, Potvin MJ, Skonieczny K. A comprehensive review of lunar-based manufacturing and construction. Prog Aerosp Sci. 2024;150:101045.[DOI]
-
4. Granier J, Cutard T, Pinet P, Le Maoult Y, Chevrel S, Sentenac T, et al. Selective laser melting of partially amorphous regolith analog for ISRU lunar applications. Acta Astronaut. 2025;226:66-77.[DOI]
-
5. Guerrero-Gonzalez FJ, Zabel P. System analysis of an ISRU production plant: Extraction of metals and oxygen from lunar regolith. Acta Astronaut. 2023;203:187-201.[DOI]
-
6. Robinot J, Rodat S, Abanades S, Paillet A, Cowley A. Review of in-situ oxygen extraction from lunar regolith with focus on solar thermal and laser vacuum pyrolysis. Acta Astronaut. 2025;234:242-259.[DOI]
-
7. Geng Z, Wu Z, Wang X, Zhang L, She W, Tan MJ. A novel 3D printing scheme for lunar construction with extremely low binder utilization. Addit Manuf. 2025;99:104657.[DOI]
-
8. Xiong G, Guo X, Yuan S, Xia M, Wang Z. The mechanical and structural properties of lunar regolith simulant based geopolymer under extreme temperature environment on the moon through experimental and simulation methods. Constr Build Mater. 2022;325:126679.[DOI]
-
9. Geng Z, Zhang L, Pan H, She W, Zhou C, Zhou H, et al. In-situ solidification of alkali-activated lunar regolith: Insights into the chemical and physical origins. J Clean Prod. 2023;391:136147.[DOI]
-
10. Sun Y, Ma S, Chen Q, Chen G, He P, Wang Y, et al. Lunar regolith simulant-derived 3D-printed geopolymers with optimized mechanical and thermal management properties. Compos Part A Appl Sci Manuf. 2025;196:108989.[DOI]
-
11. Alexiadis A, Alberini F, Meyer ME. Geopolymers from lunar and Martian soil simulants. Adv Space Res. 2017;59(1):490-495.[DOI]
-
12. Collins PJ, Edmunson J, Fiske M, Radlińska A. Materials characterization of various lunar regolith simulants for use in geopolymer lunar concrete. Adv Space Res. 2022;69(11):3941-3951.[DOI]
-
13. Davis G, Montes C, Eklund S. Preparation of lunar regolith based geopolymer cement under heat and vacuum. Adv Space Res. 2017;59(7):1872-1885.[DOI]
-
14. Xue G, Qiao G. Impacts of thermal activation on lunar regolith simulant-based precursor and resulting geopolymer: Composition, structure, solubility, and reactivity. Cem Concr Compos. 2025;155:105840.[DOI]
-
15. Geng Z, Zhang L, Wu Z, Huang J, Wang X, Tan MJ, et al. Silicate-lunar regolith composite: A vacuum self-hardened and comprehensively durable extraterrestrial construction material. Constr Build Mater. 2024;449:138517.[DOI]
-
16. Li Y, Pan P, Miao S, Feng Y. Basalt fiber reinforcement mechanism for geopolymer exposed to lunar temperature environment. Constr Build Mater. 2024;451:138845.[DOI]
-
17. Shao R, Wu C, Li J. Innovative development of geopolymer-based lunar high strength concrete for sustainable extra-terrestrial construction using large-scale regolith simulants. Constr Build Mater. 2024;450:138707.[DOI]
-
18. Wu Z, Geng Z, Pan H, Zhang L, She W. Determinants of lunar geopolymer synthesis based on experimental and statistical analysis: Towards lunar regolith’s diversity. Constr Build Mater. 2024;450:138557.[DOI]
-
19. Xue G, Qiao G. Coupling thermodynamic modelling with experimental study to reveal the evolutionary relationship of pore solutions, products, and compressive strength for lunar regolith simulant geopolymers. Compos Part B Eng. 2025;289:111949.[DOI]
-
20. Li L, Mao L, Yang J. A review of principles, analytical methods, and applications of SEM-EDS in cementitious materials characterization. Adv Mater Technol. 2025;10(7):2401175.[DOI]
-
21. Zhang W, Duan Z, Liu C, Yao Y, Liu H, Nasr A, et al. Performance optimisation and predictive modelling of rice husk ash recycled concrete under the coupled action of freeze-thaw cycles and chloride erosion: Experimental study and machine learning. Constr Build Mater. 2025;481:141467.[DOI]
-
22. Zhang W, Duan Z, Liu H, Yao Y, Zhang Z, Liu C. Salt freezing-thawing damage evolution model based on the time-dependent hydration reaction incorporating rice husk ash and recycled coarse aggregate. Constr Build Mater. 2024;440:137179.[DOI]
-
23. Li L, Ding S, Cai Y. Phase-Glue: An interpretable and efficient backscattered electron-energy dispersive X-ray spectroscopy (BSE-EDS) mapping analysis approach to dissect the phase assemblage of cementitious material systems. Expert Syst Appl. 2025;283:127923.[DOI]
-
24. Jing M, Jia H, Liu Q, Zhang K, Xu S, Zheng X, et al. Multi-objective optimization design of cement-based materials for low-carbon goals. Mater Today Commun. 2025;44:112135.[DOI]
-
25. Li F, Luo D, Niu D. Durability evaluation of concrete structure under freeze-thaw environment based on pore evolution derived from deep learning. Constr Build Mater. 2025;467:140422.[DOI]
-
26. Sun B, Ding L, Ye G, De Schutter G. Mechanical properties prediction of blast furnace slag and fly ash-based alkali-activated concrete by machine learning methods. Constr Build Mater. 2023;409:133933.[DOI]
-
27. Ding Y, Wei W, Wang J, Wang Y, Shi Y, Mei Z. Prediction of compressive strength and feature importance analysis of solid waste alkali-activated cementitious materials based on machine learning. Constr Build Mater. 2023;407:133545.[DOI]
-
28. Naser MZ. Extraterrestrial construction materials. Prog Mater Sci. 2019;105:100577.[DOI]
-
29. Debbarma S, Shi X, Torres A, Nodehi M. Fiber-reinforced lunar geopolymers synthesized using lunar regolith simulants. Acta Astronaut. 2024;214:593-608.[DOI]
-
30. Zhou S, Yang Z, Zhang R, Zhu X, Li F. Preparation and evaluation of geopolymer based on BH-2 lunar regolith simulant under lunar surface temperature and vacuum condition. Acta Astronaut. 2021;189:90-98.[DOI]
-
31. Egnaczyk TM, Hartt V WH, Mills JN, Wagner NJ. Composition-property relationships of BP-1 lunar regolith simulant geopolymers for in-situ resource utilization. Adv Space Res. 2024;73(1):885-917.[DOI]
-
32. Mills JN, Katzarova M, Wagner NJ. Comparison of lunar and Martian regolith simulant-based geopolymer cements formed by alkali-activation for in-situ resource utilization. Adv Space Res. 2022;69(1):761-777.[DOI]
-
33. Montes C, Broussard K, Gongre M, Simicevic N, Mejia J, Tham J, et al. Evaluation of lunar regolith geopolymer binder as a radioactive shielding material for space exploration applications. Adv Space Res. 2015;56(6):1212-1221.[DOI]
-
34. Zhou S, Zhu X, Lu C, Li F. Synthesis and characterization of geopolymer from lunar regolith simulant based on natural volcanic scoria. Chin J Aeronaut. 2022;35(1):144-159.[DOI]
-
35. Zhang R, Zhou S, Li F. Preparation of geopolymer based on lunar regolith simulant at in-situ lunar temperature and its durability under lunar high and cryogenic temperature. Constr Build Mater. 2022;318:126033.[DOI]
-
36. Ma Q, Zhang B. Study on the performance mechanism of alumina-alkali activator synergistic solidified lunar soil simulant. Case Stud Constr Mater. 2024;21:e03680.[DOI]
-
37. Zhan J, Xue X, Hua J, Huang L. Effects of curing temperature on early-age mechanical property and microstructure of lunar regolith simulant geopolymer. Case Stud Constr Mater. 2025;22:e04329.[DOI]
-
38. Zhou S, Lu C, Zhu X, Li F. Preparation and characterization of high-strength geopolymer based on BH-1 lunar soil simulant with low alkali content. Engineering. 2021;7(11):1631-1645.[DOI]
-
39. Chen L, Wang T, Li F, Zhou S. Preparation of geopolymer for in-situ pavement construction on the Moon utilizing minimal additives and human urine in lunar regolith simulant. Front Mater. 2024;11:1413432.[DOI]
-
40. Ma S, Jiang Y, Fu S, He P, Sun C, Duan X, et al. 3D-printed Lunar regolith simulant-based geopolymer composites with bio-inspired sandwich architectures. J Adv Ceram. 2023;12(3):510-525.[DOI]
-
41. Prater J, Kim YH. Carbon nanotube reinforced lunar-based geopolymer: Curing conditions. J Compos Sci. 2024;8(12):492.[DOI]
-
42. Sun X, Zheng X, Wang H, Wang G. Lunar regolith simulants and preparation aimed at in-situ construction technology. J Build Mater. 2025;25(1):72-81. Chinese.[DOI]
-
43. Li L, Shi Z, Wang L, Sui Y, Meng S. Experimental study on rheological properties and 3D printing of simulated lunar soil polymers. J Build Eng. 2025;104:112256.[DOI]
-
44. Yao Y, Liu C, Liu H, Chen X, Li X, Wang T, et al. Multiscale study of microstructural evolution in alkali-activated lunar regolith simulant under in-situ lunar temperatures: Insight into the reaction mechanism. J Build Eng. 2025;103:112209.[DOI]
-
45. Yao Y, Liu C, Zhang W, Liu H, Wang T, Wu Y, et al. Influence of vacuum and high-temperature on the evolution of mechanical strength and microstructure of alkali-activated lunar regolith simulant. J Build Eng. 2024;97:110709.[DOI]
-
46. Yao Y, Liu C, Zhang W, Liu H, Zhu C. Influence of in-situ lunar temperature on the pore structure and solid phases of alkali-activated lunar regolith simulant. J Build Eng. 2024;98:111162.[DOI]
-
47. Pilehvar S, Arnhof M, Pamies R, Valentini L, Kjøniksen AL. Utilization of urea as an accessible superplasticizer on the Moon for lunar geopolymer mixtures. J Clean Prod. 2020;247:119177.[DOI]
-
48. Zhang R, Zhou S, Li F, Bi Y, Zhu X. Mechanical and microstructural characterization of carbon nanofiber–reinforced geopolymer nanocomposite based on lunar regolith simulant. J Mater Civ Eng. 2022;34:04021387.[DOI]
-
49. Gu J, Ma Q. Experimental study on geopolymerization of lunar soil simulant under dry curing and sealed curing. Materials. 2024;17(6):1413.[DOI]
-
50. Lee S, van Riessen A. A review on geopolymer technology for lunar base construction. Materials. 2022;15(13):4516.[DOI]
-
51. Javed U, Shaikh FUA, Rahman AS. Synthesizing lunar regolith-geopolymer emulating lunar positive temperature regime. Planet Space Sci. 2024;244:105890.[DOI]
-
52. Geng Z, Zhang L, Wu Z, Tay YWD, Lim SG, Tan MJ. Mechanical properties and additive manufacturing of alkali-activated lunar regolith in artificial lunar environments. In: Construction 3D printing. Cham: Springer; 2024. p. 50-56.[DOI]
-
53. Yao Y, Liu C, Liu H, Zhang W. Phase assemblage characteristics and solidification reaction mechanism of alkali-activated lunar regolith simulant in thermal vacuum environment. China Civ Eng J. 2025;58(5):105-118. Chinese. Available from: https://manu36.magtech.com.cn/Jwk_tmgcxb/CN/abstract/abstract1978.shtml?utm_source
-
54. Casanova S, Espejel C, Dempster AG, Anderson RC, Caprarelli G, Saydam S. Lunar polar water resource exploration–Examination of the lunar cold trap reservoir system model and introduction of play-based exploration (PBE) techniques. Planet Space Sci. 2020;180:104742.[DOI]
-
55. Schlüter L, Cowley A. Review of techniques for In-Situ oxygen extraction on the Moon. Planet Space Sci. 2020;181:104753.[DOI]
-
56. Abdellatief M, Hassan YM, Elnabwy MT, Wong LS, Chin RJ, Mo KH. Investigation of machine learning models in predicting compressive strength for ultra-high-performance geopolymer concrete: A comparative study. Constr Build Mater. 2024;436:136884.[DOI]
-
57. Chen W, Liang L, Jiang F, Tang Z, Sun X, Yu J, et al. New opportunity: Materials genome strategy for engineered cementitious composites (ECC) design. Cem Concr Compos. 2025;159:106009.[DOI]
-
58. Sun Z, Li Y, Yang Y, Su L, Xie S. Splitting tensile strength of basalt fiber reinforced coral aggregate concrete: Optimized XGBoost models and experimental validation. Constr Build Mater. 2024;416:135133.[DOI]
-
59. Yao Y, Liu C, Liu H, Zhang W, Hu T. Deterioration mechanism understanding of recycled powder concrete under coupled sulfate attack and freeze–thaw cycles. Constr Build Mater. 2023;388:131718.[DOI]
-
60. Babiker A, Abbas YM, Khan MI, Abdel-Magid T. Optimizing compressive strength of quaternary-blended cement concrete through ensemble-instance-based machine learning. Mater Today Commun. 2024;39:109150.[DOI]
-
61. Zhang L, Zhu D, Nehdi ML, Marani A, Wang D, Zheng D, et al. Machine learning prediction of drying shrinkage for alkali-activated materials and multi-objective optimization. Mater Today Commun. 2025;45:112326.[DOI]
-
62. Hariri-Ardebili MA, Mahdavi P, Pourkamali-Anaraki F. Benchmarking AutoML solutions for concrete strength prediction: Reliability, uncertainty, and dilemma. Constr Build Mater. 2024;423:135782.[DOI]
-
63. Tran VQ, Dang VQ, Ho LS. Evaluating compressive strength of concrete made with recycled concrete aggregates using machine learning approach. Constr Build Mater. 2022;323:126578.[DOI]
-
64. Asif U, Ali Memon S. Interpretable predictive modeling, sustainability assessment, and cost analysis of cement-based composite containing secondary raw materials. Constr Build Mater. 2025;473:140924.[DOI]
-
65. Chen H, Deng T, Du T, Chen B, Skibniewski MJ, Zhang L. An RF and LSSVM–NSGA-II method for the multi-objective optimization of high-performance concrete durability. Cem Concr Compos. 2022;129:104446.[DOI]
-
66. Gandomi AH, Faramarzifar A, Rezaee PG, Asghari A, Talatahari S. New design equations for elastic modulus of concrete using multi expression programming. J Civ Eng Manag. 2015;21(6):761-774.[DOI]
Copyright
© The Author(s) 2026. This is an Open Access article licensed under a Creative Commons Attribution 4.0 International License (https://creativecommons.org/licenses/by/4.0/), which permits unrestricted use, sharing, adaptation, distribution and reproduction in any medium or format, for any purpose, even commercially, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons license, and indicate if changes were made.
Publisher’s Note
Share And Cite



