The reliance on non-renewable fossil fuels has given rise to environmental issues, necessitating an urgent energy transition to renewable energy sources [1–3]. Among various sustainable energy storage technologies, hydrogen has a significant role due to its carbon-free nature, abundance, and sustainability [4–6]. Though hydrogen has received significant attention, its commercial viability remains challenging, due to issues from environmental issue, low round trip energy efficiency to cost-effective storage [1, 7–9]. Briefly, blue hydrogen production, i.e., conventional methods, relies on the use of significant fossil fuels as the chemical feedstock, while green hydrogen production using renewables faces challenges with the conversion efficiency.
Conventional methodologies, such as the prototypical steam methane reforming (SMR), predominantly employ natural gas as the chemical feedstock. These methods, however, result in substantial carbon emissions [10]. In contrast, hydrogen production techniques with renewable resources, notably water electrolysis and biohydrogen production methods, are considered significantly more environmentally benign [11, 12]. The sustainability aspects of these three hydrogen production methods are illustrated in Fig. 1[13–19]. The current usage proportions, material sources, efficiency, advantages and drawbacks of these approaches are also compared (Table 1). It is evident that low efficiency and high costs are major issues in large-scale applications of green hydrogen production methods, i.e., water electrolysis and biohydrogen production. Within the hydrogen production process, the role of catalysts is crucial for enhancing the speed and efficiency of the reaction, representing a pivotal step in hydrogen production. The choice of catalyst consequently influences the efficiency and outcome of the hydrogen production process. The current catalyst is insufficient to meet the demands of efficient hydrogen production, necessitating further advancements.
| Approach | Usage Proportion | Raw materials | Efficiency | Advantage | Approach limitation | Catalyst limitation | Reference |
|---|---|---|---|---|---|---|---|
| Methane reforming | 95% | Fossil fuels | ±75% | Commercialized technology, high efficiency | Harmful to the environment, large emission of greenhouse gases | Coking, sintering, high cost | [20–23] |
| Water electrolysis | <5% | Water, electricity | ±60% | Sustainability, no greenhouse gases emission with water resources | High cost of the catalyst and production process, lower production efficacy than methane reforming | High cost, low efficiency | [24–26] |
| Biohydrogen production | <5% | Biomass | ±45% | Abundant resources, cost-effective | Low efficiency | Low efficiency, high cost | [27–30] |
Despite the effectiveness of human experimental work, the traditional approach to developing catalysts predominantly involved, are often time-intensive and costly. This has brought a growing reliance on computational techniques, such as density functional theory (DFT), to streamline and accelerate the process. However, the extensive computational demand of DFT has spurred interest in integrating machine learning (ML) methodologies with DFT to expedite material discovery [31–33]. The integration of DFT with ML has already demonstrated substantial achievements in various sectors, including computing, finance, and chemistry [32–34]. The fundamental principle of ML is its ability to generate predictions based on learning from historical data, often avoiding the need for explicitly defined relationships [35]. Conversely, DFT is esteemed for its precision in calculating the structural properties of catalysts [36]. Despite several articles reviewing the application of ML in the hydrogen production process, a discussion on its progress combined with DFT for catalyst designs is still limited. This review includes an evaluation and analysis of their specific applications in the hydrogen production process, illustrated through key examples. Our work also aims to highlight the benefits and impacts of using DFT and ML, intending to provide insightful guidance for future research.
DFT has emerged as a cornerstone computational technique in the field of catalysis, especially when experimental methods fail to provide knowledge in heterogeneous catalyst systems [38, 39]. DFT calculation offers valuable insights into the underlying electronic structure and reactivity of catalytic materials. This theory utilizes the electron density as the central variable, allowing for the prediction and analysis of various catalytic phenomena with remarkable accuracy and efficiency. DFT computation serves as a powerful tool for characterizing the working state of the catalysts, such as adsorption energies, reaction pathways, and surface reactivity, thereby explaining the nature of the underlying mechanisms governing catalytic processes [40]. By leveraging DFT calculations, the energetics and dynamics of catalytic reactions can be systematically explored, paving the way for designing and optimizing of novel catalyst materials tailored for specific applications.
When DFT is applied in heterogenous catalysis, catalytic descriptors are required to correlate the structure and properties of catalysts with their catalytic performance. These descriptors include a diverse array of parameters, including surface properties, electronic structure, and adsorption energies, which influence catalytic activity and selectivity. Descriptors were briefly introduced below, valid for the DFT studies of various catalytic processes and reactions [39].
(1) Binding energy refers to the energy required to break the interaction between two particles. The binding energy quantifies how strongly the reactant particle binds to the catalyst surface, which influences the kinetics and thermodynamics of the catalytic reaction [41].
(2) Electronic structure refers to the arrangement and behaviour of electrons and atoms within a solid crystal catalyst, which affects catalytic activity and selectivity [42].
(3) The d-band theory provides a framework for understanding the relationship between the electronic structure of transition metal catalysts and their catalytic activity. Accordingly, the position of the d-band centre influences the adsorption behaviour of reactant molecules on the catalyst surface [43].
(4) Bandgap refers to the energy disparity between the highest energy level of the valence band and the lowest energy level of the conduction band in a solid material [44]. It describes the general oxidative and reductive properties of the catalyst [39].
Combing the aforementioned features provides solutions to predicting the catalytic performance. However, there is a trade-off between accuracy and efficiency. Higher accuracy in DFT calculations often means more computational inputs whilst simplified models may increase efficiency but at the cost of accuracy, potentially missing critical catalysis details. Screening for these DFT results to deliver the best catalysts is also difficult.
ML algorithms are basically categorized into three large classes of learning problems [45]: supervised learning, unsupervised learning, and reinforcement learning. In supervised learning [46], the data contains labels, which indicates that both the input variables and the corresponding output variables are observed for each data point. In this paradigm, the algorithm learns a mapping or relationship between input-output pairs, which generalizes from the labelled dataset to make predictions or decisions about new, unseen data. Supervised learning is commonly used in various applications such as regression (predicting continuous numerical values) and classification (predicting class labels or categories).
Unsupervised learning algorithms [47] are trained only on input data without any labelled output. The aim is to explore and identify intrinsic patterns, hidden structures, or relationships within the dataset. Two primary tasks [48] in unsupervised learning are clustering (grouping similar data points together based on patterns) and dimension reduction (decreasing the number of features). Unsupervised learning is more flexible and adaptable to a wider range of datasets, since no labels are required. However, the deficiency of corresponding output is likely to significantly challenge the accuracy of the algorithm [49].
In reinforcement learning [50], an agent learns to make a sequence of decisions or actions while interacting with the environment and receiving feedback in the form of rewards or penalties. This learning is based on trial and error with the aim of maximizing the cumulative reward over time. Reinforcement learning finds applications in various domains, including robotics, game playing (as seen in AlphaGo and AlphaZero), autonomous vehicles, recommendation systems, and resource management, where systems need to learn to make sequential decisions in dynamic and complex environments to achieve specific goals. Other algorithms are also included in Table 2. In terms of the application of hydrogen production, while these algorithms have gained broader applicability, they are still in nascent stages and need further exploration. Moreover, the accuracy of algorithms is strongly influenced by the dataset and varies from different descriptors. Consequently, most studies involve multiple algorithms to search for suitable applications.
| Algorithms | Category | Tasks | Advantages | Limitations | Used in hydrogen production? | Reference | Possibility as future methods for hydrogen production |
|---|---|---|---|---|---|---|---|
| Linear regression | Supervised learning | Regression | Simplicity; interpretability; computational efficiency | Assumption of linearity; sensitive to outliers; homoscedasticity assumption | Yes | [51] | |
| Logistic regression | Supervised learning | Classification | Simple and interpretable; efficient for binary classification; probabilistic output | Limited to binary classification and linear relationship | No | [52] | No, the labels used in hydrogen production are far more complex than binary cases. |
| Decision trees | Supervised learning | Regression classification | Handles both numerical and categorical data; handles non-linearity and interactions; no assumptions about data distribution | Prone to overfitting; biased toward dominant classes; Unstable | Yes | [53] | |
| Random forest | Supervised learning | Regression classification | Handles missing values; reduced overfitting; high accuracy especially on complex datasets with diverse patterns | Computationally expensive; lack of interpretability; not well-suited for very high-dimensional data | Yes | [54] | |
| Support vector machines | Supervised learning | Regression classification | Effective in high-dimensional spaces; handles different types of data and captures complex relationships | Sensitivity to noise; computational complexity; difficulty in interpreting results; need for proper parameter tuning | Yes | [55] | |
| Naive bayes | Supervised learning | Classification | Simple and efficient; effective for text classification; robust to irrelevant features | Assumption of feature independence; not suitable for complex relationships; poor estimation of probabilities | No | [56] | Yes, potentially used for the classification of production features. |
| XGBoost | Supervised learning | Regression classification | Sequentially combines multiple weak predictive models to create a more robust and accurate model; supports parallel and distributed computing | Requires careful hyperparameter tuning; requires sufficient training datasets | Yes | [57] | |
| K-nearest neighbours | Supervised learning | Regression classification | No training time; handles multiclass problems; non-parametric | High memory usage for large datasets; sensitive to irrelevant features; fail to capture global patterns | Yes | [58] | |
| K-means | Unsupervised learning | Clustering | Simple and efficient; adaptability for various cluster evaluation methods; cluster centroid interpretability | Sensitivity to initial centroids; assumption of equal variance; sensitive to outliers | No | [59] | Yes, potentially used for discovery of new catalysts. |
| Artificial neural network | Supervised learning | Regression classification | Adaptability to complex patterns; benefits from parallel processing; robustness to noisy data and generalization to unseen examples; no need for manual feature engineering in some cases | Computationally expensive; need for skilled tuning; black-box nature; requires large amounts of labeled data for effective training | Yes | [60] | |
| Reinforcement learning (RL) | Reinforcement learning | Decision-making and control | Continuous learning and adaptation to changes in the environment; handles sequential decision-making | Computationally expensive; exploration-exploitation trade-off; unstable during training and sensitive to hyperparameters | No | [50] | No, RL designs a control framework aimed to address sequential decision-making tasks. |
ML can recognize trends, predict potential outcomes, classify objects, and discover unusual patterns by creating computational algorithms to learn the data by itself. These functions represent a broad range of ML capabilities, which derive insights, automate tasks, and make data-driven decisions [45]. Since DFT calculations can be used to generate large datasets of heterogeneous catalysis, such as the aforementioned catalytic descriptors, ML can efficiently evaluate catalysts and screen out the best based on sufficient training data.
The general application workflow [61] of ML in catalysis is illustrated in Fig. 2, typically invoving 10 stages as shown in the Figure. When we simplify these steps, it will be divided into the following four steps. (1) Database construction. The raw data are collected from experiments, DFT computations, or existing databases, in the form of descriptors (features). Features include properties of catalysts and experiment conditions. (2) Feature selection and transformation. Since data quality and quantity are crucial for the success of machine learning models, feature engineering is used to eliminate irrelevant and redundant descriptors that cannot contribute much to the predictive power of the model. Preprocessing modules, such as handling missing values and normalization, would then modify and transform the database to make it more suitable for analysis. (3) Model selection. Choosing an appropriate ML algorithm is based on the nature of the problem and the characteristics of the database, as well as adjusting the optimal hyperparameters. (4) Deployment and development. Given a promising performance, the selected corresponding model can be deployed into production. It will be necessary to monitor its performance over time and optimize the model as needed with new data.

DFT can finely model the structure of atoms, molecules, crystals, surfaces, and their interactions, characterized by a wide range of features. By training ML models on DFT-calculated features for various materials, researchers can capture the connection between features and performance, selectivity and stability, thereby efficiently identifying promising candidates for energy production catalysts. In contrast, traditional experimental synthesis and testing would be more expensive and time-consuming. Incidentally, other non-DFT features, including reaction conditions, can also be optimized through ML algorithms to maximize production efficiency. The integration of ML techniques with DFT computations offers a powerful framework for accelerating novel catalyst formulation discovery and advancing our understanding of production processes, ultimately contributing to the development of efficient and sustainable energy solutions.
Hydrogen production is the process of producing hydrogen by various chemical or physical methods, with the facilitation of catalysts, from water, hydrocarbons or renewable resources [62] (Fig. 3). Since the application of DFT to catalyst research, a substantial amount of experience and achievements have been obtained in DFT-based catalyst development and performance prediction. Recently, researchers have identified the formidable data processing capability of ML, which was regarded as a potential auxiliary tool to the DFT computation process. In addition to the more efficient DFT-assisted computations, these ML-based statistical data also retain certain theoretical support, thus possessing higher credibility and value.
Currently, thermal processes are the predominant approach for producing hydrogen, mainly including reforming and gasifying [63]. The categorization of these processes varies depending on the feedstock and reaction conditions. In this process, catalysts play a pivotal role in facilitating chemical reactions at high temperatures for efficient hydrogen production. They allow the reaction to occur at lower temperatures and with higher selectivity, thereby reducing energy consumption and minimizing the formation of undesired by-products. In addition, catalysts can also reduce carbon deposition and catalyst deactivation, thus helping to extend the life of the reactor system [64, 65].
Ethanol reforming to carbon monoxide and hydrogen is an attractive method for hydrogen production. For the conversion process, selectivity and activity emerge as two pivotal evaluation metrics, which are significantly influenced by catalyst performance. Artrith et al. [66]. developed two ML models in tandem with DFT to predict transition state energies and catalytic properties, respectively. Key features, including thermochemical reaction energies, bond breaking activation energies were selected from 119 calculated and published DFT data points to construct the ML models. The model shows a mean absolute error (MAE) of 0.2 eV. 248 DFT reaction energies for the hypothetical Pt-based core–shell architectures of the remaining 3d transition metals were predicted, as shown in Fig. 4(a). While the majority of compositions exhibited poorer properties than those already experimentally characterized, Cr-Pt-Pt (111), Mn-Pt-Pt (111), Co-Pt-Pt (111), and Zn-Pt-Pt (111) remains as four promising candidates.

Steam methane reforming is a well-established, highly efficient and most common method for large-scale hydrogen production. Current research on SMR catalysts mainly focuses on nickel-based or other noble and non-noble metals to address the issues of catalyst cost, activity, and long-term stability. Bimetallic catalysts with synergistic interactions show significant promise. However, the diversity of metal species and combinations poses a considerable burden on conventional experimental methods. Notably, Liu et al. developed a microkinetic ML methodology to screen bimetallic catalysts. The dataset used for ML training includes 37 transition and non-transition metals with various stoichiometric ratios. Boosting-based XGB model exhibited the best prediction accuracy and stability, with 48 promising candidates identified from a database comprising over 5000+ catalysts (Fig. 4(b)) [67]. The DFT-based microkinetic models attribute to the search of descriptors and optimal activity range and the R2 score of the ML model is above 97%. Similarly, Yu et al. performed DFT calculations and microkinetic modeling for dry methane reforming to elucidate the catalytic activity trends of eight transition metals (TMs) using information energies of adsorbed C and O as two descriptors. Unsupervised ML techniques were employed for the screening of catalytic activity combined with their stability and cost to identify 23 potential binary intermetallic compounds as potential DMR catalysts from the alloys 1482 A3B1 and 741 A1B1 (Fig. 4(c)) [68].
ML has demonstrated its potential in DFT-based catalyst research. The ML models usually exhibit excellent prediction accuracy with sufficient and accurate training datasets employed. However, factors that influence the catalyst performance can be diverse, which still requires further exploration of other more precise and easily accessible features and descriptors. In addition, the synthesisability of the catalyst combination also needs to be considered.
Despite characterized with high energy consumption, the electrolysis of water remains highly promising due to its inexhaustible raw materials and environmentally friendly products [69]. Electrocatalysts facilitate the reduction of water molecules near the cathode by decreasing the energy barrier of the HER process [64]. Precious metal catalysts, represented by platinum, offer the best performance, but are challenged by their scarcity and high cost. Consequently, current research strategies focus on improving metal utilization and other alternative non-precious metal catalysts [70].
An atomic catalyst (AC) represents a catalyst that facilitates chemical reactions through direct interaction with individual atoms anchored onto support, thereby providing enhanced control and efficiency at the atomic level [71]. Sun et al. performed DFT calculations on all potential transition metals (TM) and lanthanides covered by graphdyine-based AC in the hope of screening potential catalysts. A bag-tree algorithm based ML model was developed for the validation of performance prediction, with features including adsorption energy, adsorption trend and electronic structure employed. The high agreement with DFT calculations confirmed the feasibility of ML for electrocatalyst research (Fig. 5(a)). Umer et al. employed DFT and ML techniques to screen more than 364 catalysts in the TM-SAC design consisting of 3d/4d/5d TM monotoms and different substrates [72]. Various types of electronic, geometric, and thermodynamic were used as descriptors to construct the model. The ML model exhibited exceptional accuracy in predicting both catalyst stability and HER activity. Twenty promising catalysts were successfully identified, some better than that of commercial Pt based catalysts (Fig. 5(b)). This ML-assisted design provides great help for the theoretical and experimental research of catalysts, while further experiment validation still proves necessary.

Transition metal carbides/nitrides (MXenes) and borides (MBenes) represent a novel class of 2D materials with accordion-like structures. Their unique physicochemical properties exhibit tremendous potential in HER. Sun et al. employed DFT calculations on 110 bare MBenes and 70 randomly selected single-atom doped MBenes. The established database was used to train support vector regression (SVR)-based ML model, which successfully screened 28 potential MBenes and MXenes systems from 271 candidates. The ML model demonstrated remarkable prediction accuracy. The stability of the catalysts was further assessed using cohesive and substitution energies, ultimately facilitating the identification of the top five ideal catalysts, which show near-zero Gibbs free energies of hydrogen adsorption (Fig. 5(c)) [74]. Photocatalytic water decomposition (PWS) is another HER method for hydrogen production using solar energy. Among the studied photocatalysts, ABO3-type chalcogenide oxide semiconductor materials exhibit excellent photocatalytic performance. Tao et al. developed ML models aimed to predict the bandgap energy and hydrogen production rate of these catalysts [75]. After pre-processing nearly 30,000 combinations of 10 selected A-site cations and 24 selected B-site cations, the developed ANN model with optimal performance successfully screened out 14 candidate chalcogenide photocatalysts with excellent theoretical performance (Fig. 5(d)).
The traditional new catalyst development processes typically involve massive experimental trials during the discovery stage, as well as complex chemical production simulations. In contrast, ML-based R & D procedures are simpler and faster as they can directly establish the correlation between DFT results, or even catalytic performance and their influential features (Fig. 6(a)). It is worth mentioning that instead of using DFT calculations as the training dataset for the ML models, other research has been carried out on the direct construction of ML models to evaluate the feedstock conversion or hydrogen production efficiency of catalytic processes, thereby facilitating the prediction and optimization of catalysts. C. Kim and J. Kim, for instance, developed an ANN model to predict the carbon monoxide (CO) conversion of Pt/Cex Zr1–x O2 catalysts for water-gas shift reaction (WGSR) under different composition and operating conditions, which can directly reflect the performance of the catalysts. Two hundred fifty-one catalytic reaction data and 13 descriptors were screened from published databases for the construction of ML models. Figure 6(b) illustrates the top 10 catalysts that exhibit high CO conversion among 110 catalysts at different operating temperatures, which can help in the development and design of catalysts [76]. Biomass is considered as a promising green resource for hydrogen production due to its abundant hydrogen content. Li et al. established machine learning models for predicting hydrogen and carbon dioxide yields based on feedstock composition, process conditions, and catalyst types in the supercritical water gasification (SCWG) process [77]. ML results suggest that iron compounds may be effective catalysts for maximizing H2 and minimizing CO2 production in syngas under optimal experimental conditions.

This statistical approach offers simplicity in calculation and enhanced intuitiveness in predicting catalyst performance, particularly convenient when factors beyond the catalyst itself that affect hydrogen production performance need to be considered. It should be noted, however, that while this strategy can provide valuable insights for catalyst design to some extent, it lacks the systematic support of theoretical calculations and provides limited assistance in comprehending and advancing the underlying mechanisms of catalysts.
In addition to the catalyst design and performance prediction, ML can also be applied during catalyst synthesis to improve the performance of the obtained catalyst by optimizing relevant process parameters. Pan et al. [80] investigated ZnO nanocatalysts coated with optical plasmonic resonance nanoparticles for photoelectrochemical water decomposition. An ANN model was developed to investigate two crucial experimental parameters during the synthesis process: the time of ultraviolet (UV) treatments and the hydroquinone content, which were used for the coating and secondary growth of gold nanoparticles, respectively (Fig. 7(a)). The optimized parameters were validated experimentally, which proved instrumental in designing materials with the desired surface plasmon resonance. Yan et al. [79] employed ML model to investigate the interconnection between synthesis parameters, catalyst material properties, and hydrogen production conditions with the hydrogen production rate of element-doped graphitic carbon nitride (D-g-C3N4) for photocatalysts. Regarding synthesis parameters, Fig. 7(b) reflects the shapley additive explanation (SHAP) correlations between the processing method and temperature on the H2 yield. The results show that different synthesis methods impact the hydrogen production, which was further confirmed by the test result of temperature as the running temperature required by different methods can be more significantly different. It is worth mentioning that not all screening synthesis parameters can be effectively modeled for prediction, which is attributed to a variety of factors, data instability is one of them.

Although the exploration of the synthetic process parameters has been relatively limited, mainly due to the scarcity of accurate data available and the lack of theoretical support, the method still proved valuable assistance for future catalyst synthesis and mechanism research.
In addition to the effects on catalyst performance, the process condition during hydrogen production can also directly influence the results. The imperative for controllable and optimized process parameters proves highly beneficial for enhancing production efficiency and reducing energy consumption. For industrial hydrogen production, hydrogen yield and energy conversion efficiency stand out as the two most paramount considerations (Fig. 8).

Existing numerical simulation methods, particularly kinetic and computational fluid dynamics (CFD) modeling approaches are plagued by the complexity of the reactions and high computational costs, whereas thermodynamic models, despite their relative convenience, rely on assumptions that deviate from the real-word scenarios [81–83]. Despite sometimes being criticized as “black boxes”, ML-based models have addressed such challenges. Hosseinzadeh et al. [84]. established several ML models with different algorithms for modeling and analyzing the dark fermentation hydrogen production process from wastewater. Key parameters, including Fe, Ni, biomass ratio, pH, and hydraulic retention time (HRT) were considered as input variables to explore the correlation with H2 production. Acetate, butyrate, ethanol, iron and nickel show high importance in descending order by the permutation variable importance (PVI) procedure. Hong et al. [85] developed a hybrid DNN model coupled with a multi-objective particle swarm optimizartion (PSO) algorithm to achieve multi-objective optimizationoptimization of thermal efficiency and CO2 emission for the SMR process. The simulation revealed trade-offs between thermal efficiency and CO2 emissions. The increase in thermal efficiency with higher CO2 emissions validates the thermal efficiency advantage of direct natural gas use over hydrogen conversion, which provides insights for decision-making tailored to the specific requirements of practical circumstances.
As global challenges such as energy shortages and climate change become increasingly severe, it is urgent to accelerate the energy transition. Hydrogen energy stands out as a particularly promising candidate in this regard. This review comprehensively provides insights into the latest advancements in the application of ML to hydrogen production. In terms of catalysts, ML has proved to effectively address the issue of complexity and the time-consuming nature of DFT computational processes. ML models trained on existing DFT datasets demonstrate remarkable predictive accuracy, which greatly contributes to the development of high-performance catalysts and the research of catalytic mechanisms. ML can also be directly applied to predict catalytic performance, which shows certain advantages. In addition, ML can be used to explore and optimize catalyst synthesis conditions. The parameters of the hydrogen production process directly influence on both hydrogen production rate and energy consumption, which is also a significant focus of the recent application of ML.
The advantages conferred by ML in this context encompass but are not limited to (1) faster computation, (2) description of unknown processes, (3) examination of the interaction of factors (sensitivity), and (4) optimization of the hydrogen production process through the feedback of input parameters. Especially when integrated with DFT calculations, ML can offer fast and accurate prediction of catalyst performance and design guidance.
Despite its practical successes, the exploration of ML is still in its early stages. Numerous untapped opportunities still exist, offering substantial potential for the optimization and expansion of the ML application. For instance, the utilization of inaccurate and incomplete datasets for model generation, coupled with the time-consuming nature of feature extraction, can impact the quality of modeling and optimization processes; trained models have limited interpretability, which can be attributed to the ML modeling methodology and the intricacies inherent in the hydrogen generation process. Moreover, the optimization capability of the constructed model was not comprehensive, resulting in distinct differences in the prediction accuracy of different parameters in the same process. Therefore, further research is needed to address these limitations and a summary of some suggestions: (1) develop more powerful ML with enhanced feature recognition capabilities, such as deep learning, which automatically learns to extract features without the need for complex feature engineering; (2) advance the development of explainable AI or tools to validate trained models, facilitating the exploration of parameter correlations; (3) combination with other computing technologies can not only improve the accuracy and optimization capabilities of ML models, but also enable a more comprehensive analysis of complex hydrogen production processes; (4) reinforcement learning intelligently controls the production process, optimising it in real time in response to changes in actual reaction conditions.
Efficiently converting renewable energy into hydrogen energy can solve the instability problem in current renewable energy research and is crucial to the energy transition. At present, the application of ML machine learning in hydrogen production is still in its infancy, and more professional interdisciplinary knowledge reserves are needed to better realize the application of machine learning in hydrogen production. It should be noted that a significant portion of computational models have not been validated through empirical experiments, which may stem from the inherent differentiation between material science and computation research. Consequently, the development of integrated disciplines and talents is very necessary.
J. O. Abe, A. Popoola, E. Ajenifuja, and O. M. Popoola, "Hydrogen energy, economy and storage: Review and recommendation," International Journal of Hydrogen Energy, vol. 44, no. 29, pp. 15072–15086, 2019.
M. Yue, H. Lambert, E. Pahon, R. Roche, S. Jemei, and D. Hissel, "Hydrogen energy systems: A critical review of technologies, applications, trends and challenges," Renewable and Sustainable Energy Reviews, vol. 146, p. 111180, 2021.
J. Lai, R. Tan, H. Jiang, X. Huang, Z. Tian, B. Hong, M. Wang, J. Li, "Development of an in situ polymerized artificial layer for dendrite‐free and stable lithium metal batteries," Battery Energy, vol. 3, p.20230070, 2024.
Z. Feng, I. Eiubovi, Y. Shao, Z. Fan, R. Tan, "Review of digital twin technology applications in hydrogen energy," Chain, vol. 1, no. 1, pp. 54-74, 2024.
A. Kovač, M. Paranos, and D. Marciuš, "Hydrogen in energy transition: A review," International Journal of Hydrogen Energy, vol. 46, no. 16, pp. 10016–10035, 2021.
H. Aditiya and M. Aziz, "Prospect of hydrogen energy in Asia-Pacific: A perspective review on techno-socio-economy nexus," International Journal of Hydrogen Energy, vol. 46, no. 71, pp. 35027–35056, 2021.
H. X. Li, D. J. Edwards, M. R. Hosseini, and G. P. Costin, "A review on renewable energy transition in Australia: An updated depiction," Journal of Cleaner Production, vol. 242, p. 118475, 2020.
D. J. Davidson, "Exnovating for a renewable energy transition," Nature Energy, vol. 4, no. 4, pp. 254–256, 2019.
M. M. V. Cantarero, "Of renewable energy, energy democracy, and sustainable development: A roadmap to accelerate the energy transition in developing countries," Energy Research & Social Science, vol. 70, p. 101716, 2020.
S. Mallapaty, "How China could be carbon neutral by mid-century," Nature, vol. 586, no. 7830, pp. 482–483, 2020.
Z. Feng, G. Gupta, and M. Mamlouk, "Robust poly (p‐phenylene oxide) anion exchange membranes reinforced with pore‐filling technique for water electrolysis," Journal of Applied Polymer Science, vol. 141, no. 19, p. e55340, 2024.
Z. Feng, G. Gupta, and M. Mamlouk, "A review of anion exchange membranes prepared via Friedel-Crafts reaction for fuel cell and water electrolysis," International Journal of Hydrogen Energy, 2023.
Y. Sun, J.-P. Zhang, G. Yang, and Z.-H. Li, "Analysis of trace elements in corncob by microwave Digestion-ICP-AES," GUGNGPUXUE YU GUANGPU FENXI, vol. 27, no. 7, pp. 1424–1427, 2007.
O. Al-Juboori, F. Sher, U. Khalid, M. B. K. Niazi, and G. Z. Chen, "Electrochemical production of sustainable hydrocarbon fuels from CO2 co-electrolysis in eutectic molten melts," ACS Sustainable Chemistry & Engineering, vol. 8, no. 34, pp. 12877–12890, 2020.
O. Al-Juboori, F. Sher, A. Hazafa, M. K. Khan, and G. Z. Chen, "The effect of variable operating parameters for hydrocarbon fuel formation from CO2 by molten salts electrolysis," Journal of CO2 Utilization, vol. 40, p. 101193, 2020.
Y. Liu, J. Min, X. Feng, Y. He, J. Liu, Y. Wang, J. He, H. Do, V. Sage, G. Yang, and Y. Sun, "A review of biohydrogen productions from lignocellulosic precursor via dark fermentation: Perspective on hydrolysate composition and electron-equivalent balance," Energies, vol. 13, no. 10, p. 2451, 2020.
Y. Sun, Y. Wang, G. Yang, and Z. Sun, "Optimization of biohydrogen production using acid pretreated corn stover hydrolysate followed by nickel nanoparticle addition," International Journal of Energy Research, vol. 44, no. 3, pp. 1843–1857, 2020.
Y. Sun, J. Zhang, G. Yang, and Z. Li, "Analysis of trace elements in corn by inductively coupled plasma-atomic emission spectrometry," Food Sci, vol. 28, no. 2, pp. 236–237, 2007.
N. K. Al-Shara, F. Sher, A. Yaqoob, and G. Z. Chen, "Electrochemical investigation of novel reference electrode Ni/Ni (OH)₂ in comparison with silver and platinum inert quasi-reference electrodes for electrolysis in eutectic molten hydroxide," International Journal of Hydrogen Energy, vol. 44, no. 50, pp. 27224–27236, 2019.
M. Wang, X. Tan, J. Motuzas, J. Li, and S. Liu, "Hydrogen production by methane steam reforming using metallic nickel hollow fiber membranes," Journal of Membrane Science, vol. 620, p. 118909, 2021.
E. Meloni, M. Martino, A. Ricca, and V. Palma, "Ultracompact methane steam reforming reactor based on microwaves susceptible structured catalysts for distributed hydrogen production," International Journal of Hydrogen Energy, vol. 46, no. 26, pp. 13729–13747, 2021.
X. Zhu, X. Liu, H. -Y. Lian, J. -L. Liu, and X. -S. Li, "Plasma catalytic steam methane reforming for distributed hydrogen production," Catalysis Today, vol. 337, pp. 69–75, 2019.
H. -C. Wu, Z. Rui, and J. Y. Lin, "Hydrogen production with carbon dioxide capture by dual-phase ceramic-carbonate membrane reactor via steam reforming of methane," Journal of Membrane Science, vol. 598, p. 117780, 2020.
T. Nguyen, Z. Abdin, T. Holm, and W. Mérida, "Grid-connected hydrogen production via large-scale water electrolysis," Energy Conversion and Management, vol. 200, p. 112108, 2019.
C. Zhang, J. Greenblatt, M. Wei, J. Eichman, S. Saxena, M. Muratori, and O. J. Guerra, "Flexible grid-based electrolysis hydrogen production for fuel cell vehicles reduces costs and greenhouse gas emissions," Applied Energy, vol. 278, p. 115651, 2020.
H. Ju, S. Giddey, and S. P. Badwal, "Role of iron species as mediator in a PEM based carbon-water co-electrolysis for cost-effective hydrogen production," International Journal of Hydrogen Energy, vol. 43, no. 19, pp. 9144–9152, 2018.
E. S. Aydin, O. Yucel, and H. Sadikoglu, "Experimental study on hydrogen-rich syngas production via gasification of pine cone particles and wood pellets in a fixed bed downdraft gasifier," International Journal of Hydrogen Energy, vol. 44, no. 32, pp. 17389–17396, 2019.
S. Fail, M. Binder, R. Rauch, H. Hofbauer, A. Molino, A. Blasiand, and D. Musmarra, "Experimental investigations of hydrogen production from CO catalytic conversion of tar rich syngas by biomass gasification," Catalysis Today, vol. 277, pp. 182–191, 2016.
R. Jahromi, M. Rezaei, S. H. Samadi, and H. Jahromi, "Biomass gasification in a downdraft fixed-bed gasifier: Optimization of operating conditions," Chemical Engineering Science, vol. 231, p. 116249, 2021.
W. -X. Peng, S. -B. Ge, A. G. Ebadi, H. Hisoriev, and M. J. Esfahani, "Syngas production by catalytic co-gasification of coal-biomass blends in a circulating fluidized bed gasifier," Journal of Cleaner Production, vol. 168, pp. 1513–1517, 2017.
K. Kaur and C. V. Singh, "Amorphous TiO2 as a photocatalyst for hydrogen production: A DFT study of structural and electronic properties," Energy Procedia, vol. 29, pp. 291–299, 2012.
J. W. Goodell, S. Kumar, W. M. Lim, and D. Pattnaik, "Artificial intelligence and machine learning in finance: Identifying foundations, themes, and research clusters from bibliometric analysis," Journal of Behavioral and Experimental Finance, vol. 32, p. 100577, 2021.
A. A. Khan, A. A. Laghari, and S. A. Awan, "Machine learning in computer vision: A review," EAI Endorsed Transactions on Scalable Information Systems, vol. 8, no. 32, pp. e4–e4, 2021.
P. Geerlings and F. De Proft, "Conceptual DFT: the chemical relevance of higher response functions," Physical Chemistry Chemical Physics, vol. 10, no. 21, pp. 3028–3042, 2008.
M. I. Jordan and T. M. Mitchell, "Machine learning: Trends, perspectives, and prospects," Science, vol. 349, no. 6245, pp. 255–260, 2015.
J. Zhang, X. Wang, Y. Han, W. Li, J. Cheng, Z. Gan, and J. Gu, "The effect of supercritical water on coal pyrolysis and hydrogen production: A combined ReaxFF and DFT study," Fuel, vol. 108, pp. 682–690, 2013.
X. Wan, Z. Zhang, W. Yu, and Y. Guo, "A density-functional-theory-based and machine-learning-accelerated hybrid method for intricate system catalysis," Materials Reports: Energy, vol. 1, no. 3, p. 100046, 2021.
R. Schlögl, "Heterogeneous catalysis," Angewandte Chemie International Edition, vol. 54, no. 11, pp. 3465–3520, 2015.
L. I. Ugwu, Y. Morgan, and H. Ibrahim, "Application of density functional theory and machine learning in heterogenous-based catalytic reactions for hydrogen production," International Journal of Hydrogen Energy, vol. 47, no. 4, pp. 2245–2267, 2022.
J. K. Nørskov, F. Abild-Pedersen, F. Studt, and T. Bligaard, "Density functional theory in surface chemistry and catalysis," Proceedings of the National Academy of Sciences, vol. 108, no. 3, pp. 937–943, 2011.
H. Zhuang, A. J. Tkalych, and E. A. Carter, "Surface energy as a descriptor of catalytic activity," The Journal of Physical Chemistry C, vol. 120, no. 41, pp. 23698–23706, 2016.
H. Tao, S. Liu, J.-L. Luo, P. Choi, Q. Liu, and Z. Xu, "Descriptor of catalytic activity of metal sulfides for oxygen reduction reaction: A potential indicator for mineral flotation," Journal of Materials Chemistry A, vol. 6, no. 20, pp. 9650–9656, 2018.
C. F. Dickens, J. H. Montoya, A. R. Kulkarni, M. Bajdich, and J. K. Nørskov, "An electronic structure descriptor for oxygen reactivity at metal and metal-oxide surfaces," Surface Science, vol. 681, pp. 122–129, 2019.
A. B. Getsoian, Z. Zhai, and A. T. Bell, "Band-gap energy as a descriptor of catalytic activity for propene oxidation over mixed metal oxide catalysts," Journal of the American Chemical Society, vol. 136, no. 39, pp. 13684–13697, 2014.
G. Carleo, I. Cirac, K. Cranmer, L. Daudet, M. Schuld, N. Tishby, L. Vogt-Maranto, and L. Zdeborová, "Machine learning and the physical sciences," Reviews of Modern Physics, vol. 91, no. 4, p. 045002, 2019.
Z. -H. Zhou, "A brief introduction to weakly supervised learning," National science review, vol. 5, no. 1, pp. 44–53, 2018.
A. Glielmo, B. E. Husic, A. Rodriguez, C. Clementi, F. Noé, and A. Laio, "Unsupervised learning methods for molecular simulation data," Chemical Reviews, vol. 121, no. 16, pp. 9722–9758, 2021.
J. A. Hueffel, T. Sperger, I. Funes-Ardoiz, J. S. Ward, K. Rissanen, and F. Schoenebeck, "Accelerated dinuclear palladium catalyst identification through unsupervised machine learning," Science, vol. 374, no. 6571, pp. 1134–1140, 2021.
T. T. Nguyen, N. D. Nguyen, and S. Nahavandi, "Deep reinforcement learning for multiagent systems: A review of challenges, solutions, and applications," IEEE transactions on cybernetics, vol. 50, no. 9, pp. 3826–3839, 2020.
D. C. Montgomery, E. A. Peck, and G. G. Vining, Introduction to Linear Regression Analysis, John Wiley & Sons, 2021.
B. Charbuty and A. Abdulazeez, "Classification based on decision tree algorithm for machine learning," Journal of Applied Science and Technology Trends, vol. 2, no. 1, pp. 20–28, 2021.
S. Suthaharan and S. Suthaharan, "Support vector machine," Machine Learning Models and Algorithms for Big Data ClasSification: Thinking with Examples for Effective Learning, New York; Springer, 2016, pp. 207–235.
G. I. Webb, E. Keogh, and R. Miikkulainen, "Naïve Bayes," Encyclopedia of machine learning, vol. 15, no. 1, pp. 713–714, 2010.
T. Chen and C. Guestrin, "Xgboost: A scalable tree boosting system,"in Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 2016, pp. 785–794.
M. L. Zhang and Z. H. Zhou, "ML-KNN: A lazy learning approach to multi-label learning," Pattern Recognition, vol. 40, no. 7, pp. 2038–2048, 2007.
K. P. Sinaga and M. S. Yang, "Unsupervised K-means clustering algorithm," IEEE Access, vol. 8, pp. 80716–80727, 2020.
Y. LeCun, Y. Bengio, and G. Hinton, "Deep learning," Nature, vol. 521, no. 7553, pp. 436–444, 2015.
T. Yu, C. Wang, H. Yang, and F. Li, "Machine learning in metal-ion battery research: Advancing material prediction, characterization, and status evaluation," Journal of Energy Chemistry, vol. 90, pp. 191–204, 2024.
M. Kayfeci, A. Keçebaş, and M. Bayat, "Hydrogen production,"in Solar Hydrogen Production, Elsevier, 2019, pp. 45–83, 2019.
F. Dawood, M. Anda, and G. Shafiullah, "Hydrogen production for energy: An overview," International Journal of Hydrogen Energy, vol. 45, no. 7, pp. 3847–3869, 2020.
S. Wang, A. Lu, and C.-J. Zhong, "Hydrogen production from water electrolysis: Role of catalysts," Nano Convergence, vol. 8, pp. 1–23, 2021.
A. E. Galetti, M. F. Gomez, L. A. Arrúa, and M. C. Abello, "Hydrogen production by ethanol reforming over NiZnAl catalysts: Influence of Ce addition on carbon deposition," Applied Catalysis A: General, vol. 348, no. 1, pp. 94–102, 2008.
N. Artrith, Z. Lin, and J. G. Chen, "Predicting the activity and selectivity of bimetallic metal catalysts for ethanol reforming using machine learning," ACS Catalysis, vol. 10, no. 16, pp. 9438–9444, 2020.
Z. Liu, W. Tian, Z. Cui, and B. Liu, "A universal microkinetic-machine learning bimetallic catalyst screening method for steam methane reforming," Separation and Purification Technology, vol. 311, p. 123270, 2023.
Y. Yu, J. Yang, K. Zhu, Z. Sui, D. Chen, Y. Zhu, and X. Zhou, "High-throughput screening of alloy catalysts for dry methane reforming," ACS Catalysis, vol. 11, no. 14, pp. 8881–8894, 2021.
Z. Feng, P. O. Esteban, G. Gupta, D. A. Fulton, and M. Mamlouk, "Highly conductive partially cross-linked poly (2, 6-dimethyl-1, 4-phenylene oxide) as anion exchange membrane and ionomer for water electrolysis," International Journal of Hydrogen Energy, vol. 46, no. 75, pp. 37137–37151, 2021.
D. Strmcnik, P. P. Lopes, B. Genorio, V. R. Stamenkovic, and N. M. Markovic, "Design principles for hydrogen evolution reaction catalyst materials," Nano Energy, vol. 29, pp. 29–36, 2016.
Z. Pu, I. S. Amiinu, R. Cheng, P. Wang, C. Zhang, S. Mu, W. Zhao, F. Su, G. Zhang, S. Liao, and S. Sun, "Single-atom catalysts for electrochemical hydrogen evolution reaction: Recent advances and future perspectives," Nano-Micro Letters, vol. 12, pp. 1–29, 2020.
M. Umer, S. Umer, M. Zafari, M. Ha, R. Anand, A. Hajibabaei, A. Abbas, G. Lee, and K. S. Kim, "Machine learning assisted high-throughput screening of transition metal single atom based superb hydrogen evolution electrocatalysts," Journal of Materials Chemistry A, vol. 10, no. 12, pp. 6679–6689, 2022.
Q. Yang, Y. Y. Cai, Z. Y. Zhu, L. X. Sun, Y. S. L. Choo, Q. G. Zhang, A. M. Zhu, and Q. L. Liu, "Multiple enhancement effects of crown ether in tröger's base polymers on the performance of anion exchange membranes," ACS Applied Materials & Interfaces, vol. 12, no. 22, pp. 24806–24816, 2020.
X. Sun, J. Zheng, Y. Gao, C. Qiu, Y. Yan, Z. Yao, S. Deng, and J. Wang, "Machine-learning-accelerated screening of hydrogen evolution catalysts in MBenes materials," Applied Surface Science, vol. 526, p. 146522, 2020.
Q. Tao, T. Lu, Y. Sheng, L. Li, W. Lu, and M. Li, "Machine learning aided design of perovskite oxide materials for photocatalytic water splitting," Journal of Energy Chemistry, vol. 60, pp. 351–359, 2021.
C. Kim and J. Kim, "Machine learning‐based high‐throughput screening, strategical design and knowledge extraction of Pt/CexZr1− xO2 catalysts for water gas shift reaction," International Journal of Energy Research, vol. 46, no. 15, pp. 21293–21308, 2022.
J. Li, L. Pan, M. Suvarna, and X. Wang, "Machine learning aided supercritical water gasification for H2-rich syngas production with process optimization and catalyst screening," Chemical Engineering Journal, vol. 426, p. 131285, 2021.
C. Kim, W. Won, and J. Kim, "Early-stage evaluation of catalyst using machine learning based modeling and simulation of catalytic systems: Hydrogen production via water–gas shift over Pt catalysts," ACS Sustainable Chemistry & Engineering, vol. 10, no. 44, pp. 14417–14432, 2022.
L. Yan, S. Zhong, T. Igou, H. Gao, J. Li, and Y. Chen, "Development of machine learning models to enhance element-doped g-C3N4 photocatalyst for hydrogen production through splitting water," International Journal of Hydrogen Energy, vol. 47, no. 80, pp. 34075–34089, 2022.
F. Pan, C. -C. Wu, Y. -L. Chen, P. -Y. Kung, and Y. -H. Su, "Machine learning ensures rapid and precise selection of gold sea-urchin-like nanoparticles for desired light-to-plasmon resonance," Nanoscale, vol. 14, no. 37, pp. 13532–13541, 2022.
Ö. Ç. Mutlu and T. Zeng, "Challenges and opportunities of modeling biomass gasification in Aspen Plus: A review," Chemical Engineering & Technology, vol. 43, no. 9, pp. 1674–1689, 2020.
F. Boshagh, K. Rostami, and E. W. van Niel, "Application of kinetic models in dark fermentative hydrogen production–A critical review," International Journal of Hydrogen Energy, vol. 47, no. 52, pp. 21952–21968, 2022.
L. Mingyi, Y. Bo, X. Jingming, and C. Jing, "Thermodynamic analysis of the efficiency of high-temperature steam electrolysis system for hydrogen production," Journal of Power Sources, vol. 177, no. 2, pp. 493–499, 2008.
A. Hosseinzadeh, J. L. Zhou, A. Altaee, and D. Li, "Machine learning modeling and analysis of biohydrogen production from wastewater by dark fermentation process," Bioresource Technology, vol. 343, p. 126111, 2022.