Advancements of large language models for enhancing carbon capture technologies: A comprehensive review AITranslate
Abstract AITranslate
This paper reviews the current research status, challenges, and prospects of applying large language models (LLMs) in carbon capture technologies. The review emphasizes the importance of interdisciplinary research, integrating AI into chemistry, engineering, and environmental science to address complex challenges in carbon capture. It provides a detailed analysis of how LLMs can be utilized across various stages of carbon capture, from experimental design to industry implementation, showcasing their potential to accelerate innovation. It also reveals the use of LLMs to support gathering and analyzing sustainable information, such as carbon tax, carbon footprint, and social analysis. LLMs not only show great potential in designing and discovering materials for carbon capture technologies but also are promising to accelerate the whole industry's development through their powerful data processing and pattern recognition capabilities. In addition, the review paper also discusses challenges in the application of LLMs for carbon capture technologies and future directions and prospects.
KeyWords AITranslate
[1]International Energy Agency, World Energy Outlook 2024, 2024, Accessed: 20 March 2025 [online]. Available: https://www.cleanenergyministerial.org/resource-cesc/world-energy-outlook-2024/.
[2]A. G. Olabi, T. Wilberforce, K. Elsaid, E. T. Sayed, H. M. Maghrabie, and M. A. Abdelkareem, "Large scale application of carbon capture to processing industries–a review," Journal of Cleaner Production, vol. 362, p. 132300, 2022.
[3]S. Talei, D. Fozer, P. S. Varbanov, A. Szanyi, A. Szanyi, and P. Mizsey, "Oxyfuel combustion makes carbon capture more efficient," ACS Omega, vol. 9, no. 3, pp. 3250–3261, 2024.
[4]R. Tiwari, "Building large language models from scratch: Initial guide, medium," Accessed: 5 March 2025[online]. Available: https://medium.com/@AI-Simplified/building-large-language-models-from-scratch-a-beginners-guide-d464e54f932b.
[5]P. Shaw, J. Uszkoreit, and A. Vaswani, "Self-attention with relative position representations," arXiv preprint arXiv:1803.02155, 2018.
[6]J. Jumper, R. Evans, A. Pritzel, T. Green, M. Figurnov, O. Ronneberger, K. Tunyasuvunakool, R. Bates, A. Žídek, A. Potapenko et al., "Highly accurate protein structure prediction with AlphaFold," Nature, vol. 596, no. 7873, pp. 583–589, 2021.
[7]O. Kononova, H. Huo, T. He, Z. Rong, T. Botari, W. Sun, V. Tshitoyan, and G. Ceder, "Text-mined dataset of inorganic materials synthesis recipes," Scientific Data, vol. 6, no. 1, p. 203, 2019.
[8]S. Mysore, Z. Jensen, E. Kim, K. Huang, H.-S. Chang, E. Strubell, J. Flanigan, A. McCallum, and E. Olivetti, "The materials science procedural text corpus: Annotating materials synthesis procedures with shallow semantic structures," arXiv preprint arXiv:1905.06939, 2019.
[9]B. J. Bender, S. Gahbauer, A. Luttens, J. Lyu, C. M. Webb, R. M. Stein, E. A. Fink, T. E. Balius, J. Carlsson et al., "A practical guide to large-scale docking," Nature Protocols, vol. 16, no. 10, pp. 4799–4832, 2021.
[10]C. Rudin, "Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead," Nature Machine Intelligence, vol. 1, no. 5, pp. 206–215, 2019.
[11]Y. Xue, C. Kambhampati, Y. Cheng, N. Mishra, N. Wulandhari, and P. Deutz, "A LDA-based social media data mining framework for plastic circular economy," International Journal of Computational Intelligence Systems, vol. 17, no. 1, p. 8, 2024.
[12]A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever, "Language models are unsupervised multitask learners," OpenAI Blog, vol. 1, no. 8, p. 9, 2019.
[13]S. Alaparthi and M. Mishra, "Bidirectional encoder representations from transformers (BERT): A sentiment analysis odyssey," arXiv preprint arXiv:2007.01127, 2020.
[14]J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, "BERT: Pre-training of deep bidirectional transformers for language understanding," Proceedings of the 2019 Conference of the North American chapter of the Association for Computational Linguistics: Human Language Technologies, vol. 1, pp. 4171–4186, 2019.
[15]V. Tshitoyan, J. Dagdelen, L. Weston, A. Dunn, Z. Rong, O. Kononova, K. A. Persson, G. Ceder, and A. Jain, "Unsupervised word embeddings capture latent knowledge from materials science literature," Nature, vol. 571, no. 7763, pp. 95–98, 2019.
[16]C. Friedman, P. Kra, H. Yu, M. Krauthammer, and A. Rzhetsky, "GENIES: A natural-language processing system for the extraction of molecular pathways from journal articles," Bioinformatics, vol. 17, issue suppl_1, pp. 74–82, 2001.
[17]C. Friedman, L. Shagina, S. A. Socratous, and X. Zeng, "A WEB-based version of MedLEE: A medical language extraction and encoding system," in AMIA Annual Fall Symposium Proceedings, Washington, DC, USA, 26–30 October 1996, p. 938.
[18]H.-M. Müller, E. E. Kenny, and P. W. Sternberg, "Textpresso: an ontology-based information retrieval and extraction system for biological literature," PLoS Biology, vol. 2, no. 11, p. e309, 2004.
[19]M. C. Swain and J. M. Cole, "ChemDataExtractor: A toolkit for automated extraction of chemical information from the scientific literature," Journal of Chemical Information and Modeling, vol. 56, no. 10, pp. 1894–1904, 2016.
[20]R. Leaman, C.-H. Wei, and Z. Lu, "TmChem: A high performance approach for chemical named entity recognition and normalization," Journal of Cheminformatics, vol. 7, pp. 1–10, 2015.
[21]K. M. Hettne, R. H. Stierum, M. J. Schuemie, P. J. M. Hendriksen, B. J. A. Schijvenaars, E. M. van Mulligen, J. Kleinjans, and J. A. Kors, "A dictionary to identify small molecules and drugs in free text," Bioinformatics, vol. 25, no. 22, pp. 2983–2991, 2009.
[22]T. He, W. Sun, H. Huo, O. Kononova, Z. Rong, V. Tshitoyan, T. Botari, and G. Ceder, "Similarity of precursors in solid-state synthesis as text-mined from scientific literature," Chemistry of Materials, vol. 32, no. 18, pp. 7861–7873, 2020.
[23]T. Rocktäschel, M. Weidlich, and U. Leser, "ChemSpot: A hybrid system for chemical named entity recognition," Bioinformatics, vol. 28, no. 12, pp. 1633–1640, 2012.
[24]X. Chen, M. Li, S. Gao, R. Yan, X. Gao, and X. Zhang, "Scientific paper extractive summarization enhanced by citation graphs," arXiv preprint arXiv:2212.04214, 2022.
[25]C. Edwards, T. Lai, K. Ros, G. Honke, K. Cho, and H. Ji, "Translation between molecules and natural language," arXiv preprint arXiv:2204.11817, 2022.
[26]R. Luo, L. Sun, Y. Xia, T. Qin, S. Zhang, H. Poon, and T.-Y. Liu, "BioGPT: Generative pre-trained transformer for biomedical text generation and mining," Briefings in Bioinformatics, vol. 23, no. 6, p. bbac409, 2022.
[27]R. Taylor, M. Kardas, G. Cucurull, T. Scialom, A. Hartshorn, E. Saravia, A. Poulton, V. Kerkez, and R. Stojnic, "Galactica: A large language model for science," arXiv preprint arXiv:2211.09085, 2022.
[28]Z. Zheng, O. Zhang, C. Borgs, J. T. Chayes, and O. M. Yaghi, "ChatGPT chemistry assistant for text mining and the prediction of MOF synthesis," Journal of the American Chemical Society, vol. 145, no. 32, pp. 18048–18062, 2023.
[29]H. Park, X. Yan, R. Zhu, E. A. Huerta, S. Chaudhuri, D. Cooper, I. Foster, and E. Tajkhorshid, "A generative artificial intelligence framework based on a molecular diffusion model for the design of metal-organic frameworks for carbon capture," Communications Chemistry, vol. 7, no. 1, art. no. 21, 2024.
[30]Z. Zheng, Z. Rong, N. Rampal, C. Borgs, J. T. Chayes, and O. M. Yaghi, "A GPT‐4 reticular chemist for guiding MOF discovery," Angewandte Chemie International Edition, vol. 62, no. 46, p. e202311983, 2023.
[31]H. C. Jami, P. R. Singh, A. Kumar, B. R. Bakshi, M. Ramteke, and H. Kodamana, "CCU-Llama: A knowledge extraction LLM for carbon capture and utilization by mining scientific literature data," Industrial & Engineering Chemistry Research, vol. 63, no. 41, pp. 17585–17598, 2024.
[32]D. A. Boiko, R. MacKnight, B. Kline, and G. Gomes, "Autonomous chemical research with large language models," Nature, vol. 624, no. 7992, pp. 570–578, 2023.
[33]A. P. Anyebe, O. K. K. Yeboah, O. I. Bakinson, T. Y. Adeyinka, and F. C. Okafor, "Optimizing carbon capture efficiency through AI-driven process automation for enhancing predictive maintenance and CO2 sequestration in oil and gas facilities," American Journal of Environment and Climate, vol. 3, no. 3, pp. 44–58, 2024.
[34]X. Liu, X.-Q. Zhang, X. Chen, G.-L. Zhu, C. Yan, J.-Q. Huang, and H.-J. Peng, "A generalizable, data-driven online approach to forecast capacity degradation trajectory of lithium batteries," Journal of Energy Chemistry, vol. 68, pp. 548–555, 2022.
[35]X. Gu, C. Chen, Y. Fang, R. Mahabir, and L. Fan, "CECA: An intelligent large-language-model-enabled method for accounting embodied carbon in buildings," Building and Environment, vol. 272, p. 112694, 2025.
[36]C. dos Santos Garcia, A. Meincheim, E. R. Faria Junior, M. R. Dallagassa, D. M. V. Sato, D. R. Carvalho, E. A. P. Santos, and E. E. Scalabrin, "Process mining techniques and applications–A systematic mapping study," Expert Systems with Applications, vol. 133, pp. 260–295, 2019.
[37]M. Wrzalik, F. Faust, S. Sieber, and A. Ulges, "NetZeroFacts: Two-stage emission information extraction from company reports," in Proceedings of the Joint Workshop of the 7th Financial Technology and Natural Language Processing, the 5th Knowledge Discovery from Unstructured Data in Financial Services, and the 4th Workshop on Economics and Natural Language Processing, Torino, Italia, May 2024, pp. 70–84.
[38]R. Lorenz, J. Senoner, W. Sihn, and T. Netland, "Using process mining to improve productivity in make-to-stock manufacturing," International Journal of Production Research, vol. 59, no. 16, pp. 4869–4880, 2021.
[39]T. Wu, J. Li, J. Bao, and Q. Liu, "ProcessCarbonAgent: A large language models-empowered autonomous agent for decision-making in manufacturing carbon emission management," Journal of Manufacturing Systems, vol. 76, pp. 429–442, 2024.
[40]H. Wang, M. Zhang, Z. Chen, N. Shang, S. Yao, F. Wen, and J. Zhao, "Carbon footprint accounting driven by large language models and retrieval-augmented generation," arXiv preprint arXiv:2408.09713, 2024.
[41]K. P. Saikia, D. Mukherjee, S. Mahapatra, P. Nandy, and R. Das, "Unveiling deeper petrochemical insights: Navigating contextual question answering with the power of semantic search and LLM fine-tuning," 2023 International Conference on Computing, Communication, and Intelligent Systems (ICCCIS), Greater Noida, India, 3–4 November 2023, pp. 881–886.
[42]H. Jiang, Y. Ding, R. Chen, and C. Fan, "Carbon price forecasting with LLM-based refinement and transfer-learning," International Conference on Artificial Neural Networks, Lugano-Viganello, Switzerland, 17–20 September 2024, pp. 139–154.
[43]R. Chen, H. Jiang, T. Guo, and C. Fan, "Can large language models forecast carbon price movements? Evidence from Chinese carbon markets," Research in International Business and Finance, vol. 77, Part B, p. 102951, 2025.
[44]T. Han, R.-G. Cong, B. Yu, B. Tang, and Y.-M. Wei, "Integrating local knowledge with ChatGPT-like large-scale language models for enhanced societal comprehension of carbon neutrality," Energy and AI, vol. 18, p. 100440, 2024.
[45]S. Shan, "From correlation to causation: Understanding climate change through causal analysis and LLM interpretations," arXiv preprint arXiv:2412.16691, 2024.
[46]C. Rudin, "Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead," arXiv preprint ArXiv, 1811.10154, 2018.
[47]J. Park and B. Yang, "GIS-enabled digital twin system for sustainable evaluation of carbon emissions: A case study of Jeonju city, south Korea," Sustainability, vol. 12, no. 21, p. 9186, 2020.
[48]J. An, W. Ding, and C. Lin, "ChatGPT: Tackle the growing carbon footprint of generative AI," Nature, vol. 615, p. 586, 2023.
[49]E. Strubell, A. Ganesh, and A. McCallum, "Energy and policy considerations for modern deep learning research," Proceedings of the AAAI Conference on Artificial Intelligence 2020, vol. 34, no. 9, pp. 13693–13696, 2020.
[50]D. Patterson, J. Gonzalez, Q. Le, C. Liang, L.-M. Munguia, D. Rothchild, D. So, M. Texier, and J. Dean, "Carbon emissions and large neural network training," arXiv preprint arXiv:2104.10350, 2021.
[51]J. Osondu, "Red AI vs. green AI in education: How educational institutions and students can lead environmentally sustainable artificial intelligence practices," preprint, DOI 10.13140/RG.2.2.27929.12644.
[52]P. Jiang, C. Sonne, W. Li, F. You, and S. You, "Preventing the immense increase in the life-cycle energy and carbon footprints of llm-powered intelligent chatbots," Engineering, vol. 40, pp. 202–210, 2024.
[53]J. Wen, R. Zhang, D. Niyato, J. Kang, H. Du, Y. Zhang, and Z. Han, "Generative AI for low-carbon artificial intelligence of things with large language models," IEEE Internet of Things Magazine, vol. 8, pp. 82–91, 2025.
[54]B. Li, Y. Jiang, V. Gadepally, and D. Tiwari, "Sprout: Green generative AI with carbon-efficient LLM inference," The 2024 Conference on Empirical Methods in Natural Language Processing, Miami, Florida, USA, 12–16 November 2024.
[55]M. Lawie, "Analysing the impact of CO2 emissions from the largest artificial intelligence systems and its consequences for global warming," Preprint, December 2023, DOI 10.13140/RG.2.2.24138.95680.
[56]P. Dechamps, "The IEA world energy outlook 2022–A brief analysis and implications," European Energy & Climate Journal, vol. 11, no. 3, pp. 100–103, 2023.
[57]D. Guo, Q. Zhu, D. Yang, Z. Xie, K. Dong, W. Zhang, G. Chen, X. Bi, Y. Wu et al., "DeepSeek-coder: When the large language model meets programming - The rise of code intelligence," ArXiv preprint ArXiv 2401.14196, 2024.
[58]T. Schick, J. Dwivedi-Yu, R. Dessì, R. Raileanu, M. Lomeli, L. Zettlemoyer, N. Cancedda, and T. Scialom, "Toolformer: Language models can teach themselves to use tools," Advances in Neural Information Processing Systems, vol. 36, pp. 68539–68551, 2023.
[59]O. Topsakal and T. C. Akinci, "Creating large language model applications utilizing langchain: A primer on developing LLM apps fast," International Conference on Applied Engineering and Natural Sciences 2023, Konya, Turkey, July 2023.
[60]F. M. Megahed, Y.-J. Chen, B. M. Colosimo, M. L. G. Grasso, L. A. Jones-Farmer, S. Knoth, H. Sun, and I. Zwetsloot, "Adapting OpenAI's CLIP model for few-shot image inspection in manufacturing quality control: An expository case study with multiple application examples," ArXiv preprint ArXiv 2501.12596, 2025.
[61]Kunal, M. Rana, and J. Bansal, "The future of OpenAI tools: Opportunities and challenges for human-AI collaboration," 2023 2nd International Conference on Futuristic Technologies (INCOFT), Belagavi, Karnataka, India, 24–26 November 2023, pp. 1–6.
[62]A. Ramesh, M. Pavlov, G. Goh, S. Gray, C. Voss, A. Radford, M. Chen, and I. Sutskever, "Zero-shot text-to-image generation," ArXiv preprint ArXiv, 2102.12092, 2021.
[63]A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever, "Learning transferable visual models from natural language supervision," International Conference on Machine Learning 2021, Virtual, 18–24 July 2021.
[64]S. J. Davis, N. S. Lewis, M. Shaner, S. Aggarwal, D. Arent, I. L. Azevedo, S. M. Benson, T. Bradley, J. Brouwer, and Y.-M. Chiang, "Net-zero emissions energy systems," Science, vol. 360, no. 6396, p. eaas9793, 2018.
[65]R. Kannan, E. Panos, S. Hirschberg, and T. Kober, "A net‐zero Swiss energy system by 2050: Technological and policy options for the transition of the transportation sector," Futures & Foresight Science, vol. 4, no. 3–4, p. e126, 2022.
[66]B. Wang, Z. Cai, M. M. Karim, C. Liu, Y. Wang, "Traffic performance GPT (TP-GPT): Real-time data informed intelligent ChatBot for transportation surveillance and management," 2024 IEEE 27th International Conference on Intelligent Transportation Systems (ITSC), Edmonton, Canada, 24–27 September 2024, pp. 460–467.
Basic Information:
DOI:10.23919/CHAIN.2025.000010
Chinese Library Classification Number:
Citation Information:
This paper reviews the current research status, challenges, and prospects of applying large language models (LLMs) in carbon capture technologies. The review emphasizes the importance of interdisciplinary research, integrating AI into chemistry, engineering, and environmental science to address complex challenges in carbon capture. It provides a detailed analysis of how LLMs can be utilized across various stages of carbon capture, from experimental design to industry implementation, showcasing their potential to accelerate innovation. It also reveals the use of LLMs to support gathering and analyzing sustainable information, such as carbon tax, carbon footprint, and social analysis. LLMs not only show great potential in designing and discovering materials for carbon capture technologies but also are promising to accelerate the whole industry's development through their powerful data processing and pattern recognition capabilities. In addition, the review paper also discusses challenges in the application of LLMs for carbon capture technologies and future directions and prospects.
quote
| GB/T 7714-2015 | [1] Yangyimin Xue, Manying Liu, Kuiyuan Wang, et al. Advancements of large language models for enhancing carbon capture technologies: A comprehensive review[J]. Chain, 2025, 2(2): 131-147. DOI:10.23919/CHAIN.2025.000010. |
| MLA | [1] Yangyimin Xue, et al., "Advancements of large language models for enhancing carbon capture technologies: A comprehensive review." Chain, vol. 2, no. 2, 2025, pp. 131-147, https://doi.org/10.23919/CHAIN.2025.000010. |
| APA | [1] Yangyimin Xue, Manying Liu, Kuiyuan Wang, Yuwan Yang, Yongqiang Cheng, Xinhui Ma, & Yuanting Qiao. (2025). Advancements of large language models for enhancing carbon capture technologies: A comprehensive review. Chain, 2(2), 131-147. https://doi.org/10.23919/CHAIN.2025.000010 |
| IEEE | [1] Yangyimin Xue, Manying Liu, Kuiyuan Wang, Yuwan Yang, Yongqiang Cheng, Xinhui Ma, and Yuanting Qiao, "Advancements of large language models for enhancing carbon capture technologies: A comprehensive review," Chain, vol. 2, no. 2, pp. 131-147, 2025, doi: 10.23919/CHAIN.2025.000010. keywords: {large language models;carbon capture;artificial intelligence;machine learning} |
