Skip to main navigation Skip to search Skip to main content

Large Language Models for Automated Literature Review: An Evaluation of Reference Generation, Abstract Writing, and Review Composition

  • Xuemei Tang
  • , Xufeng Duan (Corresponding Author)
  • , Zhenguang G. Cai (Corresponding Author)

Research output: Chapter in book / Conference proceedingConference article published in proceeding or bookAcademic researchpeer-review

Abstract

Large language models (LLMs) have emerged as a potential solution to automate the complex processes involved in writing literature reviews, such as literature collection, organization, and summarization. However, it is yet unclear how good LLMs are at automating comprehensive and reliable literature reviews. This study introduces a framework to automatically evaluate the performance of LLMs in three key tasks of literature review writing: reference generation, abstract writing, and literature review composition. We introduce multidimensional evaluation metrics that assess the hallucination rates in generated references and measure the semantic coverage and factual consistency of the literature summaries and compositions against human-written counterparts. The experimental results reveal that even the most advanced models still generate hallucinated references, despite recent progress. Moreover, we observe that the performance of different models varies across disciplines when it comes to writing literature reviews. These findings highlight the need for further research and development to improve the reliability of LLMs in automating academic literature reviews. The dataset and code used in this study are publicly available in our GitHub repository.
Original languageEnglish
Title of host publicationProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
EditorsChristos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, Violet Peng
PublisherAssociation for Computational Linguistics (ACL)
Pages1602-1617
ISBN (Electronic)9798891763326
DOIs
Publication statusPublished - 4 Nov 2025
EventConference on Empirical Methods in Natural Language Processing (EMNLP 2025) - Suzhou Expo, Suzhou, China
Duration: 5 Nov 20259 Nov 2025
https://2025.emnlp.org/

Conference

ConferenceConference on Empirical Methods in Natural Language Processing (EMNLP 2025)
Abbreviated titleEMNLP 2025
Country/TerritoryChina
CitySuzhou
Period5/11/259/11/25
Internet address

Fingerprint

Dive into the research topics of 'Large Language Models for Automated Literature Review: An Evaluation of Reference Generation, Abstract Writing, and Review Composition'. Together they form a unique fingerprint.

Cite this