A Multi-Dimensional Feature Analysis for Detecting and Mitigating AI-Generated Content in Academic Journal Articles
|
An Tong, data analysis engineer, E-mail: antong@cnki.net |
|
Xue Dejun, deputy general manager and chief engineer of Tongfang CNKI; |
|
Zhou Zhe, R&D engineer |
|
Kong Xiangyu, project manager |
|
Geng Chong, deputy general manager of Central Technology Research Institute of Tongfang CNKI |
Received date: 2024-12-02
Online published: 2025-03-21
[Purpose/Significance] Investigating the characteristics of AI-generated academic journal content and establishing a multidimensional analytical framework are crucial for identifying improper use of AIGC in manuscripts, addressing challenges in peer review complexity, and upholding the quality and credibility of academic publishing. [Method/Process] This study analyzed journal articles with varying AIGC values across eight dimensions: citations, data, viewpoints and conclusions, structural framework, logical coherence, content consistency, conceptual clarity, and linguistic expression. Both explicit and implicit features of AI-generated content were systematically extracted. The practical efficacy of this approach in peer review was validated through cross-verification between AIGC-CHECK system outputs and empirically observed characteristics of AI-generated texts. A tri-level strategy was subsequently formulated, encompassing rapid preliminary screening, in-depth verification, and comprehensive countermeasures. [Result/Conclusion] Compared to relying solely on detection systems or reviewers' experience, this approach offers a more effective method to address inappropriate AIGC usage, overcomes the limitations of single review techniques, and provides practical value in enhancing review efficiency and publication quality.
An Tong , Xue Dejun , Zhou Zhe , Kong Xiangyu , Geng Chong . A Multi-Dimensional Feature Analysis for Detecting and Mitigating AI-Generated Content in Academic Journal Articles[J]. Knowledge Management Forum, 2025 , 10(2) : 99 -108 . DOI: 10.13266/j.issn.2095-5472.2025.007
表1 辨识AI生成内容的切入点、目的及分析思路Table 1 Identification of AI-generated content: key dimensions, objectives, and analytical approach |
| 序号 | 切入点 | 目的和思路 |
|---|---|---|
| 1 | 引文 | 从引文入手,分析引文和被引内容的真实性、准确性、相关性、时效性、规范性 |
| 2 | 数据 | 从数据和案例入手,分析论证过程中是否有数据和实证案例的支撑 |
| 3 | 观点和结论 | 从观点、结论、研究成果、文献综述入手,分析论文内容是否具有原创性和创新性 |
| 4 | 结构框架 | 从论文结构框架入手,分析论文大纲是否存在各类结构框架方面的问题 |
| 5 | 内在逻辑关系 | 从内在逻辑关系入手,分析稿件在逻辑连贯性、逻辑清晰度等方面的问题 |
| 6 | 内容一致性 | 从内容一致性入手,分析是否一致或重复冗余 |
| 7 | 概念阐释 | 从概念阐释入手,分析定义准确性、相似概念清晰度 |
| 8 | 语言表达 | 从语言表达入手,分析语言精准性、表达独特性、内容深刻性 |
表2 大模型生成论文样本及AIGC值Table 2 LLM-generated paper samples and corresponding AIGC values |
| AI生成论文 | 大模型 | 生成论文样本的主题 | AIGC值 |
|---|---|---|---|
| 论文1 | GPT-4o | 身体自尊与睡眠质量的联系及其中介作用 | 0.91 |
| 论文2 | Qwen2.5 | 父母角色重塑:《爸爸当家》中的女性主义视角分析 | 0.76 |
| 论文3 | Gemini-1.5 | 渔业养殖与开发的可持续管理:挑战与策略 | 0.71 |
| 论文4 | Mistral-Large | 建筑工程项目管理的计划与控制方法研究 | 0.90 |
| 论文5 | Mistral-Large | 战略管理视角下中小企业管理创新策略研究 | 0.82 |
| 论文6 | LLaMA-3.1 | 互联网+背景下技能大师工作室新样态 | 0.77 |
表3 人机协同双重验证结果Table 3 Dual verification results through human-ai collaboration |
| 切入点 | 论文1AIGC值:1 | 论文2AIGC值:0.9-1 | 论文3AIGC值:0.8-0.9 | 论文4AIGC值:0.7-0.8 | 论文5AIGC值:0.6-0.7 | 论文6AIGC值:0.5-0.6 |
|---|---|---|---|---|---|---|
| 1.引文 | 引文权威性不足,部分参考文献过时,引文数量较少,引文来源较为单一 | 引用数量偏少,权威性不足,引用过时文献,引文来源单一, 引文与正文内容匹配度一般 | 引文数量严重不足,引文权威性不足,引文与正文内容关联度低 | 引文数量较少,引文来源较单一,缺乏权威性 | 引文权威性不足,引用过时的参考文献,引用来源较单一 | 引文数量少,引文来源单一 |
| 2.数据支撑和实证案例 | 缺乏具体的实证分析和案例依据,完全无数据图表或可视化工具 | 缺乏实证数据和案例的支撑,主要通过理论阐述,无任何实际数据、统计图表或案例佐证观点 | 无任何具体的实证数据、统计分析或案例研究,论述过于概念化 | 缺乏实证分析和案例依据,未给出具体的数据来源和详细数据,无图表展示,无对比分析,论述概念化 | 缺乏实证案例和数据支撑,研究内容停留在理论层面 | 缺乏数据支撑和实证案例 |
| 3.观点/结论/研究成果/文献综述 | 观点和结论缺乏新意,只是对现有研究的总结,无作者独特见解;文献综述仅为罗列,缺乏深入分析和比较 | 观点和结论较为概念化,缺乏创新性,主要是基于已有研究的总结,未提出新的理论框架或实证发现;文献综述仅为简单罗列,缺乏深入分析和批判性评价 | 观点缺乏创新性,多为已有结论的重复,内容以概念性、描述性叙述为主;文献综述仅为简单罗列,未深入分析和比较 | 观点和结论多为概念化和描述性,观点缺乏新意,多是对已有研究的总结,未提出新的研究成果或实证发现;文献综述分析不深入 | 未提出独特见解和创新性研究成果,多是已有研究的总结;文献综述仅为罗列,未深入分析和比较 | 观点和结论多是基于现有研究的总结,缺乏新意和创新性;文献综述部分仅为罗列,未深入分析和比较 |
| 4.论文结构框架 | 部分标题重复,章节标题过于宽泛,内容层级划分不合理 | 部分章节层次重复 | 部分章节逻辑不合理,应调节顺序 | 章节标题重复、层级混乱 | 章节划分不合理,部分内容重复 | 章节层级混乱,章节划分不合理 |
| 5.内在逻辑关系 | 逻辑连贯性差,部分段落逻辑跳跃,论点与论据衔接不紧密,文中多次出现论点与论据脱节问题;段落过渡不自然 | 逻辑连贯性一般,段落间缺乏过渡,论点与论据对应性不足 | 部分内容缺乏逻辑连贯性,段逻辑跳跃,缺乏过渡,论点与论据不完全匹配, | 逻辑连贯性一般,部分段落间缺乏过渡和衔接,逻辑跳跃 | 逻辑连贯性差,逻辑跳跃,缺乏过渡和衔接,论据与论点匹配不足 | 逻辑连贯性差,段逻辑跳跃,论述同一主题或论点的内容不符,逻辑不一致 |
| 6.内容一致性 | 内容重复,冗余问题严重 | 不同章节中多次提及相似内容,内容重复冗余 | 部分章节内容重复冗余 | 部分内容在不同章节中重复冗余,被多次阐述 | 相同或相似内容被多次提及,内容重复冗余 | 相同或相似内容在不同章节中被多次提及,内容重复冗余 |
| 7.概念阐释 | 概念解释不准确,定义不严格,相似术语边界模糊 | 概念阐释不深入,缺乏清晰的界定和阐述 | 概念定义过于宽泛,关键术语阐释不足 | 关键术语的定义模糊,缺乏对相似概念和术语的关系和界限的阐释 | 核心概念阐释不精准,相似概念关系模糊 | 对相似概念和术语的关系和界限阐释得不清晰 |
| 8.语言表达 | 语言表达机械化,多为程式化学术语言,重复使用相似句式,缺乏作者个性化表达 | 语言平淡、公式化,缺乏个性化表达 | 语言表达机械化,整体语言表达偏于新闻报道或政策解读,缺乏学术严谨性和精准性 | 语言表达机械化,缺乏作者个性化表达;有语病,句子不通顺 | 语言表达机械化,使用大量结构类似的陈述句,缺乏学术个性化,句子表述冗长,缺乏层次感 | 语言表达机械化,缺乏作者个性化表达 |
安 彤:设计研究框架,进行数据分析,撰写论文;
薛德军:提出研究思路,指导论文修改;
周 哲:提供实验数据;
孔祥煜:提供实验数据;
耿 崇:提出修改意见。
| 1 |
COPE. Authorship and AI tools.[EB/OL]. [2024-10-15]. http://publicationethics.org/cope-position-statements/ai-author.
|
| 2 |
ICMJE. Recommendations[EB/OL]. [2024-10-15]. http://www.icmje.org/recommendations/.
|
| 3 |
ELSEVIER. The use of generative AI and AI-assisted technologies in writing for Elsevier[EB/OL]. [2024-10-15]. http://www.elsevier.com/about/policies-and-standards/the-use-of-generative-ai-and-ai-assisted-technologies-in-writing-for-elsevier.
|
| 4 |
|
| 5 |
Nature. Editorial policies[EB/OL]. [2024-10-15]. http://www.nature.com/nature-portfolio/editorial-policies.
|
| 6 |
Science. Science journals: editorial policies[EB/OL]. [2024-10-15]. http://www.science.org/content/page/science-journals-editorial-policies#AI.
|
| 7 |
科学技术部监督司. 负责任研究行为规范指引(2023)[EB/OL]. [2024-10-15]. http://www.most.gov.cn/kjbgz/202312/t20231221_189240.html.
(Ministry of Science and Technology,
|
| 8 |
中国科学技术信息研究所, 爱思唯尔, 施普林格·自然, 等.学术出版中AIGC使用边界指南(中文版)[EB/OL]. [2024-10-15]. http://www.istic.ac.cn/ueditor/jsp/upload/file/20230922/1695373131354098301.pdf.
(Institute of Scientific and Technical Information of China (ISTIC), Elsevier, Springer Nature, et al. Boundary guidelines for AIGC use in academic publishing (Chinese Version) [EB/OL]. [2024-10-15]. http://www.istic.ac.cn/ueditor/jsp/upload/file/20230922/1695373131354098301.pdf.)
|
| 9 |
中国科学技术信息研究所, 爱思唯尔, 施普林格·自然, 等.学术出版中AIGC使用边界指南2.0(中文版)[EB/OL]. [2024-10-15]. http://www.istic.ac.cn/ueditor/jsp/upload/file/20241009/1728463944539087271.pdf.
(Institute of Scientific and Technical Information of China (ISTIC), Elsevier, Springer Nature, et al. Boundary guidelines for AIGC use in academic publishing 2.0 (Chinese Version) [EB/OL]. [2024-10-15]. http://www.istic.ac.cn/ueditor/jsp/upload/file/20241009/1728463944539087271.pdf.)
|
| 10 |
中国科学院科研道德委员会.关于在科研活动中规范使用人工智能技术的诚信提醒[EB/OL]. [2024-10-15]. http://www.cssar.cas.cn/ztzl2015/kycx/xfjs2021/zdgf2/202409/t20240913_7362745.html.
(Chinese Academy of Sciences, Committee on Research Ethics. Ethical alert on regulating artificial intelligence technology use in scientific research activities [EB/OL]. [2024-10-15]. http://www.cssar.cas.cn/ztzl2015/kycx/xfjs2021/zdgf2/202409/t20240913_7362745.html.)
|
| 11 |
中华医学会杂志社.关于在论文写作和评审过程中使用生成式人工智能技术的有关规定[EB/OL]. [2024-10-15]. http://stm.castscs.org.cn/yw/40598.jhtml.
(Chinese Medical Association Publishing House. Official provisions on generative artificial intelligence technologies in manuscript composition and peer review procedures [EB/OL]. [2024-10-15]. http://stm.castscs.org.cn/yw/40598.jhtml.)
|
| 12 |
中山大学医学期刊联盟.关于生成性人工智能的声明[EB/OL]. [2024-10-15]. http://journal.gzzoc.com/Es/Online/ArticleList.aspx?CID=296.
(Sun Yat-sen University Medical Journals Consortium. Statement on generative artificial intelligence[EB/OL]. [2024-10-15]. http://journal.gzzoc.com/Es/Online/ArticleList.aspx?CID=296.)
|
| 13 |
Sciengine. Requirements for the use of AI technology in scientific writing[EB/OL]. [2024-10-15]. http://www.sciengine.com/policies/use-of-ai-tools.
|
| 14 |
知乎. 多所大学放宽对AI限制!英国大学宣布AI使用新规, 澳洲教育部支持AI使用[EB/OL]. [2024-10-15]. http://zhuanlan.zhihu.com/p/668344435.
(Zhihu. Multiple universities ease AI restrictions: UK universities announce new regulations on AI usage and Australian education ministry endorses AI adoption [EB/OL]. [2024-10-15]. http://zhuanlan.zhihu.com/p/668344435.)
|
| 15 |
新华网.华师大北师大传播院系联合发布AIGC学生使用指南[EB/OL]. [2024-10-31]. http://comm.ecnu.edu.cn/4b/70/c41512a609136/page.htm.
|
| 16 |
|
| 17 |
|
| 18 |
|
| 19 |
聂慧超. 学术出版交手AIGC产生“双重魔法”[N]. 中国出版传媒商报, 2024-04-05(1).
|
| 20 |
Bloomberg. AI detectors are wrongfully flagging students’ papers as cheating[EB/OL]. [2024-10-15]. http://www.bloomberg.com/news/articles/2024-01-12/ai-detectors-are-wrongfully-flagging-students-papers-as-cheating.
|
| 21 |
郭英剑. 利用人工智能的学术作弊激增, 高校如何应对[EB/OL]. [2024-10-15]. http://news.sciencenet.cn/sbhtmlnews/2024/11/381842.shtm?share_token=d25ad102-c045-4f4c-93bb-167f8f3456ff.
|
| 22 |
薛德军, 孔祥煜, 耿崇, 等. 人工智能生成中文学术论文文本检测研究[J]. 北京电子科技学院学报, 2024, 32(3): 104-112.
|
/
| 〈 |
|
〉 |