Received date: 2023-03-16
Online published: 2025-04-28
[Purpose/Significance] The basic purpose of this study is to study the use of topic characteristics in author attribution of Chinese social media texts. Word2vec is used to supplement the topic model to obtain the deficiencies of topic characteristics. At the same time, strategies are further developed to identify and screen the core topics in the topic characteristics and optimize the use of topic characteristics. So as to improve the using effect of subject features in author attribution. [Methods/Process] The research first used the LDA topic model to extract the academic topics and social topics of the candidate authors, and then used word2vec to develop a merge screening strategy to identify and represent the core topics, and finally used N-gram features and similarity calculation to achieve author attribution. [Results/Conclusion] The experimental results show that the use of core topic characteristics has a positive effect on author attribution of social texts. Meanwhile, the strategy and application of core topic characteristics proposed in this study can also optimize the effect of the use of topic-features, and the highest recognition rate will reach 83% when it is combined with stylistic-features.
Xu Meng , Jing Xie , Chunwang Li . Research on Author Attribution Based on Core Topic[J]. Knowledge Management Forum, 2023 , 8(5) : 351 -364 . DOI: 10.13266/j.issn.2095-5472.2023.030
图1 候选作者A 2013年学术文本关键词云Figure 1 Candidate author A 2013 academic text keyword cloud |
图2 候选作者A 2009年学术文本关键词云Figure 2 Candidate author A 2009 academic text keyword cloud |
图3 候选作者B 2013年学术文本关键词云Figure 3 Candidate author B 2013 academic text keyword cloud |
表1 候选作者文本内容示例Table 1 Examples of text content by candidate authors |
| 文本类型 | 文本内容 |
| 学术文本 | 随着 Internet的兴起及新的计算模式 (如面向服务的计算和普适计算等)的出现,越来越多的系统以服务组合的方式构建。服务是—种独立的计算组件,分布在网络中的各个设备,通过彼此协作提供服务…… |
| 博客文本 | 看到一篇文章谈到普适计算与云计算的区别,该文认为云计算是一个可商业实现的平台,它是包含于普适计算当中。换句话说,普适计算的概念更为广泛。本人较认同该观点,但我认为普适计算是提出了一种新的计算模式,目的还是更广泛地资源融合,以及相关技术融合;当然也产生了很多挑战…… |
表2 作者3的主题分布对比Table 2 Comparison of the topic distribution of author 3 |
| 文本类型 | 主题 | 主题词 | P |
| 社交主题 | T1 | '0.082*互联网','0.030*大脑', '0.012*人类','0.011*网络','0.010*进化' ,'0.009*系统','0.009*结构','0.008*社交','0.008*人脑','0.008*虚拟','0.007*神经系统','0.007*数据','0.006*神经学' ,'0.006*神经','0.006*功能' | 0.281 203 54 |
| T2 | '0.035*威客','0.030*互联网','0.012*知识','0.011*模式','0.009*理论','0.008*科学','0.006*发展','0.006*价值','0.005*智慧','0.004*工作','0.004*功能' '0.004*博客','0.004*进化论','0.004*人类','0.004*进化' | 0.198 570 88 | |
| T3 | '0.044*互联网','0.024*智慧','0.019*大脑','0.016*智能','0.012*系统','0.012*反射弧','0.011*人工智能','0.011*人类','0.010*云脑','0.010*神经系统','0.010*进化','0.010*建设','0.009*社会','0.009*架构','0.008*数据' | 0.055 756 76 | |
| T4 | '0.002*电平','0.002*电容','0.001*电压','0.001*开关','0.001*拓扑','0.001*输出','0.001*逆变器','0.001*单元','0.001*电流','0.001*矢量','0.001*调制','0.001*载波','0.001*三相','0.001*状态','0.001*共模' | 0.114 385 79 | |
| T5 | '0.053*大脑', '0.044*互联网', '0.014*智能', '0.013*人类', '0.010*世界' ,'0.009*发展','0.009*技术', '0.009*系统', '0.009*建设', '0.009*科技','0.007*进化' ,'0.007*模型','0.006*智慧', '0.006*数据','0.005*信息' | 0.250 082 85 | |
| 科研主题 | T1 | '0.041*调节器', '0.040*电压' ,'0.035*三相', '0.032*逆变器', '0.026*负载', '0.025*平衡','0.022*输出','0.022*电位' ,'0.016*控制','0.014*拓扑','0.014*调节','0.013*直流','0.011*三桥','0.010*采用','0.009*电流' | 0.249 974 82 |
| T2 | '0.078*电容','0.064*电压','0.056*拓扑','0.053*开关','0.051*电平','0.036*输出','0.033*单元'.'0.019*状态','0.012*逆变器','0.009*控制','0.009*箝位','0.009*飞跨','0.008*冗余','0.008*分析','0.007*数量' | 0.249 992 74 | |
| T3 | '0.003*人工智能','0.003*功能','0.003*智慧','0.002*模型','0.002*互联网','0.002*发展','0.002*开关','0.002*调制','0.001*共模','0.001*脉冲','0.001*进化','0.001*能力','0.001*测试','0.001*数据','0.001*信息' | 0.009 489 86 | |
| T4 | '0.033*电平','0.033*电压','0.031*调制','0.024*共模','0.022*矢量','0.022*逆变器','0.018*输出','0.018*载波','0.018*脉冲','0.017*电流','0.015*电容','0.013*波形','0.013*线电压','0.011*技术','0.011*三相' | 0.499 842 67 | |
| T5 | '0.038*系统', '0.028*人工智能' ,'0.027*智能' ,'0.024*人类' ,'0.018*知识' '0.015*信息','0.014*智商', '0.010*能力', '0.009*测试', '0.007*模型' ,'0.007*创新' '0.006*标准','0.005*交互', '0.005*发展', '0.005*互联网' | 0.009 486 88 |
表3 word2vec参数设置情况Table 3 Word2Vec parameter settings |
| 超参数 | 参数说明 | 设置 |
| size | 词向量的维数 | 40 |
| window | 上下文窗口的大小 | 20 |
| min-count | 词语出现的最小阈值 | 1 |
| cbow | 词向量模型参数,即是否使用cbow模型(0为使用,非0为使用skip-gram模型) | 3 |
| worker | 计算核心 | 4 |
表4 词向量训练结果示例Table 4 Examples of word vector training results |
| 词语 | 相近词 | 相似度 | 词语 | 相近词 | 相似度 |
| 计算机 | 网络 硬件 信息 专业计算机 计算机基础 计算机专业 操作系统 电脑 | 0.843 800 3 0.838 640 4 0.825 101 0 0.805 385 1 0.802 530 5 0.799 208 8 0.795 360 1 0.794 463 9 | 互联网 | 因特网 互连网 英特网 互联网络 网际网络 互联网通讯 移动互联网 | 0.879 427 7 0.877 517 3 0.864 582 5 0.861 684 6 0.822 480 3 0.822 075 4 0.804 767 3 |
表5 主题特征与核心主题特征进行作者识别结果对比Table 5 Comparison of theme characteristics and core theme characteristics with author identification results |
| 特征 | 作者 | P | R | F1 | ||||
| 主题特征+N-gram特征 | 作者1 | 0.819 3 | 0.800 0 | 0.809 5 | ||||
| 作者2 | 0.857 1 | 0.545 5 | 0.666 7 | |||||
| 作者3 | 0.411 8 | 0.875 0 | 0.560 0 | |||||
| 作者4 | 0.233 3 | 0.875 0 | 0.368 4 | |||||
| 作者5 | 1.000 0 | 1.000 0 | 1.000 0 | |||||
| 作者6 | 0.522 7 | 0.396 5 | 0.451 0 | |||||
| …… | …… | |||||||
| 综合(20名作者) | 0.667 4 | 0.660 9 | 0.688 6 | |||||
| 核心主题特征 | 作者1 | 0.951 2 | 0.917 6 | 0.934 1 | ||||
| 作者2 | 0.800 0 | 0.727 3 | 0.761 9 | |||||
| 作者3 | 0.500 0 | 0.875 0 | 0.636 4 | |||||
| 作者4 | 1.000 0 | 0.875 0 | 0.933 3 | |||||
| 作者5 | 0.909 1 | 1.000 0 | 0.952 4 | |||||
| 作者6 | 0.677 4 | 0.362 1 | 0.471 9 | |||||
| …… | …… | …… | …… | |||||
| 综合(20名作者) | 0.783 7 | 0.827 6 | 0.852 1 | |||||
表6 作者N-gram特征示例Table 6 Examples of N-gram characteristics of authors |
| 作者 | N-gram |
| 作者2 | {('软件工程', '软件工程'): 2, ('软件工程', '专业'): 2, ('专业', '必修'): 2, ('必修', '专业'): 2, ('专业', '基础课'): 2, ('涉及', '内容'): 2, ('包含', '软件'): 2, ('软件', '生命周期'): 2, ('生命周期', '阶段'): 2, ('阶段', '需要'): 2, ('需要', '知识'): 2, ('一门', '概论'): 2, ('概论', '性质'): 2, ('性质', '课程'): 2} |
表7 实验特征组合识别结果Table 7 Experimental feature combination recognition results |
| 特征 | 作者 | F1 | R | P |
| N-gram特征 | 作者1 | 0.800 0 | 0.682 4 | 0.966 7 |
| 作者2 | 0.533 3 | 0.727 3 | 0.421 1 | |
| 作者3 | 0.424 2 | 0.875 0 | 0.280 0 | |
| 作者4 | 0.482 7 | 0.875 0 | 0.333 3 | |
| 作者5 | 0.952 4 | 1.000 0 | 0.909 1 | |
| 作者6 | 0.480 0 | 0.413 8 | 0.571 4 | |
| …… | …… | …… | …… | |
| 综合(20名作者) | 0.632 6 | 0.569 0 | 0.633 6 | |
| N-gram特征+核心主题特征 | 作者1 | 0.951 2 | 0.917 6 | 0.934 1 |
| 作者2 | 0.800 0 | 0.727 2 | 0.761 9 | |
| 作者3 | 0.500 0 | 0.875 0 | 0.636 4 | |
| 作者4 | 1.000 0 | 0.875 0 | 0.933 3 | |
| 作者5 | 0.909 1 | 1.000 0 | 0.952 4 | |
| 作者6 | 0.677 4 | 0.362 1 | 0.471 9 | |
| …… | …… | …… | …… | |
| 综合(20名作者) | 0.783 7 | 0.827 6 | 0.852 1 |
孟 旭:调研及撰写论文;
谢 靖:提出论文修改意见及定稿;
李春旺:提出论文选题和论文技术路线。
| [1] |
Kalgutkar V, Kaur R, Gonzalez H, et al. Code authorship attribution: methods and challenges[J]. ACM computing surveys (CSUR), 2019, 52(1): 1-36.
|
| [2] |
Alrabaee S, Debbabi M, Wang L. CPA: accurate cross-platform binary authorship characterization using LDA[J]. IEEE transactions on information forensics and security, 2020(15): 3051-3066.
|
| [3] |
Maglogiannis I, Iliadis L, Pimenidis E. Artificial intelligence applications and innovations[J]. IFIP advances in information and communication technology, 2020(583):55-266.
|
| [4] |
刘颖,肖天久.金庸与古龙小说计量风格学研究[J].清华大学学报(哲学社会科学版), 2014,29(5):135-147,179.(LIU Y, XIAO T J. A Study of the stylistics of Jin Yong and Gu Long novels[J].Journal of Tsinghua University(philosophy and social sciences),2014,29(5):135-147,179.)
|
| [5] |
百度百科.主题[EB/OL]. [2023-04-05].https://baike.baidu.com/item/主题/2894698.(Baidu Encyclopedia. Topic[EB/OL]. [2023-04-05].https://baike.baidu.com/item/主题/2894698.)
|
| [6] |
Mendenhall T C. The characteristic curves of composition[J].Science,1887(214S):237-246.
|
| [7] |
Hoover D L. Another perspective on vocabulary richness[J].Computers and the humanities,2003(37):151-178.
|
| [8] |
De Vel O, Anderson A, Corney M, et al. Mining e-mail content for author identification forensics[J]. ACM SIGMOD record,2001,30(4):55-64.
|
| [9] |
Keselj V, Peng FC, Cercone N, et al. N-gram based author profiles for authorship attribution[EB/OL].[2023-04-05]. https://core.ac.uk/display/24680735 .
|
| [10] |
祁瑞华,杨德礼,郭旭,等.基于多层面文体特征的博客作者身份识别研究[J].情报学报, 2015,34(6):628-634. (QI R H, YANG D L, GUO X, et al. Blogger identification based on multidimensional stylistic features[J].Journal of the China Society for Scientific and Technical Information, 2015,34(6):628-634.)
|
| [11] |
祁瑞华,郭旭,刘彩虹.中文微博作者身份识别研究[J].情报学报,2017,36(1):72-78.(QI R H, GUO X, LIU C H. Authorship attribution of Chinese Microblog[J].Journal of the China Society for Scientific and Technical Information, 2017,36(1):72-78.)
|
| [12] |
Finn A, Kushmerick N. Learning to classify documents according to genre[J]. Journal of the American Society for Information Science and Technology, 2006, 57(11): 1506-1518.
|
| [13] |
Savoy J. Authorship attribution based on a probabilistic topic model[J]. Information processing & management, 2013, 49(1): 341-354.
|
| [14] |
Anwar W, Bajwa I S, Choudhary M A, et al. An empirical study on forensic analysis of Urdu text using LDA-based authorship attribution[J]. IEEE access,2019(7): 3224-3234.
|
| [15] |
Nie Y, Huang J, Li A, et al. Identifying users based on behavioral-modeling across social media sites[J].Web technologies and applications, 2014(8709):48-55.
|
| [16] |
孙学刚,陈群秀,马亮.基于主题的Web文档聚类研究[J].中文信息学报,2003(3):21-26.(SUN X G,CHEN Q L,MA L. Study on topic-based web clustering[J].Journal of Chinese information processing,2003(3):21-26.)
|
| [17] |
李湘东,张娇,袁满.基于LDA模型的科技期刊主题演化研究[J].情报杂志,2014,33(7):115-121.(LI X D, ZHANG J, YUAN M. On topic evolution of a scientific journal based on LDA model[J]. Journal of intelligence,2014,33(7):115-121.)
|
| [18] |
陈思含.基于微博的多特征情感分析方法研究[D].长春:吉林大学,2021.(CHEN S H. Research on multi-feature sentiment analysis method based on microblog[D].Changchun: Jilin University,2021.)
|
| [19] |
姚全珠,宋志理,彭程.基于LDA模型的文本分类研究[J].计算机工程与应用, 2011, 47(13): 150-153.(YAO Q Z,SONG Z L,PENG C. Research on text categorization based on LDA[J].Computer engineering and applications, 2011, 47(13): 150-153.)
|
| [20] |
王振振,何明,杜永萍.基于LDA主题模型的文本相似度计算[J].计算机科学, 2013, 40(12):229-232.(WANG Z Z, HE M,DU Y P. Text similarity computing based on topic model LDA[J].Computer science, 2013, 40(12):229-232.)
|
| [21] |
崔凯. 基于LDA的主题演化研究与实现[D]. 长沙: 国防科学技术大学,2010.
|
|
(CUI K. The research and implementation of topic evolution based on LDA [D]. Changsha: National University of Defense Technology, 2010.)
|
| [22] |
马思丹,刘东苏.基于加权Word2vec的文本分类方法研究[J].情报科学,2019,37(11):38-42.(MA S D, LIU D S. Text classification method based on weighted word2vec [J]. Information science, 2019,37(11):38-42.)
|
| [23] |
李晓,解辉,李立杰.基于Word2vec的句子语义相似度计算研究[J].计算机科学,2017, 44(9): 256-260.(LI X, JIE H,LI L J. Research on sentence semantic similarity calculation based on word2vec[J]. Computer science, 2017, 44(9): 256-260.)
|
| [24] |
唐晓波,祝黎,谢力.基于主题的微博二级好友推荐模型研究[J].图书情报工作,2014, 58(9):105-113.(TANG X B, ZHU L, XIE L. Two-level microblog friend recommendation based on topic model[J]. Library and information service,2014, 58(9):105-113.)
|
| [25] |
你好星期一.Word2vec参数[EB/OL]. [2022-12-13].
|
|
https://blog.csdn.net/DL_Iris/article/details/119175496 . (Hello on Monday. Word2vec parameter[EB/OL]. [2022-12-13].https://blog.csdn.net/DL_Iris/article/details/119175496.)
|
| [26] |
张谦,高章敏,刘嘉勇.基于Word2vec的微博短文本分类研究[J].信息网络安全, 2017(1):57-62. (ZHANG Q, GAO Z M, LIU J Y. Research of Weibo short text classfication based on word2ve[J]. Netinfo security, 2017(1):57-62.)
|
| [27] |
Johnson A, Wright D. Identifying idiolect in forensic authorship attribution: an N-gram text bite approach[J].Language and law, 2014,1(1):37-69.
|
/
| 〈 |
|
〉 |