网络谣言敏感词库的构建研究——以新浪微博谣言为例
|
夏松,讲师,博士; |
|
林荣蓉,硕士研究生。 |
收稿日期: 2019-06-18
网络出版日期: 2025-04-28
基金资助
系国家社会科学基金资助项目“基于文本挖掘的网络谣言预判研究”(14BXW033)
Construction of Sensitive Thesaurus for Network Rumors——Taking the Microblog Rumors as an Example
Received date: 2019-06-18
Online published: 2025-04-28
[目的/意义] 网络谣言严重影响网络正常信息的传播,对网络谣言进行识别有着重要的现实意义。笔者构建一个基于微博的网络谣言敏感词库,以提高网络谣言的识别精度。[方法/过程] 针对微博类社交平台短文本的特点,首先舍弃传统的分词算法,设计LBCP抽词算法,并结合位置信息和改进的TF-IDF权重来提取敏感词库的种子词集,然后通过聚类算法将种子词的近义词补充到词库中,再将常用的替代词也加入到词库中,从而得到最终的敏感词库。[结果/结论] 利用敏感词特征对谣言进行判断,在提取微博的内容特征、用户特征、传播特征以及情感分析特征的基础上,新增敏感词特征以后谣言识别率有明显提升,得到较好的识别效果。
夏松 , 林荣蓉 , 刘勘 . 网络谣言敏感词库的构建研究——以新浪微博谣言为例[J]. 知识管理论坛, 2019 , 4(5) : 267 -275 . DOI: 10.13266/j.issn.2095-5472.2019.028
[Purpose/significance] The network rumors seriously influent the spread of normal information on the internet. The purpose of this paper is to construct a sensitive lexicon on microblog rumors and to improve the recognition accuracy of the network rumors. [Method/process] According to the characteristics of microblog’s short text on social networking platforms, this paper focuses on construction of the microblog sensitive thesaurus, which is built up through LBCP algorithm and extension of multiple level words. At first, the method directly extracts words through LBCP algorithm, which considers the cohesion and polymerization of rumor words. And then, based on the core words, multiple level words are expanded to get sensitive thesaurus. [Result/conclusion] In addition to the features of the text, user characteristics, propagation characteristics, emotional analysis, and rumor features based on sensitive thesaurus are exploited. Experimental results show that the accuracy of microblog’s rumor recognition can be improved greatly based on sensitive thesaurus.
Key words: sensitive thesaurus; word embedding; feature space; network rumors
图3 计算内聚度使用的滑动窗口 |
| 央 | 视 | 已 | 经 | 报 | 道 | 此 | 事 | ...... |
| c 1 | c 2 | c 3 | c 4 | c 5 | c 6 | c 7 | c 8 | ...... |
,使得TF的权重值大于IDF的权重值);②消除了文档长度的不同对词权重的影响(公式(5)中增加分母,进行余弦归一化处理),同时通过对词频取对数来消除词频大小差异对权重计算的影响。词x改进的TF-IDF权重计算如公式(5)所示:
是词频TF的权重,tf和idf分别表示词x的TF值和IDF值,公式(5)右边的分母利用谣言微博中出现的每个词x的TF值tf i 和IDF值idf i 进行余弦归一化处理。
表1 谣言种子词集 |
| 类型 | 词集示例 |
| 种子词集 | 孩子、拐走、扩散、酬金、紧急、严重、知情、女孩、帮忙、转发、求、找、联系、死、爆炸、死亡、转转、钱、伤、黑、白血病、救、去世、农药、丢、癌症、偷、救援、专家、卖、滋、食品、导致、真相、死了、批捕、感染、必须、提醒、杀、禁用、失踪、抢救、证实、罪、打死...... |
表2 近似词集部分示例 |
| 类型 | 词集示例 |
| 近似词集 | 朋友、监控、附近、万分、抱、留意、兄弟、男人、告、姐妹、联系人、双重、关心、婴儿、达、恒天、教授、主、逸、场合、伤亡、死伤、电视、中伤、保、记、妇女、鬼子、保健、院、日本、含、果断、刷、毒素、局、发现、天地、血压、此次、余香、喝、引发、插头、真的、兰、版、海域、七、警、接力、唯一、轻、人数、新闻、消防、天津、牺牲、杭州...... |
表3 敏感词库特征对谣言判别的效果 |
| 判别模型 | 传统特征 | 传统特征+敏感词特征 | |||
| 准确率 | 召回率 | 准确率 | 召回率 | ||
| 随机森林 | 79.82% | 62.98% | 89.29 % | 85.10% | |
| GBRT | 81.44 % | 65.65% | 92.65% | 86.71% | |
| SVM | 80.71 % | 66.09% | 85.94 % | 83.22% | |
| CNN | 80.38% | 72.66% | 91.12 % | 83.54% | |
| BiLSTM | 82.68 % | 73.54% | 95.26% | 88.67% | |
| TextCNN | 81.25 % | 77.12% | 93.48 % | 86.09% | |
夏 松:设计模型, 完成实验,修改论文;
林荣蓉:采集数据,进行实验,撰写论文初稿;
刘 勘:提出研究思路,设计研究方案,修改论文与定稿。
| [1] |
徐建民,王金花,马伟瑜.利用本体关联度改进的TF-IDF特征词提取方法[J].情报科学,2011, 29(2): 279-283.
|
| [2] |
周晓. 基于互联网的情感词库扩展与优化研究[D]. 沈阳: 东北大学, 2011.
|
| [3] |
刘耕,方勇,刘嘉勇.基于关联词和扩展规则的敏感词库设计[J].四川大学学报(自然科学版), 2009, 46(3): 667-671.
|
| [4] |
徐琳宏,林鸿飞,潘宇,等.情感词汇本体的构造[J]. 情报学报, 2008,27(2): 180-185
|
| [5] |
侯丽,李姣,侯震,等.基于混合策略的公众健康领域新词识别方法研究[J].图书情报工作, 2015,59(23):115-123.
|
| [6] |
QUAN C, REN F. Construction of a blog emotion corpus for Chinese emotional expression analysis[C]//Proceedings of conference on empirical methods in natural language processing. Stroudsburg:Association for Computational Linguistics,2009:1446-1454.
|
| [7] |
PENG F, FENG F, MCCALLUM A. Chinese segmentation and new word detection using conditional random fields[C]//Proceedings of international conference on computational linguistics. Stroudsburg: Association for Computational Linguistics,2004:562-569.
|
| [8] |
周强.汉语谓词组合范畴语法词库的自动构建研究[J].中文信息学报, 2016,30(3): 196-203.
|
| [9] |
CHEN K J, MA W Y. Unknown word extraction for Chinese documents[C]// Proceedings of international conference on DBLP. Taipei: Morgan Kaufmann Publishers, 2002:169-175.
|
| [10] |
彭云,万常选,江腾蛟,等.基于语义约束LDA的商品特征和情感词提取[J].软件学报, 2017,28(3):676-693.
|
| [11] |
CHEN H, LYNCH K, BASU K, et al. Generating, integrating and activating thesauri for concept-based document retrieval[J]. IEEE intelligent systems and their applications, 1993,8(2):25-34.
|
| [12] |
YU S,CAI D,WEN J,et al. Improving pseudo-relevance feedback in web information retrieval using Web page segmentation[C]//Proceedings of the 12th international conference on World Wide Web. New York: ACM, 2003:11-18.
|
| [13] |
PNOTE J M,CROFT W B. A language modeling approach to information retrieval[C]//Proceeding of the 21st International ACM SIGIR conference on research and development in information retrieval. New York: ACM, 1998:275-281.
|
| [14] |
PEDERSEN T, KULKARNI A. Identifying similar words and contexts in natural language with sense clusters[C]// Proceedings of the 20th national conference on artificial intelligence. Pittsburgh: AAAI Press, 2010:1694-1695.
|
| [15] |
TURNEY P D, LITTMAN M L. Measuring praise and criticism: inference of semantic orientation from association[J]. ACM transactions on information systems, 2003, 21(4):315-346.
|
| [16] |
NEVIAROUSKAYA A,PRENDINGER H,ISHIZUKA M. SentiFul: a lexicon for sentiment analysis[J].IEEE transactions on affective computing,2011,2(1):22-36.
|
/
| 〈 |
|
〉 |