Text Mining-based Fake News Detection Using News And Social Media Data

Hyun, Yoonjin;Kim, Namgyu;

doi:10.7838/jsebs.2018.23.4.019

The Journal of Society for e-Business Studies (한국전자거래학회지)

Volume 23 Issue 4
/
Pages.19-39
/
2018
/
2288-3908(pISSN)
/
2765-3846(eISSN)

Society for e-Business Studies (한국전자거래학회)

DOI QR Code

Text Mining-based Fake News Detection Using News And Social Media Data

뉴스와 소셜 데이터를 활용한 텍스트 기반 가짜 뉴스 탐지 방법론

Hyun, Yoonjin (Graduate School of Business IT, Kookmin University) ;
Kim, Namgyu (School of Management Information Systems, Kookmin University)

현윤진 ;
김남규

Received : 2018.10.22
Accepted : 2018.11.18
Published : 2018.11.30

https://doi.org/10.7838/jsebs.2018.23.4.019 Citation PDF KSCI HTML

Download PDF

⟨ Previous Next ⟩

Abstract

Recently, fake news has attracted worldwide attentions regardless of the fields. The Hyundai Research Institute estimated that the amount of fake news damage reached about 30.9 trillion won per year. The government is making efforts to develop artificial intelligence source technology to detect fake news such as holding "artificial intelligence R&D challenge" competition on the title of "searching for fake news." Fact checking services are also being provided in various private sector fields. Nevertheless, in academic fields, there are also many attempts have been conducted in detecting the fake news. Typically, there are different attempts in detecting fake news such as expert-based, collective intelligence-based, artificial intelligence-based, and semantic-based. However, the more accurate the fake news manipulation is, the more difficult it is to identify the authenticity of the news by analyzing the news itself. Furthermore, the accuracy of most fake news detection models tends to be overestimated. Therefore, in this study, we first propose a method to secure the fairness of false news detection model accuracy. Secondly, we propose a method to identify the authenticity of the news using the social data broadly generated by the reaction to the news as well as the contents of the news.

최근 가짜 뉴스가 분야를 막론하고 전 세계에서 주목을 받고 있으며, 현대경제연구원에서는 이러한 가짜 뉴스로 인한 피해 규모가 연간 약 30조 900억원에 달하는 것으로 추산하였다. 정부에서는 "가짜 뉴스 찾기"를 주제로 "인공지능 R&D 챌린지" 대회를 개최하여 가짜 뉴스를 가려낼 인공지능 원천기술 개발에 대한 첫 걸음을 내딛고 있으며, 민간 차원에서도 다양한 분야에서 팩트 체크 서비스가 제공되고 있다. 학계에서도 가짜 뉴스를 탐지하기 위한 시도가 전문가 기반, 집단지성 기반, 인공지능 기반, 시맨틱 기반 등으로 활발하게 이루어지고 있다. 하지만 이러한 시도는 조작의 정밀도가 높을수록 뉴스 자체에 대한 분석만으로 진위 여부를 식별하기가 더욱 어렵다는 한계를 경험하고 있으며, 가짜 뉴스 탐지 모델의 정확도가 과평가된 경향을 보이고 있다. 따라서 본 연구에서는 가짜 뉴스 탐지 모델 정확도의 공정성을 확보하고, 뉴스의 내용뿐만 아니라 해당 뉴스에 대한 반응으로 자연적으로 발생한 광범위한 소셜 데이터를 활용하여 뉴스의 진위 여부를 판정하는 방안을 제안하고자 한다.