{"id":268,"date":"2019-01-31T13:49:09","date_gmt":"2019-01-31T13:49:09","guid":{"rendered":"https:\/\/wordpress.cs.vt.edu\/cs4984spring19\/?p=268"},"modified":"2019-01-31T13:49:10","modified_gmt":"2019-01-31T13:49:10","slug":"reflection-2-1-31-matthew-fishman","status":"publish","type":"post","link":"https:\/\/wordpress.cs.vt.edu\/cs4984spring19\/2019\/01\/31\/reflection-2-1-31-matthew-fishman\/","title":{"rendered":"[Reflection #2] \u2013 [1\/31] \u2013 [Matthew, Fishman]"},"content":{"rendered":"\n<h2 class=\"wp-block-heading\">Quick Summary<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Why is this study\nimportant? <\/strong>With the sheer quantity of news stories shared across social\nmedia platforms, news consumers must use quick heuristics to decide if a news\nsource is trustworthy. Horne et al. set out with the task of finding\ndistinguishing characteristics between real news and fake news, to help news\nconsumers be better able to distinguish the two.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>How did they do it? <\/strong>They\nstudied stylistic, complexity, and psychological features of real news, fake\nnews, and satire articles to better help them classify each from one another. Horne\net al. then used ANOVA tests on normally-distributed data sets and Wilcoxon\nrank sum tests on non-normally distributed features to find which features\ndiffer between the different categories of news. They then selected the top 4\ndistinguishing features for both the body text and title text of the articles\nto create an SVM model with a linear kernel and 5-fold cross-validation.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>What were the results? <\/strong>Their classifier achieved 71% to 91% cross-validation accuracy over a 50% baseline when distinguishing between real news and either satire or fake news. Some of the major differences they found between fake news and real news is that real news has a much more substantial body, while fake news has much longer titles with simpler words. This suggests that <strong>fake news writers are attempting to squeeze as much substance into the titles as possible.<\/strong><\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Reflection<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Again, I was not as surprised by the outcomes of this study as I had hoped.<\/strong> It seems obvious that, for example, fake news articles would have a less substantial body and use simpler words in their titles than real news. However, the lack of stop words and length of fake news titles did surprise me; I had always associated fake news with click-bait, which usually has very short titles.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">A few problems I had with the study included:<\/p>\n\n\n\n<ul class=\"wp-block-list\"><li>The data sets used. If the real enemy here is <em>fake news<\/em>, then why would the researchers use only a total of 110 fake news sources (in comparison to 308 satire news sources and 4111 real news sources). <em>No wonder the classifier had an easier time distinguishing real\u00a0news\u00a0from\u00a0the\u00a0other\u00a0two<\/em>.<\/li><li>The features extracted. The researchers could have used credibility features or different user interaction metrics like shares or clicks to better distinguish fake news from real. <em>If the study utilized more than just linguistic features, their classifier could have been much more accurate.<\/em><\/li><\/ul>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Going Forward:<\/strong><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Some improvements on the study (in addition to the dataset size and features extracted) could be to do some user research:<\/p>\n\n\n\n<ul class=\"wp-block-list\"><li>How much time do users spend reading a fake news\narticle in comparison to real news?<\/li><li>What characteristics of a news consumer are\ncorrelated with sharing, liking, or believing fake news?<\/li><li>What is the ratio of fake news articles clicked\nor shared to that of real news?<\/li><\/ul>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Questions Raised:<\/strong><\/p>\n\n\n\n<ul class=\"wp-block-list\"><li>Can we predict the type of user to be more susceptible to fake news?<\/li><li>How have fake news&#8217; linguistics changed over the years? What can we learn from this is predicting how they might change in the future?<\/li><li>Should real news sources change the format of their titles to give lazy consumers as much information as possible without needing to read the article? Or, would this hurt their baseline as their articles might not get as many clicks if all the information is in the title?<\/li><li>Should news aggregates like Facebook and Reddit be using similar classifiers to mark how potentially \u201cfake\u201d a news article is?<\/li><\/ul>\n","protected":false},"excerpt":{"rendered":"<p>Quick Summary Why is this study important? With the sheer quantity of news stories shared across social media platforms, news consumers must use quick heuristics to decide if a news source is trustworthy. Horne et al. set out with the task of finding distinguishing characteristics between real news and fake news, to help news consumers [&hellip;]<\/p>\n","protected":false},"author":232,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[1],"tags":[],"class_list":["post-268","post","type-post","status-publish","format-standard","hentry","category-uncategorized"],"jetpack_featured_media_url":"","_links":{"self":[{"href":"https:\/\/wordpress.cs.vt.edu\/cs4984spring19\/wp-json\/wp\/v2\/posts\/268","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/wordpress.cs.vt.edu\/cs4984spring19\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/wordpress.cs.vt.edu\/cs4984spring19\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/wordpress.cs.vt.edu\/cs4984spring19\/wp-json\/wp\/v2\/users\/232"}],"replies":[{"embeddable":true,"href":"https:\/\/wordpress.cs.vt.edu\/cs4984spring19\/wp-json\/wp\/v2\/comments?post=268"}],"version-history":[{"count":2,"href":"https:\/\/wordpress.cs.vt.edu\/cs4984spring19\/wp-json\/wp\/v2\/posts\/268\/revisions"}],"predecessor-version":[{"id":310,"href":"https:\/\/wordpress.cs.vt.edu\/cs4984spring19\/wp-json\/wp\/v2\/posts\/268\/revisions\/310"}],"wp:attachment":[{"href":"https:\/\/wordpress.cs.vt.edu\/cs4984spring19\/wp-json\/wp\/v2\/media?parent=268"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/wordpress.cs.vt.edu\/cs4984spring19\/wp-json\/wp\/v2\/categories?post=268"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/wordpress.cs.vt.edu\/cs4984spring19\/wp-json\/wp\/v2\/tags?post=268"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}