{"id":438,"date":"2018-03-22T04:03:14","date_gmt":"2018-03-22T04:03:14","guid":{"rendered":"http:\/\/wordpress.cs.vt.edu\/cs6724spring18\/?p=438"},"modified":"2018-03-22T04:03:14","modified_gmt":"2018-03-22T04:03:14","slug":"reflection-10-03-22-john-wenskovitch","status":"publish","type":"post","link":"https:\/\/wordpress.cs.vt.edu\/cs6724spring18\/2018\/03\/22\/reflection-10-03-22-john-wenskovitch\/","title":{"rendered":"Reflection #10 \u2013 [03\/22] \u2013 [John Wenskovitch]"},"content":{"rendered":"<p>This pair of papers describes aspects of those who ruin the Internet for the rest of us.\u00a0 Kumar\u2019s \u201cAn Army of Me\u201d paper discusses the characteristics of sockpuppets in online discussion communities (as an aside, the term \u201csockpuppet\u201d never really clicked for me until seeing its connection with \u201cpuppetmaster\u201d in the introduction of this paper).\u00a0 Looking at nine different discussion communities, the authors evaluate the posting behavior, linguistic features, and social network structure of sockpuppets, eventually using those characteristics to build a classifier which achieved moderate success in identifying sockpuppet accounts.\u00a0 Lee\u2019s \u201cUncovering Social Spammers\u201d paper uses a honeypot technique to identify social spammers (spam accounts on social networks).\u00a0 They deploy their honeypots on both MySpace and Twitter, capturing information about social spammer profiles in order to understand their characteristics, using some similar characteristics as Kumar\u2019s paper (social network structure and posting behavior).\u00a0 These authors also build classifiers for both MySpace and Twitter using the features that they uncovered with their honeypots.<\/p>\n<p>Given the discussion that we had previously when reading the Facebook papers, the first thing that jumped out at me when reading through the results of the \u201cArmy of Me\u201d paper was the <strong>small effect sizes<\/strong>, especially in the linguistics traits subsection.\u00a0 Again, these included strong p-values of p&lt;0.001 in many cases, but also showed minute differences in the rates of using words like \u201cI\u201d (0.076 vs 0.074) and \u201cyou\u201d (0.017 vs 0.015).\u00a0 Though the authors don\u2019t specifically call out their effect sizes, they do provide the means for each class and should be applauded for that.\u00a0 (They also reminded me to leave a note in my midterm report to discuss effect sizes.)<\/p>\n<p>One limitation of \u201cArmy of Me\u201d that was not discussed was the fact that all nine communities that they evaluated use Disqus as a commenting platform.\u00a0 While this made it easier for the authors to acquire their (anonymized) data for this study, <strong>there may be safety checks or other mechanisms built into Disqus that bias the characteristics of sockpuppets that appear on that platform.<\/strong>\u00a0 Some of their proposed future work, such as studying the Facebook and 4chan communities, might have made their results stronger.<\/p>\n<p>\u201cArmy of Me\u201d also reminded me of the drama from several years ago around the reddit user unidan, the \u201cexcited biologist,\u201d who was banned from the community for vote manipulation. \u00a0He used sockpuppet accounts to upvote his own posts and downvote other responses, thereby inflating his own reputation on the site.<\/p>\n<p>Besides identifying MySpace as a \u201cgrowing community\u201d in 2010, I thought that the \u201cUncovering Social Spammers\u201d paper was a mostly solid and concise piece of research.\u00a0 The use of a human-in-the-loop approach to obtain human validation of spam candidates to improve the SVM classifier appealed to the human-in-the-loop researcher in me.\u00a0 Some of the findings from their honeypot data acquisition were interesting, such as the fact that Midwesterners are popular spamming targets and that California is a popular profile location.\u00a0 <strong>I\u2019m wondering if the fact that these patterns were seen is indicative of some bias in the data collection (is the social honeypot technique biased towards picking up spammers from California?), or if there actually is a trend in spam accounts to pick California as a profile location.<\/strong>\u00a0 This wasn\u2019t particular clear to me; instead, it was just stated and then ignored.<\/p>\n<p>I really liked their use of both MySpace and Twitter, as the two different social networks enabled the collection of different features (e.g., F-F ratio for Twitter, number of friends for MySpace) in order to show that the classifier can work on multiple datasets.\u00a0 It\u2019s almost midnight and I haven\u2019t slept enough this month, but I\u2019m still puzzled by the confusion matrix that they presented in Table 1.\u00a0 Did they intend to leave variables in that table?\u00a0 If so, it doesn\u2019t really add much to the paper, as they\u2019re just describing the standard definitions of precision, recall, and false positive.\u00a0 They don\u2019t present any other confusion matrices in the paper, so it seems even more out of place.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>This pair of papers describes aspects of those who ruin the Internet for the rest of us.\u00a0 Kumar\u2019s \u201cAn Army of Me\u201d paper discusses the characteristics of sockpuppets in online discussion communities (as an aside, the term \u201csockpuppet\u201d never really clicked for me until seeing its connection with \u201cpuppetmaster\u201d in the introduction of this paper).\u00a0 [&hellip;]<\/p>\n","protected":false},"author":133,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[1],"tags":[],"class_list":["post-438","post","type-post","status-publish","format-standard","hentry","category-uncategorized"],"jetpack_featured_media_url":"","_links":{"self":[{"href":"https:\/\/wordpress.cs.vt.edu\/cs6724spring18\/wp-json\/wp\/v2\/posts\/438","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/wordpress.cs.vt.edu\/cs6724spring18\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/wordpress.cs.vt.edu\/cs6724spring18\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/wordpress.cs.vt.edu\/cs6724spring18\/wp-json\/wp\/v2\/users\/133"}],"replies":[{"embeddable":true,"href":"https:\/\/wordpress.cs.vt.edu\/cs6724spring18\/wp-json\/wp\/v2\/comments?post=438"}],"version-history":[{"count":1,"href":"https:\/\/wordpress.cs.vt.edu\/cs6724spring18\/wp-json\/wp\/v2\/posts\/438\/revisions"}],"predecessor-version":[{"id":439,"href":"https:\/\/wordpress.cs.vt.edu\/cs6724spring18\/wp-json\/wp\/v2\/posts\/438\/revisions\/439"}],"wp:attachment":[{"href":"https:\/\/wordpress.cs.vt.edu\/cs6724spring18\/wp-json\/wp\/v2\/media?parent=438"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/wordpress.cs.vt.edu\/cs6724spring18\/wp-json\/wp\/v2\/categories?post=438"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/wordpress.cs.vt.edu\/cs6724spring18\/wp-json\/wp\/v2\/tags?post=438"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}