Example: bankruptcy

Online Human-Bot Interactions: Detection, Estimation, and ...

Online Human-Bot Interactions: Detection, Estimation, and CharacterizationOnur Varol,1,*Emilio Ferrara,2 Clayton A. Davis,1 Filippo Menczer,1 Alessandro Flammini11 Center for Complex Networks and Systems Research, indiana university , Bloomington, US2 Information Sciences Institute, university of southern California, Marina del Rey, CA, USAbstractIncreasing evidence suggests that a growing amount of socialmedia content is generated by autonomous entities knownas social bots. In this work we present a framework to de-tect such entities on Twitter. We leverage more than a thou-sand features extracted from public data and meta-data aboutusers: friends, tweet content and sentiment, network patterns,and activity time series. We benchmark the classificationframework by using a publicly available dataset of Twitterbots. This training data is enriched by a manually annotatedcollection of active Twitter users that include both humansand bots of varying sophistication.

1Center for Complex Networks and Systems Research, Indiana University, Bloomington, US ... University of Southern California, Marina del Rey, CA, US Abstract Increasing evidence suggests that a growing amount of social ... study how POS tags are distributed.

Tags:

  Study, University, Southern, Indiana, Indiana university, University of southern

Information

Domain:

Source:

Link to this page:

Please notify us if you found a problem with this document:

Other abuse

Advertisement

Transcription of Online Human-Bot Interactions: Detection, Estimation, and ...

1 Online Human-Bot Interactions: Detection, Estimation, and CharacterizationOnur Varol,1,*Emilio Ferrara,2 Clayton A. Davis,1 Filippo Menczer,1 Alessandro Flammini11 Center for Complex Networks and Systems Research, indiana university , Bloomington, US2 Information Sciences Institute, university of southern California, Marina del Rey, CA, USAbstractIncreasing evidence suggests that a growing amount of socialmedia content is generated by autonomous entities knownas social bots. In this work we present a framework to de-tect such entities on Twitter. We leverage more than a thou-sand features extracted from public data and meta-data aboutusers: friends, tweet content and sentiment, network patterns,and activity time series. We benchmark the classificationframework by using a publicly available dataset of Twitterbots. This training data is enriched by a manually annotatedcollection of active Twitter users that include both humansand bots of varying sophistication.

2 Our models yield high ac-curacy and agreement with each other and can detect bots ofdifferent nature. Our estimates suggest that between 9% and15% of active Twitter accounts are bots. Characterizing tiesamong accounts, we observe that simple bots tend to interactwith bots that exhibit more human-like behaviors. Analysis ofcontent flows reveals retweet and mention strategies adoptedby bots to interact with different target groups. Using cluster-ing analysis, we characterize several subclasses of accounts,including spammers, self promoters, and accounts that postcontent from connected media are powerful tools connecting millions of peo-ple across the globe. These connections form the substratethat supports information dissemination, which ultimatelyaffects the ideas, news, and opinions to which we are ex-posed. There exist entities with both strong motivation andtechnical means to abuse Online social networks from in-dividuals aiming to artificially boost their popularity, to or-ganizations with an agenda to influence public opinion.

3 Itis not difficult to automatically target particular user groupsand promote specific content or views (Ferrara et al. 2016a;Bessi and Ferrara 2016). Reliance on social media maytherefore make us vulnerable to botsare accounts controlled by software, algo-rithmically generating content and establishing social bots perform useful functions, such as dis-semination of news and publications (Lokot and Diakopou-los 2016; Haustein et al. 2016) and coordination of vol-unteer activities (Savage, Monroy-Hernandez, and H ollerer2016). However, there is a growing record of malicious ap-plications of social bots. Some emulate human behavior tomanufacture fake grassroots political support (Ratkiewiczet al. 2011), promote terrorist propaganda and recruit-ment (Berger and Morgan 2015; Abokhodair, Yoo, and Mc-Donald 2015; Ferrara et al. 2016c), manipulate the stockmarket (Ferrara et al. 2016a), and disseminate rumors andconspiracy theories (Bessi et al.)

4 2015).A growing body of research is addressing social bot ac-tivity, its implications on the social network, and the de-tection of these accounts (Lee, Eoff, and Caverlee 2011;Boshmaf et al. 2011; Beutel et al. 2013; Yang et al. 2014;Ferrara et al. 2016a; Chavoshi, Hamooni, and Mueen 2016).The magnitude of the problem was underscored by a Twit-ter bot detection challenge recently organized by DARPA tostudy information dissemination mediated by automated ac-counts and to detect malicious activities carried out by thesebots (Subrahmanian et al. 2016).Contributions and OutlineHere we demonstrate that accounts controlled by soft-ware exhibit behaviors that reflects their intents andmodusoperandi(Bakshy et al. 2011; Das et al. 2016), and that suchbehaviors can be detected by supervised machine learningtechniques. This paper makes the following contributions: We propose a framework to extract a large collectionof features from data and meta-data about social mediausers, including friends, tweet content and sentiment, net-work patterns, and activity time series.

5 We use these fea-tures to train highly-accurate models to identify bots. Fora generic user, we produce a[0,1]score representing thelikelihood that the user is a bot. The performance of our detection system is evaluatedagainst both an existing public dataset and an additionalsample of manually-annotated Twitter accounts collectedwith a different strategy. We enrich the previously-trainedmodels using the new annotations, and investigate the ef-fects of different datasets and classification models. We classify a sample of millions of English-speaking ac-tive users. We use different models to infer thresholds inthe bot score that best discriminate between humans andbots. We estimate that the percentage of Twitter accountsexhibiting social bot behaviors is between 9% and 15%. We characterize friendship ties and information flow be-tween users that show behaviors of different nature: hu-man and bot-like. Humans tend to interact with [ ] 27 Mar 2017human-like accounts than bot-like ones, on average.

6 Reci-procity of friendship ties is higher for humans. Some botstarget users more or less randomly, others can choose tar-gets based on their intentions. Clustering analysis reveals certain specific behavioralgroups of accounts. Manual investigation of samples ex-tracted from each cluster points to three distinct botgroups: spammers, self promoters, and accounts that postcontent from connected Detection FrameworkIn the next section, we introduce a Twitter bot detectionframework ( ) thatis freely available Online . This system leverages more thanone thousand features to evaluate the extent to which a Twit-ter account exhibits similarity to the known characteristicsof social bots (Davis et al. 2016).Feature ExtractionData collected using the Twitter API are distilled in 1,150features in six different classes. The classes and types of fea-tures are reported in Table 1 and discussed extracted from user meta-data have been used to classify users and patterns be-fore (Mislove et al.)

7 2011; Ferrara et al. 2016a). We ex-tract user-based features from meta-data available throughthe Twitter API. Such features include the number of friendsand followers, the number of tweets produced by the users,profile description and Users are linked by follower-friend (fol-lowee) relations. Content travels from person to person viaretweets. Also, tweets can be addressed to specific usersvia mentions. We consider four types of links: retweeting,mentioning, being retweeted, and being mentioned. Foreach group separately, we extract features about languageuse, local time, popularity, etc. Note that, due to Twitter sAPI limits, we do not use follower/followee informationbeyond these aggregate network structure carries crucialinformation for the characterization of different types ofcommunication. In fact, the usage of network features sig-nificantly helps in tasks like political astroturf detection(Ratkiewicz et al. 2011).

8 Our system reconstructs three typesof networks: retweet, mention, and hashtag co-occurrencenetworks. Retweet and mention networks have users asnodes, with a directed link between a pair of users that fol-lows the direction of information spreading: toward the userretweeting or being mentioned. Hashtag co-occurrence net-works have undirected links between hashtag nodes whentwo hashtags occur together in a tweet. All networks areweighted according to the frequency of interactions or co-occurrences. For each network, we compute a set of fea-tures, including in- and out-strength (weighted degree) dis-tributions, density, and clustering. Note that out-degree andout-strength are measures of research suggests that the tem-poral signature of content production and consumption mayreveal important information about Online campaigns andtheir evolution (Ghosh, Surachawala, and Lerman 2011;Ferrara et al. 2016b; Chavoshi, Hamooni, and Mueen 2016).To extract this signal we measure several temporal featuresrelated to user activity, including average rates of tweet pro-duction over various time periods and distributions of timeintervals between and language recent papers havedemonstrated the importance of content and language fea-tures in revealing the nature of social media conversa-tions (Danescu-Niculescu-Mizil et al.)

9 2013; McAuley andLeskovec 2013; Mocanu et al. 2013; Botta, Moat, and Preis2015; Letchford, Moat, and Preis 2015; Das et al. 2016).For example, deceiving messages generally exhibit informallanguage and short sentences (Briscoe, Appling, and Hayes2014). Our system does not employ features capturing thequality of tweets, but collects statistics about length and en-tropy of tweet text. Additionally, we extract language fea-tures by applying thePart-of-Speech(POS) tagging tech-nique, which identifies different types of natural languagecomponents, orPOS tags. Tweets are therefore analyzed tostudy how POS tags are analysis is a powerful toolto describe the emotions conveyed by a piece of text, andmore broadly the attitude or mood of an entire conversa-tion. Sentiment extracted from social media conversationshas been used to forecast offline events including financialmarket fluctuations (Bollen, Mao, and Zeng 2011), and isknown to affect information spreading (Mitchell et al.

10 2013;Ferrara and Yang 2015). Our framework leverages sev-eral sentiment extraction techniques to generate varioussentiment features, includingarousal,valenceanddomi-nancesco res (Warriner, Kuperman, and Brysbaert 2013),happinessscore (Kloumann et al. 2012),polarizationandstrength(Wilson, Wiebe, and Hoffmann 2005), andemoti-conscore (Agarwal et al. 2011).Model EvaluationTo train our system we initially used a publicly availabledataset consisting of 15K manually verified Twitter botsidentified via ahoneypotapproach (Lee, Eoff, and Caver-lee 2011) and 16K verified human accounts. We collectedthe most recent tweets produced by those accounts using theTwitter Search API. We limited our collection to 200 publictweets from a user timeline and up to 100 of the most recentpublic tweets mentioning that user. This procedure yielded adataset of million tweets produced by manually verifiedbots and 3 million tweets produced by human benchmarked our system using several off-the-shelfalgorithms provided in thescikit-learnlibrary (Pedregosa etal.


Related search queries