ラベル StreamTwitter の投稿を表示しています。 すべての投稿を表示
ラベル StreamTwitter の投稿を表示しています。 すべての投稿を表示

2012年5月10日木曜日

情報拡散に関する研究

Information Transfer in Social Media, WWW 2012
Greg Ver Steeg, Aram Galstyan

The Role of Social Networks in Information Diffusion, WWW 2012
Eytan Bakshy, Itamar Rosenn, Cameron Marlow, Lada Adamic

Recommendations to Boost Content Spread in Social Networks, WWW 2012
Vineet Chaoji, Sayan Ranu, Rajeev Rastogi, Rushi Bhatt


Differences in the Mechanics of Information Diffusion Across Topics: Idioms, Political Hashtags, and Complex Contagion on Twitter, WWW 2011
http://www.cs.cornell.edu/home/kleinber/www11-hashtags.pdf

Limiting the Spread of Misinformation in Social Networks, WWW 2011 
Ceren Budak, Divyakant Agrawal and Amr El Abbadi

Information Credibility on Twitter, WWW, 2011 
Carlos Castillo, Marcelo Mendoza and Bárbara Poblete

Information Spreading in Contex, WWW 2011

Information diffusion in online social networks

Survey Survey on Information Diffusion, 2008

Modeling Information Diffusion in Implicit Networks

Information Diffusion Through Blogspace, WWW2004
http://people.csail.mit.edu/dln/papers/blogs/idib.pdf
We study the dynamics of information propagation in environments of low-overhead personal publishing, using a large collection of weblogs over time as our example domain. We characterize and model this collection at two levels. First, we present a macroscopic characterization of topic propagation through our corpus, formalizing the notion of long-running “chatter” topics consisting recursively of “spike” topics generated by outside world events, or more rarely, by resonances within the community. Second, we present a microscopic characterization of propagation from individual to individual, drawing on the theory of infectious diseases to model
the flow. We propose, validate, and employ an algorithm to induce the underlying propagation network from a sequence of posts, and report on the results.

2012年1月7日土曜日

Twitter + Lifelog

A location predictor based on dependencies between multiple lifelog data
http://dl.acm.org/citation.cfm?id=1867702


Towards trajectory-based experience sharing in a city
http://dl.acm.org/citation.cfm?id=2063221

Twitter as Economic Indicators

Twitter as Economic Indicators

http://dealmakersguide.com/twitter-as-an-economic-indicator/
http://professional.wsj.com/article/SB10001424052970204138204576598942105167646.html


Modeling public mood and emotion:
Twitter sentiment and socio-economic phenomena
http://arxiv.org/pdf/0911.1583

Twitter

Twitter under the microscope - First Monday

http://firstmonday.org/htbin/cgiwrap/bin/ojs/index.php/fm/article/view/2317/206
Scholars, advertisers and political activists see massive online social networks as a representation of social interactions that can be used to study the propagation of ideas, social bond dynamics and viral marketing, among others. But the linked structures of social networks do not reveal actual interactions among people. Scarcity of attention and the daily rhythms of life and work makes people default to interacting with those few that matter and that reciprocate their attention. A study of social interactions within Twitter reveals that the driver of usage is a sparse and hidden network of connections underlying the “declared” set of friends and followers.3

Twitter + Influencer (Retweet)

Want to be Retweeted? Large Scale Analytics on
Factors Impacting Retweet in Twitter Network
http://www2.parc.com/isl/members/hong/publications/socialcomputing2010.pdf

Retweeting is the key mechanism for information
diffusion in Twitter. It emerged as a simple yet powerful way of
disseminating useful information. Even though a lot of
information is shared via its social network structure in Twitter,
little is known yet about how and why certain information
spreads more widely than others. In this paper, we examine a
number of features that might affect retweetability of tweets. We
gathered content and contextual features from 74M tweets and
used this data set to identify factors that are significantly
associated with retweet rate. We also built a predictive retweet
model. We found that, amongst content features, URLs and
hashtags have strong relationships with retweetability. Amongst
contextual features, the number of followers and followees as well
as the age of the account seem to affect retweetability, while,
interestingly, the number of past tweets does not predict
retweetability of a user’s tweet. We believe that this research
would inform the design of sensemaking tools for

Identifying Influencers on Twitter
http://thenoisychannel.com/2011/04/16/identifying-influencers-on-twitter/

More Twitter Analysis: Influencers Don't Retweet
http://www.readwriteweb.com/archives/more_twitter_analysis_influencers_dont_retweet.php

Measuring User Influence in Twitter: The Million Follower Fallacy
http://an.kaist.ac.kr/~mycha/docs/icwsm2010_cha.pdf


Tweet, Tweet, Retweet: Conversational Aspects of Retweeting on Twitter
http://research.microsoft.com/pubs/102168/TweetTweetRetweet.pdf

Twitter + 災害

Twitter in Disaster Mode:
http://people.ee.ethz.ch/~hossmath/papers/woid_twimight.pdf

Recent natural disasters (earthquakes, floods, etc.) have shown
that people heavily use platforms like Twitter to communicate and organize in emergencies. However, the fixed infrastructure supporting such communications may be temporarily wiped out. In such situations, the phones’ capabilities of
infrastructure-less communication can fill in: By propagating data opportunistically (from phone to phone), tweets can
still be spread, yet at the cost of delays.
In this paper, we present Twimight and its network security extensions. Twimight is an open source Twitter client
for Android phones featured with a “disaster mode”, which
users enable upon losing connectivity. In the disaster mode,
tweets are not sent to the Twitter server but stored on the
phone, carried around as people move, and forwarded via
Bluetooth when in proximity with other phones. However,
switching from an online centralized application to a distributed and delay-tolerant service relying on opportunistic
communication requires rethinking the security architecture.
We propose security extensions to offer comparable security in the disaster mode as in the normal mode to protect
Twimight from basic attacks. We also propose a simple,
yet efficient, anti-spam scheme to avoid users from being
flooded with spam. Finally, we present a preliminary empirical performance evaluation of Twimight.
Categories and Subject Descriptors

Twitter + イベント検出 (2)

TwitterMonitor: Trend Detection over the Twitter Stream
http://queens.db.toronto.edu/~mathiou//TwitterMonitor.pdf

We present TwitterMonitor, a system that performs trend
detection over the Twitter stream. The system identifies
emerging topics (i.e. ‘trends’) on Twitter in real time and
provides meaningful analytics that synthesize an accurate
description of each topic. Users interact with the system
by ordering the identified trends using different criteria and
submitting their own description for each trend.
We discuss the motivation for trend detection over social media streams and the challenges that lie therein. We
then describe our approach to trend detection, as well as
the architecture of TwitterMonitor. Finally, we lay out our
demonstration scenario.

Twitter +

Twitter Power: Tweets as Electronic Word of Mouth
http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.155.3321

In this paper we report research results investigating microblogging as a form of electronic word-of-mouth for sharing consumer opinions concerning brands. We analyzed more than 150,000 microblog postings containing branding comments, sentiments, and opinions.We investigated the overall structure of these microblog postings, the types of expressions, and the movement in positive or negative sentiment.We compared automated methods of classifying sentiment in these microblogs with manual coding. Using a case study approach, we analyzed the range, frequency, timing, and content of tweets in a corporate account. Our research findings show that 19% of microblogs contain mention of a brand. Of the branding microblogs, nearly 20 % contained some expression of brand sentiments. Of these, more than 50 % were positive and 33 % were critical of the company or product. Our comparison of automated and manual coding showed no significant differences between the two approaches. In analyzing microblogs for structure and composition, the linguistic structure of tweets approximate the linguistic patterns of natural language expressions. We find that microblogging is an online tool for customer word of mouth communications and discuss the implications for corporations using microblogging as part of their overall marketing strategy.

Twitter + 感情分析 (金融への応用含む)

Twitter mood predicts the stock market

http://arxiv.org/abs/1010.3003

Behavioral economics tells us that emotions can profoundly affect individual behavior and decision-making. Does this also apply to societies at large, i.e., can societies experience mood states that affect their collective decision making? By extension is the public mood correlated or even predictive of economic indicators? Here we investigate whether measurements of collective mood states derived from large-scale Twitter feeds are correlated to the value of the Dow Jones Industrial Average (DJIA) over time. We analyze the text content of daily Twitter feeds by two mood tracking tools, namely OpinionFinder that measures positive vs. negative mood and Google-Profile of Mood States (GPOMS) that measures mood in terms of 6 dimensions (Calm, Alert, Sure, Vital, Kind, and Happy). We cross-validate the resulting mood time series by comparing their ability to detect the public's response to the presidential election and Thanksgiving day in 2008. A Granger causality analysis and a Self-Organizing Fuzzy Neural Network are then used to investigate the hypothesis that public mood states, as measured by the OpinionFinder and GPOMS mood time series, are predictive of changes in DJIA closing values. Our results indicate that the accuracy of DJIA predictions can be significantly improved by the inclusion of specific public mood dimensions but not others. We find an accuracy of 87.6% in predicting the daily up and down changes in the closing values of the DJIA and a reduction of the Mean Average Percentage Error by more than 6%.



Younggue Bae , Hongchul Lee, A sentiment analysis of audiences on twitter: who is the positive or negative audience of popular twitterers?, Proceedings of the 5th international conference on Convergence and hybrid information technology, September 22-24, 2011, Daejeon, Korea
http://dl.acm.org/citation.cfm?id=1944594&CFID=76892041&CFTOKEN=90820401

Automated identification of diverse sentiment types can be beneficial for many NLP systems such as review summarization and public media analysis. In some of these systems there is an option of assigning a sentiment value to a single sentence or a very short text.In this paper we propose a supervised sentiment classification framework which is based on data from Twitter, a popular microblogging service. By utilizing 50 Twitter tags and 15 smileys as sentiment labels, this framework avoids the need for labor intensive manual annotation, allowing identification and classification of diverse sentiment types of short texts. We evaluate the contribution of different feature types for sentiment classification and show that our framework successfully identifies sentiment types of untagged sentences. The quality of the sentiment identification was also confirmed by human judges. We also explore dependencies and overlap between different sentiment types represented by smileys and Twitter hashtags.
http://dl.acm.org/citation.cfm?id=1944594&CFID=76892041&CFTOKEN=90820401


Effective sentiment stream analysis with self-augmenting training and demand-driven projection
http://dl.acm.org/citation.cfm?id=2009981&CFID=76892041&CFTOKEN=90820401
How do we analyze sentiments over a set of opinionated Twitter messages? This issue has been widely studied in recent years, with a prominent approach being based on the application of classification techniques. Basically, messages are classified according to the implicit attitude of the writer with respect to a query term. A major concern, however, is that Twitter (and other media channels) follows the data stream model, and thus the classifier must operate with limited resources, including labeled data for training classification models. This imposes serious challenges for current classification techniques, since they need to be constantly fed with fresh training messages, in order to track sentiment drift and to provide up-to-date sentiment analysis.

We propose solutions to this problem. The heart of our approach is a training augmentation procedure which takes as input a small training seed, and then it automatically incorporates new relevant messages to the training data. Classification models are produced on-the-fly using association rules, which are kept up-to-date in an incremental fashion, so that at any given time the model properly reflects the sentiments in the event being analyzed. In order to track sentiment drift, training messages are projected on a demand driven basis, according to the content of the message being classified. Projecting the training data offers a series of advantages, including the ability to quickly detect trending information emerging in the stream. We performed the analysis of major events in 2010, and we show that the prediction performance remains about the same, or even increases, as the stream passes and new training messages are acquired. This result holds for different languages, even in cases where sentiment distribution changes over time, or in cases where the initial training seed is rather small. We derive lower-bounds for prediction performance, and we show that our approach is extremely effective under diverse learning scenarios, providing gains that range from 7% to 58%.


Taketoshi Ushiama , Tomoya Eguchi, An information recommendation agent on microblogging service, Proceedings of the 5th KES international conference on Agent and multi-agent systems: technologies and applications, June 29-July 01, 2011, Manchester, UK

Bernard J. Jansen , Kate Sobel , Geoff Cook, Gen X and Ys attitudes on using social media platforms for opinion sharing, Proceedings of the 28th of the international conference extended abstracts on Human factors in computing systems, April 10-15, 2010, Atlanta, Georgia, USA

Tiffany C. Chao, Data repositories: a home for microblog archives?, Proceedings of the 2011 iConference, p.655-656, February 08-11, 2011, Seattle, Washington

Wei Wu , Bin Zhang , Mari Ostendorf, Automatic generation of personalized annotation tags for Twitter users, Human Language Technologies: The 2010 Annual Conference of the North American Chapter of the Association for Computational Linguistics, p.689-692, June 02-04, 2010, Los Angeles, California

Daniel M. Romero , Wojciech Galuba , Sitaram Asur , Bernardo A. Huberman, Influence and passivity in social media, Proceedings of the 2011 European conference on Machine learning and knowledge discovery in databases, September 05-09, 2011, Athens, Greece

Younggue Bae , Hongchul Lee, A sentiment analysis of audiences on twitter: who is the positive or negative audience of popular twitterers?, Proceedings of the 5th international conference on Convergence and hybrid information technology, September 22-24, 2011, Daejeon, Korea

Kamran Massoudi , Manos Tsagkias , Maarten de Rijke , Wouter Weerkamp, Incorporating query expansion and quality indicators in searching microblog posts, Proceedings of the 33rd European conference on Advances in information retrieval, April 18-21, 2011, Dublin, Ireland

Bernard J. Jansen , Kate Sobel , Geoff Cook, Being networked and being engaged: the impact of social networking on ecommerce information behavior, Proceedings of the 2011 iConference, p.130-136, February 08-11, 2011, Seattle, Washington

Dmitry Davidov , Oren Tsur , Ari Rappoport, Enhanced sentiment learning using Twitter hashtags and smileys, Proceedings of the 23rd International Conference on Computational Linguistics: Posters, p.241-249, August 23-27, 2010, Beijing, China

Claudia Wagner , Markus Strohmaier, The wisdom in tweetonomies: acquiring latent conceptual structures from social awareness streams, Proceedings of the 3rd International Semantic Search Workshop, p.1-10, April 26-26, 2010, Raleigh, North Carolina

Matthew Michelson , Sofus A. Macskassy, Discovering users' topics of interest on twitter: a first look, Proceedings of the fourth workshop on Analytics for noisy unstructured text data, October 26-26, 2010, Toronto, ON, Canada

Surender Reddy Yerva , Zoltán Miklós , Karl Aberer, What have fruits to do with technology?: the case of Orange, Blackberry and Apple, Proceedings of the International Conference on Web Intelligence, Mining and Semantics, May 25-27, 2011, Sogndal, Norway

Jingtao Wang , Shumin Zhai , John Canny, SHRIMP: solving collision and out of vocabulary problems in mobile predictive input with motion gesture, Proceedings of the 28th international conference on Human factors in computing systems, April 10-15, 2010, Atlanta, Georgia, USA

Takeshi Sakaki , Makoto Okazaki , Yutaka Matsuo, Earthquake shakes Twitter users: real-time event detection by social sensors, Proceedings of the 19th international conference on World wide web, April 26-30, 2010, Raleigh, North Carolina, USA

Sheila Kinsella , Mengjiao Wang , John G. Breslin , Conor Hayes, Improving categorisation in social media using hyperlinks to structured data sources, Proceedings of the 8th extended semantic web conference on The semanic web: research and applications, May 29-June 02, 2011, Heraklion, Crete, Greece

Jennifer Golbeck , Justin M. Grimes , Anthony Rogers, Twitter use by the U.S. Congress, Journal of the American Society for Information Science and Technology, v.61 n.8, p.1612-1621, August 2010

Jun Huang , Mizuho Iwaihara, Realtime social sensing of support rate for microblogging, Proceedings of the 16th international conference on Database systems for advanced applications, April 22-25, 2011, Hong Kong, China

Makoto Okazaki , Yutaka Matsuo, Semantic twitter: analyzing tweets for real-time event notification, Proceedings of the 2008/2009 international conference on Social software: recent trends and developments in social software, p.63-74, March 03-04, 2008, Cork, Ireland

Thomas Heverin , Lisl Zach, Twitter for city police department information sharing, Proceedings of the 73rd ASIS&T Annual Meeting on Navigating Streams in an Information Ecosystem, October 22-27, 2010, Pittsburgh, Pennsylvania

Pedro Henrique Calais Guerra , Adriano Veloso , Wagner Meira, Jr. , Virgílio Almeida, From bias to opinion: a transfer-learning approach to real-time sentiment analysis, Proceedings of the 17th ACM SIGKDD international conference on Knowledge discovery and data mining, August 21-24, 2011, San Diego, California, USA

Miles Efron , Megan Winget, Questions are content: a taxonomy of questions in a microblogging environment, Proceedings of the 73rd ASIS&T Annual Meeting on Navigating Streams in an Information Ecosystem, October 22-27, 2010, Pittsburgh, Pennsylvania

Anlei Dong , Ruiqiang Zhang , Pranam Kolari , Jing Bai , Fernando Diaz , Yi Chang , Zhaohui Zheng , Hongyuan Zha, Time is of the essence: improving recency ranking using Twitter data, Proceedings of the 19th international conference on World wide web, April 26-30, 2010, Raleigh, North Carolina, USA

Yegin Genc , Yasuaki Sakamoto , Jeffrey V. Nickerson, Discovering context: classifying tweets through a semantic transform based on wikipedia, Proceedings of the 6th international conference on Foundations of augmented cognition: directing the future of adaptive systems, July 09-14, 2011, Orlando, FL

Evangelos Kalampokis , Michael Hausenblas , Konstantinos Tarabanis, Combining social and government open data for participatory decision-making, Proceedings of the Third IFIP WG 8.5 international conference on Electronic participation, August 29-September 01, 2011, Delft, The Netherlands

Ismael Santana Silva , Janaína Gomide , Adriano Veloso , Wagner Meira, Jr. , Renato Ferreira, Effective sentiment stream analysis with self-augmenting training and demand-driven projection, Proceedings of the 34th international ACM SIGIR conference on Research and development in Information, July 24-28, 2011, Beijing, China

Bernard J. Jansen , Zhe Liu , Courtney Weaver , Gerry Campbell , Matthew Gregg, Real time search on the web: Queries, topics, and economic value, Information Processing and Management: an International Journal, v.47 n.4, p.491-506, July, 2011

Marti Motoyama , Brendan Meeder , Kirill Levchenko , Geoffrey M. Voelker , Stefan Savage, Measuring online service availability using twitter, Proceedings of the 3rd conference on Online social networks, p.13-13, June 22-25, 2010, Boston, MA

Zi Chu , Steven Gianvecchio , Haining Wang , Sushil Jajodia, Who is tweeting on Twitter: human, bot, or cyborg?, Proceedings of the 26th Annual Computer Security Applications Conference, December 06-10, 2010, Austin, Texas

Mike Thelwall , Kevan Buckley , Georgios Paltoglou, Sentiment in Twitter events, Journal of the American Society for Information Science and Technology, v.62 n.2, p.406-418, February 2011

Laurens De Vocht , Selver Softic , Martin Ebner , Herbert Mühlburger, Semantically driven social data aggregation interfaces for Research 2.0, Proceedings of the 11th International Conference on Knowledge Management and Knowledge Technologies, September 07-09, 2011, Graz, Austria

Twitter + イベント検出

Extracting Events and Event Descriptions from Twitter
http://www.www2011india.com/proceeding/companion/p105.pdf

This paper describes methods for automatically detecting events
involving known entities from Twitter and understanding both the
events as well as the audience reaction to them. We show that NLP
techniques can be used to extract events, their main actors and the
audience reactions with encouraging results.


Earthquake Shakes Twitter Users:Real-time Event Detection by Social Sensors
http://ymatsuo.com/papers/www2010.pdf


Streaming First Story Detection with application to Twitter
http://aclweb.org/anthology/N/N10/N10-1021.pdf

With the recent rise in popularity and size of
social media, there is a growing need for systems that can extract useful information from
this amount of data. We address the problem of detecting new events from a stream of
Twitter posts. To make event detection feasible on web-scale corpora, we present an algorithm based on locality-sensitive hashing which is able overcome the limitations of traditional
approaches, while maintaining competitive results. In particular, a comparison with a stateof-the-art system on the first story detection
task shows that we achieve over an order of
magnitude speedup in processing time, while
retaining comparable performance. Event detection experiments on a collection of 160 million Twitter posts show that celebrity deaths
are the fastest spreading news on Twitter.


Detecting Controversial Events from Twitter
http://www.marcopennacchiotti.com/pro/publications/CIKM_2010.pdf
Social media provides researchers with continuously updated
information about developments of interest to large audi-
ences. This paper addresses the task of identifying contro-
versial events using Twitter as a starting point: we propose
3 models for this task and report encouraging initial results.


以下、http://hashimomau.chips.jp/weblogs/?page_id=182より抜粋

題名:Why We Twitter: Understanding Microblogging Usage and Communities
出典:Joint 9th WEBKDD and 1st SNA-KDD Workshop ’07
概要・特徴:
twitterを分析した論文(初期なのでは??)
twitterユーザの世界分布、ユーザ数の増加など
twitterユーザの行動分析
ユーザは情報収集、日常の他愛もないつぶやき、他のユーザとのコミュニケーションをしている
HITSアルゴリズムを用いてコミュニティ分析もしている
感想:この論文は引用されることが多く、twitterを研究される方は取りあえず、読んでおくと良いです!
リンク:http://ebiquity.umbc.edu/paper/html/id/367/Why-We-Twitter-Understanding-Microblogging-Usage-and-Communities
題名:TwitterRank:Finding Topic-sensitive Influential Twitterers
出典:WSDM 2010
概要・特徴:
Twitter内の有力なユーザを抽出する方法を提案
TwitterRank:ページランクのアルゴリズムを拡張して実現。実験の結果より、他のアルゴリズムと比較して良い評価を得た
Twitterユーザに“homophily”が存在することを示し、“フォロー”関係に趣向(話題)の類似性があることを証明した
感想:フォロー関係に話題の類似性があることを学術的?統計的に示した論文。技術系論文では初らしい。PageRankを用いているため、有力ユーザ(有名ユーザ)の検出には効果がありそう。しかし、どの辺がTopic-sensitiveなのか分からなかったです。
リンク:http://www.mysmu.edu/staff/jsweng/papers/TwitterRank_WSDM.pdf


題名:Predicting Elections with Twitter: What 140 Characters Reveal about Political Sentiment
出典:Proceedings of the Fourth International AAAI Conference on Weblogs and Social Media 2010
概要・特徴:
twitter×選挙(政治)の関係を示した論文
twitterの投稿数とドイツ議会選挙との関係を調査した論文(投稿数から選挙結果を予想している)
LIWC(投稿の感情を分析するツール)を用いてセンチメント分析し、当時の政治状況と考察している
感想:投稿数から選挙結果を導けるかどうか疑問。しかし、選挙とtwitterの投稿を自然言語処理を用いて導いていると論文は初?投稿(ドイツ語)→英語に変換→LIWC(投稿の感情を分析するツール)を用いている。これはドイツ語と英語の互換性が高いから出来る事。日本語じゃ無理ぽ。 ※日本だとWEBと政治の関係を示した論文が極端に少ない気がします。
リンク: http://videolectures.net/icwsm2010_sprenger_pet/
題名:Tweet the Debates Understanding Community Annotation of Uncollected Sources
出典:WSM’09
概要・特徴:
アメリカ大統領選挙中の党首討論の議論の流れと、同時期のtwitterの議論の流れを解析した研究
twitterの投稿から動画のアノテーションを試みる研究
党首討論はTV放送で中継されており、中継中の字幕放送を議論の流れとしている
twitterの投稿は、ハッシュタグを使用しているtweetなどを利用している
感想:投稿からアノテーションをする研究は興味深かったです。
リンク: http://videolectures.net/icwsm2010_sprenger_pet/
題名:Microblogging During Two Natural Hazards Events: What Twitter May Contribute to Situational Awareness
出典:CHI 2010: Crisis Informatics
概要・特徴:
マイクロブログの性質を利用して災害時における情報伝播のメカニズム?を調べた研究
実際の災害時におけるデータを使用している
the Oklahoma Grassfires of April 2009 and the Red River Floods that occurred in March and April 2009
感想:災害時にマイクロブログのリアルタイム性の高さが役つということが分かった論文。具体的な手法は覚えてないが、そのような結論だったはず。マイクロブログの特徴が社会に役に立つと思った論文でした。
リンク: http://delivery.acm.org/10.1145/1760000/1753486/p1079-vieweg.pdf?key1=1753486&key2=5717842921&coll=DL&dl=ACM&CFID=2537635&CFTOKEN=65603928
題名:Twitter mood predicts the stock market.
出典:今後どこかに投稿するらしい 2010現在
概要・特徴:
Twitterの投稿からNYダウ平均株価(DJIA)のアップダウンを87.6%の精度で予測
解析対象
We only take into account tweets that contain explicit statements of their author’s mood states.
“I feel”, “i am feeling”, “i’m feeling”, “i don’t feel”, “I’m”, “Im”, “I am”, and “makes me”.
上記解析対象をP/N(肯定/否定)判定器のOpinionFinder(OF)とセンチメント分析器のGoogle-Profile of Mood States(GPOMS)両者の差分値を示した
GPOMSは投稿を Calm, Alert, Sure, Vital, Kind, Happy に分類できる
また、両解析結果を2つの市場予想モデル?に適応している
Bivariate Granger Causality Analysis 、Self-organizing Fuzzy Neural Network (SOFNN) モデル
SOFNNモデルCalmに着目することで、 87.6%の精度を実現した

2012年1月6日金曜日

Twitter + 医療

Towards detecting influenza epidemics by analyzing Twitter messages
http://snap.stanford.edu/soma2010/papers/soma2010_16.pdf

Rapid response to a health epidemic is critical to reduce loss
of life. Existing methods mostly rely on expensive surveys
of hospitals across the country, typically with lag times of
one to two weeks for influenza reporting, and even longer
for less common diseases. In response, there have been
several recently proposed solutions to estimate a population’s health from Internet activity, most notably Google’s
Flu Trends service, which correlates search term frequency
with influenza statistics reported by the Centers for Disease
Control and Prevention (CDC). In this paper, we analyze
messages posted on the micro-blogging site Twitter.com to
determine if a similar correlation can be uncovered. We
propose several methods to identify influenza-related messages and compare a number of regression models to correlate these messages with CDC statistics. Using over 500,000
messages spanning 10 weeks, we find that our best model
achieves a correlation of .78 with CDC statistics by leveraging a document classifier to identify relevant messages.



Twitter Catches The Flu: Detecting Influenza Epidemics using Twitter
http://www.aclweb.org/anthology/D/D11/D11-1145.pdf



With the recent rise in popularity and  scale 
of social media, a growing need exists for 
systems  that can extract useful information 
from huge amounts of data. We address the 
issue  of  detecting  influenza  epidemics. 
First, the proposed system extracts influenza related tweets using Twitter API. Then, 
only tweets that mention actual influenza 
patients are extracted by the support vector 
machine (SVM) based classifier. The experiment results demonstrate the feasibility 
of the proposed approach (0.89 correlation 
to the gold standard). Especially at the outbreak and early spread (early epidemic 
stage), the proposed method shows high 
correlation  (0.97 correlation), which  outperforms the  state-of-the-art methods. This 
paper describes that Twitter texts reflect 
the real world, and that NLP techniques 
can  be applied to extract only  tweets that 
contain useful information.

TWITTER IMPROVES SEASONAL INFLUENZA PREDICTION
http://www.cs.uml.edu/~hachreka/SNEFT/images/healthinf_2012.pdf

2010年7月1日木曜日

[StreamTwitter] 2010年6月の Tweet 数解析

6月1日から6月30日までの Tweet 数がすべて取れたので、再度解析しました。前回と同様に日本語の Tweet のみを抽出しています。

結果として面白いのは、20214番目のデータ(1単位は1分)で通常のピーク数よりも3倍の Tweet 数が見られ、Bursty なデータが見て取れるということです。20214番目は、6月14日あたりということですが、これはワールドカップで、日本とカメルーンが対戦した日です。

解析ですが、NFS がボトルネックとなるので、st01 のみの4コアで計測し、3時間程度で解析が終了しました。st01-st08 まで8台あるので、入力データ及び出力データをうまくローカルディスクを利用しながら解析すれば30分台で解析できるはずです。



以下は、そのバースト時が起きた時刻を中心とした1時間のデータ。



以下、6月10日から6月30日までのグラフ。バーストが何回か起きていることが見てとれる。

2010年6月30日水曜日

[StreamTwitter] べき乗則に従う Twitter のリプライ回数の頻度

1ヶ月の Twitter の返信頻度を System S / SPADE で解析した結果、べき乗則に従うことが見てとれる。

X軸はリプライの回数。Y軸は頻度を対数にした数値。用いたデータは4月中のほぼ1ヶ月のデータ。最大返信数は197.平均は 3.874 回、中央値は 3 回。

2010年6月26日土曜日

Twitter ログ解析 4/1 から 4/26 の1分間単位の Tweet 数の変遷

Twitter ログ解析 4/1 から 4/26 の1分間単位の Tweet 数の変遷

最大値:811
最小値:1 (怪しい)
平均値:241.82
中央値:235
標準偏差:112.37







以下、任意の3日間を切り出したときのグラフ





以下、ヒストグラム。X軸は 1分間のTweet 数。

2010年5月17日月曜日

[StreamGraph] グラフアルゴリズムのライブラリ

Java 実装
  • Jung (Java Universal Network/Graph Framework) :グラフ(ネットワーク)構造の視覚化や分析を行うための、Java ベースのオープンソースライブラリ (Link)
  • JGraphT: http://www.jgrapht.org/
Python 実装
C++ 実装

2010年5月2日日曜日

Cytoscape

ネットワークの可視化ツールとして、統計言語の R の igraph や sna パッケージ、Processing による可視化などがあるが、Cytoscape はよく出来ていそうだ。元々バイオ系の可視化に作られたものだが、Twitter などの動的ネットワークの可視化にも使えるだろう。

2010年4月29日木曜日

Twitter: 1ヶ月で12億 Tweet

今年2月の Twitter トラフィック分析。1ヶ月12億 Tweet、1日単純平均4000万。更に、単純に秒間で平均すると毎秒463 Tweet といったところ。






http://royal.pingdom.com/2010/02/10/twitter-now-more-than-1-billion-tweets-per-month/




こちらは09年11月の分析。1時間当たりの Tweet 総数の最大値は10月27日米国東海岸時間8時の 184万(毎秒511)、最小値は56万程度(毎秒155)で、差は3倍程度。

2010年4月21日水曜日

Twitter ログ分析 Web インタフェース

Twitter の Streaming API を用いて収集したデータを元に、StreamCloud, StreamDS, StreamAlgo プロジェクトの基礎統計データを集めたいと思います. Streaming API は間引かれているので、絶対的な値は利用できませんが、少なくとも以下の2点の統計情報が取れればと思います。
  • Bursty 度合い: 定常状態とバースト時のトラフィックの比(相対的な比が保たれていれば良いのだが。。。)
  • 時間的周期性:  定常状態において、時間によるトラフィックの周期性があるが、最もトラフィックが時間帯と最も低いときとの違い
  • (べき乗則 (Power Law) も確認できれば, Load Shedding の戦略に役立つと思いますが。。。)
 基礎統計データの取得が主眼なので可視化する必要性はありませんが、Streaming API を用いて時々刻々取得しているかどうか確認するためにも、以下のような グラフ化ツールを用いて Web 上で見れても面白いでしょう。

Open Flash Chart : Flash でグラフを生成するオープンソースのツール。PHP が必要。

 GUI の要件としては、いつからいつまでどのくらいの間隔で取得するかどうかをユーザーが設定し(例えば、2010年3月1日0時00分~2010年3月31日23時59分の間で1時間毎, もしくは2010年3月1日の00時から24時までで5分毎、など)、再描画ボタンでグラフの再処理リクエストをサーバー側に送信。
 サーバー側の PHP スクリプトでは、パラメータで設定された情報を元に、Twitter のログデータから件数を取得(または事前に件数をあるファイル又は DB に書き込んでおいてもよい)。X 軸を時間、Y 軸をデータ件数として、Open Flash Chart のオブジェクトにセットし、グラフを生成。

2009年12月18日金曜日

老木君学会発表@インターネットアーキテクチャ研究会



インターネットアーキテクチャ研究会にて老木君の学会発表。会場は、東京タワーの目の前にある機械振興会館でした。

発表は非常に良かったと思いますし、こうやってひとつの仕事を論文としてまとめ上げる、世の中に公表することは大切なことだと思います。本当にお疲れ様でした!



 私はずっと学生時代より情報処理学会関連の学会にしか参加したことがなかったので、電子情報通信学会(通称: 信学会)に参加したのは初めてでした。

ひとつの収穫としては、Live E! プロジェクトの従事者たちとお知り合いになれたことでしょうか。Live E ! プロジェクトは、気象データや CO2 濃度を収集するセンサーを日本、世界中にばらまき、センサーネットワークを築き、高等教育や気象観測、防災などに役立てることを目指したプロジェクトです。東工大にも西7号館と西8号館の屋上に設置されたそうで、我々のストリームコンピューティングのターゲットアプリケーションになりそうです。

Live E! プロジェクトでは、非常に多くのセンサーをばらまくのが目標のようですが、1センサーの設置コストが最低でも6万ぐらいかかり、運用、保守が必要になります。私自身としては、このような形態は、結局、スケールしないのではないだろうかと思っています。例えば、ゲリラ豪雨などの突発的、かつ局所的なイベントに対しては、非常に密なセンサーネットワークの構築が必要ですが、このようなコストのかかるものだと検知できる範囲も限られるでしょう。

このような専用型センサーではなく、個人が既に持っている携帯電話や、既存のインフラを活用した方が、よりスケールに可能になるのは当然でしょう。例えば、自動車が携帯電話網とつながる ITS (Intelligent Transport System) があと数年で普及してと言われていますが、ワイパーの動きを検知して、降雨状態を把握するなどのアイデアが出されており、自動車会社と携帯会社が本気になれば、そのようなインフラがあっという間にできてしまいます。また、StreamTwitter のようにマイクロブログのような集合知を利用することで、より安価にかつスケーラビリティ高く実現できます。

ただし、やはり、Live E! のような、より精緻な気象観測を提供するような装置も結局、必要であり、相補的に使っていくのではないでしょうか。