Jumat, 15 Juli 2011

More Trends of the Greek Debt Crisis

Here are some more results on mentions of various Concepts being discussed in Greek Blogs about the Greek Debt Crisis. Using Text Analytics, thousands of Greek Blogs are being annotated on a daily basis with the purpose of identifying the frequency with which several aspects of the Greek Debt crisis are discussed.

First let's have a look at the trend line of the Indignant Citizens Movement :






We can see that there is a clear down-trend in the number of Blog Mentions. This is also supported by a very significant reduction of the total Tweets found for this subject.


Next let's see how the trend of the mentions of "Greek default" looks like in the past month :



 We notice a severe spike beginning from July 12th because several Blogs and News sites were having mentions on a possible "Selective Default" which could happen to Greece.  

Interestingly, the trend on mentions of a US Default is also rising in Greek Blogs but is found with a much smaller frequency :





Jumat, 17 Juni 2011

The Greek Debt crisis - Some Trends

Several friends and blog readers ask me very frequently on what i think about Greece and the problems that Greece has on a Social and Economic level. Since this is not a blog about Politics or the Economy i will try to give my point of view with some analytics added. 

 It is always interesting to know how people feel and what do they think about the economy,their future, the politicians and how the general sentiment is. Also of great importance is the trend of all opinions and/or sentiment as this is recorded in Blog posts and other Social Media sources.

Here are some examples from data that i collect on a daily basis, several times a day from Greek blogs. Hundreds of Concepts are annotated within thousands of Blogs entries and collected for further analysis.


The results that i will show here are for :

- the latest Government Reform

- words that communicate Negative Sentiment.

-The "Indignants Movement" : Citizens that do not agree with the practices of both 2 largest Greek political parties  during the past 30 years and spending cuts directed by the IMF.

- Debt Crisis

Let us begin with the trend of "Government Reform" which at the time of writing (17/06/11 - Note that date format is  DD/MM/YY) has just happened. Here is the trend of mentions :





Notice how during the previous days not many mentions were captured and how much the trend increases until June 17th were the reform took place.


Next, let's look at entries that communicate "economic default" and their trend :





Again notice how on previous days mentions of Greek default start to rise (starting from June 3rd) and gradually the trend appears to fade out (French and German leaders said they will back up Greek debt on June 17th). It was no surprise that on June 8th and 9th (yet more) Greeks rushed in Banks to withdraw their money.

Here is the Trend of "The Indignants" movement :



Notice dates May 29th-30th, June 5th, Jun 12th-13th. All of these dates are Sundays (or close to Sundays) which is the day that most people gather in Syntagma square to express their anger for the IMF and Government practices. The trend however appears to be falling but  this may well be changing in the next days. Time will tell.

How about the words that communicate Negative Sentiment? Here is the trend :





Negative sentiment words appear to be somewhat rising after 31/05 but are coming down to previous levels.

FYI,  words that frequently occur with the concept "Politicians" are : "leaders", "cheats", "traitors".


More on the next post.

Kamis, 16 Juni 2011

Apple Products on Twitter - A Text Analytics example


My presentation on the 7th annual text analytics summit was a tutorial in one of the methodologies one could use to analyze unstructured text. The sample consisted of 365000 tweets that contained keywords of Apple products and concepts such as iPad, iPhone, iPod, Apple Store, Mac, Steve Jobs and the goal was to get an understanding of what people where tweeting about each product or concept.

The first step is to use a text analysis toolkit (i used GATE) to annotate the tweets and identify which concepts and keywords occur within the tweets. But this is not always easy. Take the word Mac for example. According to the context, Mac could be a computer type,  a burger type, the MAC beauty products or Mac Arthur airport. So when a query sent to Twitter API that contains the word  Mac we end up with lots of erroneous information.

So one of the things that have to be done to ensure good results  is word sense disambiguation. We know for example that if a tweet contains a word such as fries, lettuce and/or salad then quite likely the word Mac that was also found within this tweet was about the Big Mac (even though the word Big may not be present). If we find the word Arthur next to the word Mac then the tweet is about the Mac Arthur airport, etc. Here is GATE in action, identifying different keywords and concepts in Tweets :







Now we can see which concepts and keywords appear frequently in Re-Tweets ('USER' denotes that a '@' was present in the Tweet, 'URL' that a URL link was found in the Tweet,etc)





 We can also see which words frequently occur with iPhone5 :

















Selasa, 26 April 2011

Event Detection: Analytics becoming more personal

Sentiment Analysis is a hot technology at the moment. Marketers are interested in the perception that consumers have about  a specific brand, product or service as this is found in unstructured text. Some people claim that Sentiment Analysis does not meet their expectations but also that it is not straightforward for a company to find the "right" solution.  Comparing different Sentiment Analysis solutions could prove a difficult task.

Marketers and Decision makers need insights with which they can make better decisions - They need both Reports and Intelligence. Therefore the question that always follows the finding that "Your product has a 35% negative sentiment in the past 10 days" is "Why".  Social Media Monitoring tools must also  provide actionable Intelligence. 

All this is important information as it shows why your Brand / Product / Service could be losing customers. You monitor what is being said, identify whether a negative or positive Sentiment Trend is declining or rising and take necessary actions accordingly.

One of the questions i often get is what other applications can emerge from using Text Analytics and Data Mining. With Text Analytics and Data Mining we can find behavior patterns on many levels and -assuming that information such as Tweets will keep coming- the understanding of consumers can  go to the next -and sometimes more personal- level.

One of these applications is Event Detection. I am not aware if Event Detection is provided by any tool at the moment but i believe that this type of analysis could become a next major source of consumer insights. But what exactly is "Event Detection"?

Since we are able to have a computer automatically identify whether a phrase contains positive, negative or neutral sentiment, perhaps we could use Text Analytics and Machine Learning to detect that a specific event has occurred to an individual from the Tweets that someone posted such as "i've just returned from holidays". But that's not all. We can mine for patterns of consumer behavior given the fact that an event has occurred. And that potential knowledge from such an analysis could be very powerful. Because apart from the emotions that a product / service / person generates, the same applies for events happening in our lives. These events and the emotions they create can sometimes change our lives and also drive our decisions. A logical next step is to collect several behavioral Data and use Data Mining to analyze this information.

I will discuss an example of using Event Detection towards the end of my presentation on the 7th Annual Text Analytics Summit in Boston this May along with the reasons for such an analysis being important and i am looking forward to the reactions.

The fact is that with more insights, privacy issues arise even more and I get an increasing number of people asking me about privacy. I was also interviewed by a major British newspaper last month on what companies can learn by applying "Super Crunching" on Tweets. I tried to show both worlds of "Super Crunching" but the truth is that consumer insights become more personal as companies understand the value of structured and (more recently) unstructured information. 

Selasa, 15 Maret 2011

Social Media Data and what analysts can do with it

It is worth looking at what having our lives "digitalized" means since all of the information currently generated from usage of Social Media is available for analysis : "Collective Intelligence" and "Behavior Mining" are terms that are becoming increasingly known.


But what exactly is Social Media Data? Here are some examples :

  • The number of followers you have on Twitter and number of friends on FaceBook.
  • The number of links you provide, groups you join, retweets you make and how often you talk with other friends / followers.
  • The number of re-tweets, FaceBook "likes", comments and views that a blog post generates.
  • The personal information you provide (such as Twitter Bio)
  • The concepts being discussed in Tweets and FaceBook walls.



By applying Predictive Analytics to all of this information an impressive number of applications arises such as :

- Analysis of your Twitter Bio and words that are contained in your Tweets. For example we can identify what do people stating in their Bio being "Computer Geeks" discuss more frequently (in terms of Electronic Brands, technology trends etc). (See more here)

- Analyze thousands of Twitter accounts and find words that could make a difference in your follower count. (It appears that you should  keep things positive -at least most of the time-. See why here).

- Identify best practices on how to use Social Media  : When to post your new blog post, which words and concepts to avoid writing about and ultimately what concepts (such as Personal Branding) you should focus on. ( See more here).

- Understand consumer behavior : What people liked, how they feel and what they would like to see in upcoming products and/or experiences. See this example on how different aspects of consumer behavior in shopping malls is "mined".

Note that these are just some examples. The list goes on.

There is no doubt that new exciting Social Media apps will become available. This in turn will produce even more Social Media data (such as ones that contain location information). Being able to combine Data Mining and Text Mining techniques to extract insights from Social Media Data will become a very  important skill to have.

Rabu, 16 Februari 2011

7th Annual Text Analytics Summit



I would like to say a few words about an upcoming major event for all of those interested in Text Analytics and its various uses in Social Media, Marketing and Business Intelligence. Starting on May 18th, the annual Text Analytics summit (the only conference dedicated completely to Text Mining) will take place in Boston, MA with a total of 28 speakers presenting material on applications of Text Analytics including :




  • Social Media Analytics
  • Sentiment Analysis
  • Voice of the Customer
  • Marketing
  • Semantics


Well-known names in the industry will be there ( Seth Grimes, Tom Anderson, Gregory Piatetsky-Shapiro, Ronen Feldman) as well as experts from companies such as SAS, IBM, Forrester Research, Attensity, Adobe, J.D. Power&Associates, Clarabridge and others.   


My presentation will be about Behavior Mining in Social Media using Text Analytics and i will be giving a step-by-step tutorial  on the analysis of data originating from Twitter regarding a major Electronics Brand in USA. More specifically :

  • I will show how Tweets can be transformed and then analyzed using various statistical NLP techniques and software. 
  • Discuss the various problems that are found when one wants to analyze Text data
  • Discuss and introduce new ways of seeking for valuable information and extracting insights when it comes to Mining Behavior.

Although the Case Study will be using data from Twitter, the techniques shown can be applied to any other Text Data such as those found in FaceBook, Blog posts, User comments, etc.

I am looking forward to seeing the work done by others, learning about successful applications of Text Analytics and the knowledge gained and also seeing the issues that professionals come across and how these are faced by them. 

Kamis, 10 Februari 2011

Forex Trading with R : Part 2

In the previous post the first steps were given for building the basis for trading forex. Now it is time to build the actual classifiers that  can give us future buy / hold / sell signals.

Assuming that everything is in working order and the instructions given in  the previous post were followed we can start building these classifiers.

First let's train a Neural Network. The following command trains a Neural Network and then applies the trained model on our test data and outputs the predictions for buy/sell/hold signals :

set.seed(134)
nn <- nnet(class~.,traindata, size = 3, rang = 0.1,decay = 0.001, maxit = 3000,trace="F")
table(actual=testdata$class,predicted=predict(nn,newdata=testdata,type="class"))

Note that a seed number was used.  You should either try different seed numbers (so that network weights are re-initialized) or omit the set.seed() directive. You should also experiment with other Neural Net parameters such as the number of iterations (maxit), the learning decay (decay), etc.


The confusion matrix shows us the necessary information for calculating TP, FP, TN,FN rates for each class (ie for each signal type).


Similarly we can train and test a Random Forest :


 rf.model<-randomForest(class~.,data=traindata,nodesize=40,importance=FALSE,mtry=3,ntree=100)
table(actual=testdata$class,predicted=predict(rf.model,newdata=testdata,type="class"))



Now let's train an SVM for our data. We can issue the following command : 

###train SVM
sv<-svm(class~.,traindata,gamma=0.01,cost=5,kernel="radial")

To see how the classifier did on the test set, we enter :

table(actual=testdata$class,predicted=predict(sv,newdata=testdata,type="class"))


Next we can try to optimize parameters of the SVM classifier as follows :


#find optimal values of Gamma and Cost for an RBF- SVM classifier
tuned <- tune(svm, class~., data = traindata,ranges = list(gamma = c(0.0001,0.001,0.05,0.1,0.2,0.3), cost = c(1,5,10,20,50,100,120,130)),tunecontrol = tune.control(sampling = "cross"),cross=10)


tuned


The first command uses 10-fold cross validation to identify the best gamma and cost parameters among some predetermined values. We then issue the command tuned to see which combination of parameters  gives us the lowest classification error. Knowing these parameters we can then use these parameters to train an SVM classifier and see how this model performs (as was shown previously).

Be aware of the following key points :


  1. Three sets of data should be used : Training, Test and Validation. The Validation set should not be a part of the optimization (=finding the best algorithm parameters) process.
  2. Make sure that you create classifiers for several time periods. Test the performance of any classifier according to the percentage of available data you use for training / testing / validation and the number of periods you use for the sliding window.
  3. Make also sure that once you have chosen your model, you use a correct way to test your system by simulating buy / hold / sell signals and taking under consideration all associated trading costs.