Senin, 13 Februari 2012

Text Analytics for Telecommunications - Part 2

In the previous post we have seen the problems that a highly inflected language creates and also a very basic example of Competitive Intelligence. The Case Study that i will present in the forthcoming European Text Analytics Summit is about the analysis of Telco Subscriber conversations on FaceBook and Twitter that involve Telenor, MT:S and VIP Mobile located in Serbia.

It is time to see what Topics are found in subscriber conversations. Each Telco has its own FaceBook page which contains posts and comments generated by page curators and subscribers. Each post and comment also generates "Likes" and "Shares". Several types of analysis can be performed to find out :

1) What kind of Topics are discussed in posts and comments of each Telco FaceBook page?
2) What is the sentiment?
3) Which posts (and comments) tend to be liked and shared (=generate Interest and reactions)?



For each FaceBook page post, an identifier is added to the post text which designates the origination page (either MT:S, Telenor or VIP Mobile) of the post. Prior the analysis of  a FaceBook Post which says "We want more promotions" we need to be aware that this text originated -for example- from the MT:S FaceBook Page and not Telenor's.

Identifying the topics discussed in Telco subscribers posts and comments has a number of benefits. We gain a better understanding in the areas that a Telco should focus on. If we find that the topic of INTERNET is on the top  list of discussions repeatedly then this is where a Telco should pay attention. If we find Network mentions  to be associated with a competitor Telco repeatedly (which most likely is not good) we can choose the right time to air commercials implying that we are constantly working for better Network coverage.  We can also identify how much "buzz' was created from a new marketing campaign or  a new phone offer and the sentiment associated with it.

Let's have a look at the first type of analysis, namely Topic Detection. Using Information Extraction we  identify the Topics mentioned in thousands of FaceBook posts and comments for a particular period. Here are the results :




In terms of user engagement on FaceBook, MT:S is the winner since in absolute numbers, its FaceBook page contains more posts and comments than the other two FaceBook Telco Pages. Notice the frequencies of other topics found such as SMILEYs, INTERNET, PROMOCIJA (= Promotion), MREZA (=Network), POSTPAID and ANDROID.

With the chart shown above we become aware of the distribution with which Topics are discussed on all three FaceBook Pages. We do not know what is being discussed for now but we know that subscribers  talked more about Internet, then Promotions (=PROMOCIJA), then Network (=MREZA) and so on.

Let's look at  what Topics exist with PROMOCIJA (=Promotion). In other words, which other Topics are found in FaceBook Posts when Promotions are mentioned? Here are the results :




Most posts collected that discuss about Promotions are actually posts found on the MT:S FaceBook page. Notice also that the presence of Topic HOCU (=I want) which tells us that subscribers simply state that they want new promotions. Here is what the picture looks like for topic INTERNET :




So Telenor is found more frequently in INTERNET mentions. However caution is required since we do not know if all of  these Topic distributions found and their associations with specific Telcos can be attributed to pure chance or not.

It is very important to be confident enough to communicate to any Telco that being associated with Network mentions or any other Topic  is -or is not- simply a random event.

Rabu, 25 Januari 2012

Text Analytics for Telecommunications - Part 1

As discussed in the previous post, performing Text Analytics for a language for which no tools exist is not an easy task. The Case Study which i will present in the European Text Analytics Summit is about analyzing and understanding thousands of Non-English FaceBook posts and Tweets for Telco Brands and their Topics, leading to what is known as Competitive Intelligence.


The Telcos used for the Case Study  are Telenor, MT:S and VIP Mobile which are located in Serbia. The analysis aims to identify  the perception of Customers for each of the  three Companies mentioned and understand the Positive and Negative elements of each Telco as this is captured from the Voice of the Customers - Subscribers.


By analyzing several thousands of Tweets and FaceBook posts and comments we can have a first glimpse of Competitive Intelligence. For example when we wish to identify which words frequently occur with mentions about postpaid packages this is what we find  :




Red boxes show Telco Brands - notice "mts" and "mtsa" which point to the same Telco, namely mt:s.  Blue boxes indicate similar words that should be merged.  From a first look at the results above we see that : 

a) mt:s is found more frequently when users mention PostPaid packages.

b) Telenor and VIP Mobile are not found as frequently as MT:S in PostPaid package conversations.

c) We see several  problems from insufficient pre-processing : Kredit and Kredita (=credit) should merge into one word, the same applies for telefona - telefon, internet - interneta and mts - mtsa.



Notice that we can perform the same High-level analysis for several Telco Topics such as Network, Billing, Customer Care, Promotions, Questions of subscribers and so on. The next task is to identify the reason(s) why MT:S was found to have more mentions about PostPaid packages. Note that at this point we do not know why this is so : It could be the fact that MT:S prices of prepaid packages are high, very cheap or something else is happening that needs to be identified.


The Serbian Language poses extra work because it is a highly inflected language : Even the ending  of  Brand names change according to the usage.  Consider the following :

U mts-u (at mts)
Sa mts-om (With mts)
Bez mts-a (Without mts)


It is evident that a highly inflected language explodes our feature space and for this reason R can come to the rescue with some success. We can use R for changing several synonyms to one word, removing (Serbian) stop words, removing URLs and performing several other pre-processing steps that are necessary prior to an extensive analysis. More on the next post.

Senin, 09 Januari 2012

Case Study : Competitive Intelligence for Telecommunications

Telcos are a good example of a fast moving business environment and a good candidate for using Competitive Intelligence analysis from Social Media sources. The Case Study involves three major Telcos located in an Eastern European Country and shows the results from the analysis of thousands of Tweets and FaceBook wall posts to understand the following :


- How subscribers perceive each Telco Brand? 

- Which information do subscribers tend to Re-Tweet and "Like" on FaceBook Wall Posts? 

- Which words and Topics are commonly found with Intense feelings / thoughts?

- Which topics are mostly discussed when subscribers compare two or more Telco operators?

- What do subscribers discuss about  Network Quality and Speed, Billing, Promotions, Marketing Events, Customer Care, TV Commercials etc.

- How do they prioritize these topics and which of them are interesting and why?  

- What do subscribers talk about in general (i.e without any Telco Brand being mentioned) regarding Internet speed, Charges and what would they expect to see more?

I will present the Case Study mentioned  above in the forthcoming 9th Annual European Text Analytics Summit in April in London - UK. The Case Study is an example of application of Text Analytics to a language for which currently no tools exist and thus all difficulties and possible solutions will also be discussed. Examples will be also given on analyzing information to different conceptual levels and how this technique provides even more insights in consumer behavior.

The following tools were used for the analysis : 

- GATE to annotate all Topics that occur within Telco conversations (such as "sms", "internet", "dropped call", "network","promotion") and for setting up Conceptual Levels.

- R for pre-processing Text and performing Text Classification, Topic Detection and Cluster Analysis.

- WEKA  for Feature Selection and Text Classification.

- Finally,  Java is used to manage the information that is generated from GATE such as  understanding how subscribers prioritize various Telco Concepts and Topics and also identify important phrases and/or words that frequently occur when these Topics are being discussed.  

Senin, 28 November 2011

New Insights from Text Analytics

Text Analytics has gained the attention it deserves in the past few years. Sentiment Analysis is perhaps the most frequently discussed type of analysis but there will be always new ways to analyze and gain insights from text data.  

Examples of new types of analysis -and they have a vast potential- are in my opinion two :  Sequence Detection and Concept Mining. I am not aware whether  these types of analysis are currently being implemented by any Text Mining practitioner at the moment and if there is one, feel free to add your comments below.


So what is Sequence Detection and Concept Mining ? Some examples :

Suppose that you receive several similar e-mails sent from customers as the one seen below :

"I have been trying repeatedly to solve my billing problem through customer care. I first talked with someone called  Mrs Jane Doe. She said she should transfer my call to another representative from the sales department. Yet another rep from the sales department informed me that i should be talking with the Billing department instead. Unfortunately my bad experience of being transferred through various representatives was not over because the Billing department informed me that i should speak to the......"

Currently Text Analytics software will identify key elements of the above text but a very important piece of information goes unnoticed. It is the sequence of events which takes place :

 (Jane Doe => Sales Dept =>Billing  Dept =>...)


Being able to detect the sequence of events is an important element in understanding customer interaction. In our example above, imagine the possibility of detecting similar sequences through thousands of e-mails or call center transcripts and running a sentiment analysis, a process which then could correlate sentiment with specific event sequences.

Next,  is the usage of Concept Mining (this is just a phrase i coined for this post) : Being able to analyze information to different conceptual levels. A very powerful technique indeed and let's see why this is so.

People that have attended the 7th annual Text Analytics Summit in Boston had the opportunity to listen to several presentations regarding Semantics. The discussions between experts from the Semantics Panel and the attendees revealed that people could not find Semantics practical for several reasons. Yet, in Semantics lies the power of being able to find patterns on different conceptual levels.

As a -very basic- example, if we use Information Extraction to annotate -say- the Tweets containing mentions of American Telcos we can tag each one as a more general category called TELCOS. We can also tag individual prepaid packages as a more general category called PREPAID_PACKAGES. By doing that we can then search for patterns in a more general conceptual level than searching for patterns only at a Telco Brand level or a specific Telco's prepaid package. As an example we can run  a sentiment analysis on all prepaid packages mentions,  identify patterns of negative or positive sentiment and see which Telco is the winner of positive sentiment at a conceptual level.

The possibilities are endless.

Senin, 19 September 2011

Big Data : Case Studies, Best Practices and Why America should care

We know that Knowledge is Power. Due to Data Explosion more Data Scientists will be needed and being a Data Scientist becomes increasingly a "cool" profession. Needless to say that America should be preparing for the increased need for Predictive Analytics professionals in Research and Businesses.

Being able to collect, analyze and extract knowledge from a huge amount of Data is not only about Businesses being able to make the right decisions but also critical for a Country as a whole. The more efficient and fast this cycle is, the better for the Country that puts Analytics to work.

This Blog post is actually about the words and phrases being used for this post : All words and phrases on the title of the post (and the introductory text) were carefully selected to produce specific thoughts which can be broken down in three parts :

  •  Being a Data Scientist has high value. 
  • "Case Studies" and "Best Practices" communicate to readers successful applications and knowledge worthwhile reading.
  • "America should". This phrase obviously creates specific emotions and feelings to Americans.

"Case Study" and "Best Practices" were phrases found to be commonly associated with posts of high visibility. You might also get many views if you create a post which proves that whatever concept you are writing about is the right thing to do (for example write a post that clearly demonstrates yet another reason to use Social Media and have this post shown to Social Media Professionals).  Regarding our example : It is very probable (and logical) for Data Miners to look at and then re-tweet (or otherwise share) information which is a "proof" about Data Mining being useful  and also a "cool" profession. The higher concept / motive which works behind the scenes is that "I am doing the right job and this post proves it".

You might also get many views by submitting a post which disproves well-accepted concepts or posts that demonstrate the difficulties that well-accepted concepts face : For example, if you were a Data Scientist or a BI Professional, you would be inclined to read a post titled "Big Data is a Big Hype".  Whether you will re-tweet or share the post is of course under your discretion. At this point it should be noted that there is a big difference between number of clicks of a post and the number of shares it got (by Retweeting it, Liking it, etc) because sharing a post means that this post is considered worthwhile to read.

All of the above (and much more) have been found by analyzing thousands of Blog posts along with their number of clicks and shares they got (either by RT's , FaceBook "Likes", etc) and this is what i will be presenting in Text Analytics World in New York this October. It was also very interesting to see that some findings are in tandem with findings discussed by Joseph Carrabis during the Text Analytics Summit 2011 in Boston back in May. 

Of course it is not suggested  that   by using specific words and phrases you are guaranteed a successful post being re-tweeted from thousands of people and there are many reasons for this which i will not get into here. Additionally, Text Analytics cannot infer the higher meaning and concepts suggested within Text and this problem deserves a post on its own. This analysis however identifies concepts and/or phrases that point Bloggers and Marketers to look at a specific direction and with this knowledge to have increased probabilities for a successful Web presence. Again, this is an example of true Social Media Intelligence. Not (just) Reports.

So, in case that this post title immediately got your attention from other posts, you've just had a little taste of Predictive Analytics in action.

Jumat, 09 September 2011

Do Social Media Monitoring tools provide True Intelligence?

Having recently read a report from WebLiquid one of the interesting facts to consider is that around 70% of Marketers replied finding the insights gleaned from Social Media Monitoring tools "Somewhat Valuable".  Slightly more than 20% of them found these insights "Extremely valuable". The report also shows that most Marketers plan to invest more in SMM tools with few of them retreating from any further investment.

This is Big News. 70% of Marketers finding insights gleaned from SMM tools "Somewhat" valuable is not a good thing and perhaps there are reasons for this. It would be very interesting to know what do Marketers consider Insights, how they prioritize those Insights and how easily they can act once they have those insights . The problem can be summarized in one sentence:


- Marketers do not want (just) Reports.


There is a lot of useful information provided by many Social Media Monitoring tools  : The number of mentions of a Brand (or Product or Service) per channel, which users talk frequently about your Brand  (and which of them are considered influential). Sentiment Analysis provides Marketers with the perception of a Brand but also the perception about competitive Brands leading to what is known as Competitive Intelligence. Perhaps Social Media Monitoring platforms have many types of metrics still to offer : For example, a potentially useful metric could be the ability to identify Consumer  Intentions ("I will definitely buy...") and how these intentions differentiate - such as "I would buy 'ABC' if it was cheaper" or "I would buy 'ABC' if i hadn't  purchased 'XYZ' already".


Notice that SMM tools provide metrics  : Number of mentions per channel, Top influential users, percentage of positive / negative / neutral sentiment and sentiment intensity, how mentions of a new product disperse through different social media channels, etc. 


But what is considered Intelligence in Social Media? Would someone identify as intelligence the fact that during the past 2 months there was an increase in specific Brand Mentions on Twitter but not on YouTube? Or is it Intelligence when we notice that there has been a decline in positive sentiment about a product?  All of this information is Reporting and Feedback. It is not meant that this is not useful information :  It is important to know what is happening and why.

So what True Intelligence is all about?

True Intelligence is about knowing how to successfully Promote and Market a Brand, Product or Service. To do that a Marketer wants to know the Best Practices : With Social Media Reports, Marketers know what is happening (a decline in positive mentions on our new smartphone) and why this is happening (a potential hardware problem). Social Media Analytics can identify the right strategies to make things happen. True Social Media Intelligence is about knowing which parameters (channels, number of mentions) are important in achieving a result. Is it important to have a product associated with intense (positive) sentiment? Or could it be more important to have a Product being highly associated with Rumors?

There is still a long way to go in terms of Insights from Social Media Monitoring tools. There are many processes and parameters that will eventually used for deriving more Insights and better Strategies. The answer to true Social Media Intelligence is the use of Predictive Analytics (Data and Text Mining)  applied to Social Data : One area that is currently untouched by most Social Media Monitoring tools.

Jumat, 15 Juli 2011

More Trends of the Greek Debt Crisis

Here are some more results on mentions of various Concepts being discussed in Greek Blogs about the Greek Debt Crisis. Using Text Analytics, thousands of Greek Blogs are being annotated on a daily basis with the purpose of identifying the frequency with which several aspects of the Greek Debt crisis are discussed.

First let's have a look at the trend line of the Indignant Citizens Movement :






We can see that there is a clear down-trend in the number of Blog Mentions. This is also supported by a very significant reduction of the total Tweets found for this subject.


Next let's see how the trend of the mentions of "Greek default" looks like in the past month :



 We notice a severe spike beginning from July 12th because several Blogs and News sites were having mentions on a possible "Selective Default" which could happen to Greece.  

Interestingly, the trend on mentions of a US Default is also rising in Greek Blogs but is found with a much smaller frequency :