Showing posts with label IBM. Show all posts
Showing posts with label IBM. Show all posts

Wednesday, July 27, 2011

Predictive Analytics Conference!

In my previous post, I had talked about "Debunking Myths about Watson" and now I am going to talk about Predictive Analytics World. So is there a connection? Well, one connection is Dr. David Ferruci who is an IBM Fellow and the principal investigator for the Watson/Jeopardy is one of the keynote speaker at Predictive Analytics conference. He also led the team who developed UIMA. Also, the other connection is that I had the opportunity to know Eric Siegal for some time, who is an expert in predictive analytics, founding chair of the Predictive Analytics World and is also a former computer science professor at Columbia University. Eric is very knowledgeable, practical and articulate about this area and really knows how to simplify this topic. The last reason is that line is blurring between so many of these technologies and we will continue to see overlap -- the effect will be that we will also see more and more reorganization in future enterprises. There is a always a difference between how vendors build categories for different technologies and how businesses view them.  

Predictive Analytics by definition is a business intelligence technology that produces a predictive score for each customer or a prospect. Some people in the industry consider analytics as a different discipline than business intelligence -- because conventionally BI is more about what happened in the past or what is happening now but analytics is why it happened and what will happen in future. Predictive analytics (one of the key areas of analytics among areas like data mining, forecasting, optimization, text analytics etc.) is becoming increasingly important as marketing is the main business driver behind this discipline. It was already being used for many years in the applications for fraud detection, credit scoring and insurance pricing. But now there is almost an explosion of this discipline in the areas of direct  marketing, customer retention, product recommendations, behavior-based advertising, email targeting and leads scoring. The academic term for predictive analytics is "machine learning". As Mark Twain said, "The art of prophesy is very difficult, especially with respect to future" but the good news is that  predictive analytics doesn't need to be very accurate to provide value. It can help you answer questions like "People who buy life insurance are probably more likely to buy a luxury sedan."  Probably, we know this already but it matters when you are dealing with large volumes of data about customers and their interactions with products and services. Also, if you know in advance, which of your customers are likely to leave you, you can take measures to hit only those customers with right campaigns to retain them. Business Intelligence doesn't get more actionable than that! Forrester research expects the growth to double within five years as ROI is very high.

You might consider going to Predictive Analytics World, October 16-21 in New York City (pawcon.com/nyc) which will give you deep dive in the Predictive Analytics. There is also a new conference Text Analytics World (tawgo.com/nyc), co-located with PAW NYC. You can get a 15% discount on the 2 Day Conference Pass by using this code:  PMNY11

Wednesday, June 8, 2011

Debunking Myths about IBM's Watson!


There are many instances of televised technological feats of human race which have managed to leave a lasting impression on us - In my opinion, probably, Apollo 11 landing on Moon on July 20th 1969 will always retain the number one spot - in terms of impact value at a given time. I don’t know who is the close second but IBM Watson’s televised Jeopardy challenge when it bested Brad Rutter, the biggest all-time money winner on Jeopardy!, and Ken Jennings, the record holder for the longest championship streak, will always have a place in history. Organizing these kind of challenges is not new to IBM - who can forget when Deep Blue, a chess playing computer developed by IBM, beat world champion Gary Kasparov in a controversial match on 11th May 1997.




A lot of great things have been written about Watson, named after the IBM's first President but also influenced by the name of Sherlock Holmes' assistant Dr. Watson, since the televised challenge. Though, at the same time, you will also find many tweets ridiculing, hopefully in a good humour, that it needed lessons in geography after the (in)famous "Toronto" answer. But if you understand even little bit of information retrieval, natural language processing, machine learning, knoweldge representation etc.. then you would have realized what an amazing accomplishment this is. Recently, I had the opportunity to meet Aditya Kalyanpur, one of the team members of the DeepQA project which built Watson. Rome wasn't built in a day! Similarly it took more than four years and roughly twenty five brilliant technologists to build Watson. There are many unknown facts which we will come to know in due course of time. for e.g. Did you know that from September 2010 through December 2010 Watson played 55 games against Tournament of Champion Jeopardy! players and won 71% of the games. These players represent some of the best Jeopardy! players in the world.

Aditya managed to debunk few myths about Watson and also highlighted the approach taken by the team to develop the software. I would like to share them with you:

Three prominent myths:

• Watson answers a question by changing the query to structured query; then it queries a structured knowledge base. This is not true at all. This approach is taken by traditional QA systems and is very brittle with poor domain coverage.

• Watson identifies the question type and generates candidates from precompiled list of instances. This is false. Watson relies on a radically novel open domain type “coercion” technique.

• It uses either structured or unstructured data analytics. This is not true either. Watson integrates information from both unstructured and structured data analytics.


Some of the notable things about the DeepQA project:

• It does deep analysis of a question by breaking it down into relevant components like the focus, key entities and relationships, classifying it into broad classes, requirements of special handling etc..In one such strategy the system identifies independent facts within a given question, poses these as new questions to the underlying QA system, and generates answer candidates supported by the facts. Candidates that have support from multiple independent facts reinforce each other to boost system confidence in them

• Hypotheses, which in the QA case, are potential answers to the question, are generated by the Search component, which retrieves content relevant to a given question from the large volume of local knowledge resources Watson can access. The sources of information for Watson include encyclopedias, dictionaries, thesauri, newswire articles, and literary works. Watson also used databases, taxonomies, and ontologies, for example, DBPedia, YAGO, WordNet.

• Candidate Generation component identifies potential answers to the question from the retrieved content. A variety of answer scoring algorithms are then applied using DeepQA’s pervasive probabilistic framework. Other than linguistic processing, taxonomic, geospatial, temporal, popularity and source reliability are some of the evidence dimensions used by the scorers to constrain or support the right answer.

UIMA-AS was used for orchestrating the overall processing and an in-memory implementation of Sesame was used for storing RDF data. Initially, it used to take 1-2 hours to answer a question but using a massively parallel architecture and exploiting more than 2800 P7 cores, QA time was reduced to a few seconds.

• Machine learning and Monte Carlo methods used by the strategy components to estimate and optimize the win probabilities for the various players in a particular game state



Thursday, April 7, 2011

Elsevier Challenge: Another Big Step towards Open Data Movement

Nobody was happy after hearing that Data.gov, along with a number of other data-related sites of the government such as USAspending.gov and Apps.gov, are slated to be shut down due to budget cuts. The current annual budget of $37 million will be reduced to $2 million. It wasn't long ago when I had written very enthusiastically about the open data movement. In general, despite the fate of data.gov, the open data initiative is still going strong. Today, we have almost 25 cities in US who have opendata. Sanfrancisco's datasf.org is another success story which has almost 60 applications built by the developers. What we really need, in this context, is more participation from the commercial world!

On a similar note - Today, I was contacted by Elsevier, one of the largest publisher of medical and scientific literature in the world about their open data initiative. I am more than happy to write about this great initiative from a well known commercial enterprise.They just announced its first worldwide challenge called “Apps for Science,” powered by ChallengePost. The challenge is designed to bring together developers and a community of 15M researchers to collaborate more efficiently via new and innovative apps.

This database comprises of more than 25% of world’s academic and scientific articles for the challenge. Now, developers can access Elsevier’s data catalog and APIs from its SciVerse Suite, a content discovery platform + developer network w/ 10M+ articles, an abstract database with 41,000,000 records and more. This video explains SciVerse Suite better:






Why should we care for it?

  • If you care about a cure for cancer, AIDS or any other deadly disease then it is an important step. We need to help scientists and support them with the best tools and information possible because better communication is equal to greater knowledge share which will result into innovation breakthroughs.
  • It will free up approximately 12 hours per week previously spent on collecting and organizing research (according to a 2007 Outsell survey of 6,300 knowledge workers).
What’s in it for developers?

You can build & host tools (free or fee-based) for a captive audience of 15M researchers & 10k research institutions on Elsevier’s Application Marketplace.


Judges include known names like Jeff Jonas and James Handler among others. 

Thursday, November 11, 2010

Predictive Analytics for social data: Is there a role for Semantics?

Recently, Gartner, the leading technology analyst company, came out with its predictions for key technologies for 2011. The list comprises of cloud computing, mobile applications, social collaboration, next generation analytics, social analytics and many others - shouldn’t surprise you if work in information technology. The only thing which was not clear to me that why next generation analytics and social analytics are in two different categories when social collaboration is already emphasized as a part of a roadmap for large enterprises. I am sure they must be having their own valid reasons to do so. But the point is that overall it is getting a bit confusing about the various terms which are being used in context of analytics. What I mean here is what is the difference between just analytics versus predictive analytics versus forecasting versus predictive modeling versus optimization versus data mining versus advanced analytics. To many people, it sounds same! Also, it will depend upon who you ask this question. If you are a vendor or a consultant then try explaining it to a decision maker in an enterprise - basically, try not to get into that conversation. Many of these disciplines are more than a decade old as something like predictive modeling has been used in credit scoring for years. Also, academically, there is not much difference between classic techniques used in data mining and in statistics. Though, data mining has evolved to deal better with real life messy data. Unfortunately, unlike analytics, statistics could never become a hot topic but maybe it is about to change. In short, analytics or predictive analytics is the umbrella term or the new term - maybe the buzz word. It seems there is new surge of interest in predictive analytics because it is about the future outcomes in context of business intelligence. There is a difference between insights and gaining foresight!

You can always question that companies were always worried about the future outcomes so what has changed now if many of the methods were available before also. Probably, more data is available now, and there has been advancement and simplification of  tools/techniques - you can hire a good business analyst to do the job instead of someone with a doctorate in statistics. Predictive analytics enables you to develop mathematical models to help you better understand the variables driving success. Predictive analytics relies on formulas that compare past successes and failures, and then uses those formulas to predict future outcomes. Also, if you consider the fact that IBM has spent almost eleven billion dollar in the last five years acquiring software companies, like SPSS and Unica, for its analytics consulting organization then it starts making more sense.

On top of that, social analytics is a new kid on the block and there is new buzz that it is going to play a significant role in predictive analytics. I do believe that it will become true gradually but it is  not going to happen as quickly as we are being made to believe. What are the challenges in it? From the process perspective, predictive analytics is about understanding the prediction variables to the business problem, selecting the relevant statistical technique, validating the model with the test data and finally applying/adjusting the model iteratively with the production data. Do you think that it should be very different in context of social data? First of all, just because a company is doing brand monitoring (there are just too many companies), it doesn't make it a social analytics company. What I mean here is that if they are just following converations about an entity and don’t have much semantic intelligence in their software. If you look at most of the common examples of so called social media helping in predictions, they are about topics about election predictions (as recently claimed by Facebook political team) or about how a new product like iPad is perceived - the outcome is most of the time boolean i.e success or failure. In my opinion, these are interesting examples but much more is expected to do predictive analytics from social data. Maybe, if you consider social analytics and all associated prediction with it as a seperate or standalone discipline then it is good enough - but then you don't have an integrated view from an enterprise perspective. Maybe, in some cases, you don't need to integrate social data with enterprise data and still can get some value. But, you still need to build a repeatable predictive model using the social media data. And before you build the predictive model, you need to do true semantic analysis of the social data. Build some kind of normalized social data model to work with enterprise data for predictive analytics. IBM has come out with a new offerring where it claims to have enhanced SPSS software with social analytics and it can do predictive analytics for your business needs. It also offers semantic network analysis of the social data. I am not aware how well it works or if it can work across different business domains without too much business analysis or customization.

It is not an easy task to build a repeatable predictive analytics with a everchanging large volume of social data. The quality of meaningful data is also very important in this context. I see the ambiguity of social data as one of the biggest hurdle. You will have to do lots of preprocessing and deal with many new attributes being added in the social context. Every company is unique and we will see predictive analytics manifesting itself differently for each of them. Though, I do believe that we will see many companies building very high number of predictive models which will take social data into consideration for their business needs - maybe a predictive model per product and total turn around time of a week or less from problem definition to scoring. That brings up another question - even if real-time social data is present, can you take real time actions? If not then what is the true value of real time data in context of predictive analytics? You really need a very different level of infrastruture to take advantage of it.
 
This whole integration of social media analytics with predictive analytics should be owned by business - not by IT.  Infact, in most of the cases, it should be owned by marketing departments because they value/understand social data more than any other department and they are also most qualified to define prediction goals from a social context, predictive behaviors, pragmatic tradeoffs and evaluating results.

It seems there are some good opportunities in this space. We still have a long way to go. A lot needs to be understood before making any claims in the industry. There are definetely some good examples in the research community like SOMA, a forecasting model developed by a researcher at the University of Maryland, Terror Organization Portal. It analyses a wide range of information about politics, business and society in Lebanon to predict, with surprising accuracy, rocket attacks by the country’s Hizbullah militia on Israel. By the middle of 2010 SOMA was sucking up data from more than 200 sources, many of them newspaper websites. There is another example of two researchers at HP Labs who have established that they can use tweets to predict how well a movie will do - the results turned out to be fairly accurate. What we don't understand that can these examples be generalized for the indusry adaption? It will be good to know more examples from the industry.

Tuesday, February 9, 2010

New Patent from IBM for Semantic Web!

IBM has always been a leader in patent filing. Do you know that just in 2008, it filed more than 4000 patents which is three times more than its nearest rival HP. Recently, IBM filed a patent to improve traditional tag clouds by using semantic technology. Basically, the idea is that since tags are single words and users can't have description and context with a tag, it is very limiting and value of tag diminishes as the tag space grows. For e.g. picture tagged as "dog" will not show up when the user searches for content associated with tag "puppy." You can think of hundreds of similar examples. So, with the help of ontologies and associations, you can have more meaningful, descriptive and understandable tags. For example, once this tag cloud is represented in an ontology form then a "German Shepherd" can be classified as a type of dog with attributes like eye color, fur color etc. and relationships like "owned by". You can also specify that Puppy is a yound dog in this context.




The method, explained in the attached filing,  comprises of receiving a tag cloud which includes tags that hyperlink to web content. It will seperate the tag into different linguistics categories, assigning a weight to each tag, and grouping the tags into clusters, whereas tags in a cluster are associated with a context. The server will have components like linguistic analyzer, semantic domain analyzer, taxonomy builder, attribute analyzer, relationship analyzer and ontology generator. Like any onology generator, the process will be iterative in nature. As a result of this, it will eventually lead to more accurate searches for the content you are looking for.

All of it makes sense to me but I wonder if there are risks associated with companies in future, who might try to accomplish similar goals using different flavors of technology, without infringing on this patent. I will let patent lawyers figure this out in future.