Recently, Gartner, the leading technology analyst company, came out with its predictions for key technologies for 2011. The list comprises of cloud computing, mobile applications, social collaboration, next generation analytics, social analytics and many others - shouldn’t surprise you if work in information technology. The only thing which was not clear to me that why next generation analytics and social analytics are in two different categories when social collaboration is already emphasized as a part of a roadmap for large enterprises. I am sure they must be having their own valid reasons to do so. But the point is that overall it is getting a bit confusing about the various terms which are being used in context of analytics. What I mean here is what is the difference between just analytics versus predictive analytics versus forecasting versus predictive modeling versus optimization versus data mining versus advanced analytics. To many people, it sounds same! Also, it will depend upon who you ask this question. If you are a vendor or a consultant then try explaining it to a decision maker in an enterprise - basically, try not to get into that conversation. Many of these disciplines are more than a decade old as something like predictive modeling has been used in credit scoring for years. Also, academically, there is not much difference between classic techniques used in data mining and in statistics. Though, data mining has evolved to deal better with real life messy data. Unfortunately, unlike analytics, statistics could never become a hot topic but maybe it is about to change. In short, analytics or predictive analytics is the umbrella term or the new term - maybe the buzz word. It seems there is new surge of interest in predictive analytics because it is about the future outcomes in context of business intelligence. There is a difference between insights and gaining foresight!
You can always question that companies were always worried about the future outcomes so what has changed now if many of the methods were available before also. Probably, more data is available now, and there has been advancement and simplification of tools/techniques - you can hire a good business analyst to do the job instead of someone with a doctorate in statistics. Predictive analytics enables you to develop mathematical models to help you better understand the variables driving success. Predictive analytics relies on formulas that compare past successes and failures, and then uses those formulas to predict future outcomes. Also, if you consider the fact that IBM has spent almost eleven billion dollar in the last five years acquiring software companies, like SPSS and Unica, for its analytics consulting organization then it starts making more sense.
On top of that, social analytics is a new kid on the block and there is new buzz that it is going to play a significant role in predictive analytics. I do believe that it will become true gradually but it is not going to happen as quickly as we are being made to believe. What are the challenges in it? From the process perspective, predictive analytics is about understanding the prediction variables to the business problem, selecting the relevant statistical technique, validating the model with the test data and finally applying/adjusting the model iteratively with the production data. Do you think that it should be very different in context of social data? First of all, just because a company is doing brand monitoring (there are just too many companies), it doesn't make it a social analytics company. What I mean here is that if they are just following converations about an entity and don’t have much semantic intelligence in their software. If you look at most of the common examples of so called social media helping in predictions, they are about topics about election predictions (as recently claimed by Facebook political team) or about how a new product like iPad is perceived - the outcome is most of the time boolean i.e success or failure. In my opinion, these are interesting examples but much more is expected to do predictive analytics from social data. Maybe, if you consider social analytics and all associated prediction with it as a seperate or standalone discipline then it is good enough - but then you don't have an integrated view from an enterprise perspective. Maybe, in some cases, you don't need to integrate social data with enterprise data and still can get some value. But, you still need to build a repeatable predictive model using the social media data. And before you build the predictive model, you need to do true semantic analysis of the social data. Build some kind of normalized social data model to work with enterprise data for predictive analytics. IBM has come out with a new offerring where it claims to have enhanced SPSS software with social analytics and it can do predictive analytics for your business needs. It also offers semantic network analysis of the social data. I am not aware how well it works or if it can work across different business domains without too much business analysis or customization.
It is not an easy task to build a repeatable predictive analytics with a everchanging large volume of social data. The quality of meaningful data is also very important in this context. I see the ambiguity of social data as one of the biggest hurdle. You will have to do lots of preprocessing and deal with many new attributes being added in the social context. Every company is unique and we will see predictive analytics manifesting itself differently for each of them. Though, I do believe that we will see many companies building very high number of predictive models which will take social data into consideration for their business needs - maybe a predictive model per product and total turn around time of a week or less from problem definition to scoring. That brings up another question - even if real-time social data is present, can you take real time actions? If not then what is the true value of real time data in context of predictive analytics? You really need a very different level of infrastruture to take advantage of it.
This whole integration of social media analytics with predictive analytics should be owned by business - not by IT. Infact, in most of the cases, it should be owned by marketing departments because they value/understand social data more than any other department and they are also most qualified to define prediction goals from a social context, predictive behaviors, pragmatic tradeoffs and evaluating results.
It seems there are some good opportunities in this space. We still have a long way to go. A lot needs to be understood before making any claims in the industry. There are definetely some good examples in the research community like SOMA, a forecasting model developed by a researcher at the University of Maryland, Terror Organization Portal. It analyses a wide range of information about politics, business and society in Lebanon to predict, with surprising accuracy, rocket attacks by the country’s Hizbullah militia on Israel. By the middle of 2010 SOMA was sucking up data from more than 200 sources, many of them newspaper websites. There is another example of two researchers at HP Labs who have established that they can use tweets to predict how well a movie will do - the results turned out to be fairly accurate. What we don't understand that can these examples be generalized for the indusry adaption? It will be good to know more examples from the industry.
Showing posts with label semantic. Show all posts
Showing posts with label semantic. Show all posts
Thursday, November 11, 2010
Tuesday, March 23, 2010
XBRL Implementation Status: Doing a Reality Check
It was only few months back when I wrote the post - Semantic Technology, Financial Reporting and Toxic Assets. The post was more about how XBRL and the Semantic technology can play a big role in addressing issues related with toxic assets in the world of financial derivatives. It is becoming more clear now that XBRL is almost a movement which is going to have deep impact on how information between businesses, regulators and investors across the Internet will be communicated in the next decade. XBRL is not only considered the most revolutionary change in financial reporting since the first general ledger but also one of the most successful semantic web format. So what is the status of XBRL implementations? You can check this interview with Eric E. Cohen, co-founder XBRL, about the recent updates. You can also find some good information at Charles Hoffman's blog or http://xbrlplanet.org/. I also liked the section on XBRL in David Siegel's Pull where he talks about the history of XBRL.Some of the points which I found interesting are:
- XBRL is almost ten years old but the real adaption has started in the last few years only
- Regulators are one of the first group to adopt it. Regulators who are adopting XBRL are : Capital market regulators, tax offices, banking regulators, national statistic offices, corporate registrars
- Corporate America has already been complying with a mandate from the Securities and Exchange Commission for nearly a year to “tag” financial data in XBRL
- There have been issues in many projects but no significanlty failed XBRL project has been reported
- June 15th, 500 public companies did XBRL filings. Another 50 have started doing this even though they didn't have to file. 2009 taxonomies came late so 2008 taxonomies were used
- June 15th 2010, another 1500 will start filing in XBRL
- June 15th , 2011, another 10,000 will do
- One pending bill in Congress would direct all federal agencies to adopt XBRL for all requests for government bailout funds and all required reports on how those funds are used
- Edgar Online gets $12 million from Bain Capital for XBRL efforts
- Using XBRL, FDIC reduced the time to report information from 45 days to 2 days
- European Parliament is the largest government body who has expressed interest in XBRL
- The Securities and Exchange Commission launched its XBRL information portal, which can be found at http://xbrl.sec.gov/
- The investment in person hours that it took to create either the IFRS or the US GAAP taxonomies dwarfed the total hours needed to create XBRL itself
- Over 90% of Spanish banks now report in XBRL
- Holland and Newzealand are already accepting tax returns and other government required documents in XBRL
- Nevada is one of the first US state who is trying to use XBRL in many of its operations
- XBRL taxonomies may not be interoperable. For e.g, the US GAAP and the IFRS taxonomies are all used for financial reporting but are significantly different
- MIX (Microfinance Information Exchange) which collects information from more than thousand microfinance institutions is using XBRL. It is the first non-profit to use XBRL
- Data aggregators and distributors have not embraced XBRL in any significant way
- Taxonomy extensions is one of the hardest issue to solve as it reduces standardization and make interoperability very difficult
- You can find some common errors in XBRL implementations in this excellent article
- It is observed that few technologists have accounting domain knowledge and few accounting experts have technology domain knowledge in XBRL projects
- XBRL International, Inc. (XII), has released “XBRL: Towards a Diverse Ecosystem," a discussion document seeking public comment on the future business requirements and technical roadmap for the XBRL business information standard. The document may be downloaded here.
Tuesday, March 2, 2010
Semantic Search momentum continues: Netbase gets 9 million dollar funding
It was only six months back when Netbase, a semantic search company focused on Health and Consumer segments, went through lot of negative publicity after Techcrunch published this story titled "Netbase thinks you can get rid of Jews with Alcohol and Salt." NetBase Solutions’ HealthBase, a semantic search engine that aggregates medical content from millions of authoritative health sites including WebMD, Wikipedia, and PubMed was the center of this story when many users found some glaring glitches. It wasn't over for Netbase despite heavy criticism as it is evident from this news which say that it has raised $9 million in venture capital funding, completing its Series C round with Altos Ventures and Thomvest Ventures Inc --obviously, the VCs are seeing something which the media couldn't. In my opinion, these technologies are still a long way from being perfect and we will continue to see 20% to 40% errors - in many cases, it is dependent on the domain/content and the technology needs to be tuned for that.
The marketing message of Netbase is not very different from what you hear from many companies in the semantic search, natural language or text analytics space but they do have global brands like PandG and Elsevier as customers. The company also claims that by April about sixty percent of nurses around the world will be looking up information in its resources with the help of NetBase’s HealthBase Insight Discovery tool.
Overall, it is good news for semantic search companies who are trying to raise money as VCs believe that they have a good future and demand will only grow. I am sure they are aware of imperfections but can still see the value it can deliver to solve certain problems in a particular domain.
The marketing message of Netbase is not very different from what you hear from many companies in the semantic search, natural language or text analytics space but they do have global brands like PandG and Elsevier as customers. The company also claims that by April about sixty percent of nurses around the world will be looking up information in its resources with the help of NetBase’s HealthBase Insight Discovery tool.
Overall, it is good news for semantic search companies who are trying to raise money as VCs believe that they have a good future and demand will only grow. I am sure they are aware of imperfections but can still see the value it can deliver to solve certain problems in a particular domain.
Friday, February 26, 2010
Bill and Melinda Gates Foundation: Philanthropy goes Semantic
Bill and Melinda Gates Foundation is undoubtedly one of the most transparent and largest privately owned philanthropic organization in the world. The foundation's work ranges from providing vaccines to prevent childhood diseases to cutting edge research to improve agricultural yields and prevent malaria, among other things. Though, philanthropy is philanthropy and there is no big and small in that world as anybody who is doing it genuinely is a unique entity - you only need to travel outside US, Western Europe and few other developed nations to understand how desperately it is needed in rest of the world. Still, I believe that this particular foundation is very visionary in all its pursuits and focus areas if you follow them closely. Today, I was pleasantly surprised when I was contacted by Fenton communications who handles communications for them to write about this interesting press release on my blog. This is the bare minimum I can do as I understand the importance of what they are trying to accomplish - I have also been fortunate to work with few philanthropic organizations and I understand how valuable this concept can be for them.
The Bill and Melinda Gates Foundation is funding a new digital-media hub call ViewChange.org. The hub will use semantic technology to create a platform that combines the video sharing power of YouTube with the open information of Wikipedia and the mission of your favorite advocacy organization.
ViewChange.org is being created by a social change organization, making it one of the first time a non-profit is on the leading edge when it comes to technological innovation. They’re partnering with Zemanta, FreeBase and OpenCalais - the three well known companies in the semantics world.
Actor Danny Glover announced the launch of the project today via email and in a video.
ViewChange.org is using the power of semantic technology to make videos, articles, blogs, and actions readily available to people working in global development. While watching high-impact video stories on the site, viewers can choose to dig deeper by exploring up-to-date details on which organizations are involved, links to related content, and lists of relevant actions they can take.
For example, imagine you are watching a short documentary about clean water issues in India. As the video plays, adjacent windows will dynamically generate links to actions and media directly related to each scene. These could include organizations involved in clean water and sanitation, action campaigns related to water issues, relevant videos from YouTube, articles from research organizations, and the latest updates from news services and blogs.
This is a great cause and it will be immense help to anybody who is doing anything noble for any society. Please spread the word in any way you can.
The Bill and Melinda Gates Foundation is funding a new digital-media hub call ViewChange.org. The hub will use semantic technology to create a platform that combines the video sharing power of YouTube with the open information of Wikipedia and the mission of your favorite advocacy organization.
ViewChange.org is being created by a social change organization, making it one of the first time a non-profit is on the leading edge when it comes to technological innovation. They’re partnering with Zemanta, FreeBase and OpenCalais - the three well known companies in the semantics world.
Actor Danny Glover announced the launch of the project today via email and in a video.
ViewChange.org is using the power of semantic technology to make videos, articles, blogs, and actions readily available to people working in global development. While watching high-impact video stories on the site, viewers can choose to dig deeper by exploring up-to-date details on which organizations are involved, links to related content, and lists of relevant actions they can take.
For example, imagine you are watching a short documentary about clean water issues in India. As the video plays, adjacent windows will dynamically generate links to actions and media directly related to each scene. These could include organizations involved in clean water and sanitation, action campaigns related to water issues, relevant videos from YouTube, articles from research organizations, and the latest updates from news services and blogs.
This is a great cause and it will be immense help to anybody who is doing anything noble for any society. Please spread the word in any way you can.
Thursday, February 25, 2010
Where does my money go? A begining for an open government in US, UK and rest of the world
Once James Madison, the political philospher and the fourth president of America, said "If men were angels then no government would be necessary." Though, there are no clear historical records available about the first official government, democracy or parliament, despite various claims by few old civilizations, but I guess that men figured out much before James Madison that they can never be angels and they will always need a goverment to live happily and prosper. I am sure people just didn't want any government but also hoped and craved for a smarter, open and transparent government. But it seems nobody could define it clearly in last so many generations what openness and transparency really means for a government. The good news is that all of it is changing! Surprisingly, for the first time "data" is taking the lead in defining an open government - maybe because it is measurable and never lies.
The two big initiatives, data.gov and data.gov.uk (still in Beta), were launched by US and UK government respectively in May 2009 and January 2010. Infact, Prime minister Gordon Brown of UK asked Berners-Lee to look at access to government data in June, after Barack Obama's administration launched an open source data site. Well, UK already ranks number three in the OECD (Organisation for Economic Cooperation and Development) study behind Austria and Portugal in the sophistication of its e-services so making it data available was the next logical step. At high level, the goals of both the initiatives are same as they want to make the government data available online to general public for improved access; creative use of that data outside the walls of the government; public participation, collaboration and feedback; identify unexpected and insightful data relationships - insights that would normally take several decades and hundreds or thousands of brilliant socialist scientists, statisticians, psychologists, focus groups and public policy experts to simply suspect. Both these initiatives are great because of the intention behind them but lets take a closer look at the state of the initiatives.
The data.gov started with forty seven data sets but already has thousands of datasets from eighty-one US agencies. It links directly to data files in various formats including CSV, XML, Excel, and KML. A lot seems to be lacking though:
I am still not sure why Linked data (Semantic Technology) approach was not taken from the begining. Overall, data.gov in US has still a long way to go before its goals are met. Ideally, it will be great to see more impressive applications which uses data from different sources and gives you an insight about a specific problem. Nevertheless, it is moving forward - it is also understandable that managing and simplifying the process of publishing humungous data from so many agencies is a herculean task.
Now, when I look at data.gov.uk then I have to say that I am just simply impressed considering the progress they have made in six-seven months. Kudos to the team, along with Sir Tim Berners Lee, who has been working on it. They just used the semantic technology, basically linked data, approach from the begining. It also has a modern design with a very developer friendly approach. Combining data and creating mashups from different sources in this context is not an easy task but semantic technologies have made it possible. Overall, data.gov.uk's approach is simple and clear - they have used open standards, open source and open data. The website has quality and elegance written all over it even though it is in beta. They also need to figure out many things like modelling various datasets behind the scenes, encourage more participation and many other things you can think of in a project of this complexity. But the results are showing! The top ten application as rated in this telegraph article are impressive. You can see the screen shot of one of the application called "Where does my money go." Or click here to access the prototype.
If you are really keen to go deeper in the approach of UK govenment in this implementation then I will encourage you to read this document - Putting the frontline first:Smarter Government. UK might have followed US, Australia, New Zealand in implementing open and smarter government but it seems at least now that their template will be followed by rest of the world. It will be interesting to see if/when countries like India, China and Russia will follow this trend!
"This is very much the beginning. Hopefully, this is the tip of the iceberg. There is a whole lot more to do." But what a beginning! These were the words of Sir Tim Berners Lee when he launched the beta site for data.gov.uk.
The two big initiatives, data.gov and data.gov.uk (still in Beta), were launched by US and UK government respectively in May 2009 and January 2010. Infact, Prime minister Gordon Brown of UK asked Berners-Lee to look at access to government data in June, after Barack Obama's administration launched an open source data site. Well, UK already ranks number three in the OECD (Organisation for Economic Cooperation and Development) study behind Austria and Portugal in the sophistication of its e-services so making it data available was the next logical step. At high level, the goals of both the initiatives are same as they want to make the government data available online to general public for improved access; creative use of that data outside the walls of the government; public participation, collaboration and feedback; identify unexpected and insightful data relationships - insights that would normally take several decades and hundreds or thousands of brilliant socialist scientists, statisticians, psychologists, focus groups and public policy experts to simply suspect. Both these initiatives are great because of the intention behind them but lets take a closer look at the state of the initiatives.
The data.gov started with forty seven data sets but already has thousands of datasets from eighty-one US agencies. It links directly to data files in various formats including CSV, XML, Excel, and KML. A lot seems to be lacking though:
- It makes little effort to highlight or promote any projects that uses the data from the site
- The focus is more on a repository
- What you do with the data is not very clear
- The website needs lot of work in terms of clarity and user experience
- It is still not developer friendly and needs to develop an ecosystem
- There's no basic demographic data like population from the Census Bureau
- Browse and search functionality seems to be missing
- One of the application is about the amount of money received by the government for corporate and personal income taxes projecte through 2014 - click on the link to access it.
- If you want to know how knowledgeable is your state - click on the link to access it.
I am still not sure why Linked data (Semantic Technology) approach was not taken from the begining. Overall, data.gov in US has still a long way to go before its goals are met. Ideally, it will be great to see more impressive applications which uses data from different sources and gives you an insight about a specific problem. Nevertheless, it is moving forward - it is also understandable that managing and simplifying the process of publishing humungous data from so many agencies is a herculean task.
Now, when I look at data.gov.uk then I have to say that I am just simply impressed considering the progress they have made in six-seven months. Kudos to the team, along with Sir Tim Berners Lee, who has been working on it. They just used the semantic technology, basically linked data, approach from the begining. It also has a modern design with a very developer friendly approach. Combining data and creating mashups from different sources in this context is not an easy task but semantic technologies have made it possible. Overall, data.gov.uk's approach is simple and clear - they have used open standards, open source and open data. The website has quality and elegance written all over it even though it is in beta. They also need to figure out many things like modelling various datasets behind the scenes, encourage more participation and many other things you can think of in a project of this complexity. But the results are showing! The top ten application as rated in this telegraph article are impressive. You can see the screen shot of one of the application called "Where does my money go." Or click here to access the prototype.
If you are really keen to go deeper in the approach of UK govenment in this implementation then I will encourage you to read this document - Putting the frontline first:Smarter Government. UK might have followed US, Australia, New Zealand in implementing open and smarter government but it seems at least now that their template will be followed by rest of the world. It will be interesting to see if/when countries like India, China and Russia will follow this trend!
"This is very much the beginning. Hopefully, this is the tip of the iceberg. There is a whole lot more to do." But what a beginning! These were the words of Sir Tim Berners Lee when he launched the beta site for data.gov.uk.
Monday, February 22, 2010
After Best Buy, now Tesco adopts Semantic Technology!
It wasn't longtime back when I wrote how Best Buy is using Semantic technology to define a new trend. Now, it is Tesco , largest British retailer, who is following similar steps. Simply put, Tesco is to UK the same way Walmart is to US. Though, both of them get into each other's territory as it was almost an year back when Tesco launched Fresh and Easy,a chain of 10,000 square foot convenience stores in US to compete with Walmart.
Tesco has also been known as a company with a unique ability to manage vast reams of data and translate it into sales. It also uses information gathered from Dunnhumby, a British data mining firm of which it has majority control, to manage every aspect of its business, from creating new shop formats to arranging store layouts to developing private-label products and targeted sales promotions.
Tesco has always known to be pioneer in e-commerce and it is one of the most visited online supermarket sites in UK. They have started experimenting with RDFa in their website which seems like a logical next step. I won't be surprised if they adapt Goodrelations ontology also in near future.
Tesco has also been known as a company with a unique ability to manage vast reams of data and translate it into sales. It also uses information gathered from Dunnhumby, a British data mining firm of which it has majority control, to manage every aspect of its business, from creating new shop formats to arranging store layouts to developing private-label products and targeted sales promotions.
Tesco has always known to be pioneer in e-commerce and it is one of the most visited online supermarket sites in UK. They have started experimenting with RDFa in their website which seems like a logical next step. I won't be surprised if they adapt Goodrelations ontology also in near future.
Thursday, February 18, 2010
Cognition Technologies to power Microsoft's Bing now!
According to this news, Bing will also be powered by Cognition Technologies as Microsoft has licensed its technology for its search application. It is an interesting news because Congnition Technologies was always compared with Powerset who Microsoft acquired for $100 million plus in June of 2008. It is not replacing Powerset technology but adding more power to its capabilities. What is the relevance behind this non-exclusive licensing?
Though, Bing has made lot of improvements in the last year, to become more relevant in the search space and grab more share of the search market, it still needed to add more semantic capabilities in its arsenal. Bing is perceived to be good in following things:

Precision and Recall (the two criteria to measure the accuracy/relevancy of search results) is always a debatable topic if you ask any search company. Eveybody claims different results! As far as capability to undersatnd semantics/meaning of query is concerned, that is also debatable as you get mixed results if you test different queries with different search engines. Different techniques are used to really understand the meaning of a query by all semantic engines. But in case of Google, it gets its context from the humungous amount of data it indexes - they are processing so much data that they have lot of context around things like acronyms etc.. Suddenly, the search engine seems smart, like it achieved that semantic understanding, but it hasn't really. There have been instances where I have seen google nailing it quite accurately! So what is left to measure the effectiveness of search engine? Probably, this is the reason, the site traffic ends up becoming the only crietria for effectiveness or success. But, it takes years to build traffic and every improvement in the search engine counts.
I believe that Congition's real value propostion in this licensing is its advanced semantic map (think of it as a combination of dictionary, thesaurus, ontology etc.) which has millions of semantic connections that are comprised of semantic contexts, meaning representations, taxonomy and word meaning distinctions. Bing should benefit from this in making its query parsing more powerful and in improving relevance!
Though, Bing has made lot of improvements in the last year, to become more relevant in the search space and grab more share of the search market, it still needed to add more semantic capabilities in its arsenal. Bing is perceived to be good in following things:
- Clean interface and a good user experience
- Excellent information aggregation capability
- Extracting concepts and summarization
- Very useful for advanced search on Wikipedia and Freebase

Precision and Recall (the two criteria to measure the accuracy/relevancy of search results) is always a debatable topic if you ask any search company. Eveybody claims different results! As far as capability to undersatnd semantics/meaning of query is concerned, that is also debatable as you get mixed results if you test different queries with different search engines. Different techniques are used to really understand the meaning of a query by all semantic engines. But in case of Google, it gets its context from the humungous amount of data it indexes - they are processing so much data that they have lot of context around things like acronyms etc.. Suddenly, the search engine seems smart, like it achieved that semantic understanding, but it hasn't really. There have been instances where I have seen google nailing it quite accurately! So what is left to measure the effectiveness of search engine? Probably, this is the reason, the site traffic ends up becoming the only crietria for effectiveness or success. But, it takes years to build traffic and every improvement in the search engine counts.
I believe that Congition's real value propostion in this licensing is its advanced semantic map (think of it as a combination of dictionary, thesaurus, ontology etc.) which has millions of semantic connections that are comprised of semantic contexts, meaning representations, taxonomy and word meaning distinctions. Bing should benefit from this in making its query parsing more powerful and in improving relevance!
Tuesday, February 16, 2010
HP lists Semantic Technology among the top 10 BI trends for 2010!
It is always good to know that HP has listed Semantic Technology among the top 10 BI trends for 2010 in this new white paper. Ideally, the technology enthusiasts would have liked more details about the semantic technolgy as a trend in this white paper but please keep in mind that this is a business white paper. Meant for executives who will probably scan through it! It is great to have it atleast mentioned there! The key to acceptance of any technology by business is always simplified marketing message and validation by a brand like HP.
It is also good to have it categorized under BI umbrella by a global leader! HP has become a serious player in the IT services sector after its $13.2 billion acquisition of EDS. So far, the acquisition has been working out very well as it grew its 4th quarter revenue of 2009 by 8% to $8.9 billion. Even though HP services only represents 30% of total company revenues, services created 40% of the total operating profit dollars. When technology markets mature, the revenue and margins start becoming more service-centric. Software is another high-margin business for HP. Semantic Technology market is not matured so it will have both product and services opportunity in the long run. Hopefully, HP will proactively help in its widespread acceptance. Maybe, there will be more opportunity for tools and products vendors in semantic technology space who can develop alliances with HP.
Semantic Technology is not new to HP - the most significant thing that HP has done is that it provided Jena. It’s an open source framework for building semantic Web applications. It incorporates RDF and owl APIs; it also includes a rules-based inference engine, it includes in-memory and persistent storage for the data. It includes the SPARQL query engine, and it’s by far the most popularly used framework for developing applications.
In the end, it is always difficult to say whether semantic technology will become a trend in 2010 or later because most of the companies buy/adapt products and services based on their business cycles, not on the vendors' products and services roadmap.
It is also good to have it categorized under BI umbrella by a global leader! HP has become a serious player in the IT services sector after its $13.2 billion acquisition of EDS. So far, the acquisition has been working out very well as it grew its 4th quarter revenue of 2009 by 8% to $8.9 billion. Even though HP services only represents 30% of total company revenues, services created 40% of the total operating profit dollars. When technology markets mature, the revenue and margins start becoming more service-centric. Software is another high-margin business for HP. Semantic Technology market is not matured so it will have both product and services opportunity in the long run. Hopefully, HP will proactively help in its widespread acceptance. Maybe, there will be more opportunity for tools and products vendors in semantic technology space who can develop alliances with HP.
Semantic Technology is not new to HP - the most significant thing that HP has done is that it provided Jena. It’s an open source framework for building semantic Web applications. It incorporates RDF and owl APIs; it also includes a rules-based inference engine, it includes in-memory and persistent storage for the data. It includes the SPARQL query engine, and it’s by far the most popularly used framework for developing applications.
In the end, it is always difficult to say whether semantic technology will become a trend in 2010 or later because most of the companies buy/adapt products and services based on their business cycles, not on the vendors' products and services roadmap.
Tuesday, February 9, 2010
New Patent from IBM for Semantic Web!
IBM has always been a leader in patent filing. Do you know that just in 2008, it filed more than 4000 patents which is three times more than its nearest rival HP. Recently, IBM filed a patent to improve traditional tag clouds by using semantic technology. Basically, the idea is that since tags are single words and users can't have description and context with a tag, it is very limiting and value of tag diminishes as the tag space grows. For e.g. picture tagged as "dog" will not show up when the user searches for content associated with tag "puppy." You can think of hundreds of similar examples. So, with the help of ontologies and associations, you can have more meaningful, descriptive and understandable tags. For example, once this tag cloud is represented in an ontology form then a "German Shepherd" can be classified as a type of dog with attributes like eye color, fur color etc. and relationships like "owned by". You can also specify that Puppy is a yound dog in this context.
The method, explained in the attached filing, comprises of receiving a tag cloud which includes tags that hyperlink to web content. It will seperate the tag into different linguistics categories, assigning a weight to each tag, and grouping the tags into clusters, whereas tags in a cluster are associated with a context. The server will have components like linguistic analyzer, semantic domain analyzer, taxonomy builder, attribute analyzer, relationship analyzer and ontology generator. Like any onology generator, the process will be iterative in nature. As a result of this, it will eventually lead to more accurate searches for the content you are looking for.
All of it makes sense to me but I wonder if there are risks associated with companies in future, who might try to accomplish similar goals using different flavors of technology, without infringing on this patent. I will let patent lawyers figure this out in future.
The method, explained in the attached filing, comprises of receiving a tag cloud which includes tags that hyperlink to web content. It will seperate the tag into different linguistics categories, assigning a weight to each tag, and grouping the tags into clusters, whereas tags in a cluster are associated with a context. The server will have components like linguistic analyzer, semantic domain analyzer, taxonomy builder, attribute analyzer, relationship analyzer and ontology generator. Like any onology generator, the process will be iterative in nature. As a result of this, it will eventually lead to more accurate searches for the content you are looking for.
All of it makes sense to me but I wonder if there are risks associated with companies in future, who might try to accomplish similar goals using different flavors of technology, without infringing on this patent. I will let patent lawyers figure this out in future.
Saturday, February 6, 2010
Your new Virtual Assistant on iphone: Can it deliver?
It always surprises me how little is known about Siri, the virtual assistant, outside the Semantic Technology circle which is still very small community. Intuitively, everyone understands what a virtual assistant can do or deliver but most of the people still think that these technologies are not really ready for serious use. For those of you who don't know, Siri was born out of SRI's CALO Project, the largest Artificial Intelligence project in U.S. history. (CALO stands for Cognitive Assistant that Learns and Organizes). Made possible by a $150 million DARPA (Defense Advanced Research Projects Agency) investment, the CALO Project included 25 research organizations and institutions and spanned 5 years. Siri is bringing the benefits of this technology to the public in the first mainstream consumer application of a virtual personal assistant. The Siri application is just released for iphones and you can download it for free.
Siri - The Personal Assistant in your Phone from Tom Gruber on Vimeo.
It uses Semantic Technology for intelligent mash-ups that automatically make connections, takes action and communicates information based on dimensions such as personal data, theme or task awareness, time and location awareness much the same way a real live virtual assistant working the Internet could. It is almost like a very personalized semantic search focused on concierge-oriented tasks but can also take voice as input.
I downloaded the application on my iphone and it did manage to surprise me pleasantly. This technology has come a long way! I did face few issues as sometimes it struggled to understand what I am trying to say. Actually, it uses Nuance technology for voice recognition.I don't know how it will work outside United States - basically, lets say if you are in India, China or Russia then can this virtual assistant deliver the same help as it can do in US?
All of us need virtual assistant at some point in our professional and personal life - more when we think that time spent on some tasks is not worth our time. Ofcourse, we won't mind as long as this virtual assistant is working free for us and giving us some extra input. Will we trust it enough to do high-value transactions just based on its advice? Maybe buying a movie ticket is not a big risk and the prices are almost standard. But beyond that? It really comes down to confidence we need to develop gradually in our virtual assistant so that we can start delegating more and more tasks to Siri or other similar virtual assistants in future. Nothing different from level of complexity of tasks which we assign to out real life assistant - the trust needs to be earned everyday.
Release of Siri on iphone is still a very important milestone in a new era of AI and semantic technology-based applications.Can Siri build a formidable user base and become synonymous with the consumer Internet? It still has a long way to go and company will have to learn many things from its user base. Can it make enough money by just charging its affiliate network and giving the application to consumers for free? Time will tell this. I still think that the company can eventually make more money by making it more focused for business-oriented professionals. It just needs to think more about applying its technology to solve unique use cases. In the business scenario, it will still have to compete with other flavors of virtual assistants likeTimesvr, which is pronounced Time Saver. It is virtual assistant of different kind - an offshore-based online service which provides a task based (as opposed to assistant based) solution, where each task goes into a queue and is handled by whatever assistant (a real person) is available and qualified. It works surprisingly nicely and economically for few. Well, you can argue, it is not really fair to compare a top-notch AI/Semantic Tech. startup with an offshore-based service whose business model is basically labor arbitrage. But this is the reality of this flat world!You have to compete at every level and in all the fronts to make business sense!
Siri - The Personal Assistant in your Phone from Tom Gruber on Vimeo.
It uses Semantic Technology for intelligent mash-ups that automatically make connections, takes action and communicates information based on dimensions such as personal data, theme or task awareness, time and location awareness much the same way a real live virtual assistant working the Internet could. It is almost like a very personalized semantic search focused on concierge-oriented tasks but can also take voice as input.
I downloaded the application on my iphone and it did manage to surprise me pleasantly. This technology has come a long way! I did face few issues as sometimes it struggled to understand what I am trying to say. Actually, it uses Nuance technology for voice recognition.I don't know how it will work outside United States - basically, lets say if you are in India, China or Russia then can this virtual assistant deliver the same help as it can do in US?
All of us need virtual assistant at some point in our professional and personal life - more when we think that time spent on some tasks is not worth our time. Ofcourse, we won't mind as long as this virtual assistant is working free for us and giving us some extra input. Will we trust it enough to do high-value transactions just based on its advice? Maybe buying a movie ticket is not a big risk and the prices are almost standard. But beyond that? It really comes down to confidence we need to develop gradually in our virtual assistant so that we can start delegating more and more tasks to Siri or other similar virtual assistants in future. Nothing different from level of complexity of tasks which we assign to out real life assistant - the trust needs to be earned everyday.
Release of Siri on iphone is still a very important milestone in a new era of AI and semantic technology-based applications.Can Siri build a formidable user base and become synonymous with the consumer Internet? It still has a long way to go and company will have to learn many things from its user base. Can it make enough money by just charging its affiliate network and giving the application to consumers for free? Time will tell this. I still think that the company can eventually make more money by making it more focused for business-oriented professionals. It just needs to think more about applying its technology to solve unique use cases. In the business scenario, it will still have to compete with other flavors of virtual assistants likeTimesvr, which is pronounced Time Saver. It is virtual assistant of different kind - an offshore-based online service which provides a task based (as opposed to assistant based) solution, where each task goes into a queue and is handled by whatever assistant (a real person) is available and qualified. It works surprisingly nicely and economically for few. Well, you can argue, it is not really fair to compare a top-notch AI/Semantic Tech. startup with an offshore-based service whose business model is basically labor arbitrage. But this is the reality of this flat world!You have to compete at every level and in all the fronts to make business sense!
Wednesday, February 3, 2010
Measuring Semantics!
A lot has been written and discussed about the power and benefits of semantic technology. But still when it comes down to quantifying the benefits of semantic technology, you will find very few case studies where it is measured or some kind of ROI analysis has been done. It is always good to know about this case study by Telefonica which is one of the largest fixed-line and mobile telecommunications companies in the world: third largest in terms of number of customers only behind China Mobile and Vodafone, and in the top five in market value. They used semantic technology for their tariff calculation which is a very complex acitivity for an operator of their size. They could achieve 80% reduction in working hours and 75% reduction in errors for activities related with tariff calculation. If Semantic Technology has to become mainstream in the enterprises then we will continue to need more case studies like this.
Monday, January 25, 2010
Semantic Search: Finding Stuff and Creating more Businesses in this Flat World!
If you want more good jobs then spawn more Steve Jobs" says Thomas Friedman, the author of "The World is Flat " in this new article in NY times. Not everyone is a genius like Steve Jobs, can put 10,000 hours or is a part of 1955 club which were the key characterstics of these successful enterpreneurs as pointed by Malcom Galdwell in his interesting book "Outliers." Not every company becomes Apple also. Well, what US really needs is more of small businesses than ever which are one of the driving force behind its economy. There are some real facts about the small businesses in US:
We need more enterpeneurs than ever who can identify opportunities worldwide and develop it into profitable ventures. Where do you start? How do you get the information? How do I know about the gaps in the markets for particular products and services? What should be the focus area? Has it been done before? Who can partner with me? And there are so many questions you can think of. I undertand that you don't start every business just by searching on the web as there are other important things like personal contacts, capital, your network, your own experience, trade associations etc. etc.. But the search on the web is increasingly becoming the major starting point for many of these activities. It is more relevant than ever to do the research because everything you want to do can be outsourced; can be imported; maybe already exists somewhere; or demand is going to go away and you are blissfully unaware. Not that it is easy to figure this out but atleast we should have more resources than just the big three search engines - Google, Yahoo and Bing.
Try to find or research that information on these engines and you will know it is so hard. All three of them have done good jobs in the last many years but they can't continue to be all things to all people in all the contexts. Too much of emphasis has been on ranking also. I personally like google but it definetely falls short as far as exploratory and interactivity aspect is concerned. Bing, which calls itself decision engine, has shown some very good improvements in last one year. Also, somehow the big three have ended up promoting a marketing view of search on the web and they thrive on the tension created between SEO consultants/advertisers and them. It seems that media and analyst community are also too concerned with glorifying every percentage gain by Bing over Yahoo as you can see in this news and may others. Maybe, "its just not about search, its about business" - probably, Michael Corleane (from movie Godfather) would have said if he worked for Google in this era.
In general, I have seen a very narrow view of the search on the web from a school of thought which believes that whatever could be done in search will be just confined to these three as far as web is concerned - they think that new entrants will make some noise initially and then go away quietly. I completey disagree. In my opinion, it is no different from the view in eighteenth century when there was a school of thought which believed that whatever human beings can think of or can invent has already been done - there is no scope of anything new. Sounds ridiculous if you evaluate the progress mankind has made since then!
A lot has been written about the benefits about the Semantic Search and how it is better than the key-word based search. I have also written in one of my previous article about "semantics" in semantic search. Recently, Seth Grimes also compiled a very good article about types of semantic search. So I am not going to talk about what semantic search is but more about the opportunities for the new breed of semantic search engines.
Occassionally, you do see articles like "semantic search engines which will change the world which lists new breed of semantic search engines - some call them google killers. I always wonder why these semantic search have not been able to make measurable impact yet. Some of these engines have very good technology also though the list doesn't include many others which are out there. The internal details of most of these engines are still proprietary and they combine a natural-language processing with various flavours of semantics. One more semantic engine which I like is TipTop which mines Twitter database and does sentiment analysis also. But there are more than six Twitter search-based products in the market as you can see in this review - product from Tiptop is not even mentioned here while the other products may not be doing semantic search. Now, even Bing has a product which searches Twitter. So what can be the next step for these new semantic search engines?
In my opinion, the big issue is the lack of focus for some of these "semantic search" startups and their obsession to boil the ocean. Many of them waste lot of time comparing themselves with google. Some of them are also trying to do very similar things. Market will continue to get crowded with semantic search engines in next few years but there is a risk that many companies with excellent technologies will get lost in the crowd. Very soon, every search engine will start calling itself a semantic search engine, the same way every SAAS offering is a Cloud offering nowadays. The other issue is lack of understanding from business and users about what "semanticity" in search engines really means. Ideally, Semantic search engines should have some aspects of natural language, contextual (focus on disambiguating queries), ontologies and reasoning. The hard part is always developing the understanding how much of work is required to customize the technology to incorporate all these aspects so that it is relevant for a particular domain.
There is a great opportunity for these new breed of semantic search engines to rethink about their strategy. They should also not try to be all things to all people. They really need to carve out a space for themselves in specific segments. If they go after enterprises, they will face stiff competion from the big three in the enterprise - Sharepoint/Fast, Autonomy and Endeca who have customers in hundreds and have evolved over the years. Among them, Autonomy has done a great job in e-discovery space by following a vertical strategy. Even companies like Marklogic with its powerful XML server can solve many search related problems for unstructured content. In my opinion, semantic search startups can always continue to tweek their algorithms and enhance semantics but simultaneously the focus should be on verticalization, branding, strategy, positioning etc.. Application-centric or vertical strategy will be better for them as opposed to platform-centric strategy. They can also think about merging what is there on the web with the enterprise data/content to give extended BI inxights which is still a new area to develop powerful applications. Though I can count few small companies in this area also and even big ones like Business Objects after acquistion of Inxight. Still , there is ample oportunity to think creatively and develop useful analytics-based applications for enterprise.
- Employ just over half of the country’s private sector workforce
- Hire 40 percent of high tech workers, such as scientists, engineers and computer workers
- Include 52 percent home-based businesses and two percent franchises
- Represent 97.3 percent of all the exporters of goods
- Represent 99.7 percent of all employer firms
- Generate a majority of the innovations that come from United States companies
We need more enterpeneurs than ever who can identify opportunities worldwide and develop it into profitable ventures. Where do you start? How do you get the information? How do I know about the gaps in the markets for particular products and services? What should be the focus area? Has it been done before? Who can partner with me? And there are so many questions you can think of. I undertand that you don't start every business just by searching on the web as there are other important things like personal contacts, capital, your network, your own experience, trade associations etc. etc.. But the search on the web is increasingly becoming the major starting point for many of these activities. It is more relevant than ever to do the research because everything you want to do can be outsourced; can be imported; maybe already exists somewhere; or demand is going to go away and you are blissfully unaware. Not that it is easy to figure this out but atleast we should have more resources than just the big three search engines - Google, Yahoo and Bing.
Try to find or research that information on these engines and you will know it is so hard. All three of them have done good jobs in the last many years but they can't continue to be all things to all people in all the contexts. Too much of emphasis has been on ranking also. I personally like google but it definetely falls short as far as exploratory and interactivity aspect is concerned. Bing, which calls itself decision engine, has shown some very good improvements in last one year. Also, somehow the big three have ended up promoting a marketing view of search on the web and they thrive on the tension created between SEO consultants/advertisers and them. It seems that media and analyst community are also too concerned with glorifying every percentage gain by Bing over Yahoo as you can see in this news and may others. Maybe, "its just not about search, its about business" - probably, Michael Corleane (from movie Godfather) would have said if he worked for Google in this era.
In general, I have seen a very narrow view of the search on the web from a school of thought which believes that whatever could be done in search will be just confined to these three as far as web is concerned - they think that new entrants will make some noise initially and then go away quietly. I completey disagree. In my opinion, it is no different from the view in eighteenth century when there was a school of thought which believed that whatever human beings can think of or can invent has already been done - there is no scope of anything new. Sounds ridiculous if you evaluate the progress mankind has made since then!
A lot has been written about the benefits about the Semantic Search and how it is better than the key-word based search. I have also written in one of my previous article about "semantics" in semantic search. Recently, Seth Grimes also compiled a very good article about types of semantic search. So I am not going to talk about what semantic search is but more about the opportunities for the new breed of semantic search engines.
In my opinion, the big issue is the lack of focus for some of these "semantic search" startups and their obsession to boil the ocean. Many of them waste lot of time comparing themselves with google. Some of them are also trying to do very similar things. Market will continue to get crowded with semantic search engines in next few years but there is a risk that many companies with excellent technologies will get lost in the crowd. Very soon, every search engine will start calling itself a semantic search engine, the same way every SAAS offering is a Cloud offering nowadays. The other issue is lack of understanding from business and users about what "semanticity" in search engines really means. Ideally, Semantic search engines should have some aspects of natural language, contextual (focus on disambiguating queries), ontologies and reasoning. The hard part is always developing the understanding how much of work is required to customize the technology to incorporate all these aspects so that it is relevant for a particular domain.
Recently, Financial times launched Newssift (still in Beta) which is a business news semantic search engine which indexes thousands of news sources worldwide. They have used Endeca technology for faceted search and sentiment analysis is provided by Lexalytics. It can be a useful tool but more can be done in this area. If any of these new breed of semantic search engines can correlate data from historical sources and the one which is acquired from multiple sources to: identify patterns and indicate important events then it can be a killer application on Wall Street. Unstructured data is already being leveraged in electronic trading strategies but the adaption is not so fast. Generating alpha from the stream of unstructured data is not an easy task but a great opportunity.
Another set of innovative companies I want to mention in this context are Bintro ,Trialx and Echo Nest. Bintro matches you to what you are looking for like employment, partnerships, investment and joint ventures. TrialX is a free service that matches participants to relevant clinical trials based on their personal health information. TrialX uses a comprehensive database of 25,000+ clinical trials approved by the Food and Drug Administration (FDA) in the United States. Echo Nest helps you find your audience with targetted music production. It claims that it can understand every music writer on the web (bloggers, review sites etc.) and helps you find the writer most likely to review your music. Again, very useful way to leverage semantic technology to solve problems in a particular domain!
In the end, I believe that there is a scope for hundreds of similar applications for semantic search engines and they can happily coexist with Google, Yahoo and Bing. We will also see very interesting changes once the data web evolves and semantic markup starts becoming more prevalent.
Another set of innovative companies I want to mention in this context are Bintro ,Trialx and Echo Nest. Bintro matches you to what you are looking for like employment, partnerships, investment and joint ventures. TrialX is a free service that matches participants to relevant clinical trials based on their personal health information. TrialX uses a comprehensive database of 25,000+ clinical trials approved by the Food and Drug Administration (FDA) in the United States. Echo Nest helps you find your audience with targetted music production. It claims that it can understand every music writer on the web (bloggers, review sites etc.) and helps you find the writer most likely to review your music. Again, very useful way to leverage semantic technology to solve problems in a particular domain!
In the end, I believe that there is a scope for hundreds of similar applications for semantic search engines and they can happily coexist with Google, Yahoo and Bing. We will also see very interesting changes once the data web evolves and semantic markup starts becoming more prevalent.
Tuesday, January 12, 2010
Christmas Bomber: Connecting The Dots!
It is now official that it was the fault of software which almost got 289 people killed in the bungled Christmas day bombing. Couple of facts have emerged:
- The suspect, Umar Farouk Abdulmutallab, was added to a catch-all terrorism-related database when his father reported concerns about his son's radicalizations and associations. Though his name was not on flight watchlists
- A misspelling of Mr. Abdulmutallab's name initially resulted in the State Department believing he did not have a valid U.S. visa. It seems that his visa could not be revoked earlier because of it
- According to remarks by the President Obama:
- The intelligence community did not agrresively follow up on and priortize stream of intelligence related to possible attack.
- a failure to connect the dots of intelligence that existed across our intelligence community and which, together, could have revealed that Abdulmutallab was planning an attack.
- In sum, the U.S. government had the information -- scattered throughout the system -- to potentially uncover this plot and disrupt the attack. Rather than a failure to collect or share intelligence, this was a failure to connect and understand the intelligence that we already had.
It is useless to play the blaming game at this stage as written in this Economist article but what it means that we have to rely more on the intelligence of the software before a terrorist, with all valid documents, tries to board the plane. Yes, there are ways to detect the device at the airport as explained in the Scientific American article - adding unpredictable or layered security screening in future but that will always be a very costly solution if we have to implement in all the airports in this world. I am still not clear about this report released by National Research Council in 2008 which says that data mining is not the most effective way to smoke out terrorists. Yes, there can be issues of false positives which really means that a non-match can be declared as a match but it is always not the case as evident in the Christmas bomber's case.
To me, what really stands out is how/why we fail to connect the dots." as the suspect's name was in an international database indicating "a significant terrorist connection". It is clear that there is a strong need for superior knowledge discovery, database integration, cross-database search and the ability to correalte biographic information with terrorism-related information. I can't imagine doing any of these things without taking semantic technology into consideration. Infact, it should be one of the biggest drivers for any new initiatives in this context!
You might have heard story of David Headley (whose earlier name was Daood Gilani) - he is named as the key architect behind the Mumbai/India attacks in Nomberber 2008 in which 173 people died and 308 were injured. The residents of Mumbai were not as lucky as the passengers on the flight with Umar Farouk Abdulmutallab. David Headley is an American with a Pakistani father and also served as an agent for the Drug Enforcement Agency after being caught twice doing drug dealings. He was an operative of the Pakistan-based terror group Lashkar-e-Taiba. After his arrest by U.S. authorities, Indian officials discovered that he was given a long-term business visa for India. It is also alleged that he was already on a watch list which Indian authorties were not aware of. Indian authorities also say Headley traveled seamlessly between borders and stayed in various hotels in the same city while scouting for targets. What is more shocking is that he came back to India after the Mumbai attacks! This is another case of big failure to connect the dots! There can be many of these in the future also!
Semantic technologies can really help in connecting these dots because OWL/RDF can help build views in a more natural data graph format that is highly expressive and strongly deterministic. It is also more applicable in scenarios like this which places more premium on adaptiveness, agility, flexibility and grounded unambiguous level of truth. It is very useful when you really care to see end-to-end picture of how things are logicaly connected. The consistency can still be maintained while changing and asserting new facts! Inferencing is also a powerful capability which can unearth many new facts.
It is understandable that there are very complex protocols and policies involved in sharing of databases between various agencies around the world but not having the right technology shouldn't be an excuse. Because Semantic technology can be a very good solution to this problem.
I would really like to know your opinion about this. If you have new ideas/thoughts or you are aware of existing work being done in this area then please comment or write directly to me.
Thursday, January 7, 2010
"Pull" by David Siegel: Book Review
I just finished the final pages of "Pull" while watching the People's choice awards on TV. Who could have thought few years back that there will be a category for most popular "Web Celebrity" award which will be won by Ashton Kutcher for his more than one million followers on Twitter. Such is the power of web technology! Are you curious about the power of Semantic Web technology? Read "Pull" by David Siegel.
First of all, I would like to acknowledge the courage of David Siegel to write a business book on a difficult topic like this. The book is more about the power of Semantic Web to transform your business and is meant for business managers and enterpreneurs. It is not easy to write a "business book" on a topic like Semantic Web which has more sceptics than believers. In general, I have found that most of the good business books are more about analyzing the past and there are very few which are visionary or predict the future. In this complex world, it is so hard to see the future even beyond five years from now! So don't expect perfection! David makes a very good attempt in this direction.
Other than explaining the benefits of pull vs the push approach as practiced in most of the businesses, you will find some very useful information and statistics. The book is more conceptual in nature and talks about the recent developments in the world of semantic web and also takes you to the decade between 2020 to 2030. Some of you might question his timing also, if you are in a mood to question everything, but that is not that important from my perspective. Eventually the market forces dictate everything so why should we worry about it? You can always argue that he has become over enthusiastic about certain topics and is almost Utopian in its approach at some places but keep in mind that many of the concepts in this book are about a distant future. You may wish for more details or hope that he should have covered more domains but then he had to draw a line somewhere to make it readable for everyone. Overall, he has done a good job in doing gap analysis between the present state and the future state across various verticals but don't expect that you are going to get the perfect technology and process roadmap to achieve future state - the book is not about making incremental improvements but to find a completely new learning curve.
In the end, this is a thought provoking book which you should read with an open mind. It will definetely make you think! Without giving away too much about this book, I just want you to know that I really enjoyed reading it. It is a passionate work by someone who has put two years of life writing about a subject he believes in. The level of effort he has put in research also shows. Apart from that, for $18.45 on Amazon.com, this book is value for money. A real bargain! I will recommend all of you to read it.
First of all, I would like to acknowledge the courage of David Siegel to write a business book on a difficult topic like this. The book is more about the power of Semantic Web to transform your business and is meant for business managers and enterpreneurs. It is not easy to write a "business book" on a topic like Semantic Web which has more sceptics than believers. In general, I have found that most of the good business books are more about analyzing the past and there are very few which are visionary or predict the future. In this complex world, it is so hard to see the future even beyond five years from now! So don't expect perfection! David makes a very good attempt in this direction.
Other than explaining the benefits of pull vs the push approach as practiced in most of the businesses, you will find some very useful information and statistics. The book is more conceptual in nature and talks about the recent developments in the world of semantic web and also takes you to the decade between 2020 to 2030. Some of you might question his timing also, if you are in a mood to question everything, but that is not that important from my perspective. Eventually the market forces dictate everything so why should we worry about it? You can always argue that he has become over enthusiastic about certain topics and is almost Utopian in its approach at some places but keep in mind that many of the concepts in this book are about a distant future. You may wish for more details or hope that he should have covered more domains but then he had to draw a line somewhere to make it readable for everyone. Overall, he has done a good job in doing gap analysis between the present state and the future state across various verticals but don't expect that you are going to get the perfect technology and process roadmap to achieve future state - the book is not about making incremental improvements but to find a completely new learning curve.
In the end, this is a thought provoking book which you should read with an open mind. It will definetely make you think! Without giving away too much about this book, I just want you to know that I really enjoyed reading it. It is a passionate work by someone who has put two years of life writing about a subject he believes in. The level of effort he has put in research also shows. Apart from that, for $18.45 on Amazon.com, this book is value for money. A real bargain! I will recommend all of you to read it.
Wednesday, December 23, 2009
Mobile Internet report: some interesting insights
In this age of information overload, it is very easy to miss this comprehensive report by Mary Meeker of Morgan Stanley. It is a solid and very detailed work! Basically, the theme of the report is about the growth of mobile which will be much bigger than any of us can imagine. This new technology cycle is compared with what Windows 3 did for the PC in 1990 and Netscape browser did for desktop internet in 1995. And how the mobile internet has potential to create/destroy more wealth than prior computing cycles!
The report is well supported by solid research and numbers to back the analysis and forecasts. It took me some time to go through it but it was worth it as it gave me some interesting insights. In general, if it comes to mobile, we don't need to read a report to predict the future - we just have to look around and see what/how everyone is using these smart devices. I will skip the obvious like the spark created by iphone, 100k+ apps etc. but there are still few interesting facts such as:
The report is well supported by solid research and numbers to back the analysis and forecasts. It took me some time to go through it but it was worth it as it gave me some interesting insights. In general, if it comes to mobile, we don't need to read a report to predict the future - we just have to look around and see what/how everyone is using these smart devices. I will skip the obvious like the spark created by iphone, 100k+ apps etc. but there are still few interesting facts such as:
- 5 trends converging - 3G + social networking + video+ VOIP + impressive mobile devices
- 57 m iphone +163% growth, 125k developers worldwide, 2B+ downloads
- Many consumers are finding that their online usage rises dramatically when they have 24*7 mobile access to cloud based stuff
- Growth/monetization roadmap for mobile is provided by Japan
- Physical products are gaining share in mobile ecommerce - 20% in Japan but less than 1% in US
- China leads world in virtual good monetization
- Google (Android) has the best chance to to serve as a more open counter-balance to apple. Opera leading transformation of mobile browsers
- Professional content repository is still open - amazon, iphone, netflix, hulu?
- Open mobile web potentially more attractive to developers - Apple may believe its platform management is prudent way to ensure high quality content but it also runs the risk of stifling development and innovation
- Clouds will be providing real infrastructure for mobile applications
- Location aware ads will be better targetted
- People are more willing to pay for content on mobile than desktop
- AT&T - 50* mobile traffic growth in last three years despite only 40% subscriber growth over the same period
- Shift in type of applications from games, lifestyle, utilities, enterptainment etc. to business oriented though among the top 100 applications on iphone, business oriented applications are just five
- Daily usage for productivity based applications is probably less than 5-6 %
- Browser still remains the weakest link in these devices. There is no integrated/personalized experience.
- The principles of end user interaction have not been established
- Semantic search will become more relevant as users will have less tolerance for too many results.
- Semantic and location specific advertising will have a role to play
- Cloud computing will be one of the biggest enabler of writing semantic applications for mobile devices as computing power will no longer be an issue for even smaller shops
- Semantic web powered commerce can also be triggered
- Agent technology should find better business cases
Thursday, December 10, 2009
Online retail : How Best Buy is using Semantic Technology to define a new trend
This downturn has changed the behavior of consumers considerably in the retail sector. Though, the retail industry, just in US, has been more than four trillion dollars in last few years but it has been affected by this economy in last two years. Despite of this trend, 80% of retailers feel that the online retail channel continues to be better suited to withstand an economic slowdown better than other channels. It is validated by recent observations in the retail sector:
2. Jay also reported a 30 % percent (!) increase in traffic on the BestBuy stores pages, e.g. http://stores.bestbuy.com/1895
3. Yahoo observes a 15% increase in the Click-through-Rate (CTR). Nick Cox from Yahoo also recently reported that augmented search results, e.g. those with GoodRelations / RDFa in Yahoo get a 15 % higher Click-through-Rate (CTR).
So in short: While better visibility in traditional search engines is of course not the main intended effect by adding GoodRelations & RDFa rich mark-up, I think these findings are so substantial that any SEO / SEM consultant should apply it - now!
Black Friday spending rose 0.5%, ($54 million), to $10.7 billion, this year from last year. - Online sales have been up 17% (Thurs. to Sun.) over the same period last year.
- Cyber Monday has been up 11%, more than they did a year ago
- Online retail has been growing more than 20-25% in countries like Germany, Italy, UK
- Forrester Research projects that online retail sales will increase by 8 percent to $44.7 billion this holiday season
The channel is slowly maturing and with many of the easy wins now maximised, further progress will be much slower. In industries where consumer shifts as small as 1 percent can severely dent the profitability of brand, retailers now need to think more strategically about maximising revenue online. Textbook theory tells you that changes in the relationship between how much consumers are willing to pay, on the one hand, and their perception of the value they are receiving, on the other, underpins behavioral changes. There is also a strong trend where consumere are learning to live without expensive products and they are no longer willing to pay easily for premium brands. According to this article, the retailers need to be aware of following key trends:
- As acquiring new customers becomes more of a challenge, retailers should switch more marketing budget to maintaining existing customers and driving repeat business
- They must clearly communicate why customers should shop with them, and what extra benefits can be gained from doing so
- Providing clear, accurate and detailed information on products, prices and additional charges is a key
- Deep knowledge of your competitors’ online offerings coupled with sophisticated testing of different customer acquisition strategies will be crucial to stay ahead of the market
Till now, SEO, Search Engine Optimization, has been a successful strategy applied by online retailers to increase the traffic to their website but it has its own limitations - the links to the website doesn't communicate clearly why customers should shop with them and also doesn't provide clear information accurate and detailed information on products, prices and many other relevant things. GoodRelations ontology, which is just an year old, can fulfill those gaps and give retailers that extra advantage. It is a standardized vocabulary for product, price, and company data that can (1) be embedded into existing static and dynamic Web pages and that (2) can be processed by other computers. This increases the visibility of your products and services in the latest generation of search engines. I have explained some of the concepts of microformats and RDFa in one of my previous blogs but I will highly recommend you to visit website maintained by Martin Hepp, the creator of Good Relations Ontology. Martin has done a great job in coming up with a practical application of Semantic Technology which can deliver value. The best part is that adaption will eventually increase as the learning curve is so simple.
In his talk at the Search Engine Strategies 2009 conference in Chicago, Jay Myers, Lead Web Development Engineer for Best Buy, Co., Inc., reported very surprising effects of adding GoodRelations and RDFa to their products pages:
1. GoodRelations + RDFa improved the rank of the respective pages in Google tremendously. In fact, if you try the query "BestBuy Ferris Bueller" on Google, then the page comes on rank # 1 ahead of the much more established page . This indicates a strong effect of GoodRelations + RDFa on Google's appreciation of a page. It is particularly surprising since the age of a domain has now a huge influence on ranking in Google - older ones get a much higher ranking. In this case, the semantically augmented one is just eight weeks old but it is still ranked higher!
2. Jay also reported a 30 % percent (!) increase in traffic on the BestBuy stores pages, e.g. http://stores.bestbuy.com/1895
3. Yahoo observes a 15% increase in the Click-through-Rate (CTR). Nick Cox from Yahoo also recently reported that augmented search results, e.g. those with GoodRelations / RDFa in Yahoo get a 15 % higher Click-through-Rate (CTR).
So in short: While better visibility in traditional search engines is of course not the main intended effect by adding GoodRelations & RDFa rich mark-up, I think these findings are so substantial that any SEO / SEM consultant should apply it - now!
Kudos to Best Buy for showing leadership in early adaption of this technology! I don't know whether it is a tipping point or not but I do recognize that it is a very positive step in adaption of semantic technology by the retail industry! In the end, the most important thing is to give SEO experts and those who pay them an incentive to add rich meta-data now. I hope to see many other online retailers to join the bandwagon and take full advantage of this simple technology.
Tuesday, December 8, 2009
Microsoft Semantic Engine!
One of the highlights of the PDC09, an event focused on the technical strategy of the Microsoft developer platform, was the overview and demonstration of the Microsoft Semantic Engine. The Semantic Engine unifies search, structured querying, and analytics over structured and unstructured data.You can read some more details about PDC at the CTO's blog. It goes beyond existing components like Lucene by supporting both text and non-text, such as audio, video, and images. The key points are:
- Micorsoft has been working and investing heavily on this technology for the last two years
- The "Microsoft Semantic Engine" name is just a place holder
- It is not a W3C SemanticWeb(tm) approach but one which melds the unique capabilities of unsupervised machine learning (hierarchical clustering), information retrieval models (higher-dimensional vector spaces), pluggable and trainable classifiers (SVMs, Naive Bayesian, Maximum Entropy, Decision Tree, etc.), and personalized filtering and ranking.
- One of the goals is to make search semantically enhanced. Clustering the results based on Semantics is a key differentiator.
- You can expect to see the Microsoft Semantic Engine in one of the upcoming SQL Server Betas.
- This is a good move by Microsoft and was long overdue because there is a solid business case for integrating this technology in the enterprise.
- I view it more as a flavor of text extraction/analytics technology - infact, Msft has said it very clearly that it is not using the semantic web technology approach
- Having Microsoft in the Semantic technology space is a good thing for the Semantic Technology indutsry. This is a great validation from one of the most successful leaders in the software space
- They are not the only one who is trying this approach. As far as I understand, the goal behind Inxight's acquisition by Business Objects (now SAP) was the same one. I am not aware exactly how that integration has worked out
- I am sure Micosroft will come out with a clear message regarding the "Semantic Engine's" positioning in the enterprise in comparison to Fast Search Engine (Now part of Sharepoint division).
- It will intersting to understand if Powerset (now Bing) technology, Microsoft's semantic search, was used in this effort
- More details are needed to understand how it will work with disparate data sources
- Microsoft has been underestimated for too long as far as their search strategy is concerned. I agree that Google has a big lead in terms of number of users on the web but I really think that they have made all the right moves, both for web as well as enterprises,at least in last 2-3 years as far as search space is concerned. There are three great moves:
- Acquistion of Fast search and integrating it with Sharepoint (MOSS)
- Acquistion of Powerset and launch of Bing
- Plan to introduce Semantic Engine and integrating it with SQL serve
- Probably, services and product-based companies in the Semantic space, need to go back and revise their marketing message.
- In the end, it just validates my initial thoughts that Semantic technology is a superset of Semantic Web Technology. Microsoft's approach or Semantic Web technology (W3C) approach are two different approaches to solve Semantic related problems in the enterprise. This is the reason, I named my blog as Semantic Technology blog.
Tuesday, November 24, 2009
Finally, a National Semantic Technology Roadmap!
Probably, you might have assumed it that I am talking about the US but actually it is envisioned by the Malaysian government on their National Information technology Council (NITC) web site. More details would have helped but I will still give lot of credit to Malaysian government to start thinking in this direction and having the courage to recognize Semantic Technology on their official technology web site under the Ministry of Science, Technology and Innovation. Among other great things, the Malaysian governement also has the credit of creating one of the tallest towers in the world called Petrona Twin Towers. You have to be there to believe it and these towers have also become the epitome of astonishing growth in the last two decades in Malaysia.
While in US, sometimes, we still have to run sessions in conferences like - "Is Semantic Technology for Real? "
Friday, November 20, 2009
The Death of Taxonomies! Is there a role for Semantic Technologies?
- Since social computing is being widely adapted by companies so there is no fixed way of categorizing things i.e days of single, heirarchical taxonomies are far behind. People will tag content in any way they want.
- You should be able to get to information the way you want, which may be different from your colleague's approach.
- Text mining and auto-tagging software is gradually improving, and extracted terms can be applied as metadata. Metadata needs to be very fluid - cloud like. See the fig. above
- Metadata architects should really understand the domain
- What will be the future of existing large taxonomies?
- Many companies have already invested in more than one content management system like Sharepoint, Documentum or Opentext. There is a good opportunity for semantic web technologies, basically OWL and RDF, which can unify the taxonomies of different content management systems and provide a single data model to retrieve the content through.
- We should always be very clear about what the word "Metadata" means in any context. It means different things to different people and has been overused. A metadata in a document-centric world is entirely different from say a metadata in a word processing program.
- Semantic Web is really an umbrella language for metadata. More thought needs to be given to how this new trend of adding tags/metadata on the fly can be leveraged to add more value to the semantic web or linked data cloud.
Friday, October 16, 2009
Is it GE or Google as proxy for US economy? The role of Search in being a leading indicator!
We know that the best way to understand the state of US economy is by looking at GDP . It wasn't long ago, probably two to three years back, when GE was unanimously considered as proxy for US economy. It was always most diversified company as its businesses spanned from entertainment, medical devices, energy, locomotives, equipment, infrastructure services to finance. It always remained among the leaders in revenue and market capitalization. Unfortunately, it has fallen out of favor in last two years because of its average performance across the board in most of its businesses. It is no longer a leader in market cap. also. The result of all of this is that GE is not considered as a proxy for US economy or their sales pipeline as one of the leading indicator of economy.
I will leave it to economists to debate what is a true economic indicator but recently, Google has started claiming that its search patterns are saying that economy is recovering. Basically, Google's chief economist says that he can tell the direction of economy from American's search habits.
For e.g Google is seeing following trends:
You can always disagree with Google's claim and argue that a search for "property prices in a neighbourhood of Las Vegas" shouldn't be interpreted as recovery sign. Maybe, the person is trying to to sell the house! In the end, it is very difficult to interpet the true motive from few key words in the search. After all, economic forecast is a serious business and econometric modeling it is a very complex thing!
I think that we will continue to see this trend in future where "search on the web" will continue to indicate at least partially where the economy is heading. As search on the web moves from "keyword search" to more advanced "semantic search", we will have better idea about the intention of the query.
If we can analyze the "query intent" then we can build better models also. We also have to take into consideration the search patterns from Yahoo, Bing and other search engines to have a complete picture. I will still not dismiss off leading indicators from GE as it remains a great company despite a plunge in its stock value.
I will leave it to economists to debate what is a true economic indicator but recently, Google has started claiming that its search patterns are saying that economy is recovering. Basically, Google's chief economist says that he can tell the direction of economy from American's search habits.
For e.g Google is seeing following trends:
- Decrease in searches for unemployment benefits
- Increase in searches for homes and real estate agents
- They also showed an early spike in government "cash for clunkers" program.
You can always disagree with Google's claim and argue that a search for "property prices in a neighbourhood of Las Vegas" shouldn't be interpreted as recovery sign. Maybe, the person is trying to to sell the house! In the end, it is very difficult to interpet the true motive from few key words in the search. After all, economic forecast is a serious business and econometric modeling it is a very complex thing!
I think that we will continue to see this trend in future where "search on the web" will continue to indicate at least partially where the economy is heading. As search on the web moves from "keyword search" to more advanced "semantic search", we will have better idea about the intention of the query.
If we can analyze the "query intent" then we can build better models also. We also have to take into consideration the search patterns from Yahoo, Bing and other search engines to have a complete picture. I will still not dismiss off leading indicators from GE as it remains a great company despite a plunge in its stock value.
Subscribe to:
Posts (Atom)










