Showing posts with label Semantic Web. Show all posts
Showing posts with label Semantic Web. Show all posts

Thursday, February 25, 2010

Where does my money go? A begining for an open government in US, UK and rest of the world

Once James Madison, the political philospher and the fourth president of America, said "If men were angels then no government would be necessary." Though, there are no clear historical records available about the first official government, democracy or parliament, despite various claims by few old civilizations, but I guess that men figured out much before James Madison that they can never be angels and they will always need a goverment to live happily and prosper. I am sure people just didn't want any government but also hoped and craved for a smarter, open and transparent government. But it seems nobody could define it clearly in last so many generations what openness and transparency really means for a government. The good news is that all of it is changing! Surprisingly, for the first time "data" is taking the lead in defining an open government - maybe because it is measurable and never lies.

The two big initiatives, data.gov and data.gov.uk (still in Beta), were launched by US and UK government respectively in May 2009 and January 2010. Infact, Prime minister Gordon Brown of UK asked Berners-Lee to look at access to government data in June, after Barack Obama's administration launched an open source data site. Well, UK already ranks number three in the OECD (Organisation for Economic Cooperation and Development) study behind Austria and Portugal in the sophistication of its e-services so making it data available was the next logical step. At high level, the goals of both the initiatives are same as they want to make the government data available online to general public for improved access; creative use of that data outside the walls of the government; public participation, collaboration and feedback; identify unexpected and insightful data relationships - insights that would normally take several decades and hundreds or thousands of brilliant socialist scientists, statisticians, psychologists, focus groups and public policy experts to simply suspect. Both these initiatives are great because of the intention behind them but lets take a closer look at the state of the initiatives.


The data.gov started with forty seven data sets but already has thousands of datasets from eighty-one US agencies. It links directly to data files in various formats including CSV, XML, Excel, and KML. A lot seems to be lacking though:
  • It makes little effort to highlight or promote any projects that uses the data from the site
  • The focus is more on a repository
  • What you do with the data is not very clear
  • The website needs lot of work in terms of clarity and user experience
  • It is still not developer friendly and needs to develop an ecosystem
  • There's no basic demographic data like population from the Census Bureau
  • Browse and search functionality seems to be missing
The Sunlight labs, a DC based non-profit organization, is working on some projects to take advantage of this data but it seems you need to have larger developer community doing the same. I have also seen some good applications from the team of  James Hendler of Rensselaer Polytechnic Institute, USA. His team is converting the data sets into RDF and  taking advantage of semantic technolgies to build few applications. For e.g:
  • One of the application is about the amount of money received by the government for corporate and personal income taxes projecte through 2014 - click on the link to access it. 
  •  If you want to know how knowledgeable is your state - click on the link to access it.

I am still not sure why Linked data (Semantic Technology) approach was not taken from the begining. Overall, data.gov in US has still a long way to go before its goals are met. Ideally, it will be great to see more impressive applications which uses data from different sources and gives you an insight about a specific problem. Nevertheless, it is moving forward - it is also understandable that managing and simplifying the process of publishing humungous data from so many agencies is a herculean task.

Now, when I look at data.gov.uk then I have to say that I am just simply impressed considering the progress they have made in six-seven months. Kudos to the team, along with Sir Tim Berners Lee, who has been working on it. They just used the semantic technology, basically linked data, approach from the begining. It also has a modern design with a very developer friendly approach. Combining data and creating mashups from different sources in this context is not an easy task but semantic technologies have made it possible. Overall, data.gov.uk's approach is simple and clear - they have used open standards, open source and open data. The website has quality and elegance written all over it even though it is in beta. They also need to figure out many things like modelling various datasets behind the scenes, encourage more participation and many other things you can think of in a project of this complexity. But the results are showing! The top ten application as rated in this telegraph article are impressive. You can see the screen shot of  one of the application called "Where does my money go." Or click here to access the prototype.







 If you are really keen to go deeper in the approach of UK govenment in this implementation then I will encourage you to read this document - Putting the frontline first:Smarter Government. UK might have followed US, Australia, New Zealand in implementing open and smarter government but it seems at least now that their template will be followed by rest of the world. It will be interesting to see if/when countries like India, China and Russia will follow this trend!

"This is very much the beginning. Hopefully, this is the tip of the iceberg. There is a whole lot more to do." But what a beginning! These were the words of Sir Tim Berners Lee when he launched the beta site for data.gov.uk.

Thursday, January 7, 2010

"Pull" by David Siegel: Book Review

I just finished the final pages of "Pull" while watching the People's choice awards on TV. Who could have thought few years back that there will be a category for most popular "Web Celebrity" award which will be won by Ashton Kutcher for his more than one million followers on Twitter. Such is the power of web technology! Are you curious about the power of Semantic Web technology? Read "Pull" by David Siegel.


First of all, I would like to acknowledge the courage of David Siegel to write a business book on a difficult topic like this. The book is more about the power of Semantic Web to transform your business and is meant for business managers and enterpreneurs. It is not easy to write a "business book" on a topic like Semantic Web which has more sceptics than believers. In general, I have found that most of the good business books are more about analyzing the past and there are very few which are visionary or predict the future. In this complex world, it is so hard to see the future even beyond five years from now! So don't expect perfection! David makes a very good attempt in this direction.

Other than explaining the benefits of pull vs the push approach as practiced in most of the businesses, you will find some very useful information and statistics. The book is more conceptual in nature and talks about the recent developments in the world of semantic web and also takes you to the decade between 2020 to 2030. Some of you might question his timing also, if you are in a mood to question everything, but that is not that important from my perspective. Eventually the market forces dictate everything so why should we worry about it? You can always argue that he has become over enthusiastic about certain topics and is almost Utopian in its approach at some places but keep in mind that many of the concepts in this book are about a distant future. You may wish for more details or hope that he should have covered more domains but then he had to draw a line somewhere to make it readable for everyone. Overall, he has done a good job in doing gap analysis between the present state and the future state across various verticals but don't expect that you are going to get the perfect technology and process roadmap to achieve future state - the book is not about making incremental improvements but to find a completely new learning curve.

In the end, this is a thought provoking book which you should read with an open mind. It will definetely make you think! Without giving away too much about this book, I just want you to know that I really enjoyed reading it. It is a passionate work by someone who has put two years of life writing about a subject he believes in. The level of effort he has put in research also shows. Apart from that, for $18.45 on Amazon.com, this book is value for money. A real bargain! I will recommend all of you to read it.

Wednesday, December 23, 2009

Mobile Internet report: some interesting insights

In this age of information overload, it is very easy to miss this comprehensive report by Mary Meeker of Morgan Stanley. It is a solid and very detailed work! Basically, the theme of the report is about the growth of mobile which will be much bigger than any of us can imagine.  This new technology cycle is compared with what Windows 3 did for the PC in 1990 and Netscape browser did for desktop internet in 1995. And how the mobile internet has potential to create/destroy more wealth than prior computing cycles!




The report is well supported by solid research and numbers to back the analysis and forecasts. It took me some time to go through it but it was worth it as it gave me some interesting insights. In general, if it comes to mobile, we don't need to read a report to predict the future - we just have to look around and see what/how everyone is using these smart devices. I will skip the obvious like the spark created by iphone, 100k+ apps etc. but there are still few interesting facts such as:
  • 5 trends converging - 3G + social networking + video+ VOIP + impressive mobile devices
  • 57 m iphone +163% growth, 125k developers worldwide, 2B+ downloads
  • Many consumers are finding that their online usage rises dramatically when they have 24*7 mobile access to cloud based stuff
  • Growth/monetization roadmap for mobile is provided by Japan
  • Physical products are gaining share in mobile ecommerce - 20% in Japan but less than 1% in US
  • China leads world in virtual good monetization
  • Google (Android) has the best chance to to serve as a more open counter-balance to apple. Opera leading transformation of mobile browsers
  • Professional content repository is still open - amazon, iphone, netflix, hulu?
  • Open mobile web potentially more attractive to developers - Apple may believe its platform management is prudent way to ensure high quality content but it also runs the risk of stifling development and innovation
  • Clouds will be providing real infrastructure for mobile applications
  • Location aware ads will be better targetted
  • People are more willing to pay for content on mobile than desktop
  • AT&T - 50* mobile traffic growth in last three years despite only 40% subscriber growth over the same period
  • Shift in type of applications from games, lifestyle, utilities, enterptainment etc. to business oriented though among the top 100 applications on iphone, business oriented applications are just five
  • Daily usage for productivity based applications is probably less than 5-6 %
I wish there was little more information about the user profile but I still consider it the best work by Mary Meeker after her famous Internet report in 1994. Since new business models are often created during technology changes so there is an opportunity for eveyone inlcuding Semantic Technology to take advantage of this booming market because:
  •  Browser still remains the weakest link in these devices. There is no integrated/personalized experience.
  • The principles of end user interaction have not been established
  • Semantic search will become more relevant as users will have less tolerance for too many results.
  • Semantic and location specific advertising will have a role to play
  • Cloud computing will be one of the biggest enabler of writing semantic applications for mobile devices as computing power will no longer be an issue for even smaller shops
  • Semantic web powered commerce can also be triggered
  • Agent technology should find better business cases
Obviously, any analyst report will not be specific about the type of future applications/services which will reside in these mobile devices. That job is for innovators to figure out!
Reblog this post [with Zemanta]

Tuesday, December 8, 2009

Microsoft Semantic Engine!

One of the highlights of the PDC09, an event focused on the technical strategy of the Microsoft developer platform, was the overview and demonstration of the Microsoft Semantic Engine. The Semantic Engine unifies search, structured querying, and analytics over structured and unstructured data.You can read some more details about PDC at the CTO's blog. It goes beyond existing components like Lucene by supporting both text and non-text, such as audio, video, and images. The key points are:

  • Micorsoft has been working and investing heavily on this technology for the last two years
  • The "Microsoft Semantic Engine" name is just a place holder
  • It is not a W3C SemanticWeb(tm) approach but one which melds the unique capabilities of unsupervised machine learning (hierarchical clustering), information retrieval models (higher-dimensional vector spaces), pluggable and trainable classifiers (SVMs, Naive Bayesian, Maximum Entropy, Decision Tree, etc.), and personalized filtering and ranking.
  • One of the goals is to make search semantically enhanced. Clustering the results based on Semantics is a key differentiator.
  • You can expect to see the Microsoft Semantic Engine in one of the upcoming SQL Server Betas.
This demo/presentation will explain the approach in more detail. My initial thoughts are:

  • This is a good move by Microsoft and was long overdue because there is a solid business case for integrating this technology in the enterprise.
  • I view it more as a flavor of text extraction/analytics technology - infact, Msft has said it very clearly that it is not using the semantic web technology approach
  • Having Microsoft in the Semantic technology space is a good thing for the Semantic Technology indutsry. This is a great validation from one of the most successful leaders in the software space
  • They are not the only one who is trying this approach. As far as I understand, the goal behind Inxight's acquisition by Business Objects (now SAP) was the same one. I am not aware exactly how that integration has worked out
  • I am sure Micosroft will come out with a clear message regarding the "Semantic Engine's" positioning in the enterprise in comparison to Fast Search Engine (Now part of Sharepoint division).
  • It will intersting to understand if Powerset (now Bing) technology, Microsoft's semantic search, was used in this effort
  • More details are needed to understand how it will work with disparate data sources
  • Microsoft has been underestimated for too long as far as their search strategy is concerned. I agree that Google has a big lead in terms of number of users on the web but I really think that they have made all the right moves, both for web as well as enterprises,at least  in last 2-3 years as far as search space is concerned. There are three great moves:
    • Acquistion of Fast search and integrating it with Sharepoint (MOSS)
    • Acquistion of Powerset and launch of Bing
    • Plan to introduce Semantic Engine and integrating it with SQL serve 
  • Probably,  services and product-based companies in the Semantic space, need to go back and revise their marketing message.
  • In the end, it just validates my initial thoughts that Semantic technology is a superset of Semantic Web Technology. Microsoft's approach or Semantic Web technology (W3C) approach are two different approaches to solve Semantic related problems in the enterprise. This is the reason, I named my blog as Semantic Technology blog.
Reblog this post [with Zemanta]

Thursday, December 3, 2009

Semantic Technology and iPhone!

Recently, I came across this intersting article about "Inside the Apple Economy"  - http://bit.ly/2FnBPt . Please read about some of the unknown names who have made a killing in last two years just by exploting this new ecosystem. I believe that Semantic web, linked data and Semantic technologies in general for the enteprise will also be creating an ecosystem which will be exploited by many people. But I only realized this Thanksgiving weekend during a NY city visit along with few friends of mine that there can be a corelation between iphone and Semantic Technology.

We had not made any dinner reservations as we were not sure about timings and were confident of finding an interesting cusine in the NY city, practically a paradise for your palate, which would work for everyone in the group. I was also counting on my Zagat membership which I have always found to be a very trusted source for restaurant reviews. We soon realized that everytime we agreed on a particular cuisine and restaurant after looking up ratings on Zagat on my iphone, the restuarant was always few miles away and it was very difficult to find two cabs on a crowded weekend. So we decided to use "Around Me" - an interesting application on iphone which can list critical services around you — banks, coffee shops, bars, gas stations, hospitals, movie theaters, restaurants and so on. Using geolocation, the app orders each service by its proximity to you — how many yards away — and, like other apps aimed at the traveler, maps out a route from here to there, if requested.





We kept looking up restaurants in "around me" but we couldn't find ratings and reviews for those restaurants. So I ended up toggling between these two applications, during our long walk, for a while till everybody agreed on a Turkish cusine.

I was wondering if these two services shared there data using Semantic Technology, basically Linked data, then life would have been so simple. All I had to do was to ask - "Give me all restaurants with minimum 4 star rating within 400 yards of my position." Hopefully, some food for thought for enterpreneurs who want to exploit the iphone economy. As somebody rightly said, "The food that enters the mind must be watched as closely as the food that enters the body."

Thursday, November 12, 2009

Semantics, Metal Reserves and how long will they last!

The picture below tells you about the state of planet earth's metal reserves!



Basically, this diagram says that there are fifty eight years left when all the metal resources will be exhausted  if the world continues to consume at today's rate. There is nothing to worry about because I am sure that geologists on this planet will figure something out much before that.  I am saying this because many of us worried about the state of oil reserves also in last two decades after many pessimistic predictions - more than half of the time people were "Crying Wolf." In any case, the point I want to talk more about is the role of semantic technology in helping the geologists around the world.

I came across an interesting paper in this context. The key points explaining why semantics is becoming important in geology are:
  • Geology, perhaps more than any other science, has advanced with the aid of pictures. The pictures, which are geological maps shows the distribution of different rock types. The current efforts to improve them have spawned a global e-science initiative to review the semantics of geology.

  • One example of issue of semantics in context of geology is - The US department of agriculture and Forest service and the British Columbia Ministry of Lands, Parks and Environment will describe and classify the same geological phenomenan in a completely different way

  • When working at relatively large scales, geologists are always dependent on data and information gathered and reported by other geologists regarding smaller scale features. Semantics is very critical here also

  • One of the consistent problems in geology is the omni-present "map boundary fault" - an apparent geological continuity along the border of adjoining maps which is nothing more than difference in nomenclature or semantics

  • Research is being done to prototype vocabularies by developing web services which can be used to store, compare, and rank on similarity, decsriptions of models (concepts) and instances (physical entitites and events) of mineral deposits, landslides and landslide hazards.
To go through the full research papaer please go here.
Reblog this post [with Zemanta]

Semantic Advertising or not : Diamonds are forever!

Most of us know the origins of the  "De Beers", one of the largest diamond company out of South Africa, and how its founder got his start by renting water pumps to miners during the diamond rush. It is also believed that the founders  merged their interests in various diamond mines into a single entity that would be powerful enough to control production and perpetuate the illusion of scarcity of diamonds. Rest is history! One thing at least I didn't know was how De Beers is used as an example, as shown below, of  "Bad" ad placements. Basically, it highlights the limitation of existing ad technologies like contextual advertising and behavioral targetting in understanding the meaning of a web page.





Semantic Advertising or Semantic Targeting, which is more about targeting the relevant ad based on the meaning and sentiment of the page has gained lot of traction in last few years. I will recommend everyone to read this article by Scott Brinker which explains four different kinds of Semantic advertising in a very clear way.

In theory, the value of semantic advertising is great and companies like peer39, expert system and few others have gained lots of visibility. The opportunity is large as publishing industry is going through the biggest turmoil in decades because of this new generation who doesn't believe in paying for the content. So a relevant or highly targeted advertisement will end up being a major source of revenue for publishers.

In the end, semantic advertising is still a niche space and not a whole lot is understood by the industry about the effectiveness of this approach. There are not enough demos to show the value of semantic advertisement and publishers still need to be convinced more about the benefits of switching to semantic advertising. It will probably take some more time for semantic advertising to become completetely mainstream but I don't think that it will have any bearing on future of De Beers despite the "flaw" in ad placement - as there is a saying that "it is better to have a diamond with a flaw then to have a pebble without one." "Diamonds are forever" and they will always remain a scarce and sought after resource at least in next few decades.
Reblog this post [with Zemanta]

Thursday, November 5, 2009

Linked Open Data : Can we learn anything from failure of many B2B Exchanges?

"You can't stop an idea whose time has come!" Linked data seems to be that idea in the grand vision of Semantic Web! The growth of Linked data cloud always reminds you that we are in exponential times! The Semantic technology community, including me, believes that Linked Data is the best thing happened to the Semantic Web vision. It makes sense and it is the next step for web and can also contribute significantly, if done right, to the evolution of this civilization. The possibilities are endless! We also have one of the best brains behind this initiative.  Whenever I hear Metcalfe's law (by Bob Metcalfe) , in context of Linked data, which states that "the value of the network is proportional to the square of the connected users of the system" then it always reminds me of those early B2B days during late 90s when the world was about to change. Metcalfe law was very popular during those times also! B2B exchanges were considered the pillars of the new economy and their valuations made us lose our sleep. It seemed to all of us that everything was going to be re-defined and we wanted to be part of it. Well, it didn't happen exactly as we were led to believe! Despite the differences between B2B and Linkedata like one is trasactional and other is knowledge oriented, privately owned vs free, different technologies, different era etc., there are some similarities in both of them. Lets do some introspection so that we don't repeat some of those mistakes!

B2B (Business to Business) exchanges is an entity which brings multiple buyers and sellers to a  marketplace where all kinds of commodity, financial instruments, intellectual property and various other goods can be eletronically traded - web is used as the medium in most of the cases. It has following characterstics:
  • The perceived value followed Metcalfe's law
  • Enterprenuers could fundamentally re-invent how work gets done
  • No longer comparison between big-small and so on
  • Shift of power from producers to consumers who are in control of everything
  • Standardized marketplace and standardized contracts
  • Markets operated at fraction of physical word cost
  • Global reach and one stop shopping
  • Neutrality, transparency, self-regulation, market efficiency, confidentiality and anonymity were other virtues of this marketplace
  • The winner, of a particluar vetical B2B , takes all 
It did have some good principles but a very large number of the B2B exchanges failed within few years. There were few fundamental reasons for their failure:

  • The internet enterpreneurs didn't really understand their place in the overall marketplace
  • What they offered to business was something that already existed - at least in some form
  • Companies had always done business with other businesses and most of these businesses negotiate to lower the price. Adapting those existing processes to an internet format didn't really create anything new or different in the field. Unfortunately, many companies felt that switching to an internet-based sites controlled by third parties was a riskier bet than staying with the current partners and vendors
Few of these exchanges figured it out early and survived by either addressing these issues or by being creative about it - like some of them just focused on small businesses. ECN (Electronic Communication Network) are another success story in the world of B2B exchanges as they allowed a more efficient price discovery mechanism for stocks and currencies.  In the end, B2B exchanges were all about attracting buyers and sellers or producers and consumers which they couldn't do. In a similar way, the success of Linked data cloud will depend on creating a marketplace which should be able to attract producers and consumers or buyers or sellers. The technical design prinicples by Sir Tim Berners Lee are just great and well thought of. We are also aware of the existing issues in the Linked data Cloud like quality of data, what is out there, disambiguations issues, trust of the source, frequency of the update, how to get started, should we always publish in RDF, how do you erase inaccuracy and many others things like this. We had very similar issues, maybe in a different flavor, when Web 1.0 was developing and we have come a long way. We might say that HTML was much easier to adapt then RDF but the success or wide adaption of the original web was not only simplicity of HTML and HTTP protocol but because it was a great sales, marketing and ecommerce tool. In context of Linked data cloud, lets accept the fact that most of the businesses will not be  interested in "the future," they will always be interested in "their future." Their first responsibility will always be their share holders.

Companies need very strong incentives or business case to publish their data to this "Cloud" and write applications to consume the collective intelligence of the nodes in this cloud. Initiatives like exposing government data (UK and US at this point), dbpedia, scientific, map oriented data, music, individual research projects and so many other examples which can be good and interesting from exploration standpoint are already underway. We will also see more invovlement from non-profits and intelligence communities at some point. It is a great effort but still not enough! The Cloud will continue to grow and will add billions of more RDF triples but we need to proactively involve the corporate world. Technology enterpreneurs for Semantic Web or Linked data can't do it alone. They have to partner with the business community and the enterprise. The value proposition of Linked data needs to be articulated to the CXO community and their participation needs to be encouraged. We have to start having more conversations like the "benefits of publishing data to the cloud for the enterprise." Basically, much more on business development, business case analysis, education, marketing then technology alone - technology will always remain important and core. I had written an earlier blog about a new category name for Semantic technology which talks about some aspects. Once we have their attention and buy in then the ecosystem of new tools, applications, security infrastructure, new architecture, developers, APIs, SEO, creative sales and marketing ideas, consulting, outsourcing, selling and buying and many either good things will start maturing. It will be very important to do so in next few years to sustain this momentum of Linked data cloud. It will also attract significant investments in this effort! Don't we think that there are dollars to be spent on the Linked data cloud in the  worldwide IT spending of more than three trillion dollars in 2009. We should be thinking or advocating about the "guidelines" for  publishing to the Linked data Cloud like:

  • The benefits of Linked data and possible use cases.
  • How are they impacting their supply chain? 
  • What are the advantages in reciprocity of links? 
  • What should be the process of identifying the data which needs to be published
  • How will it help their revenue and help in their relationships with partners and customers?
  • What are the security issues? What kind of security tools or approaches which are out there?
  • How will it mitigate risks for the company?
  • What will be the possible compliance and legal issues for them?
  • What data access choices can they have - free access, partial access, sign up , qualification, payment?
  • What are the data packing options like - depth, breath, granularity, freshness etc..
  • Can they have handle on who is accessing this data, frequency of access, redistribution issues?
  • Can they create virtual networks of data clouds for their partners and business customers?
  • Can they promote their data?
  • Can they increase the findability of their data?
  • How can they do it in a phased manner?
  • If they use RDFa then are they part of Linked data?
  • Are they accountable for consistency, organization, correctness of their data?
I can think of so many more questions. A very clear policies and procedures aproach needs to be articulated which will lead to a governance model for publishing to Linked data cloud within an enterprise. Probably, we will have to learn more from the experience of present initiatives like publishing data from the UK government to the cloud. During the early days of web, if any company didn't have a web presence then it was perceived that they were missing the action, opportunity or future revenue. At some point, today's businesses should start looking at Linked data cloud with the same lens.
Reblog this post [with Zemanta]




Wednesday, October 28, 2009

Book of Odds: A Semantic Knowledgebase of Probability

When I think about quotes related to probability, I can't decide which one of these is more profound or I like more:
1) "Probability is the very guide of life" by Marcus Tullius Cicero
2) “Once you eliminate the impossible, whatever remains, no matter how improbable, must be the truth.” by Sherlock  Holmes
3) "Million to one odds happen eight times a day in New York" by Penn Jillette

I am always tempted to choose number three because it is more funny and  if you have lived, worked or visited NewYork city then you can relate to it. But this blog is not about quotes related to probability! It is about Book of Odds: World's first reference on the odds of everyday life or a Knowledgebase of Probability built using Semantic Technologies. It is not another search engine or "yet another knowledgbase" as I explained in one of my earlier blogs. It tells you odds about things which you care about like "odds of surviving in a plane crash" or  "odds of having a heart attack if you are of certain age and ethnicity". Odds, Probability and Chances are used interchangeably in this context -- it is basically the ratio of favourable outcomes to total outcomes. The world had  references like google, wikipedia, Webster Dictionary, encyclopedias and so many other for  almost everyhting but didn't have "Book Of Odds" to understand the probability of an event or an occurence - so the founders had a good business case to start this initiative.

It seems that the project started almost three years back and millions of data points were collected by the team of very talented researchers. Once data have been collected, analyzed, and the odds have been computed, every Odds Statement then undergoes a rigorous quality review, is modeled ontologically (or in simple words by using semantic technologies explained throughout this blog), and finds its place in underlying semantic database of probabilities. The Semantic middleware is provided by Cambridge Semantics, a Semantic Technology company, mentioned in one of my earlier blogs, which structures the data, adds metadata and other information that helps users find the relevant odds statements using simple search. I consider it a very creative use of applying semantic technologies to the problems which could not have been solved using traditional database management systems. This proves again that a typical DBMS is not the answer to all the storage and query related problems! The semantic technology also allows it to find interesting and unexpected connections. It is not only a fun site but aims to be a respected new reference source to increase the tolerance for uncertainity and help people to deal with tough decisions.

To start using Book of Odds, you can browse by topics like health and illness; accidents and death; relationship and society; and daily life and activities. It is understood that it will always remain work in progress as this planet is becoming more complex and eventful everyday. There will always be  new areas to add and grow this reference.

My limited intearction with the website tells me that content is of high quality. I will recommend readers to at least give it a try and encourage this team. The most important thing for "Book of Odds Inc." is to have more people registered and start using this knowledgebase. I will "not" keep on pressurizing the founders at this stage by asking about business models and revenue stream.What business model did google have for many years? I am sure that they know that these questions are going to come up repeatedly after an year of their launch.  If  they will keep on producing high quality content or probabilistic ratios in this case and keep it meaningful for people then good things will happen. They do need to think creatively about advertising and marketing campaigns. They shouldn't be perceived as another google killer! Atleast they were smart enough to not use "engine" in their name. A business model will follow if they can attract more users and keep them interested.  My only recommendation to the founders will be to start thinking about combining their work or knowledgebase with various forms of predictive modeling or business analytics in the business world. Who knows better than companies in this economy that  "By looking at the past we can come to a greater understanding of the world and what is forever an uncertain future!"

Thursday, October 15, 2009

Semantic Technology, Financial Reporting and the Toxic Assets!

Financial markets, traditionally the earliest adapter of any new technology relative to other industries, has been a laggard as far as Semantic Technology is concerned. It seems that the turmoil in the capital markets in last two years has managed to dampen enthusiasm for new technologies in capital markets and banking industry. All of it is about to change! The two obvious reasons are: we are coming out of recession and  new regulations regarding financial reporting in XBRL. I believe that the third reason is the inherent limitation of XBRL as far as Semantics is concerned!


As we know, XBRL, solves two significant problems for companies who prepare financial statements along with analysts, investors, regulators, financial publishers and data aggregators:
  • The first problem is that preparing a financial statement for printing, for a Web site, and for filing today means that a company could typically enter information three times
  • The second problem is that today (if the report is not in XBRL), extracting specified detailed information from a financial statement - for e.g we still can't ask questions like "Give me depreciation expense from 2003 a financial report."
The basic idea behind XBRL is to provide grammer and syntax behind financial reporting so that it can be extracted, analyzed and queried. Although XBRL has been around for 10 years now, the adoption and acceptance has only begun to significantly accelerate during 2007 with the support of  SEC. Since the year 2009, the filing has become mandatory for largest five hundred US corporations and other companies will follow in a phased manner from 2010 onwards. Market is already flooded with XBRL products , services and tools. Most of these products and tools help in one or more of following things : creation, viewing, analysis, taxonomy creation, custom document creation and various other automation features. XBRL can be stored in RDBMS as well as XML databases like Marklogic.


So what is the problem? Why do we need Semantic Technology in this context? While XBRL allows for more accurate consumption and interpretation of financial information, there is still a need to connect to the authentic source of the document and to recombine the XBRL content with other data sources. The fundamental issue here is that XBRL document working with other dat source doesn't understand anything about the semantics of data. There is just no meaning associated with the nesting of tags. The limitation of XBRL becomes more obvious when you have to use/analyze/query XBRL reports along with other sources of data which is not XBRL compliant.

If you read this  article in Wall Street Jornal on Toxic assets then it will make you think more clearly about importance of "semantics" in reporting in the world of derivatives.. The key points are:

  1. Ever since humans started trading, lending and investing beyond the confines of the family and the tribe, we have depended on legally authenticated written statements to get the facts about things of value
  2. Derivatives are the root of the credit crunch. Why? Unlike all other property paper, derivatives are not required by law to be recorded, continually tracked and tied to the assets they represent. Nobody knows precisely how many there are, where they are, and who is finally accountable for them.
  3. Every financial deal must be firmly tethered to the real performance of the asset from which it originated.
  4. All documents and the assets and transactions they represent or are derived from must be recorded in publicly accessible registries
  5. Governments can encourage assets to be leveraged, transformed, combined, recombined and repackaged into any number of tranches, provided the process intends to improve the value of the original asset
  6. Financial institutions will have to serve society and fully report what they own and what they owe -- just like the rest of us -- so that we get the facts necessary to find our way out of the current maze
  7. Governments can no longer tolerate the use of opaque and confusing language in drafting financial instruments. Clarity and precision are indispensable for the creation of credit and capital through paper.
XBRL, by itself can't fulfill all of these requirements as we need to corelate/link/resolve various reports to the source data - a very important thing in the world of derivatives reporting. You need Semantic Technology for that! We need to represent XBRL in RDF or OWL representation. I will recommend my readers to read another rebuilding public trust - a nice article on the same topic. The author is also talking about services which can allow the financial data in XBRL to be combined with data from other industry and government sectors — basically, transforming the way we explore information.

There are various techniques to convert an XBRL document to RDF. I will not go into those details in this blog. One example -  GoodMorningResearch.com machine automates XBRL tagging of Excel data in RDF format with one-click Save As XBRL functionality.

I believe that  long-term (probably very long-term) vision of XBRL reports should be to publish it as RDF triples and make it a part of Linkedata cloud. This will help in achieving all linkages, transparency and verification as far as financial reporting is concened. I would like sceptics to know that by April 2009, more than 600 XBRL reports, approx. 1,3 million RDF triplets,  were already part of Linked data cloud.  But at the same time, you need lot more governance, regulations and process behind this effort to get real value. Also, there has to be some kind of incentives for financial organizations to do this.

Tuesday, October 13, 2009

Applying Semantic Technology to cure cancer!

All of us know that in the world of molecular biology, significant research efforts are moving from wet labs to the dry labs. And by dry labs, I mean bioinformatics. It is a relatively a new discipline but a lot of work has been done in last few years.  In a nutshell, bioinformatics is about building an information system using biological data. Semantic Technology is playing a big role in transforming this biological data into connectable information which is giving scientists completely new insight about biological mysteries.

Recently, scientists at The University of Texas, M. D. Anderson Cancer Center got funding to derive meaningful information from an ocean of data about the aberrant genetics that drive human cancers. The work will be primariy focused on parsing the multiple genetic pathways that fuel more than 20 types of cancer. Medical Scientists are going to use semantic technology to solve this puzzle.


For full news:
 http://www.eurekalert.org/pub_releases/2009-10/uotm-mda100709.php

Friday, September 25, 2009

Google adds semantic web support for video search!

Google, who has been behind Yahoo in semantic web efforts, is trying to catch up fast. Google announced support for enhanced markup for video search. This will allow webmasters to include important information, such as titles and descriptions, in machine-readable HTML along with the JavaScript or Flash videos themselves. It will use structured data open standards such as microformats and RDFa to give users more detailed previews of the information contained on the web page. Yahoo! searchMonkey has done something very similar for the content.



Reblog this post [with Zemanta]

Tim Berners Lee video about his journey and vision!

This interesting video was published this year.

http://www.ted.com/talks/tim_berners_lee_on_the_next_web.html

He talks about his journey of last years from regular web to the semantic web.
Reblog this post [with Zemanta]

Wednesday, September 23, 2009

How Open Calais initiative is helping Semantic Technology!

Open Calais initiative, started by Thomson-Reuters, is one of the most interesting things which has happened in favour of semantic web vision. Reuters acquired this technology as part of their ClearForest, one of the leading vendors in the text analytics space, acquisition. It is stated that the service could quickly become the largest repository of metadata (in the form of named entites and facts) on the Web if it stored the resulting metadata from each request. Open Calais is the "metadata extraction service" ; it is a Web service that allows you to automatically annotate content and extract information like facts and named entities (people, places, and organizations, and much more) from unstructured text. Calais uses linguistic parsing (also known as entity extraction) in a service enables way to producr RDF triples and Semantic Web data models.
Open Calais opens the door to the possibility of lowering the barrier enough for everyday users to publish semantic content. It finally does what critics say to be the greatest obstacle to the Semantic Web: Taking the metadata burden from the end-user by providing an automatic meta-tagging tool. Open Calais initiative will also be one of the biggest enabler of the Linked data initiative.
Recently, CNET has joined OpenCalais initiative as one of the first commercial media companies to publish core data assets for public, programmatic use on the open semantic Web. CNET will leverage OpenCalais' connection to the rapidly expanding 'Linked Data cloud' to allow its original content -- such as tech product reviews on laptops, TVs, smart phones, and digital cameras; news articles and blog posts from its CNET News editorial staff; and parts of its core technology product catalog - to be available for public use.
Reblog this post [with Zemanta]

Tuesday, September 22, 2009

Example of Semantic Technology in action - Yahoo Search Monkey

Its is a misconception when people question the viability of semantic technologies by saying that "it isn't possible to convert all the data to RDF format?" In reality , it is not true at all. Again, it is important to keep in mind that we are not talking about semantic web, we are talking about semantic technologies. It can be best explained by talking about how Yahoo Search Monkey is using RDFa (RDF in Attributes) to enhance search results more useful and visually appealing. It is just one of the simple examples but it can hopefully answer the sceptics. It can be a boon to small businesses who can drive more traffic to their web sites. SearchMonkey looks for special data inside websites, based on a standard called RDFa. Your website should include this data so it is available to Yahoo as they crawl and index your site. This way your business information is available to any developers who build SearchMonkey apps, and you will show up with enhanced results as this gets adopted over time. The SearchMonkey platform has three main components: - "Site owners share structured data with Yahoo!, using semantic markup (microformats, RDF), standardized XML feeds, APIs (OpenSearch or other web services), and page extraction. -Third party developers build SearchMonkey applications. -Consumers customize their search experience." RDFa is a way to encode data within HTMLand XHTML pages which helps people and machines to embed structured data within HTML and XHTML pages. The underlying representation of RDFa is RDF because it is flexible enough to let publishers build an devolve their own vocabularies. You can see Search Monkey in action by clicking on this link http://www.yelp.com/search?find_desc=nobu&ns=1&rpp=10&find_loc=San+Francisco%2C+CA#find_loc=new%20york
You can see the enhanced quality of the result "nobu new york restaurant". You will probably realize that you have many using Semantic technologies without even realizing it.

Tuesday, July 14, 2009

What do we mean by "Semantics" in Semantic Search?

I have often been asked this question in last few months by clients and other people who are interested in finding more about the activities in the world of search. Is it same as natural language search? Or it is just applicable for semantic data expressed in RDF and OWL? So companies who are developing semantic search are basically trying to throw google out? How relevant is it for an enterprise. These are the standard questions which I often get asked. At high level, semantic search engines is about interpreting meaning of the "query" in a smart way. Most of todays search engines today are primarily based on "key words" focused though the trend has been changing since early 2008. If you ask a question: q: What is the time difference between US and Tanzania ( a country in Africa)? average search engine : Will give you all documents having key words 'time difference', 'US' and 'Tanzania'. This is the case with millions of web sites who are using "plain vanilla" search. good search engine (having some semantic capabilities) : Will probably understand the intention of the query and give you links to the websites which talks about time difference between countries. There are still various inconsistencies in the answers if you test any of the popular search engines out there. You just need to ask the same question in a different way to validate the 'semanticity" of the engine. A Real Semantic Search Engine : Will give you a time difference by getting the right content or it will tell you that it doesn't know the answer. This is very difficult to attain on a consistent basis by by any technology out there. So most of the discussion will focus on good semantic search engines. Even though, the popular search engines like google, yahoo, bing (Microsoft), ask.com etc. claim that they have semantic abilities, they will still be in the category of a good search engine. They are hybrid of "key word" and "semantic" based search. None of these engines are purely semantic in nature. Google is very clear about its messaging - they are not going to replace key word search with semantic search capabilities. Most of the search engines today are really based on how the query is phrased. They do a poor job in understanding the meaning of the query if is phrased in different way. Semantic search engines should do disambiguation very cleverly - If someone is talking about "furniture", the semantic engine should even cosnider documents with and "tables" and "chairs" in it. Basically, it should understand the context. A good semantic search engine takes the burden from the user to answer a query even if it is asked in very different way. It is also very important to understand that these engines are not getting their content from semantic web. They are still working with the same set of millions of document in the regular web. They haven't touched semantic web for your query needs. The "semantic" aspect in these search engines are only related with their ability to interpret the query. Swoogle is the only semantic search engine which only fetches the content from the semantic web - basically data written in RDF format. The need to improve the "semanticity" of search engines will grow exponentially. It is very difficult to measure it though. There are primarily two categories : Pure play semantic search engines like Hakia,Sensebot, Congnition Search, Exalead, Powerset and many others. And the other category is Google,Yahoo, Microsoft and Ask.com of the world who are integrating semantic search algorithms in their core search technology. The line is blurring between these two categories. There is not just one approach to semantic search. Most semantic search engines mix and match them in various ways to yield a unique search experience for their users. There are at least four approaches to semantic search. Different semantic search engines may use one or more of these approaches. The point of semantic search is to use meaning to improve the user's search experience. For example, one approach is to use contextual analysis to help to disambiguate queries. Another approach focuses on reasoning. Given a set of facts that are represented in the system, additional facts can be inferred from them. A number of semantic search engines emphasize natural language understanding. These engines process the content they index and the queries people submit to try to identify the intent of the information. They use the syntax of the sentence and rules to identify people, places, organizations, and so forth. Powerset makes extensive use of natural language understanding. The fourth approach uses an ontology to represent knowledge about a domain and expand queries. On this approach, when a user enters a query for a word like "sofa sets," the system adds terms from its ontology (e.g., "furniture" because a sofaset is a kind of furniture) to make the search more focused as well as more broad. This approach is used by a large number of semantic search systems. Google and Yahoo will continue to have an edge because of existing vistors to their site and their continued investments in these technologies. Despite claims by various vendors, semantic search engines have a long way to go. There is a big gap between the existing information out there versus the tools which can actually get that information.

Sunday, July 12, 2009

Semantic Conference 2009

I was fortunate to be one of the attendees at the Semantic Conference 2009 held in Fairmont hotel, San Jose, Ca. I have to admit that I was very impressed to witness a huge crowd, probably 1300+, during these recessionary times. I was told by one of the organizers that the number of attendees has been growing considerably since the inception of this conference back in 2005. It was also very encouraging to see more than 20% of the attendees from the business side that just validates the fact that semantic technology is a no longer confined to R&D labs and intellectual discussions. There is a strong interest to apply this technology to derive maximum business value from it. Some notable sessions were:
  • Keynote from Reuters about state of Semantic technology and where is the money
  • Semantic Search discussion between C-level executives from Microsoft, Google and Yahoo
  • At least two sessions from the VC community expressing strong interest in the business side of semantic technologies
I had no doubt, even before going to this conference, that semantic technology can solve some problems which is very difficult to solve using the traditional or well-accepted technologies in the enterprise. I was was very disappointed to see almost lack of representation from even a single well known name from the Wall Street. It is well-known fact that financial industry has been an early adapter of most of the new technology but it seemed this wasn't the case in this context. Maybe, the turmoil in the financial industry is taking its toll on everything. On a brighter side, I saw a large number from the health care and pharma companies who had presented many use cases and application of this technology to solve some real business problems. I was also surprised to see a large convoy from Europe who presented many government-sponsored projects in the semantic technology. It seems that Europe has finally decided to take a very pro-active and leadership role in advocating semantic technology!
To sum it all, I left the conference with a smile on my face and with added enthusiasm for the semantic technology. There is no doubt left in my mind that it is just a question of time when this technology will be a mainstream technology and business benefits will be very real.
More to follow in the next blog ..