Showing posts with label Linked Data. Show all posts
Showing posts with label Linked Data. Show all posts

Thursday, April 7, 2011

Elsevier Challenge: Another Big Step towards Open Data Movement

Nobody was happy after hearing that Data.gov, along with a number of other data-related sites of the government such as USAspending.gov and Apps.gov, are slated to be shut down due to budget cuts. The current annual budget of $37 million will be reduced to $2 million. It wasn't long ago when I had written very enthusiastically about the open data movement. In general, despite the fate of data.gov, the open data initiative is still going strong. Today, we have almost 25 cities in US who have opendata. Sanfrancisco's datasf.org is another success story which has almost 60 applications built by the developers. What we really need, in this context, is more participation from the commercial world!

On a similar note - Today, I was contacted by Elsevier, one of the largest publisher of medical and scientific literature in the world about their open data initiative. I am more than happy to write about this great initiative from a well known commercial enterprise.They just announced its first worldwide challenge called “Apps for Science,” powered by ChallengePost. The challenge is designed to bring together developers and a community of 15M researchers to collaborate more efficiently via new and innovative apps.

This database comprises of more than 25% of world’s academic and scientific articles for the challenge. Now, developers can access Elsevier’s data catalog and APIs from its SciVerse Suite, a content discovery platform + developer network w/ 10M+ articles, an abstract database with 41,000,000 records and more. This video explains SciVerse Suite better:






Why should we care for it?

  • If you care about a cure for cancer, AIDS or any other deadly disease then it is an important step. We need to help scientists and support them with the best tools and information possible because better communication is equal to greater knowledge share which will result into innovation breakthroughs.
  • It will free up approximately 12 hours per week previously spent on collecting and organizing research (according to a 2007 Outsell survey of 6,300 knowledge workers).
What’s in it for developers?

You can build & host tools (free or fee-based) for a captive audience of 15M researchers & 10k research institutions on Elsevier’s Application Marketplace.


Judges include known names like Jeff Jonas and James Handler among others. 

Thursday, February 25, 2010

Where does my money go? A begining for an open government in US, UK and rest of the world

Once James Madison, the political philospher and the fourth president of America, said "If men were angels then no government would be necessary." Though, there are no clear historical records available about the first official government, democracy or parliament, despite various claims by few old civilizations, but I guess that men figured out much before James Madison that they can never be angels and they will always need a goverment to live happily and prosper. I am sure people just didn't want any government but also hoped and craved for a smarter, open and transparent government. But it seems nobody could define it clearly in last so many generations what openness and transparency really means for a government. The good news is that all of it is changing! Surprisingly, for the first time "data" is taking the lead in defining an open government - maybe because it is measurable and never lies.

The two big initiatives, data.gov and data.gov.uk (still in Beta), were launched by US and UK government respectively in May 2009 and January 2010. Infact, Prime minister Gordon Brown of UK asked Berners-Lee to look at access to government data in June, after Barack Obama's administration launched an open source data site. Well, UK already ranks number three in the OECD (Organisation for Economic Cooperation and Development) study behind Austria and Portugal in the sophistication of its e-services so making it data available was the next logical step. At high level, the goals of both the initiatives are same as they want to make the government data available online to general public for improved access; creative use of that data outside the walls of the government; public participation, collaboration and feedback; identify unexpected and insightful data relationships - insights that would normally take several decades and hundreds or thousands of brilliant socialist scientists, statisticians, psychologists, focus groups and public policy experts to simply suspect. Both these initiatives are great because of the intention behind them but lets take a closer look at the state of the initiatives.


The data.gov started with forty seven data sets but already has thousands of datasets from eighty-one US agencies. It links directly to data files in various formats including CSV, XML, Excel, and KML. A lot seems to be lacking though:
  • It makes little effort to highlight or promote any projects that uses the data from the site
  • The focus is more on a repository
  • What you do with the data is not very clear
  • The website needs lot of work in terms of clarity and user experience
  • It is still not developer friendly and needs to develop an ecosystem
  • There's no basic demographic data like population from the Census Bureau
  • Browse and search functionality seems to be missing
The Sunlight labs, a DC based non-profit organization, is working on some projects to take advantage of this data but it seems you need to have larger developer community doing the same. I have also seen some good applications from the team of  James Hendler of Rensselaer Polytechnic Institute, USA. His team is converting the data sets into RDF and  taking advantage of semantic technolgies to build few applications. For e.g:
  • One of the application is about the amount of money received by the government for corporate and personal income taxes projecte through 2014 - click on the link to access it. 
  •  If you want to know how knowledgeable is your state - click on the link to access it.

I am still not sure why Linked data (Semantic Technology) approach was not taken from the begining. Overall, data.gov in US has still a long way to go before its goals are met. Ideally, it will be great to see more impressive applications which uses data from different sources and gives you an insight about a specific problem. Nevertheless, it is moving forward - it is also understandable that managing and simplifying the process of publishing humungous data from so many agencies is a herculean task.

Now, when I look at data.gov.uk then I have to say that I am just simply impressed considering the progress they have made in six-seven months. Kudos to the team, along with Sir Tim Berners Lee, who has been working on it. They just used the semantic technology, basically linked data, approach from the begining. It also has a modern design with a very developer friendly approach. Combining data and creating mashups from different sources in this context is not an easy task but semantic technologies have made it possible. Overall, data.gov.uk's approach is simple and clear - they have used open standards, open source and open data. The website has quality and elegance written all over it even though it is in beta. They also need to figure out many things like modelling various datasets behind the scenes, encourage more participation and many other things you can think of in a project of this complexity. But the results are showing! The top ten application as rated in this telegraph article are impressive. You can see the screen shot of  one of the application called "Where does my money go." Or click here to access the prototype.







 If you are really keen to go deeper in the approach of UK govenment in this implementation then I will encourage you to read this document - Putting the frontline first:Smarter Government. UK might have followed US, Australia, New Zealand in implementing open and smarter government but it seems at least now that their template will be followed by rest of the world. It will be interesting to see if/when countries like India, China and Russia will follow this trend!

"This is very much the beginning. Hopefully, this is the tip of the iceberg. There is a whole lot more to do." But what a beginning! These were the words of Sir Tim Berners Lee when he launched the beta site for data.gov.uk.

Thursday, November 5, 2009

Linked Open Data : Can we learn anything from failure of many B2B Exchanges?

"You can't stop an idea whose time has come!" Linked data seems to be that idea in the grand vision of Semantic Web! The growth of Linked data cloud always reminds you that we are in exponential times! The Semantic technology community, including me, believes that Linked Data is the best thing happened to the Semantic Web vision. It makes sense and it is the next step for web and can also contribute significantly, if done right, to the evolution of this civilization. The possibilities are endless! We also have one of the best brains behind this initiative.  Whenever I hear Metcalfe's law (by Bob Metcalfe) , in context of Linked data, which states that "the value of the network is proportional to the square of the connected users of the system" then it always reminds me of those early B2B days during late 90s when the world was about to change. Metcalfe law was very popular during those times also! B2B exchanges were considered the pillars of the new economy and their valuations made us lose our sleep. It seemed to all of us that everything was going to be re-defined and we wanted to be part of it. Well, it didn't happen exactly as we were led to believe! Despite the differences between B2B and Linkedata like one is trasactional and other is knowledge oriented, privately owned vs free, different technologies, different era etc., there are some similarities in both of them. Lets do some introspection so that we don't repeat some of those mistakes!

B2B (Business to Business) exchanges is an entity which brings multiple buyers and sellers to a  marketplace where all kinds of commodity, financial instruments, intellectual property and various other goods can be eletronically traded - web is used as the medium in most of the cases. It has following characterstics:
  • The perceived value followed Metcalfe's law
  • Enterprenuers could fundamentally re-invent how work gets done
  • No longer comparison between big-small and so on
  • Shift of power from producers to consumers who are in control of everything
  • Standardized marketplace and standardized contracts
  • Markets operated at fraction of physical word cost
  • Global reach and one stop shopping
  • Neutrality, transparency, self-regulation, market efficiency, confidentiality and anonymity were other virtues of this marketplace
  • The winner, of a particluar vetical B2B , takes all 
It did have some good principles but a very large number of the B2B exchanges failed within few years. There were few fundamental reasons for their failure:

  • The internet enterpreneurs didn't really understand their place in the overall marketplace
  • What they offered to business was something that already existed - at least in some form
  • Companies had always done business with other businesses and most of these businesses negotiate to lower the price. Adapting those existing processes to an internet format didn't really create anything new or different in the field. Unfortunately, many companies felt that switching to an internet-based sites controlled by third parties was a riskier bet than staying with the current partners and vendors
Few of these exchanges figured it out early and survived by either addressing these issues or by being creative about it - like some of them just focused on small businesses. ECN (Electronic Communication Network) are another success story in the world of B2B exchanges as they allowed a more efficient price discovery mechanism for stocks and currencies.  In the end, B2B exchanges were all about attracting buyers and sellers or producers and consumers which they couldn't do. In a similar way, the success of Linked data cloud will depend on creating a marketplace which should be able to attract producers and consumers or buyers or sellers. The technical design prinicples by Sir Tim Berners Lee are just great and well thought of. We are also aware of the existing issues in the Linked data Cloud like quality of data, what is out there, disambiguations issues, trust of the source, frequency of the update, how to get started, should we always publish in RDF, how do you erase inaccuracy and many others things like this. We had very similar issues, maybe in a different flavor, when Web 1.0 was developing and we have come a long way. We might say that HTML was much easier to adapt then RDF but the success or wide adaption of the original web was not only simplicity of HTML and HTTP protocol but because it was a great sales, marketing and ecommerce tool. In context of Linked data cloud, lets accept the fact that most of the businesses will not be  interested in "the future," they will always be interested in "their future." Their first responsibility will always be their share holders.

Companies need very strong incentives or business case to publish their data to this "Cloud" and write applications to consume the collective intelligence of the nodes in this cloud. Initiatives like exposing government data (UK and US at this point), dbpedia, scientific, map oriented data, music, individual research projects and so many other examples which can be good and interesting from exploration standpoint are already underway. We will also see more invovlement from non-profits and intelligence communities at some point. It is a great effort but still not enough! The Cloud will continue to grow and will add billions of more RDF triples but we need to proactively involve the corporate world. Technology enterpreneurs for Semantic Web or Linked data can't do it alone. They have to partner with the business community and the enterprise. The value proposition of Linked data needs to be articulated to the CXO community and their participation needs to be encouraged. We have to start having more conversations like the "benefits of publishing data to the cloud for the enterprise." Basically, much more on business development, business case analysis, education, marketing then technology alone - technology will always remain important and core. I had written an earlier blog about a new category name for Semantic technology which talks about some aspects. Once we have their attention and buy in then the ecosystem of new tools, applications, security infrastructure, new architecture, developers, APIs, SEO, creative sales and marketing ideas, consulting, outsourcing, selling and buying and many either good things will start maturing. It will be very important to do so in next few years to sustain this momentum of Linked data cloud. It will also attract significant investments in this effort! Don't we think that there are dollars to be spent on the Linked data cloud in the  worldwide IT spending of more than three trillion dollars in 2009. We should be thinking or advocating about the "guidelines" for  publishing to the Linked data Cloud like:

  • The benefits of Linked data and possible use cases.
  • How are they impacting their supply chain? 
  • What are the advantages in reciprocity of links? 
  • What should be the process of identifying the data which needs to be published
  • How will it help their revenue and help in their relationships with partners and customers?
  • What are the security issues? What kind of security tools or approaches which are out there?
  • How will it mitigate risks for the company?
  • What will be the possible compliance and legal issues for them?
  • What data access choices can they have - free access, partial access, sign up , qualification, payment?
  • What are the data packing options like - depth, breath, granularity, freshness etc..
  • Can they have handle on who is accessing this data, frequency of access, redistribution issues?
  • Can they create virtual networks of data clouds for their partners and business customers?
  • Can they promote their data?
  • Can they increase the findability of their data?
  • How can they do it in a phased manner?
  • If they use RDFa then are they part of Linked data?
  • Are they accountable for consistency, organization, correctness of their data?
I can think of so many more questions. A very clear policies and procedures aproach needs to be articulated which will lead to a governance model for publishing to Linked data cloud within an enterprise. Probably, we will have to learn more from the experience of present initiatives like publishing data from the UK government to the cloud. During the early days of web, if any company didn't have a web presence then it was perceived that they were missing the action, opportunity or future revenue. At some point, today's businesses should start looking at Linked data cloud with the same lens.
Reblog this post [with Zemanta]




Wednesday, September 23, 2009

How Open Calais initiative is helping Semantic Technology!

Open Calais initiative, started by Thomson-Reuters, is one of the most interesting things which has happened in favour of semantic web vision. Reuters acquired this technology as part of their ClearForest, one of the leading vendors in the text analytics space, acquisition. It is stated that the service could quickly become the largest repository of metadata (in the form of named entites and facts) on the Web if it stored the resulting metadata from each request. Open Calais is the "metadata extraction service" ; it is a Web service that allows you to automatically annotate content and extract information like facts and named entities (people, places, and organizations, and much more) from unstructured text. Calais uses linguistic parsing (also known as entity extraction) in a service enables way to producr RDF triples and Semantic Web data models.
Open Calais opens the door to the possibility of lowering the barrier enough for everyday users to publish semantic content. It finally does what critics say to be the greatest obstacle to the Semantic Web: Taking the metadata burden from the end-user by providing an automatic meta-tagging tool. Open Calais initiative will also be one of the biggest enabler of the Linked data initiative.
Recently, CNET has joined OpenCalais initiative as one of the first commercial media companies to publish core data assets for public, programmatic use on the open semantic Web. CNET will leverage OpenCalais' connection to the rapidly expanding 'Linked Data cloud' to allow its original content -- such as tech product reviews on laptops, TVs, smart phones, and digital cameras; news articles and blog posts from its CNET News editorial staff; and parts of its core technology product catalog - to be available for public use.
Reblog this post [with Zemanta]

Tuesday, September 22, 2009

Example of Semantic Technology in action - Yahoo Search Monkey

Its is a misconception when people question the viability of semantic technologies by saying that "it isn't possible to convert all the data to RDF format?" In reality , it is not true at all. Again, it is important to keep in mind that we are not talking about semantic web, we are talking about semantic technologies. It can be best explained by talking about how Yahoo Search Monkey is using RDFa (RDF in Attributes) to enhance search results more useful and visually appealing. It is just one of the simple examples but it can hopefully answer the sceptics. It can be a boon to small businesses who can drive more traffic to their web sites. SearchMonkey looks for special data inside websites, based on a standard called RDFa. Your website should include this data so it is available to Yahoo as they crawl and index your site. This way your business information is available to any developers who build SearchMonkey apps, and you will show up with enhanced results as this gets adopted over time. The SearchMonkey platform has three main components: - "Site owners share structured data with Yahoo!, using semantic markup (microformats, RDF), standardized XML feeds, APIs (OpenSearch or other web services), and page extraction. -Third party developers build SearchMonkey applications. -Consumers customize their search experience." RDFa is a way to encode data within HTMLand XHTML pages which helps people and machines to embed structured data within HTML and XHTML pages. The underlying representation of RDFa is RDF because it is flexible enough to let publishers build an devolve their own vocabularies. You can see Search Monkey in action by clicking on this link http://www.yelp.com/search?find_desc=nobu&ns=1&rpp=10&find_loc=San+Francisco%2C+CA#find_loc=new%20york
You can see the enhanced quality of the result "nobu new york restaurant". You will probably realize that you have many using Semantic technologies without even realizing it.