Why Internet search has emerged as such a powerful tool? One of the reasons is the haphazard growth of knowledge on the Internet. Domain name system of Internet is not intended to classify the knowledge created on Internet. For example, .Org domain is supposed to be used by non-profit entities, which has nothing to do with the content. Similarly, .Com domain is expected to be used for commercial businesses and so on. Internet domain system allows us to virtually mimic the hierarchical structure of organizations. Therefore, once you own MyCompany.Com domain, you can establish multiple sub-domains like Subsidiary1.MyCompany.Com, Subsidiary2.MyCompany.Com and sub-sub-domains. There are root domain servers that classify this naming scheme in a hierarchical tree like structure, which allows you to quickly map from a domain name to an IP address that computers can use to communicate with each other. Well, this was a quick introduction to Internet domain name system to clarify two things:
1. Internet domain naming system has a hierarchical tree like structure
2. Domains are classified into categories
3. Our search of IP addresses from domain names is through the domain name tree
Domain naming system does not have any relationship with the content. Therefore, All-About-Birds.Org content may be related to cooking. This creates tons of opportunities for search engines to index all this knowledge, represent this knowledge in a search-able format and let people search through the knowledge in an easy manner. Most of these searches are based on pattern matches but there is some element of semantic search. A semantic search is based on the meaning of the content. For example, a search of "black African sunflower" is about African sunflower not African people. Due to nascent linguistic capabilities, modern day search engines are able to differentiate such basic semantic differences. However, search engines are likely to be confused if your earlier search referred to a particular brand of jeans for instance.
Clearly, there has to be a better way to semantically organize all the Internet content. In fact, semantic web is considered to be the next big thing after the Internet. Several people are working on using XML for semantic representation of content.
So here is a million dollar idea for you. Build a utility program that website owners can allow people to use on their website. This utility program will go through the content of each page and semantically organize the information as best as it can using cues from the XML schemas as well as some natural language processing capabilities. It will give content owners or web page owners a chance to review the semantic organization of the content and make appropriate changes as needed. Once that is done, the utility program will send that semantically indexed data to a group of semantic root servers to allow people to do more intelligent searches than is currently possible. Semantic root servers will expand and grow as more knowledge is added to them.
Such a search engine will not need as much processing power as the current generation of search engines that crawl through the website content and index that content on their own servers with links to the originating website. Basically, each web server will be responsible for semantically organizing their content and linking that into the root knowledge server. The root knowledge server will be responsible for maintaining the root trees of knowledge in a way similar to the current domain servers. However, such knowledge trees will be evolutionary and organic instead of being static and hard coded like the current root servers of Internet domain name system.
This is not the perfect machine-readable semantic web that computers can process without human intervention. But this is mid-way between where we are now and a perfect semantic web concept.
Subscribe to:
Post Comments (Atom)

No comments:
Post a Comment