Monday, 25 November 2013

Paraphrase original text

RWSL Assignment Part two of the Continuous Assessment – original Paraphrase text The original text is taken from IEEE Journal. Authors: Jiawei Han and Kevin Chen-Chuan Chang Article title: Data mining for Web intelligence Published in: Computer, volume.35, Issue no.11, pp.64 -70 Issue Date: Nov 2002 ISSN: 0018-9162 Digital Object Identifier: 10.1109/MC.2002.1046977 INSPEC Accession Number: 7469291 Sub-section title: Mining Web search-engine data Mining Web search-engine data An index-based Web search engine crawls the Web, indexes Web pages, and builds and stores huge keyword-based indices that help locate sets of Web pages that contain specific keywords. By using a set of tightly constrained keywords and phrases, an experienced user can quickly locate relevant documents. However, current keyword-based search engines suffer from several deficiencies. First, a topic of any breadth can easily contain hundreds of thousands of documents. This can lead to a search engine returning a huge number of document entries, many of which are only marginally relevant to the topic or contain only poor-quality materials. Second, many highly relevant documents may not contain keywords that explicitly define the topic, a phenomenon known as the polysemy problem. For example, the keyword data mining may turn up many Web pages related to other mining industries, yet fail to identify relevant papers on knowledge discovery, statistical analysis, or machine learning because they did not contain the data mining keyword. Based on these observations, we believe data mining should be integrated with the Web search engine service to enhance the quality of Web searches. To do so, we can start by enlarging the set of search keywords to include a set of keyword synonyms. For example, a search for the keyword data mining can include a few synonyms so that an index-based Web search engine can perform a parallel search that will obtain a larger set of documents than the search for the keywords alone would return. The search engine then can search the set of relevant Web documents obtained so far to select a smaller set of highly relevant and authoritative documents to present to the user. Web-linkage and Web-dynamics analysis thus provide the basis for discovering high-quality documents. Reference Jiawei Han; Chang, K.C.-C., "Data mining for Web intelligence," Computer , vol.35, no.11, pp.64,70, Nov 2002, doi: 10.1109/MC.2002.1046977 URL: http://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=1046977&isnumber=22439 [Accessed 17th November 2013]

No comments:

Post a Comment