Monday, 23 December 2013

literature review and summary writing: Data Analytics and Predictive Analytics, Inferential Statistics and CRISP-DM Methodology

ABSTRACT
Data Analytics as a science that deals with data has evolved immensely over the years. It has grown to become very vital in the processing and analysis of data for businesses and organisations.  Its applications are in both private and public sectors as well as in governmental, institutional and commercial areas. Using the findings of the data analysis, informed decisions can be made as well as making future predictions. This paper looks at Data analytics from Data mining perspective taking a closer look at the processes involved in the data mining analysis, in line with a recognised standard for data mining projects. There are various process models used by data analyst for carrying out data mining projects and some companies have developed their own process to match their data mining tools. An example is the process called Sample Explore Modify Model and Analyse (SEMMA) developed by the SAS Institute to match their data mining tool called SAS Enterprise Miner (SAS, 2013). Although there are different process models for data analysis and knowledge Discovery in Databases (KDD) the process reviewed in this paper is the generally recognised standard process for data mining projects which is the Cross Industry Standard Process for Data Mining (CRISP-DM). Even if the CRISP-DM process is not all encompassing, it does not really fulfil the criteria for a methodology and sometimes not all the processes proposed in the model are really needed in a project, it contains details of very useful guidelines that can be followed for any successful data mining project. There is no analysis of the various disciplines involved in the data analytics process like machine learning, visualization and artificial intelligence in this paper
Keywords
Data analytics, predictive analytics, data analysis, data mining, knowledge discovery in databases, CRISP-DM, KDD, CRISP, process model for data mining

1.     INTRODUCTION

For a layman it might be difficult to tell the difference between data analytics and data mining. While data analytics encompasses all the various disciplines from statistics, visualization, neurocomputing, machine learning and artificial intelligence to databases, data mining is just one of these many different methodologies that make up data analytics. Data mining on its own also employs various methodologies from areas like database, mathematics, statistics and pattern recognition in its analysis, following a sequence of processes. It is not mandatory to follow the predefined processes in the CRISP-DM. Other process models can also be used to accomplish the data mining task. On the other hand the data analytics profession is an area that is still evolving at a very fast pace and it is the same for data mining that is a part of it. More and more professionals are getting involved in it both within and outside the data science profession. In order to make it easy for the different parties involved and to foster better understanding and communication for the data analysts, the CRISP-DM reference model was developed and agreed on as the standard for all data mining projects (Chapman et al., 2000).
This paper takes a general look at data analytics and a more detailed look at data mining using CRISP-DM. the first section looks at data analytics and the different type of data analytics: Descriptive, Prescriptive and Predictive
It then looks at data mining and the stages in data mining before looking at Data mining processes models. This is followed by the CRISP-DM process reference model and finally the conclusion

2.     Data analytics

Data Analytics as a science that deals with data has evolved immensely over the years. It has evolved to become very vital in the processing and analysis of data for businesses and organisations and has provided huge advantages for both businesses and organisations in being able to use the finding of the analysis not just to make informed decisions but also to be able to predict future situations based 
A simple definition of Data Analytics is that it is a branch of science that deals with gathering huge data set, sometimes called Big Data, exploring and analysing the data, transforming the data and modelling the data so that hidden patterns and associations or correlations can be discovered from the data. The outcome or knowledge discovered from the data analysis can then be used to make informed decisions and to predict future outcomes and future expectations.  It is a multi-disciplinary process that incorporates different methodologies from various expert domains like Visualization, Artificial Intelligence, Machine Learning, Database, High performance computing, Statistics and Data Mining (Hilborn, 2013; Wikipedia, 2013).
Fayyad et al. (1996) in his definition referred to these multi-processes simply as Knowledge Discovery in Database (KDD) and stated that Data Mining and Statistics are just two particular steps in the entire process. Data Mining and Statistics use specific algorithms to extract patterns from the data while the other additional steps in the KDD process (Data gathering, data exploring, data analysis, data transformation and data modelling) help to infer knowledge from the data (Fayyad et al., 1996).
The data lifecycle both for Data Analytics or KDD follow more or less the same process as shown in the figure 1 below
 

 
The acquire stage involves selecting and gathering the raw data. The organize stage involves cleaning the data and doing the necessary pre-processing and storing the data. The analyse stage involves applying algorithms and methodologies to model the data while the decide stage will involve extracting the knowledge, making the decision and visualizing the outcome (Oracle, 2013). The data analytics and KDD processes are extensions of these phases
Data Analytics can be categorized roughly into 3 main areas which are Descriptive Analytics, Prescriptive Analytics and Predictive Analytics

2.1     Descriptive analytics

Descriptive analytics uses the analysis of historical data to model or classify data into groups or clusters by identifying not just single relationships in the data but rather many different relationships within and between data. It can be used for example to categorize or cluster customers into groups of high spenders, low spenders and product preferences, based on the issue being analyzed. For example a target group can be determined for the new products to be introduced (Rose Business Technologies).
Descriptive Analytics tends to provide solutions to analytic questions like, what happened and what did it happen, using the provided data as input. This aspect of Analytics generally strives to identify business opportunities and likely related problems (Dursun Delen and Haluk Demirkan, 2013).

2.2     Prescriptive Analytics

Prescriptive Analytics uses a combination of mathematical algorithms on the data to find best possible alternative actions or best alternative decisions that will improve the given analytic challenge (Dursun Delen and Haluk Demirkan, 2013).
 Prescriptive Analytics predicts future outcomes and also provides alternative measures or options for example, it shows the options on how to better utilize advantages or rather how to mitigate risk showing the effects of both decisions. Prescriptive analytics can model outcomes using both internal and external factors at the same time to provide decision options and the impact of the various options (Rose Business Technologies). 

2.3     Predictive Analytics

Predictive Analytics uses mathematical methods on the data to find out the hidden patterns, tendencies and relationship that can be inferred from the data. This aspect of Analytics provides solution to analytic questions like, what will happen and why it will happen. It is the outcome of this aspect of analytics that is used to make future predictions (Dursun Delen and Haluk Demirkan, 2013).
Predictive Analytics includes lots of other areas of specialties like the Machine Learning, Data Mining, Game Theory, Modelling and Statistical Methods. In a nutshell it uses both historical and current data to analyze the given situations and to predict the future. The main areas of predictive analytics that is most commonly used are the Decision Analysis and Optimization, Transaction profiling and Predictive modelling. In Decision Analysis and Optimization Prescriptive Analytics brings out the patterns in the given data for example customer data, and this patterns can then be used to predict the customer behavior for the present or future situation (Rose Business Technologies).  Statistics and Data mining are two of the most important methodologies used in the KDD process for patterns finding (Fayyad et al., 1996).

2.4     Statistics

Statistics is the branch of science that deals with collecting and analyzing numerical data and then using the result of the analysis to make decisions, to resolve issues or to create new products and processes (Montgomery and Runger, 2003). There are two areas of statistics: Descriptive statistics and Inferential Statistics

2.4.1     Descriptive Statistics

Descriptive Statistics is the part of statistics that, like the name implies, describes data and present a summary of the data showing the Central Tendency of the sample data or the distribution of the sample data. The Central Tendency will show if it is a normal distribution or not. The main types of Central Tendencies are the Mean, Median and the Mode of the sample data. The Mean represents the average which is the sum of the values of the variable divided by the number of items. The Median is the middle value of all the items if they are arranged in descending or ascending order while the Mode is the item that has got the most counts or that is the most frequency occurring item.  When any of the above Central Tendency is used depends of the nature of the data. If the data set includes Outliers, which is a numerical value, that is clearly different from the rest of the data set, then Median is used because it is less sensitive to outliers. The Mean is very sensitive to outliers and the Mode is mostly used to represent categorical data for example on a bar chart but it does not produce unique results like the Mean or the Median (Montgomery and Runger, 2003).

2.4.2     Inferential Statistics

 Inferential Statistics on the other hand is the area of statistics that concerns taking a sample from a large number objects (population, customers or product) and using the analysis of this sample to infer outcomes or make predictions that can be applied to the entire objects (Marshall and Jonker, 2011; Montgomery and Runger, 2003).
It involves using an appropriate sampling method to obtain the sample that will be representative of the entire objects in question. Using statistical techniques it is possible to calculate the required size of the samples that will be needed to carry out an inferential analysis. Inferential statistics involves different statistical test like the hypothesis test and the t-test using certain conventions like the Confidence Level of 95% or the Significance Level of 0.05. These values though arbitrary, are set by statistician and used as a convention based on best practice over many years (Marshall and Jonker, 2011). These values are predefined in the various data mining tools and they can also be adjusted as needed during any data mining project.
However it is worth stating at this point that stand-alone projects in statistics do not have any general standard process model like the CRISP-DM standard process that can act as a reference when doing statistical projects

3.     Data Mining

When talking about Data Analytics the first discipline that comes to mind is Data Mining. Data mining as defined by Daciuk and Jou (2013) is the process whereby Statistical and Mathematical Methods as well as patterns recognition Methodologies are used to analyse large sets of data with the aim of finding patterns, tendencies and mutual relationships in the data (Tim Daciuk and Stephan Jou, 2011; Fayyad et al., 1996). In other words Data Mining is a complex process that involves a significant amount of computing resources in observing data and extracting hidden but useful and meaningful patterns from data using different methodologies and statistical algorithms. The result of the analysis can then be used to make decisions for current situations or used to make predictions or forecast future situations for the given analytical issue (Dursun Delen and Haluk Demirkan, 2013).
Although lots of the processes can be done using sophisticated automated Data Mining tools it requires a human with the domain knowledge to be able to utilize these tools and make accurate decisions and predictions. The basic stages of the Data Mining process of which there are available in most of the Data Mining tools are:
Exploring the data, fitting models to the data, comparing the models to know which model is the best for the given analytic issue and presenting the report.

3.1     Exploring the data

This involves studying the raw data and finding out if there might be for example missing values for some variables or not, if the values provided values are misleading or not representative enough, if values need to be replaced or modified before the doing the data analysis. This is generally regarded as the cleaning part, the data understanding and data preparation part of the process. It will also determine the type of Model to be used on the data. For example if the modelling has to be done with the provided data with the missing values then the Decision Tree Model will be the suitable Model  because they can best handle missing values in variables (Chapman et al., 2000; Fayyad et al., 1996).

3.2     Fitting models to the data

Based on the goal of the analysis and the function that the model should serve, this is the stage where the specific model, containing the parameters to be determined from the data, is fitted to the data. It can be classification if there are predefined classes (classifications) also known as supervised classifications, to predict the categorical class variables. It can be clustering model if there are no predefined classes or classifications and this is also known as unsupervised classification. It can also be Regression Model which can predict a variable by mapping the variable to a real value. Also at this stage it can be decided on what additional algorithm or rules to use on the data for example there are Association rules that can be used to find the datasets that occur together frequently (Fayyad et al., 1996).

3.3     Comparing the models and representation the outcomes

Based on the Data Mining tool used the models can be compared to see which one has the best answer to the analytic problem. This can be done by comparing the error outcome of each model and then choosing a representation model that best suits the situation and that can best display the outcome in a human readable form. Models for representation include Decision Tree model, linear model like the Regression model or non-linear models like the Neural Network Model. Other models include the Nearest Neighbour and the Bayesian network model (Fayyad et al., 1996). The process model used during the entire project can be a major determinant of the output type, quality and deployment method

4.     Process models

All data mining projects basically follow a sequence of processes, mostly predefined. Some analytics use their own process, some use the specific company’s process and others choose from a variety of available processes like SEMMA, KDD process, and CRISP-DM. The most commonly used standard is the Cross Industry Standard for Data Mining (CRISP-DM)according to Piatetsky-Shapiro (2007) on his website from his survey in 2007 as shown on the figure below taken from his website

Figure 2: data mining methodology poll (Piatetsky-Shapiro (2007)

4.1     Cross Industry Standard for Data Mining (CRISP-DM) – the origin

CRISP-DM is a reference model for Data Mining projects first propounded in 1996 by a group of companies who were using the Data Mining technology. The companies were the then Daimler Benz, an Automobile giant, now known as DaimlerChystler AG in Germany. The second company was the SPSS Incorporated, a computer software company in the United States of America (USA), it was acquired by IBM since 2009 and is now called IBM SPSS. The third company was the NCR Incorporation, a computer hardware and electronics company based in the USA. The fourth company was Teradata Corporation, a computer company also based in the USA. The fifth company was the OHRA Insurance Company based in the Netherlands. These companies came up with this idea to find a model that will be standard for all Data Mining related processes. They invented the acronym, CRISP-DM, got the European Commission funding for the project and finally came out with the end result which is the documentation of the standard procedure that should serve as a reference model for Data Mining processes (SPSS, Chapman et al. 2000).

4.2     What is CRISP-DM

CRISP-DM is a reference model that proposes a sequence of processes that should be followed in a typical Data Mining project. As a reference model all the proposed processes might not be needed depending on the nature of a particular project but CRISP-DM is the recognized process model standardized for use in all data mining projects (Marbán et al., 2009)
As shown in the figure 3 below, it comprises of 6 processes that entails all the processes involved in the Data Mining project lifecycle starting from the Business understanding stage and goes on sequentially until the final deployment stage. It is meant to serve as a guide for professionals in the Data Mining profession when doing any data mining project.

Figure 3: CRISP-DM process lifecycle (IBM SPSS Modeler CRISP-DM, 2011).

4.3     CRISP-DM processes

4.3.1     Business understanding

Business understanding phase of CRISP-DM  is the stage where the analyst tries to understand the objectives of the project and what the business goals are. This includes gathering as much information as possible regarding the proposed project, look at the current solutions if any and describe the problems area and how it can be resolved using data mining. Assess the situation to know the available data that can be used for the analysis taking into consideration the time factor, the financial factor. Try to figure out any risk that might be involved, make contingency plans and take note of the resources available for the project. During this phase create a data mining plan stating the data mining problem to be resolved like clustering or classification or prediction providing actual number in the initial sketch where possible. At the end of this phase create a data mining plan that should show all the phases of the project from start to finish, the estimated time for each phase and the available resources and risk for each phase (IBM SPSS Modeler CRISP-DM, 2011).

4.3.2     Data understanding

The data understanding phase involves collecting required data from available sources, studying and exploring the data. Using available methods describe and summarize the data so as to have a clear overview of the data and be able to figure out some details about the quality of the data. The data summary should show the attributes of the variables, the variable types and the values displayed using tables and/or graphs and charts. Do a quality check to discover the variables with missing values or wrong inputs and conclude the phase with a data quality report (IBM SPSS Modeler CRISP-DM, 2011).

4.3.3     Data preparation

Based on the reports from the previous phases the data can be prepared for the intended analysis type by cleaning and formatting the data if necessary. This phase involves selecting the required data from the explored data, adding additional data if needed, replacing the missing values, adding new attributes and splitting the data into test , validation and training data set as needed (IBM SPSS Modeler CRISP-DM, 2011).
A point to note here is that although the data mining tools provide this option for partitioning the data into 3 different sections: test; validation and training, most analysts think the test partition is a waste of resources and prefer to partition the data equally into validation and training so as to have enough data for the analysis. The consequence of this action though comes later in the process when comparing the models to know the best model that best resolves the problem.  If more data is assigned to the Training partition it will result in a better and stable prediction but the model assessment will not be stable in the model assessment stage. On the other hand if less data is assigned to the training partition it will result in a less stable prediction but more stable model assessment. This is an area that will require some future research in data mining tools.

4.3.4     Modelling

The modelling phase involves fitting models to the data by using the required model and setting the parameters as needed. The modelling stage can be done a couple of times by testing and trying out different models techniques with different parameters. The test can be repeating by tuning the parameters slightly as the case may be, monitor the result and draw some initial conclusions on the models. Assess the results of the various tests, compare the models using available techniques and find the best result that meets the requirement of the analytic problem (IBM SPSS Modeler CRISP-DM, 2011).

4.3.5     Evaluation

The evaluation phase involves assessing the modelling results to ensure that it aligns with the predefined business and data mining goals. Document the findings and state the conclusions and any assumptions or new issues raised from the findings. Make sure the findings are presented in human understandable forms. Summarize the activities and the decisions for all the phases and if necessary go back and review the phases making adjustments if need be (IBM SPSS Modeler CRISP-DM, 2011).

4.3.6     Deployment

This is the last phase of the project where the final results of the entire process is deployed into the production environment by way of improvement or informed decision or prediction depending on the original goal of the analysis. If necessary create a deployment plan for this stage, note all the requirements and state any further actions that need to be taken or monitored or any future requirements. Write a final report summarizing the findings of the project and any recommendations while taking into consideration the intended audience and structure the report to the meet the level of the audience. Present the project report if required (IBM SPSS Modeler CRISP-DM, 2011).

5.     Conclusion

This paper touched on data analytics in general as an area of data science that is multi-disciplinary, involving several other methodologies from Statistics, machine learning, visualization, artificial intelligence and data mining. Data analytics can be subdivided into 3 different types: descriptive, prescriptive and predictive analytics. Statistics is one of the methodologies in data analytics. It comprises of Descriptive statistics that describes data by providing a summary of the data in the form of the central tendency measures: mean, median and mode and Inferential statistics that uses different statistical tests to resolve problems and predict future outcomes. There is no cross-industry standard for standalone projects in statistics like the CRISP-DM. Data mining is a part of data analytics that uses statistical and mathematical methods to explore, analyse and model data, using the findings to make predictions and decisions. Data mining projects follow the standard process proposed in the reference model CRISP-DM which details the different phases of a data mining project, how it should be done and what the outcome of each phase should strive to be.

5.1     Further research issues

The data partition phase of the data mining process involves some ambiguities. A further research is needed in this area to resolve the problem of data partitioning in the data preparation stage. How either all 3 partitions can be used in the analysis without compromising the model assessment or how the 3 partitioning system can be reduced to 2 partitions, eliminating the test partition?
Just like the CRISP-DM there could be an industry standard for standalone statistics projects too: a CRISP for Statistical Inference, aka CRISP-SI for example or something similar that can server as reference when doing statistics prediction projects
CRISP-DM is a well detailed process model and most people refer to it as methodology although it does not fulfil the criteria of a methodology in that it does not provide how things can be done. When using a different process model like SEMMA in SAS Enterprise Mining tool the process is easy to follow because the tool is developed to support the process? A research on how to make CRISP customizable in such a way that all data mining tool can provide features that support the processes
Since there are different disciplines in the data analytics and knowledge discovery process CRISP can be extended to other areas too just like data mining so that there is a CRISP for every methodology involved in data analysis and knowledge discovery process  

6.     References

SAS Enterprise Miner – SEMMA. SAS Institute., 2013, [Accessed 21st December 2013]. URL:  http://www.sas.com/technologies/analytics/datamining/
Pete Chapman, Julian Clinton, Randy Kerber, Thomas Khabaza, Thomas Reinartz, Colin Shearer, and Rüdiger Wirth, August 2000, CRISP-DM 1.0 Step-by-step data mining guide. [Accessed 22th December 2013], URL: ftp://ftp.software.ibm.com/software/analytics/spss/support/Modeler/Documentation/14/UserManual/CRISP-DM.pdf
Hilborn, Don Leo, Healthcare Informational Data Analytics (December 3, 2013). Available at SSRN: http://ssrn.com/abstract=2362781, [Accessed 23rd December 2013].
Wikipedia, 2013, Analytics, [Last updated: 21 December 2013], Available at: http://en.wikipedia.org/wiki/Analytics, [Accessed 23rd December  2013].
Usama Fayyad, Gregory Piatetsky-Shapiro, and Padhraic Smyth. 1996. The KDD process for extracting useful knowledge from volumes of data. Commun. ACM 39, 11 (November 1996), 27-34. DOI=10.1145/240455.240464 http://doi.acm.org/10.1145/240455.240464
MySQL and Hadoop – Big Data Integration, Oracle, 2013, The Data Lifecycle photographer unknown, URL: < http://www.mysql.com/why-mysql/white-papers/mysql_wp_enterprise_ready.php>.
Oracle, 2013, MySQL and Hadoop – Big Data Integration, URL: http://www.mysql.com/why-mysql/white-papers/mysql_wp_enterprise_ready.php
Rose Business Technologies. 2012. Rose Business Technologies. [ONLINE] Available at: http://www.rosebt.com/1/post/2012/08/predictive-descriptive-prescriptive-analytics.html. [Accessed 22 December 13].
Dursun Delen and Haluk Demirkan, 2013. Data, information and analytics as services, [3. Analytics-as-a-service], Decision Support Systems 55, 1 (April 2013), 359-363. DOI=10.1016/j.dss.2012.05.044 http://dx.doi.org/10.1016/j.dss.2012.05.044
Douglas C.Montgomery and George C. Runger, 2003. Applied Statistics and Probability for Engineers. 3rd edition. United States of America: John Wiley & Sons, Inc.
Gill Marshall, Leon Jonker, An introduction to inferential statistics: A review and practical guide, Radiography, Volume 17, Issue 1, February 2011, Pages e1-e6, ISSN 1078-8174, URL: http://dx.doi.org/10.1016/j.radi.2009.12.006{Marshall, 2011, An introduction to inferential statistics: A review and practical guide}
Tim Daciuk and Stephan Jou, 2011, An introduction to data mining and predictive analytics, In Proceedings of the 2011 Conference of the Center for Advanced Studies on Collaborative Research (CASCON '11), Marin Litoiu, Eleni Stroulia, and Stephen MacKay (Eds.). IBM Corp., Riverton, NJ, USA, 323-324.
Gregory Piatetsky-Shapiro, KDnuggets : Polls : Data Mining Methodology (Aug 2007), [Accessed: 24th December 2013 ], URL:http://www.kdnuggets.com/polls/2007/data_mining_methodology.htm.
Óscar Marbán, Gonzalo Mariscal and Javier Segovia (2009), A Data Mining & Knowledge Discovery Process Model, In Data Mining and Knowledge Discovery in Real Life Applications. Book edited by: Julio Ponce and Adem Karahoca, ISBN 978-3-902613-53-0, pp. 438-453, February 2009, I-Tech, Vienna, Austria.
IBM SPSS Modeler CRISP-DM Guide, 2011, [Accessed: 20th December 2013]. URL: ftp://public.dhe.ibm.com/software/analytics/spss/documentation/modeler/14.2/en/CRISP_DM.pdf
Nadali, A., E. N. Kakhky, and H. E. Nosratabadi. "Evaluating the success level of data mining projects based on CRISP-DM methodology by a Fuzzy expert system." Electronics Computer Technology (ICECT), 2011 3rd International Conference on. Vol. 6. IEEE, 2011.

Monday, 25 November 2013

RWSL Assignment Part two of the Continuous Assessment -Summarizing


Original text summarized is in the post titled "original summary text"

This is a summary of the Introduction section from the conference article “Don't match twice: redundancy-free similarity computation with MapReduce” by Lars Kolb, Andreas Thor, and Erhard Rahm (2013). 
 

They stated that redundancy is an outstanding issue associated with comparing object pairs and finding similarities. Applications that use large amounts of data rely heavily on the Pair-wise-Similarity-Computation (PSC) operation which is a resource intensive operation and PSC uses the MapReduce (MR) model to find similarities between objects that can be strings, documents or data entities like in a database. They noted that using PSC without in-depth understanding of it results in a multiple set of pairs of objects which is very costly and ineffective. To avoid this, they noted, is possible by grouping the objects into clusters and then performing the similarity search operation on the clusters following the MR model. The MR model generates signatures like keys or tokens for each object, before grouping the objects into clusters of small sizes each. The signatures are then used to identify the clusters during the computation and the objects in the same cluster are compared to each other. They argued that although this can reduce redundancy it has disadvantages because, first, creating the signatures can be very challenging and secondly the expected high standard result can be compromised by the type of data input. Moreover although creating small clusters is good, it might not include other objects that are of the same category. A cluster with location or zip code as the signature, they remarked, might not group similar objects into the same cluster because of their varied location or zip code. Most implementations use a combination of signatures to represent objects in a cluster, undermining the fact that there are still redundancy issues with the quality of the data. Such computations reduces the effectiveness of these applications which is why they propose an "optimization approach" to eradicate redundancy 

 

Reference:

Lars Kolb, Andreas Thor, and Erhard Rahm, (2013) ‘Don’t match twice: redundancy-free similarity computation with MapReduce’ In Proceedings of the Second Workshop on Data Analytics in the Cloud (DanaC '13), ACM, New York, NY, USA, 1-5. Available: http://doi.acm.org/10.1145/2486767.2486768
[Accessed 25 November 2013]

RWSL Assignment Part two of the Continuous Assessment - Paraphrasing


This is a paraphrase of the Sub-section title: Mining Web search-engine data

In the journal article Data mining for Web intelligence  by Jiawei Han and Kevin Chen-Chuan Chang
More details of the source is in the post “Paraphrase original text ”

There are lots of problems associated with the currently available search engines when using keywords to search for topics. One problem is that a search engine can find and return a vast number of results containing a large number of records that may have little or no relevance to the intended topic. This is because any topic can very easily be spread out and be linked to or from a vast number of records. Another issue is the problem whereby a document that is of high relevance may not contain the appropriate keywords that will precisely and clearly represent its relevance and classification. This is called polysemy (Jiawei and Chang, 2002). Taking these and other similar factors into consideration, data mining should be combined with the web search engines in order to improve the standards of the services provided by search engines. A technology called “Web-linkage and Web-dynamics analysis”, an index-based search engine (Jiawei and Chang, 2002) can be used to achieve this goal. The technology scans the websites on the Web for keywords and creates a large index database using the keywords and the synonyms to the keywords. When carrying out a search these keywords and the synonyms will then be used to find Websites that have them on their websites. In order words, the index-based search engine is searching for a larger set of records using keywords and synonyms while the individual search engines will search those records to obtain the most appropriate records that best match the search criteria. This will make it easy for power users to be able to find documents that are relevant to their intended topic by using a combination of keywords and synonyms

 
Reference:
Jiawei Han and Kevin Chen-Chuan Chang, "Data mining for Web intelligence," Computer, vol.35, no.11, pp.64, 70, Nov 2002, Available: http://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=1046977&isnumber=22439
[Accessed: 17 November 2013]

RWSL Assignment Part two of the Continuous Assessment - Quotation


The original text is taken from the Journal, IT Professional, with details in the post "Quotation original source"

Sub-section title: Data challenges

Data mining is a process that includes the collecting of data, processing the data, cleaning the data, transforming and modelling the data in order to find patterns in the data that can be used to make decisions or predict future trends (Wikipedia, 2013).

 Predictive analytics is the process of making future prediction by using analytic techniques from various disciplines like Data Mining, Statistics, Modelling, Machine Learning, to analyse data (Wikipedia, 2013).

The use of data mining and predictive analytics technology in almost all areas of life has increased significantly over the years. Its applications areas has extended to areas like politics, manufacturing, education and application in general social areas like social security and safety (Colleen McCue, 2006). The data mining and predictive analysis technology has got the intelligence and ability, not just to process large data sets and look for patterns, but also to use the results to predict future trends or to make informed decisions.  This widespread application of the technology notwithstanding, Data Analysts still faces a lot of challenges in accessing the right data for the required analysis or extracting the specific data required from the excessive data that may be available (Colleen McCue, 2006). According to McCue:

The challenges associated with public safety and intelligence data transcend the oft-cited “stove pipes” that limit access and functional integration of data resources. Fundamental issues associated with the fact that most (if not all) public safety and intelligence data resources were collected for reasons other than analysis seriously limit the analyst’s ability to extract meaningful output from these data. Moreover, it is difficult (if not impossible) to anticipate the array of data resources that analysts will encounter during the course of their work. Incident data, narrative reports, financial transactions, telephone records, and Internet activity represent only a few of the many and varied information resources used in crime and intelligence analysis… .(McCue, 2006)
The amount of data available to the Data Analyst from which they have to extract the relevant information that relevant to the analysis being carried out is a lot. Clearly these factors impede the work of the Analyst. These shortcomings not withstanding Analysts are still able to achieve effective and very useful results

References
Wikipedia, Data Analysis 2013, Available: http://en.wikipedia.org/wiki/Data_analytics 
[Accessed 21 November 2013]

Wikipedia, Predictive Analytics, 2013, Available: http://en.wikipedia.org/wiki/Predictive_analytics  
[Accessed 21 November 2013]

McCue, Colleen, 'Data Mining and Predictive Analytics in Public Safety and Security', IT Professional, vol.8, no.4, pp.12, 18, July-Aug. 2006

 

Quotation original source

RWSL Assignment Part two of the Continuous Assessment - Quotation
The original text is taken from the Journal IT Professional
Author: Colleen McCue
Article title: Data Mining and Predictive Analytics in Public Safety and Security IT Professional (Volume: 8 , Issue: 4 ) Page(s): 12 - 18
Date of Publication: July-Aug. 2006;
ISSN: 1520-9202
Digital Object Identifier: 10.1109/MITP.2006.84 INSPEC Accession Number: 9137302
Sponsored by: IEEE Computer Society
Sub-section title: Data challenges

Data Analytics is a process that includes the collecting of data, processing the data, cleaning the data, transforming and modelling the data in order to find patterns in the data that can be used to make decisions or predict future trends (Data Analysis, Wikipedia). Predictive analytics is the process of making future prediction by using analytic techniques from various disciplines like Data Mining, Statistics, Modelling, Machine Learning, to analyse data (Predictive Analytics, Wikipedia). The use of data mining and predictive analytics technology in almost all areas of life has increased significantly over the years. Its applications areas has extended from the mostly business sphere to include other areas like politics, manufacturing, education and application in general social areas like social security and safety (Colleen McCue, 2006). The data mining and predictive analysis technology has got the intelligence and ability, not just to process large data sets and look for patterns, but also to use the results to predict future trends or to make informed decisions. The technology is applied a lot in sensitive areas like the public safety and security. This widespread application of the technology notwithstanding, Data Analysts still faces a lot of challenges like how to access the right data for the required analysis or how extract the specific data required from the excessive data that may be available.

As Colleen McCue wrote,

The challenges associated with public safety and intelligence data transcend the oft-cited “stove pipes” that limit access and functional integration of data resources. Fundamental issues associated with the fact that most (if not all) public safety and intelligence data resources were collected for reasons other than analysis seriously limit the analyst’s ability to extract meaningful output from these data. Moreover, it is difficult (if not impossible) to anticipate the array of data resources that analysts will encounter during the course of their work. Incident data, narrative reports, financial transactions, telephone records, and Internet activity represent only a few of the many and varied information resources used in crime and intelligence analysis… .



Clearly these factors impede the work of the Analyst but Analysts still work hard enough to achieve effective results that are not affected but these shortcomings

References Wikipedia, Data Analysis. http://en.wikipedia.org/wiki/Data_analytics [Accessed 21 November 2013] Wikipedia, Predictive Analytics. http://en.wikipedia.org/wiki/Predictive_analytics [Accessed 21 November 2013] McCue, Colleen, "Data Mining and Predictive Analytics in Public Safety and Security," IT Professional , vol.8, no.4, pp.12,18, July-Aug. 2006

Paraphrase original text

RWSL Assignment Part two of the Continuous Assessment – original Paraphrase text The original text is taken from IEEE Journal. Authors: Jiawei Han and Kevin Chen-Chuan Chang Article title: Data mining for Web intelligence Published in: Computer, volume.35, Issue no.11, pp.64 -70 Issue Date: Nov 2002 ISSN: 0018-9162 Digital Object Identifier: 10.1109/MC.2002.1046977 INSPEC Accession Number: 7469291 Sub-section title: Mining Web search-engine data Mining Web search-engine data An index-based Web search engine crawls the Web, indexes Web pages, and builds and stores huge keyword-based indices that help locate sets of Web pages that contain specific keywords. By using a set of tightly constrained keywords and phrases, an experienced user can quickly locate relevant documents. However, current keyword-based search engines suffer from several deficiencies. First, a topic of any breadth can easily contain hundreds of thousands of documents. This can lead to a search engine returning a huge number of document entries, many of which are only marginally relevant to the topic or contain only poor-quality materials. Second, many highly relevant documents may not contain keywords that explicitly define the topic, a phenomenon known as the polysemy problem. For example, the keyword data mining may turn up many Web pages related to other mining industries, yet fail to identify relevant papers on knowledge discovery, statistical analysis, or machine learning because they did not contain the data mining keyword. Based on these observations, we believe data mining should be integrated with the Web search engine service to enhance the quality of Web searches. To do so, we can start by enlarging the set of search keywords to include a set of keyword synonyms. For example, a search for the keyword data mining can include a few synonyms so that an index-based Web search engine can perform a parallel search that will obtain a larger set of documents than the search for the keywords alone would return. The search engine then can search the set of relevant Web documents obtained so far to select a smaller set of highly relevant and authoritative documents to present to the user. Web-linkage and Web-dynamics analysis thus provide the basis for discovering high-quality documents. Reference Jiawei Han; Chang, K.C.-C., "Data mining for Web intelligence," Computer , vol.35, no.11, pp.64,70, Nov 2002, doi: 10.1109/MC.2002.1046977 URL: http://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=1046977&isnumber=22439 [Accessed 17th November 2013]

Summary original text

RWSL Assignment Part two of the Continuous Assessment –Summary original text ISBN 978-1-4503-2202-7 Lars Kolb, Andreas Thor, and Erhard Rahm. 2013. Don't match twice: redundancy-free similarity computation with MapReduce. In Proceedings of the Second Workshop on Data Analytics in the Cloud (DanaC '13). ACM, New York, NY, USA, 1-5. DOI=10.1145/2486767.2486768 http://doi.acm.org/10.1145/2486767.2486768 ISBN: 978-1-4503-2202-7 Original text 1. INTRODUCTION Pair-wise similarity computation (PSC) is an important aspect of many data-intensive applications, e.g., identifying similar documents for clustering [10], efficient set-similarity joins in databases [16], or identifying duplicates (entity resolution) [8]. PSC usually is an expensive operation because it is inherently of O(n2) complexity and typically involves complex (string) similarity functions. Therefore, it particularly benefits from the parallel MapReduce (MR) model and we observe an increasing number of MapReduce-based PSC implementations [4, 15, 3, 6, 14, 11]. The na¨ıve approach for PSC examines the complete Cartesian product of object pairs. The resulting quadratic complexity is intolerable for large datasets even when using MR. A common approach is pruning the search space to avoid processing pairs with presumably low similarity. This is achieved by grouping objects into (possibly overlapping) clusters and restricting similarity computation to objects of the same data cluster. To this end, many MR implementations follow a similar strategy: For each object (e.g., document, string, or entity) one or more signatures (e.g., terms, tokens, or blocking keys) are generated. A signature identifies a particular cluster and objects are assigned to all clusters of their signatures. The map phase emits a (key=signature, value=object) pair for each signature. The MR framework then groups all pairs based on their key (signature) and thus groups together objects of the same cluster. The actual similarity (match) computation is performed within the reduce phase, i.e., all objects of the same clusters are compared with each other. The creation of appropriate signatures for objects is difficult because it has to balance between efficiency and data quality. On the one hand, cluster sizes should be as small as possible to reduce the number of pairs and thus increase efficiency. On the other hand, small cluster sizes tend to miss similar object pairs especially for dirty (web) data. For example, if customer objects are clustered by their location, wrong or missing zip codes may place very similar objects into different clusters. Many approaches such as document clustering [10], entity resolution [8], or sequence alignment [14] therefore make use of multiple signatures per object to ensure that similar objects are still grouped together for comparison even in the presence of data quality issues. For example, entity resolution approaches frequently apply standard blocking [2] for generating blocking keys (signatures) based on the values of one or several entity attributes. Blocking keys for finding duplicates customers in enterprise databases can be the first three letters of the customer’s name or zip code. Due to common data quality issues, e.g., missing or false zip code, it is of crucial importance to utilize several blocking keys (multi-pass blocking) to achieve sufficient match pair completeness and thus match quality compared to single-pass blocking. The sketched na¨ıve MR-based implementation for PSC is unaware of redundancy introduced by multiple signatures per object. If an object pair shares more than one signature, it will be redundantly compared by several reduce tasks that are likely to be executed on different nodes. This unnecessary computation obviously deteriorates the run-time efficiency. Consequently, eliminating redundant pair comparison is a promising optimization approach. Lars Kolb, Andreas Thor, and Erhard Rahm, (2013) ‘Don’t match twice: redundancy-free similarity computation with MapReduce’ In Proceedings of the Second Workshop on Data Analytics in the Cloud (DanaC '13), ACM, New York, NY, USA, 1-5. DOI=10.1145/2486767.2486768 URL: http://doi.acm.org/10.1145/2486767.2486768 [Accessed 25 November 2013]

Thursday, 21 November 2013

RWSL Assignment Part two of the Continuous Assessment - Quotation


RWSL Assignment Part two of the Continuous Assessment - Quotation

The original text is taken from the Journal IT Professional

Author: Colleen McCue
Article title: Data Mining and Predictive Analytics in Public Safety and Security
IT Professional (Volume: 8 ,  Issue: 4 ) Page(s): 12 - 18
Date of Publication: July-Aug. 2006; ISSN:  1520-9202
Digital Object Identifier: 10.1109/MITP.2006.84
INSPEC Accession Number: 9137302
Sponsored by: IEEE Computer Society
Sub-section title: Data challenges

Data Analytics is a process that includes the collecting of data, processing the data, cleaning the data, transforming and modelling the data in order to find patterns in the data that can be used to make decisions or predict future trends (Data Analysis, Wikipedia).

Predictive analytics is the process of making future prediction by using analytic techniques from various disciplines like Data Mining, Statistics, Modelling, Machine Learning, to analyse data (Predictive Analytics, Wikipedia).

The use of data mining and predictive analytics technology in almost all areas of life has increased significantly over the years. Its applications areas has extended from the mostly business sphere to include other areas like politics, manufacturing, education and application in general social areas like social security and safety (Colleen McCue, 2006). The data mining and predictive analysis technology has got the intelligence and ability, not just to process large data sets and look for patterns, but also to use the results to predict future trends or to make informed decisions.  The technology is applied a lot in sensitive areas like the public safety and security. This widespread application of the technology notwithstanding, Data Analysts still faces a lot of challenges like how to access the right data for the required analysis or how extract the specific data required from the excessive data that may be available. As Colleen McCue wrote,

The challenges associated with public safety and intelligence data transcend the oft-cited “stove pipes” that limit access and functional integration of data resources. Fundamental issues associated with the fact that most (if not all) public safety and intelligence data resources were collected for reasons other than analysis seriously limit the analyst’s ability to extract meaningful output from these data. Moreover, it is difficult (if not impossible) to anticipate the array of data resources that analysts will encounter during the course of their work. Incident data, narrative reports, financial transactions, telephone records, and Internet activity represent only a few of the many and varied information resources used in crime and intelligence analysis… .

Clearly these factors impede the work of the Analyst but Analysts still work hard enough to achieve effective results that are not affected but these shortcomings

References
1. Wikipedia, Data Analysis. http://en.wikipedia.org/wiki/Data_analytics 
[Accessed 21 November 2013]

2. Wikipedia, Predictive Analytics. http://en.wikipedia.org/wiki/Predictive_analytics  
[Accessed 21 November 2013]

3. McCue, Colleen, "Data Mining and Predictive Analytics in Public Safety and Security," IT Professional , vol.8, no.4, pp.12,18, July-Aug. 2006

 

Wednesday, 20 November 2013

Quotation

http://blog.apastyle.org/apastyle/2013/06/block-quotations-in-apa-style.html

RWSL Assignment Part two of the Continuous Assessment - Paraphrasing

RWSL Assignment Part two of the Continuous Assessment - Paraphrasing
 
The original text is taken from IEEE Journal.
 
Authors: Jiawei Han and Kevin Chen-Chuan Chang
Article title: Data mining for Web intelligence
Published in: Computer, volume.35, Issue no.11, pp.64 -70
Issue Date: Nov 2002
ISSN:  0018-9162
Digital Object Identifier: 10.1109/MC.2002.1046977
INSPEC Accession Number: 7469291
Sub-section title: Mining Web search-engine data
There are lots of problems associated with the currently available search engines when searching for topics using keywords. One problem is that a search engine can find and return a vast number of results containing a large number of records that may have little or no relevance to the intended topic. The reason for this being that any topic can very easily spread out and be linked to or from a vast number of records. Another issue is the polysemy problem whereby a document that is of high relevance may not contain the appropriate keywords that will precisely and clearly represent its relevance and classification. In order to improve the standards of the services provided by search engines, considering the above study, data mining should be combined with the web search engines. An technology called “Web-linkage and Web-dynamics analysis” is an index-based search engine (Jiawei and Chang, 2002) can be used to achieve this goal. The technology scans the websites on the WWW for keywords and creates a large index database. The keywords are then extended to include further sets of keyword synonyms. When carrying out a search these keywords and the synonyms will then be used to find Websites that have them on their websites. In order words the index-based search engine is searching for a larger set of records using keywords and synonyms while the individual search engines will search the records to obtain the most appropriate records that matches the search criteria. So an experienced user can enter a carefully selected keyword or group of words to easily find documents that are relevant to the intended topic
 
Reference:
Jiawei Han and Kevin Chen-Chuan Chang, "Data mining for Web intelligence," Computer, vol.35, no.11, pp.64, 70, Nov 2002
DOI: 10.1109/MC.2002.1046977
[Accessed: 17 November 2013]