Monday, 25 November 2013

RWSL Assignment Part two of the Continuous Assessment -Summarizing


Original text summarized is in the post titled "original summary text"

This is a summary of the Introduction section from the conference article “Don't match twice: redundancy-free similarity computation with MapReduce” by Lars Kolb, Andreas Thor, and Erhard Rahm (2013). 
 

They stated that redundancy is an outstanding issue associated with comparing object pairs and finding similarities. Applications that use large amounts of data rely heavily on the Pair-wise-Similarity-Computation (PSC) operation which is a resource intensive operation and PSC uses the MapReduce (MR) model to find similarities between objects that can be strings, documents or data entities like in a database. They noted that using PSC without in-depth understanding of it results in a multiple set of pairs of objects which is very costly and ineffective. To avoid this, they noted, is possible by grouping the objects into clusters and then performing the similarity search operation on the clusters following the MR model. The MR model generates signatures like keys or tokens for each object, before grouping the objects into clusters of small sizes each. The signatures are then used to identify the clusters during the computation and the objects in the same cluster are compared to each other. They argued that although this can reduce redundancy it has disadvantages because, first, creating the signatures can be very challenging and secondly the expected high standard result can be compromised by the type of data input. Moreover although creating small clusters is good, it might not include other objects that are of the same category. A cluster with location or zip code as the signature, they remarked, might not group similar objects into the same cluster because of their varied location or zip code. Most implementations use a combination of signatures to represent objects in a cluster, undermining the fact that there are still redundancy issues with the quality of the data. Such computations reduces the effectiveness of these applications which is why they propose an "optimization approach" to eradicate redundancy 

 

Reference:

Lars Kolb, Andreas Thor, and Erhard Rahm, (2013) ‘Don’t match twice: redundancy-free similarity computation with MapReduce’ In Proceedings of the Second Workshop on Data Analytics in the Cloud (DanaC '13), ACM, New York, NY, USA, 1-5. Available: http://doi.acm.org/10.1145/2486767.2486768
[Accessed 25 November 2013]

RWSL Assignment Part two of the Continuous Assessment - Paraphrasing


This is a paraphrase of the Sub-section title: Mining Web search-engine data

In the journal article Data mining for Web intelligence  by Jiawei Han and Kevin Chen-Chuan Chang
More details of the source is in the post “Paraphrase original text ”

There are lots of problems associated with the currently available search engines when using keywords to search for topics. One problem is that a search engine can find and return a vast number of results containing a large number of records that may have little or no relevance to the intended topic. This is because any topic can very easily be spread out and be linked to or from a vast number of records. Another issue is the problem whereby a document that is of high relevance may not contain the appropriate keywords that will precisely and clearly represent its relevance and classification. This is called polysemy (Jiawei and Chang, 2002). Taking these and other similar factors into consideration, data mining should be combined with the web search engines in order to improve the standards of the services provided by search engines. A technology called “Web-linkage and Web-dynamics analysis”, an index-based search engine (Jiawei and Chang, 2002) can be used to achieve this goal. The technology scans the websites on the Web for keywords and creates a large index database using the keywords and the synonyms to the keywords. When carrying out a search these keywords and the synonyms will then be used to find Websites that have them on their websites. In order words, the index-based search engine is searching for a larger set of records using keywords and synonyms while the individual search engines will search those records to obtain the most appropriate records that best match the search criteria. This will make it easy for power users to be able to find documents that are relevant to their intended topic by using a combination of keywords and synonyms

 
Reference:
Jiawei Han and Kevin Chen-Chuan Chang, "Data mining for Web intelligence," Computer, vol.35, no.11, pp.64, 70, Nov 2002, Available: http://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=1046977&isnumber=22439
[Accessed: 17 November 2013]

RWSL Assignment Part two of the Continuous Assessment - Quotation


The original text is taken from the Journal, IT Professional, with details in the post "Quotation original source"

Sub-section title: Data challenges

Data mining is a process that includes the collecting of data, processing the data, cleaning the data, transforming and modelling the data in order to find patterns in the data that can be used to make decisions or predict future trends (Wikipedia, 2013).

 Predictive analytics is the process of making future prediction by using analytic techniques from various disciplines like Data Mining, Statistics, Modelling, Machine Learning, to analyse data (Wikipedia, 2013).

The use of data mining and predictive analytics technology in almost all areas of life has increased significantly over the years. Its applications areas has extended to areas like politics, manufacturing, education and application in general social areas like social security and safety (Colleen McCue, 2006). The data mining and predictive analysis technology has got the intelligence and ability, not just to process large data sets and look for patterns, but also to use the results to predict future trends or to make informed decisions.  This widespread application of the technology notwithstanding, Data Analysts still faces a lot of challenges in accessing the right data for the required analysis or extracting the specific data required from the excessive data that may be available (Colleen McCue, 2006). According to McCue:

The challenges associated with public safety and intelligence data transcend the oft-cited “stove pipes” that limit access and functional integration of data resources. Fundamental issues associated with the fact that most (if not all) public safety and intelligence data resources were collected for reasons other than analysis seriously limit the analyst’s ability to extract meaningful output from these data. Moreover, it is difficult (if not impossible) to anticipate the array of data resources that analysts will encounter during the course of their work. Incident data, narrative reports, financial transactions, telephone records, and Internet activity represent only a few of the many and varied information resources used in crime and intelligence analysis… .(McCue, 2006)
The amount of data available to the Data Analyst from which they have to extract the relevant information that relevant to the analysis being carried out is a lot. Clearly these factors impede the work of the Analyst. These shortcomings not withstanding Analysts are still able to achieve effective and very useful results

References
Wikipedia, Data Analysis 2013, Available: http://en.wikipedia.org/wiki/Data_analytics 
[Accessed 21 November 2013]

Wikipedia, Predictive Analytics, 2013, Available: http://en.wikipedia.org/wiki/Predictive_analytics  
[Accessed 21 November 2013]

McCue, Colleen, 'Data Mining and Predictive Analytics in Public Safety and Security', IT Professional, vol.8, no.4, pp.12, 18, July-Aug. 2006

 

Quotation original source

RWSL Assignment Part two of the Continuous Assessment - Quotation
The original text is taken from the Journal IT Professional
Author: Colleen McCue
Article title: Data Mining and Predictive Analytics in Public Safety and Security IT Professional (Volume: 8 , Issue: 4 ) Page(s): 12 - 18
Date of Publication: July-Aug. 2006;
ISSN: 1520-9202
Digital Object Identifier: 10.1109/MITP.2006.84 INSPEC Accession Number: 9137302
Sponsored by: IEEE Computer Society
Sub-section title: Data challenges

Data Analytics is a process that includes the collecting of data, processing the data, cleaning the data, transforming and modelling the data in order to find patterns in the data that can be used to make decisions or predict future trends (Data Analysis, Wikipedia). Predictive analytics is the process of making future prediction by using analytic techniques from various disciplines like Data Mining, Statistics, Modelling, Machine Learning, to analyse data (Predictive Analytics, Wikipedia). The use of data mining and predictive analytics technology in almost all areas of life has increased significantly over the years. Its applications areas has extended from the mostly business sphere to include other areas like politics, manufacturing, education and application in general social areas like social security and safety (Colleen McCue, 2006). The data mining and predictive analysis technology has got the intelligence and ability, not just to process large data sets and look for patterns, but also to use the results to predict future trends or to make informed decisions. The technology is applied a lot in sensitive areas like the public safety and security. This widespread application of the technology notwithstanding, Data Analysts still faces a lot of challenges like how to access the right data for the required analysis or how extract the specific data required from the excessive data that may be available.

As Colleen McCue wrote,

The challenges associated with public safety and intelligence data transcend the oft-cited “stove pipes” that limit access and functional integration of data resources. Fundamental issues associated with the fact that most (if not all) public safety and intelligence data resources were collected for reasons other than analysis seriously limit the analyst’s ability to extract meaningful output from these data. Moreover, it is difficult (if not impossible) to anticipate the array of data resources that analysts will encounter during the course of their work. Incident data, narrative reports, financial transactions, telephone records, and Internet activity represent only a few of the many and varied information resources used in crime and intelligence analysis… .



Clearly these factors impede the work of the Analyst but Analysts still work hard enough to achieve effective results that are not affected but these shortcomings

References Wikipedia, Data Analysis. http://en.wikipedia.org/wiki/Data_analytics [Accessed 21 November 2013] Wikipedia, Predictive Analytics. http://en.wikipedia.org/wiki/Predictive_analytics [Accessed 21 November 2013] McCue, Colleen, "Data Mining and Predictive Analytics in Public Safety and Security," IT Professional , vol.8, no.4, pp.12,18, July-Aug. 2006

Paraphrase original text

RWSL Assignment Part two of the Continuous Assessment – original Paraphrase text The original text is taken from IEEE Journal. Authors: Jiawei Han and Kevin Chen-Chuan Chang Article title: Data mining for Web intelligence Published in: Computer, volume.35, Issue no.11, pp.64 -70 Issue Date: Nov 2002 ISSN: 0018-9162 Digital Object Identifier: 10.1109/MC.2002.1046977 INSPEC Accession Number: 7469291 Sub-section title: Mining Web search-engine data Mining Web search-engine data An index-based Web search engine crawls the Web, indexes Web pages, and builds and stores huge keyword-based indices that help locate sets of Web pages that contain specific keywords. By using a set of tightly constrained keywords and phrases, an experienced user can quickly locate relevant documents. However, current keyword-based search engines suffer from several deficiencies. First, a topic of any breadth can easily contain hundreds of thousands of documents. This can lead to a search engine returning a huge number of document entries, many of which are only marginally relevant to the topic or contain only poor-quality materials. Second, many highly relevant documents may not contain keywords that explicitly define the topic, a phenomenon known as the polysemy problem. For example, the keyword data mining may turn up many Web pages related to other mining industries, yet fail to identify relevant papers on knowledge discovery, statistical analysis, or machine learning because they did not contain the data mining keyword. Based on these observations, we believe data mining should be integrated with the Web search engine service to enhance the quality of Web searches. To do so, we can start by enlarging the set of search keywords to include a set of keyword synonyms. For example, a search for the keyword data mining can include a few synonyms so that an index-based Web search engine can perform a parallel search that will obtain a larger set of documents than the search for the keywords alone would return. The search engine then can search the set of relevant Web documents obtained so far to select a smaller set of highly relevant and authoritative documents to present to the user. Web-linkage and Web-dynamics analysis thus provide the basis for discovering high-quality documents. Reference Jiawei Han; Chang, K.C.-C., "Data mining for Web intelligence," Computer , vol.35, no.11, pp.64,70, Nov 2002, doi: 10.1109/MC.2002.1046977 URL: http://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=1046977&isnumber=22439 [Accessed 17th November 2013]

Summary original text

RWSL Assignment Part two of the Continuous Assessment –Summary original text ISBN 978-1-4503-2202-7 Lars Kolb, Andreas Thor, and Erhard Rahm. 2013. Don't match twice: redundancy-free similarity computation with MapReduce. In Proceedings of the Second Workshop on Data Analytics in the Cloud (DanaC '13). ACM, New York, NY, USA, 1-5. DOI=10.1145/2486767.2486768 http://doi.acm.org/10.1145/2486767.2486768 ISBN: 978-1-4503-2202-7 Original text 1. INTRODUCTION Pair-wise similarity computation (PSC) is an important aspect of many data-intensive applications, e.g., identifying similar documents for clustering [10], efficient set-similarity joins in databases [16], or identifying duplicates (entity resolution) [8]. PSC usually is an expensive operation because it is inherently of O(n2) complexity and typically involves complex (string) similarity functions. Therefore, it particularly benefits from the parallel MapReduce (MR) model and we observe an increasing number of MapReduce-based PSC implementations [4, 15, 3, 6, 14, 11]. The na¨Ä±ve approach for PSC examines the complete Cartesian product of object pairs. The resulting quadratic complexity is intolerable for large datasets even when using MR. A common approach is pruning the search space to avoid processing pairs with presumably low similarity. This is achieved by grouping objects into (possibly overlapping) clusters and restricting similarity computation to objects of the same data cluster. To this end, many MR implementations follow a similar strategy: For each object (e.g., document, string, or entity) one or more signatures (e.g., terms, tokens, or blocking keys) are generated. A signature identifies a particular cluster and objects are assigned to all clusters of their signatures. The map phase emits a (key=signature, value=object) pair for each signature. The MR framework then groups all pairs based on their key (signature) and thus groups together objects of the same cluster. The actual similarity (match) computation is performed within the reduce phase, i.e., all objects of the same clusters are compared with each other. The creation of appropriate signatures for objects is difficult because it has to balance between efficiency and data quality. On the one hand, cluster sizes should be as small as possible to reduce the number of pairs and thus increase efficiency. On the other hand, small cluster sizes tend to miss similar object pairs especially for dirty (web) data. For example, if customer objects are clustered by their location, wrong or missing zip codes may place very similar objects into different clusters. Many approaches such as document clustering [10], entity resolution [8], or sequence alignment [14] therefore make use of multiple signatures per object to ensure that similar objects are still grouped together for comparison even in the presence of data quality issues. For example, entity resolution approaches frequently apply standard blocking [2] for generating blocking keys (signatures) based on the values of one or several entity attributes. Blocking keys for finding duplicates customers in enterprise databases can be the first three letters of the customer’s name or zip code. Due to common data quality issues, e.g., missing or false zip code, it is of crucial importance to utilize several blocking keys (multi-pass blocking) to achieve sufficient match pair completeness and thus match quality compared to single-pass blocking. The sketched na¨Ä±ve MR-based implementation for PSC is unaware of redundancy introduced by multiple signatures per object. If an object pair shares more than one signature, it will be redundantly compared by several reduce tasks that are likely to be executed on different nodes. This unnecessary computation obviously deteriorates the run-time efficiency. Consequently, eliminating redundant pair comparison is a promising optimization approach. Lars Kolb, Andreas Thor, and Erhard Rahm, (2013) ‘Don’t match twice: redundancy-free similarity computation with MapReduce’ In Proceedings of the Second Workshop on Data Analytics in the Cloud (DanaC '13), ACM, New York, NY, USA, 1-5. DOI=10.1145/2486767.2486768 URL: http://doi.acm.org/10.1145/2486767.2486768 [Accessed 25 November 2013]

Thursday, 21 November 2013

RWSL Assignment Part two of the Continuous Assessment - Quotation


RWSL Assignment Part two of the Continuous Assessment - Quotation

The original text is taken from the Journal IT Professional

Author: Colleen McCue
Article title: Data Mining and Predictive Analytics in Public Safety and Security
IT Professional (Volume: 8 ,  Issue: 4 ) Page(s): 12 - 18
Date of Publication: July-Aug. 2006; ISSN:  1520-9202
Digital Object Identifier: 10.1109/MITP.2006.84
INSPEC Accession Number: 9137302
Sponsored by: IEEE Computer Society
Sub-section title: Data challenges

Data Analytics is a process that includes the collecting of data, processing the data, cleaning the data, transforming and modelling the data in order to find patterns in the data that can be used to make decisions or predict future trends (Data Analysis, Wikipedia).

Predictive analytics is the process of making future prediction by using analytic techniques from various disciplines like Data Mining, Statistics, Modelling, Machine Learning, to analyse data (Predictive Analytics, Wikipedia).

The use of data mining and predictive analytics technology in almost all areas of life has increased significantly over the years. Its applications areas has extended from the mostly business sphere to include other areas like politics, manufacturing, education and application in general social areas like social security and safety (Colleen McCue, 2006). The data mining and predictive analysis technology has got the intelligence and ability, not just to process large data sets and look for patterns, but also to use the results to predict future trends or to make informed decisions.  The technology is applied a lot in sensitive areas like the public safety and security. This widespread application of the technology notwithstanding, Data Analysts still faces a lot of challenges like how to access the right data for the required analysis or how extract the specific data required from the excessive data that may be available. As Colleen McCue wrote,

The challenges associated with public safety and intelligence data transcend the oft-cited “stove pipes” that limit access and functional integration of data resources. Fundamental issues associated with the fact that most (if not all) public safety and intelligence data resources were collected for reasons other than analysis seriously limit the analyst’s ability to extract meaningful output from these data. Moreover, it is difficult (if not impossible) to anticipate the array of data resources that analysts will encounter during the course of their work. Incident data, narrative reports, financial transactions, telephone records, and Internet activity represent only a few of the many and varied information resources used in crime and intelligence analysis… .

Clearly these factors impede the work of the Analyst but Analysts still work hard enough to achieve effective results that are not affected but these shortcomings

References
1. Wikipedia, Data Analysis. http://en.wikipedia.org/wiki/Data_analytics 
[Accessed 21 November 2013]

2. Wikipedia, Predictive Analytics. http://en.wikipedia.org/wiki/Predictive_analytics  
[Accessed 21 November 2013]

3. McCue, Colleen, "Data Mining and Predictive Analytics in Public Safety and Security," IT Professional , vol.8, no.4, pp.12,18, July-Aug. 2006

 

Wednesday, 20 November 2013

Quotation

http://blog.apastyle.org/apastyle/2013/06/block-quotations-in-apa-style.html

RWSL Assignment Part two of the Continuous Assessment - Paraphrasing

RWSL Assignment Part two of the Continuous Assessment - Paraphrasing
 
The original text is taken from IEEE Journal.
 
Authors: Jiawei Han and Kevin Chen-Chuan Chang
Article title: Data mining for Web intelligence
Published in: Computer, volume.35, Issue no.11, pp.64 -70
Issue Date: Nov 2002
ISSN:  0018-9162
Digital Object Identifier: 10.1109/MC.2002.1046977
INSPEC Accession Number: 7469291
Sub-section title: Mining Web search-engine data
There are lots of problems associated with the currently available search engines when searching for topics using keywords. One problem is that a search engine can find and return a vast number of results containing a large number of records that may have little or no relevance to the intended topic. The reason for this being that any topic can very easily spread out and be linked to or from a vast number of records. Another issue is the polysemy problem whereby a document that is of high relevance may not contain the appropriate keywords that will precisely and clearly represent its relevance and classification. In order to improve the standards of the services provided by search engines, considering the above study, data mining should be combined with the web search engines. An technology called “Web-linkage and Web-dynamics analysis” is an index-based search engine (Jiawei and Chang, 2002) can be used to achieve this goal. The technology scans the websites on the WWW for keywords and creates a large index database. The keywords are then extended to include further sets of keyword synonyms. When carrying out a search these keywords and the synonyms will then be used to find Websites that have them on their websites. In order words the index-based search engine is searching for a larger set of records using keywords and synonyms while the individual search engines will search the records to obtain the most appropriate records that matches the search criteria. So an experienced user can enter a carefully selected keyword or group of words to easily find documents that are relevant to the intended topic
 
Reference:
Jiawei Han and Kevin Chen-Chuan Chang, "Data mining for Web intelligence," Computer, vol.35, no.11, pp.64, 70, Nov 2002
DOI: 10.1109/MC.2002.1046977
[Accessed: 17 November 2013]

Sunday, 17 November 2013

Summarizing

Tips on how to summarize:
A pdf file on how to summarize as of 23 november 2013

1. It has to include all the main points,
2. It has to be concise
3. It has to maintain a good paragraph structure
4. Topic sentence should identify itself as a summary - it should include the Title/Autor/Speaker of  original material
5. Supporting sentences are the main points and should follow the order of information as original
6. Paraphrasing and no direct quotes or copy and paste from original
7. Be objective and not add you own opinion
8. Concluding sentences
9. Should be about one third the lenght of the original material
Source: Youtube video on how to summarize
Accessed: 17.11.2013



1. Prediction from the Title about what it is
2. Ask Question - what information should be there?
3. Compare - compare prediction with questions
4. Visualize - try to picture it
5. Summary - use the keywords from the visualization to form the summary

Source: Youtube video on how to summarize
Accessed: 17.11.2013

Paraphrasing

Use Synonyms - similar words
Use Antonyms - using the opposite of the original word
Use phrasal verbs - use phrase to represent the original word
General verbs - verbs that represent a phrase
Phrases - phrase that represent a word
Quotation - the original sentence in quote

source:http://www.youtube.com/watch?v=sgMJ16WUEPg
Accessed: 17.11.2013

Saturday, 2 November 2013

Final Sumarry

Kingsley Gaius Ufumwen

Data Analytics (DT228B)
Research Writing and Scientific Literature
Dublin Institute of Technology

Assignment part 1: Search Strategy
Date: 2.11.2013

Search Strategy

Finding a topic
Finding a topic was very difficult as computer science as a whole and data analytics in particular is a very broad area of study.  A topic of interest might not be the topic can be very challenging. It might not be a topic where one is already skilled or it might be a topic that one is particularly interested but it might not be the areas of study where one wishes to contribute. It might also not be the area that needs further development based on one’s intention or need of the society or of a particular society if one happens to have a particular focus.

Somehow by comparing the pros and the cons and usefulness of my intention to my target environment I decided on the topic:
Utilization of Data Analytics for improvements in the developing world

First task was to get some comprehensive overview on this research area by trying to acquire some relevant and trusted information on the topic by searching for research papers and company white papers, journals and articles, conferences, podcasts, books and websites on the topic

Overview of Data Analytics and Developing World
Data Analytics - Simply put, Data Analytics is the process of gathering data and then inspecting the data, cleaning the data, transforming the data and modelling the data in order to be able to find patterns or discover some knowledge from the data that can be used to make future predictions or simply used to make good decisions

Developing World – developing world are countries or nations with low standard of living, less developed industries and low human development index compared to other countries
Developing country, Available: http://en.wikipedia.org/wiki/Developing_world
[Accessed: 02.11.2013]

Initial research to gather more information
There are lots of websites and links that one can get from searching any topic. The challenge is to be able to find the good and relevant websites that are not fake and then to be able to filter the relevant information from the search result. Assuming one is on the correct website or search engine filtering the search result can be based on the specific topic and more specific criteria relevant to the topic sentence. To achieve this it is necessary to create some keywords or words combination and review the search results

Creating Keywords from the topic
Before creating the keywords it was necessary to first split the topic into concepts that will be a shorter combination of words. The concept can be of every possible variety of words combinations. It somehow seems like applying statistical probability theory but every words combination can produce a unique search result that will be of vital importance in the research.
 More information can be found in my post titled “keywords”

 Some of the words combinations are:
·         Data analytic utilization in development
·         Using Data Analytics
·         Development in underdeveloped countries
·         Development using data analytics
·         Developing with data analytics
·         Data analytics in rural development
·         ICT Development in poor countries
·         Data analysis and countries development
·         Data analytics and development
·         Data analytics advantages
·         Advantages of data analytics
·         Benefits of data analytics
·         Uses of data analytics
·         Pros of data analytics
·         Data analytics in government
·         Data analytics in economic


Finding synonyms
Apart from the keywords or words combination concept it is also important to try to find synonyms to the words because this will help further in the searching process. The reason for this is because most search engines and databases might have the same information stored under different names or stored under a different title using a word that is just the synonym of what we are looking for.

 I have more information of on this in the post titled “possible synonyms”
Some of the synonyms:
·         Developing world; developing countries; developing nations
·         Under-developed countries; underdeveloped
·         Third world nations; third world countries
·         Data Analytics; analytics solution; business analytics; business intelligence;
·         Utilization; application; employment; implementation; exertion; usage; operation; use;
·         Development; growth; advancement; evolution; expansion; improvement; progress; developing; advancing; increasing;  progress; progression; augmentation; boost; evolving;
·         Analysis; study; search; investigation; inquiry

Brainstorming the topic
In the post titled “Brainstorming the topic” I tried to further decompose the topic into further smaller chunks and into other possible areas of study and of relevance using PESLTE analysis method and a self-composed acronym - MMECT(manufacturing, Medical, Education, Commercial and Transportation)

What is a PESTLE Analysis?
PESTLE is an acronym that stands for "political, economic, social, technological, legal and environmental." PESTLE analysis is used to identify various external political, economic, social, technological, legal and environmental factors
 Gregory Hamel, (Demand Media), Reason to Use SWOT & PESTLE Analysis
Available:
http://smallbusiness.chron.com/reason-use-swot-pestle-analysis-40810.html,
[Accessed:29th October 2013]

Political: political data analysis; data analytics in politics; data analytics and politics; data analysis and politics; data analysis in politics; data analysis in political development; development and data analytics/analytics; political trends and data analysis; data analysis in government; data analysis in institutions; the use of data analysis in political institutions/parastatals/organizations/organisations/ departments

Economics: economic data analysis; economic Analytics; economic analysis in ICT; economic development and data analysis; big data and economic analysis

Social: analyzing/analysing social network data; role of data analytics in social media; analysing/analyzing data from social media/social network; developing the social network with data analytics; improving the social network with data analytics; implementing data analytics in social network/social media; applying data analysis in social network; big data and social network

Technological: technical data analysis; growth of data analytics in technology; role of data analytics in technology; using data analytics for further development; exploring data analytics in technology

Legal: legal data analysis; role of data analytics in legal systems; using data analytics to combat crimes; using data analytics to capture criminals; using data analytics to prevent crimes; using data analytics to control crimes; using data analytics in criminal investigations; using data as evidence in court

Environmental: big data in environmental development; maps and data analytics; big data and the environment; environmental data analysis

manufacturing: data analytics in manufacturing

Medical: data analytics in medicine; application of data analytics in medical science; bioinformatics and data analysis

Educational: introducing data analysis in educational syllabus; using data analysis to improve learning/education; promoting research in data analytics; research in data analytics

Commercial: improving trade using data analysis; developing export and import using data analysis;  commercial use of data analysis

Transportation: planning road network using big data; development using statistical data; controlling traffic using data analysis; development infrastructure using data analysis

Identifying available resources
After creating the concepts with possible keywords and word combinations the next step is to try to identify available resources where the information can be found. Details of this step can be found in the post “Journal sources and ACM Digital Library”, “Journal review sheet”, “Conferences” and “Databases”

A list of some Available resources
The DIT Library both the physical shelves and the online resources, Journals, conferences, books, white papers, articles, podcasts, blogs; Forums, news magazines

The DIT library also provide a service by creating an Athens account for student and with this created account a student can access most of the Digital libraries and have full access to the Journals

A list of the sites I accessed includes:
Search engines like Google scholar
Databases like ACM, IEEE,
Digital libraries

More details can be found n the post “Journal sources and ACM Digital Library”, “Journal review sheet”, “Conferences” and “Databases”

The Review sheet
The review list is on the post titled “Journal review sheet”

Journals list
A complete list of the Journals is in the post “Journal review sheet”. Some of the Journals I found in my search are:

1.       International Journal of Data Analysis Techniques and Strategies

ISSN: 17558050, 17558069

This Journal explains the topic in-depth with real world examples and in a way that non Analysts can also understand

2.       Advances in Data Analysis and Classification

ISSN: 18625347, 18625355

This Journal also explains the topic in a variety of ways especially from a statistically perspective

3.       Educational modelling language: modelling reusable, interoperable, rich and personalised units of learning.

ISSN: 1467-8535

The Authors of this Journal are Rob Koper and  Jocelyn Manderveld and it was first published in August of 2004. It is a series from the British Journal of Educational Technology with 9000 citations. The Journal weights in on the importance of introducing Technologies into educational system

4.       Journal: ACM transactions on management information systems

ISSN:2158-656X EISSN:2158-6578

This Journal deals with the development of Analytical technologies that can be used to analyse business data in a way that will provide knowledge of how to improve on business and market processes so as to be able to achieve better operational efficiency and customer satisfaction

5.       Bringing Analytics to the Masses

ISSN : 0018-9162

This Journal explains the possibilities available to the ordinary, non-technology oriented people. It shows that there are ready-made technology packages with in-built, so-to-speak, analytics capabilities

 

More information on the Journals can be found in the post “Journal review sheet”
Conferences list
1.       Proceedings of the sixth Australasian conference on Data mining and analytics - Volume 70
ISBN: 978-1-920682-51-4
2.       Proceedings of the 2006 ACM/IEEE conference on Supercomputing
ISBN:0-7695-2700-0
3.       Behavior Informatics and Analytics: Let Behavior Talk
E-ISBN : 978-0-7695-3503-6 Print ISBN: 978-0-7695-3503-6
4.       Analytic estimation of subsample spatial shift using the phases of multidimensional analytic signals
ISSN: 1057-7149, EISSN: 1941-0042
5.       PALAPA - Distributed power generation for the development of underdeveloped villages in Indonesia
ISSN : 2155-6822 Print ISBN: 978-1-4577-0753-7

Full details on the conferences can be found in the post “Conferences”

Databases
Some of the databases I accessed are:
1.       Academic Search Premier (EBSCO): http://www.ebscohost.com/academic/academic-search-premier

2.       JSTOR:  http://www.jstor.org/


4.       Google Scholar:  http://scholar.google.com/

5.       Microsoft Academic Search:  http://academic.research.microsoft.com/

A list of the databases can be found in the post Databases

Additional resources
Additional resources I found useful was the DIT library. There are lot of journals available also online. Other materials that were of relevance were research articles available on most of the research databases, podcast and even videos on youtube from some of the authors of the journals are also available. There are also blogs like http://ws-dl.blogspot.ie/