2020
Vol 5, No 1 (2020): Relational Anomaly Detection in Knowledge Graphs: A Review
Abstract
Knowledge graphs (KGs) represent entities and relations in structured form, enabling intelligent reasoning and semantic search. However, large-scale KGs often contain anomalous or inconsistent relational patterns caused by noise, incompleteness, or malicious insertion. Relational anomaly detection aims to identify such irregular structures considering both entities and relationships. This paper reviews the foundations, methods, datasets, and challenges of relational anomaly detection in knowledge graphs. We categorize approaches into statistical, embedding-based, graph neural network, rule-based, and hybrid techniques. Comparative analysis highlights strengths and limitations across scalability, interpretability, and detection accuracy. Applications in fraud detection, cybersecurity, biomedical discovery, and data quality management are also discussed. Finally, open challenges such as evolving graphs, explainability, and privacy-aware anomaly detection are outlined. The review provides a consolidated understanding for researchers and practitioners working in graph mining and knowledge engineering.
Keywords: Knowledge graph, relational anomaly detection, graph mining, link prediction, graph neural networks, outlier detection, semantic data quality
Vol 5, No 1 (2020): Privacy-Preserving Knowledge Discovery: Techniques, Frameworks, and Emerging Challenges
Abstract
Knowledge discovery from data has become an essential process in domains like healthcare, finance, education, and smart cities. However, the increasing volume of personal and sensitive information in datasets raises serious privacy concerns. Privacy-preserving knowledge discovery (PPKD) aims to extract useful patterns and models without exposing confidential information of individuals or organizations. This paper reviews the foundations, techniques, and applications of privacy-preserving knowledge discovery. It discusses major approaches such as anonymization, cryptographic computation, differential privacy, federated learning, and secure multi-party mining. A comparative analysis is presented to evaluate trade-offs between data utility and privacy guarantees. The paper also highlights challenges like scalability, regulatory compliance, adversarial attacks, and interpretability. Emerging research directions including privacy-aware deep learning and decentralized data mining are also explored. The study shows that achieving strong privacy with high analytical accuracy remains a complex but critical goal for trustworthy data-driven systems.
Keywords: Privacy-preserving data mining, knowledge discovery, differential privacy, secure multi-party computation, federated learning, data anonymization
Vol 5, No 1 (2020): Neuro-Symbolic Data Mining for Hybrid Reasoning
Abstract
Neuro-symbolic data mining has emerged as a promising paradigm that integrates neural learning with symbolic reasoning to enable hybrid intelligence systems. Traditional data mining techniques based purely on statistical or neural models often lack interpretability and logical consistency, while symbolic approaches struggle with scalability and noisy data. Neuro-symbolic methods combine the strengths of both paradigms, allowing models to learn from data while also respecting structured knowledge and logical constraints. This paper reviews the concepts, architectures, algorithms, and applications of neuro-symbolic data mining for hybrid reasoning. It discusses knowledge representation, neural symbolic integration strategies, reasoning mechanisms, and learning frameworks. A comparative analysis of existing approaches is presented along with challenges such as explainability, knowledge acquisition, and computational complexity. The study concludes that neuro-symbolic data mining can significantly enhance decision-making systems in domains such as healthcare, robotics, engineering analytics, and intelligent automation.
Keywords: Neuro-symbolic learning, hybrid reasoning, knowledge mining, explainable AI, logical constraints, intelligent data mining
2019
Vol 4, No 2 (2019): Mining Dynamic Social and Information Networks: A Review
Abstract
Dynamic social and information networks are rapidly evolving structures where nodes and links change over time. Examples include online social media platforms, citation networks, communication graphs, and knowledge networks. Mining such networks is challenging because traditional static network analysis fails to capture temporal behavior, evolving communities, and dynamic influence patterns. This paper presents a comprehensive review of mining techniques for dynamic social and information networks, including temporal graph modeling, community evolution detection, link prediction, influence analysis, and anomaly detection. We discuss key algorithms, evaluation metrics, datasets, and applications in real-world scenarios such as social media analysis, recommendation systems, fraud detection, and information diffusion modeling. Challenges like scalability, noise, privacy, and interpretability are also highlighted. The review concludes with emerging research directions including deep dynamic graph learning and real-time network mining.
Keywords: Dynamic networks, social network mining, temporal graphs, community evolution, link prediction, information diffusion, dynamic graph learning
Vol 4, No 2 (2019): Big Data Mining Using Dataset Partition for Social Welfare
Abstract
Data exist all around the world. No one can measure the exact amount of data. As we know, we have a lot of data and we are struggling to store and analyze it. There is a solution to analyze the large volume of electronically stored data. The Technology in known as Big Data. These technique is capable to analysis the complex and real time data as well as complex hidden data, no matter what is the size of data. Today the population growth in very high and these leads to unemployment problem. Many people have tremendous skill but they don't get opportunity to show their skills. Partitioning Dataset and nearest data searching algorithm are used to get the useful data from the large data set. Here we highlight, how to mine the nearby data of employees for the users using Partitioning Dataset and nearest data searching algorithm.
Keywords: Big data, Data Mining, MapReduce, Hadoop, Big Data Analysis
Vol 4, No 2 (2019): A Review: Big data Analysis framework
Abstract
Big data is about high volume and often high velocity data streams with highly diverse data types. To make maximum utilization of massive data, there should be some solid framework to manage the data. In this paper, various architectures have been studied for review. The universal architecture for big data is applicable for common massive data application. In many situations, you may need to react to the current state of data. There is also another class of architecture, real-time big data analysis architecture. Real-time big data architecture is to analyze the stream of online data. Streaming computing is designed to handle a continuous stream of a large amount of unstructured data. Batch mode architecture analyzes the data which are stored earlier. We have also studied RUBA architecture which deals with analysis of real-time data. There is some system which provides recommendation regarding some product or service based on experience of pervious customer or consumer. These types of system collects large amount of data and then analyze them called recommendation system. The architecture of a recommendation system, consumer behavior analysis has been presented here.
Keywords: Big data, big data architecture, big data framework, Real-time big data
Vol 4, No 2 (2019): Mental Health Disease using Classification Approach in Data Mining
Abstract
Data mining is the best method for finding out the Prediction of data through various sources. It Provide the information about the methodology we are presenting. Through data mining techniques, useful evidence can be collected from such source, which can help fitness seekers to get immediate support for their fitness related problems. This paper presents investigation on small data mining method mainly in mental disease dataset. Three various algorithms of data mining i.e. KNN, SVM, Naïve Bayes are applied on a disease dataset, to analyze the performance of the classifiers. The predictive rate is evaluated using four evaluation parameters i.e. Correctness, exactness, recollection and measuring point. The implementation is performed in Matlab 7.15 tool shows that Naïve Bayes outdoes as compared to remaining classifiers. This gives performance measure in various techniques in a single dataset and also comparison with previous result.
Keywords: KNN, SVM, Global, Local
Vol 4, No 2 (2019): Evaluation of Customer Ratings on Restaurant by Clustering Techniques using R
Abstract
In today’s modern times food and lifestyle has become integral part of human system. People today aspire for good day at work and sumptuous and delicious food to eat at the end of the day. Hospitality sector has come up in a big way and in this business serving good food with great ambience has become mandatory to attract customers. So, in this paper we try to collect data from a restaurant in Bangalore and evaluate its popularity based on ratings given by customers. In this paper we use K- Means clustering techniques to cluster the popular restaurants. This analysis would help people to choose restaurants for better food and ambience
Keywords: Food, Clustering, Data Science, Hospitality, K-Means
Vol 4, No 1 (2019): Truth Discovery in Big Data Social Media Application
Abstract
In this system first one is “misinformation spread” where a significant number of sources are contributing to false claims, making the identification of truthful claims difficult. For example, on, Instagram, rumors, Twitter scams, and influence bots are common examples of sources colluding, either intentionally or unintentionally, to spread misinformation and obscure the truth. The challenge is “data sparsity” or the “long-tail phenomenon” where a majority of sources only contribute a small number of claims, providing insufficient evidence to determine those sources’ trustworthiness. For example, in the Twitter datasets that we collected during real-world events, more than 90only contributed to a single claim. Third, many current solutions are not scalable to large-scale social sensing events because of the centralized nature of their truth discovery algorithms. We are going develop a Scalable and Robust Truth Discovery (SRTD) scheme to address the above all challenges. In this, the SRTD scheme jointly quantifies both the reliability of sources and the credibility of claims using a principled approach.
Keywords: Database Management, Query Processing, Scalable and Robust Truth Discovery (SRTD), Truth Discovery
Vol 4, No 1 (2019): Secure Health Monitoring and Predictive System Using Cloud Computing
Abstract
Security, in information technology is the shield of statistical knowledge and IT assets against centralized and outmost, vicious and unexpected hazards. Security is demanding for enterprises and organizations of all sizes and in all industries. Weak security can result in compromised systems or data, either by a vicious hazard actor or an unintentional internal hazard. Now-a-days healthcare business is growing tremendously due to the increase in elderly population and decline in birthrate. A healthcare becomes a big concern due to lack of opportunity of expert doctors. Due to this concern there is a paradigm shift from need based health monitoring to preventive health monitoring service. Keeping in view this scheme we are proposing a health care system which will be unified with cloud computing. That will make system adept of generating EMR i.e. Electronic Medical Records of patients which will play a favorable role for patient’s diagnostic and speedy gain process as well as for medical practicing doctors who need vast medical cases for their own study purpose. This system will keep track of patient’s health in a timely manner and generate a alert when the patient’s vital parameters crosses the normal value. This paper focuses on a health care monitoring with security
Keywords: Security, Navigation, Global Positioning System, Cloud, Geo-location, Augmented Reality
Vol 4, No 1 (2019): Crime Rate Prediction using Data Mining Algorithms
Abstract
Crime analysis and prevention is a systematic approach for identifying and analyzing patterns and trends in crime. Our system can predict regions which have high probability for crime occurrence and can visualize crime prone areas. With the increasing advent of computerized systems, crime data analysts can help the Law enforcement officers to speed up the process of solving crimes. Using the concept of data mining we can extract previously unknown, useful information from an unstructured data. Here we have an approach between computer science and criminal justice to develop a data mining procedure the can help solve crimes faster. Instead of focusing on causes of crime occurrence like criminal background of offender, political enmity etc. We are focusing mainly on crime factors of each day. Crime is one of the most important social problems in the country, affecting public safety, children development, and adult socioeconomic status. Understanding what factors cause higher crime is critical for policy makers in their efforts to reduce crime and increase citizens’ life quality. We tackle a fundamental problem in our paper: crime rate inference at the neighborhood level. Traditional approaches have used demographics and geographical influences to estimate crime rates in a region.
Keywords: Data Mining, Linear regression, Crime rate analysis
Vol 4, No 1 (2019): Role Mining for Context-Aware Recommendation Using Cluster on E-Commerce Clothing
Abstract
Recommendation systems have been trying to utilize context information to recommend services that better meet the needs of the consumers. However, current service recommendation techniques are mainly based on individual intelligence or the local knowledge of users. The main objective of this paper is to develop a web application which will provide clothing recommendations based on the group of users having the same roles
Keywords: Recommendation, context-aware, role mining, cluster.
Vol 4, No 1 (2019): A Comparative Study about Improving Accuracy in Big Data
Abstract
Big data is one of the most popular and most significant used terms to describe the epidemic growth and availability of data in the modern age. It describes any multitudinous amount of both structured and unstructured data that has the potential to be mined for information. Due to the voluminous amount of data, it becomes more complex to perform in effective manner. Nowadays, it is not provide accurate result. If data collection (or) analysis may not have enough accuracy, then it will critically affect the decision based on an analytics. This paper provides a survey of recent work on accuracy approaches. In particular, the paper considers Classification techniques, Filtering techniques, Prediction method, Mining method and Estimation Method. Our results discussed the comparison of improving accuracy and different level of accuracy.
Keywords: Big Data Analytics, Accuracy, K- Means++, CSO, Ranking Approach, Load Balancing Technique;
2018
Vol 3, No 2 (2018): Knowledge Graph Completion Using Deep Learning
Abstract
Knowledge graphs (KGs) are structured representations of entities and their relationships, widely used in domains like natural language processing, recommendation systems, and semantic search. Despite their utility, knowledge graphs are often incomplete, containing missing entities or relations, which can reduce their effectiveness. Knowledge graph completion (KGC) aims to infer these missing links and enhance the graph’s coverage. Recently, deep learning approaches have shown significant improvements in KGC by effectively modeling complex patterns and semantic information within large-scale graphs. This paper provides a comprehensive review of KGC methods leveraging deep learning, including embedding-based models, graph neural networks (GNNs), and hybrid architectures. Comparative analyses, challenges, and future research directions are discussed. The paper also presents key datasets, evaluation metrics, and illustrative results through tables and figures to highlight trends and performance in KGC research.
Keywords: Knowledge Graph Completion, Deep Learning, Graph Neural Networks, Embedding Models, Link Prediction, Semantic Knowledge Graphs
Vol 3, No 2 (2018): “A Study of Credit Risk Assessment for Vehicle Loan Applications Using Data Mining Technique”
Abstract
Data mining techniques are greatly used in the banking industry which helps them compete in the market and provide the right product to the right customer with less risk. Credit risks which account for the risk of loss and loan defaults are the major source of risk encountered by banking industry. Data mining techniques like classification and prediction can be applied to overcome this to a great extent. In this paper we introduce an effective prediction model for the bankers that help them predict the credible customers who have applied for loan. Decision Tree Induction Data Mining Algorithm is applied to predict the attributes relevant for credibility. A prototype of the model is described in this paper which can be used by the organizations in making the right decision to approve or reject the loan request of the customers.
Keywords:- NBFC, Finance Industry, Data Mining; Risk Management; Credit Scoring; Non-Performing Assets; Default Detection; Non Performing Loans Decision Tree; Credit Risk Assessment; Classification; Prediction
Vol 3, No 2 (2018): Identification of Black Spot Analysis
Abstract
The region where the road accident occurs frequently are called as Black spot .Identifying black spot of accident zones in the developing country plays an important role in the development of an economic, industrial, social and cultural fields. Due to the increasing the number of vehicle quantity, heavy traffic occurs which results into the large number of road accidents .The major factors for accidents are driver, vehicle condition, health problem, road environment such as obstruction visibility, bad shoulder, trees and poles on the shoulder. A study work has been made in this paper regarding the black spot analysis through various research papers. It deals with the process of finding the accident black spot through various factors between the two or specified region. We can find the accident zone through the cluster and regression analysis. Through the analysis of black spot region, we can take preventive measures to avoid the road accident in the particular region by placing the danger board or any other safety measures.
Keywords— Black spot; Accident; Cluster; regression.
Vol 3, No 2 (2018): Big Data: Trend, Technology and Applications
Abstract
Big Data is an excessive amount of imprecise data in assortment of formats generated from variety of sources with rapid speed. It is most buzzed terms among researcher, industry and academia. Big Data is not only limited to data perspective but it has been emerged as a stream that includes associated technologies, tools and real word applications. The objective of this paper is to provide a simple, comprehensive and brief introduction of Big Data to the beginners in subject. In this paper, we provide an overview of Hadoop and its sub-projects and a concise review of various developed technologies for Big Data. We also discuss some recent trends and eminent applications in Big Data. Though this paper does not touch each and every dimension of Big Data as it is not possible to make it in a single paper but essential aspects are covered, which may benefit to the people new in Big Data world.
Keywords: Big Data; Hadoop; Map Reduce; Yarn; Technology; Eco System
Vol 3, No 2 (2018): Clinical Data Mining and Knowledge Discovery
Abstract
Applying data mining in the medical field is a very challenging undertaking due to the idiosyncrasies of the medical profession. By introducing status and progress of data mining and analyzing characteristics of medical data, a mathematical model of medical data mining based on computation intelligence such as data preprocessing, statistical techniques, artificial neural network, decision tree and manifold learning have been introduced. This paper focusing on data mining on clinical datasets, examining of high risk population with hypertension, cancer and child obesity etc., the main objective of data mining in clinical data warehouse had been an appropriate and sufficiently sensitive method to analyze the outcomes of effective treatments and reduce costs. Medical diagnosis is extremely important but complicated task that should be performed accurately and efficiently. In contrast, the slight difference could change the balance between life or death. Towards this direction, in this study a context-aware approach is proposed, aiming to provide medical supervisors with a series of applications and personalized services targeted to exploit the multi parameter analysis for the assessment of medical –related risk factors targeting in the reduction of diseases.
Keywords: Data Mining, Clinical data, Regression Analysis, Artificial Neural Network, Decision Tree
Vol 3, No 1 (2018): Real Time Sentiment Analysis Based On Social Media Mining
Abstract
Social media is a platform where people from any corner of the world share their opinions with each other thus current era is called as social media era. Social networking sites such as Facebook and Twitter plays an important role in information retrieval and web data analysis. In a survey it is found that Twitter produces more than 500 millions of tweets each day which is about 8 TB of data which can be mined and sentiments analysis can be carried out. The purpose of mining and exacting opinions is to discover and categorize the positive and negative sentiments of society. So as there is a huge repository of data, clustering is efficient and quick way to study people’s expression and get a conclusion. This paper surveys the different mining techniques to carry out opinions and sentiments analysis and represent it to the best way based on subjectivity and polarity.
Keywords: Social Media, Opinion Mining, Sentiments Analysis, Clustering, Expressions, Social Networking.
Vol 3, No 1 (2018): Survey on Community Question Answering System to Solve the Lexical Gap
Abstract
Web search engines give a ranked list of related documents based on users keywords which depends on various aspects like popularity measures, keyword match, and frequency of accessing documents in which users have to check every specific document for getting the desired information and it causes information retrieval a prolonged process. Community Question Answering (CQA) system focus to deliver users short and precise answers instead of irrelevant documents. CQA is a specialized application which deals with information retrieval which has an ability to retrieve the right answers to questions posed in natural language. Natural Language Processing (NLP) techniques used to process a question, then searches for the required information regarding user questions to determine the answer accurately. This survey mainly focus on different approaches to solve the problems arising due to the lexical gap and also rank the accurate answers in question and answering blogs such as community question answering websites.
Keywords: Community Question Answering system, Information Retrieval, Lexical Gap, Natural Language Processing.
Vol 3, No 1 (2018): Temporal Information Identification and Normalization from Multilingual Social Media
Abstract
Social media generates massive amount of multilingual real time data. On Many occasion, Social media have proved that it can be fastest media to break event compare to any other traditional news media e.g. News channel, Internet news portal. In this work, the identification and the normalization of temporal information have been done on tweets. Each tweet can represent or describe event as opposed to previous work which rely of bursty keyword detection. In the Identification of temporal entity, three approach: Conditional Random Field, Support Vector Machine and Rule base approach used. In the Normalization of temporal entity, Rule base approach used. The rule base approach proceeds texts as inputs and from various rules, it identify temporal information. The conditional approach is trained using tagged data and then using generated model tagging on test dataset done. Morphological, Syntactic and gazetteers features are defined in Conditional random field. The use of chunking in Conditional Random Filed is not effect in identification of temporal entity. In support vector machine approach, training is completed on tagged dataset and then using generated model file temporal entity tagged. The result on Hindi temporal named entity identification terms of Precision 0.88, Recall 0.81, F1 - measure 0.85 for rule base approach; Precision 0.93, Recall 0.74, F1 - measure 0.83 for conditional random field; Precision 0.73, Recall 0.69, F1 - measure 0.71 for support vector machine and its normalization in terms of Precision 0.85, Recall 0.80 and F1 measure 0.83. Rule base approach gives good result compare to other two approach for temporal entity identification. Temporal expression identification can be helpful for generation tweet calendar, Question/Answering over social media and event summarization.
Keywords: Hindi temporal, temporal expression identification
Vol 3, No 1 (2018): Study of the Impact of Big Data on Research and Knowledge Mining for Indian Data Processing Units
Abstract
The massive repository of terabytes of data is generated each day from modern information systems and digital technologies of ours surroundings such as Internet of Things and cloud computing. Analysis of these huge data requires a lot of efforts at multiple levels to extract knowledge for the purpose of decision making to do our daily practices. Therefore, big data analysis is a current area of research and development for the SOP of Indian data processing units like CRA, digi-bank, gene-bank etc. The basic objective of this paper is to explore the potential impact of big data challenges, open research issues and various tools associated with it. As a result, this research paper provides a platform to explore big data at numerous stages. Additionally, it opens a new horizon for researchers to develop the solution, based on the challenges and open research issues.
Keywords: Big data analytics, Hadoop, Massive data, structured data, Unstructured Data, Drill, and Storm
Vol 3, No 1 (2018): Making Earth More Green and Healthy using Data mining Techniques
Abstract
Trees are vital. As the biggest plants on the planet, they give us oxygen, store carbon, stabilize the soil and give life to the world's wildlife. A study from 2017 reveals the information that more than 150 acres lost every minute of every day, and 78 million acres lost every year!, yet there is no proper technical solution to identify the places of deforestation and to plant trees on those places. This may lead to a dangerous future. But on the other hand the growth of technology in our day-to-day life is immeasurable and uncontrollable. So the technology can be used to control this scenario. This project is all about afforestation, which is one of the important things to save our future generations. This project is an innovative idea for making earth more green and healthy. Of course, it is a long term project of finding the places with fewer trees and sowing seeds on those places either manually or using drone. Those seeds are selected based on the place and the current monsoon on that place so that trees could grow well by adopting themselves to the surroundings. In this paper we present a better solution to prevent the loss of trees.
Keywords: Data mining techniques, Deforestation, Afforestation
2017
Vol 2, No 2 (2017): A Framework for Handling Evolving Data Streams
Abstract
With recent advancement in technology need for analysis of such unbounded streams is increasing day by day. Data mining process helps to excavate useful knowledge from rapidly generated raw data streams. In context with the continuously generated data, mining data streams is emerging challenging task in which several issues like limited space, limited time, accuracy, handling evolving data need to be considered. In this paper the main method of research is clustering which is focused to handle evolving data streams. Most of the previously proposed methods inherit the drawbacks of k means method and fail to handle the issues. A hybrid data mining approach encompassing windowing, grid and density clustering and divide and merge method is proposed in this paper. A dynamic data stream clustering algorithm (DDS) is used in which a dynamic density threshold is designed to accommodate the changing density of grids with time in data stream. At last divide and merge approach is used to handle varying data points and further refine the quality of result obtained.
Keywords: Data streams; Data mining; Concept drift; Clustering; Threshold
Vol 2, No 2 (2017): Study and Analysis on Data Mining Technologies for Diabetes Mellitus Prediction
Abstract
Data mining now-a-days plays a dynamic role in prediction of Disease. Data Mining is frequently defined as the process of determining patterns, correlations, trends or relationships by searching through a huge amount of data stored in repositories, databases, and data warehouses. Diabetes mellitus is an enduring disease and a foremost public health challenge allover. Diabetes affected over 246 million people worldwide with a common of them being women. According to the WHO report, by 2025 this number is projected to increase to over 380 million. This paper provides a survey and analysis of data mining methods that have been commonly applied to diabetes data analysis and prediction of the disease.
Keywords: Data mining, Diabetes mellitus, Classification, C4.5, Naïve Bayes, CART, Bayes Network, Pima Indian Diabetes Data(PIDD)