EDBT 2026 Demo / reviewers in the wild / expert
Richi Nayak
dblp:99/1071
· DBLP profile ↗
77ranked-venue papers in the field
11as first author
12since 2021 · last 2025
0000-0002-9954-0159ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 26 (4 first)Data Mining & Knowledge Discovery · 23 (3 first)Other / Interdisciplinary · 13 (2 first)Database Systems & Data Management · 12 (2 first)Knowledge Engineering, Semantic Web & Information Systems · 2Big Data, Cloud & Distributed Data Systems · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Chatting with organisational data: a generative AI approach applied to scientific reports for information seekingabstractAbstract The world of data is vast and complex, harbouring valuable insights and patterns that can drive decision-making in various fields. However, accessing useful information from raw data (e.g. mining project reports) can be a formidable task, often requiring specialised skills and tools. The emergence of generative artificial intelligence (AI) has opened up an intriguing and novel means of engaging with data conversation. This article delves into the novel concept of chatting with organisational data using generative AI. We present an innovative solution that combines a generative AI chatbot (e.g. ChatGPT-QAM) with a language model (e.g. BERT) based extractive question-answer model (BERT-QAM) to generate responses based on given contexts. We use an answer verification model to resolve any disagreements between the responses of ChatGPT-QAM and BERT-QAM. We use the context filtering model to enhance responses by considering valid contexts. Our solution is tested on mining project reports made available by the Geological Survey of Queensland. Through this case study, we highlight several challenges that must be addressed to utilise this approach effectively. The case study shows that the concept of chatting with organisational data can revolutionise how we interact with complex scientific reports, which contain a mix of tables, text and images, to find valuable insights. Md. Abul Bashar, Richi Nayak |
Knowl. Inf. Syst. | 2 |
| 2025 | Relative Entropy-based Regularized Non-negative Matrix Factorization for Attributed Graph ClusteringabstractAttributed graph clustering is a fundamental task in network mining, essential for uncovering valuable insights in various applications. However, the heterogeneity of information from structural and attribute spaces poses significant challenges in achieving consistent and meaningful clustering. To address this, we propose Relative Entropy-based Regularized Non-negative Matrix Factorization (RENMF), a novel approach that integrates structural and attribute information through advanced matrix factorization techniques. RENMF employs Symmetric NMF and Projective NMF to extract community membership distributions from the structural and attribute spaces, respectively. By treating these distributions as homogeneous, RENMF preserves distinct, denoised information from both spaces while considering their heterogeneous complementary information. We introduce Relative Entropy (RE) as a novel regularization term to facilitate interaction between these spaces, aiming to maximize consistency between the discovered latent distributions. In this interaction, we leverage the asymmetric property of RE to emphasize attributes as essential complementary information for structural clustering. The RENMF model is solved using a new iterative multiplicative update rule, with convergence theoretically proven. We evaluate RENMF’s effectiveness through extensive experiments on 10 real-world networks, comparing it to 11 state-of-the-art clustering methods. The results demonstrate RENMF’s superiority in ground truth matching and key quality metrics, outperforming existing methods. Kamal Berahmand, Mehrnoush Mohammadi, Razieh Sheikhpour, Mahdi Jalili, Richi Nayak, Hassan Khosravi |
ACM Trans. Knowl. Discov. Data | 5 |
| 2024 | Joint Representation Learning with Generative Adversarial Imputation Network for Improved Classification of Longitudinal DataabstractAbstract Generative adversarial networks (GANs) have demonstrated their effectiveness in generating temporal data to fill in missing values, enhancing the classification performance of time series data. Longitudinal datasets encompass multivariate time series data with additional static features that contribute to sample variability over time. These datasets often encounter missing values due to factors such as irregular sampling. However, existing GAN-based imputation methods that address this type of data missingness often overlook the impact of static features on temporal observations and classification outcomes. This paper presents a novel method, fusion-aided imputer-classifier GAN (FaIC-GAN), tailored for longitudinal data classification. FaIC-GAN simultaneously leverages partially observed temporal data and static features to enhance imputation and classification learning. We present four multimodal fusion strategies that effectively extract correlated information from both static and temporal modalities. Our extensive experiments reveal that FaIC-GAN successfully exploits partially observed temporal data and static features, resulting in improved classification accuracy compared to unimodal models. Our post-additive and attention-based multimodal fusion approaches within the FaIC-GAN model consistently rank among the top three methods for classification. Sharon Torao-Pingi, Duoyi Zhang, Md. Abul Bashar, Richi Nayak |
Data Sci. Eng. | 4 |
| 2024 | Conditional Generative Adversarial Network for Early Classification of Longitudinal Datasets Using an Imputation ApproachabstractEarly classification of longitudinal data remains an active area of research today. The complexity of these datasets and the high rates of missing data caused by irregular sampling present data-level challenges for the Early Longitudinal Data Classification (ELDC) problem. Coupled with the algorithmic challenge of optimising the opposing objectives of early classification (i.e., earliness and accuracy), ELDC becomes a non-trivial task. Inspired by the generative power and utility of the Generative Adversarial Network (GAN), we propose a novel context-conditional, longitudinal early classifier GAN (LEC-GAN). This model utilises informative missingness, static features and earlier observations to improve the ELDC objective. It achieves this by incorporating ELDC as an auxiliary task within an imputation optimization process. Our experiments on several datasets demonstrate that LEC-GAN outperforms all relevant baselines in terms of F1 scores while increasing the earliness of prediction. Sharon Torao-Pingi, Richi Nayak, Md. Abul Bashar |
ACM Trans. Knowl. Discov. Data | 2 |
| 2023 | Enhanced Topic Modeling with Multi-modal Representation Learning
Duoyi Zhang, Yue Wang 0130, Md. Abul Bashar, Richi Nayak |
PAKDD (1) | 4 |
| 2023 | GAN-IE: Generative Adversarial Network for Information Extraction with Limited Annotated Data
Ahmed Shoeb Talukder, Richi Nayak, Md. Abul Bashar |
WISE | 2 |
| 2022 | Learning Inter- and Intra-Manifolds for Matrix Factorization-Based Multi-Aspect Data ClusteringabstractClustering on the data with multiple aspects, such as multi-view or multi-type relational data, has become popular in recent years due to their wide applicability. The approach using manifold learning with the Non-negative Matrix Factorization (NMF) framework, that learns the accurate low-rank representation of the multi-dimensional data, has shown effectiveness. We propose to include the inter-manifold in the NMF framework, utilizing the distance information of data points of different data types (or views) to learn the diverse manifold for data clustering. Empirical analysis reveals that the proposed method can find partial representations of various interrelated types and select useful features during clustering. Results on several datasets demonstrate that the proposed method outperforms the state-of-the-art multi-aspect data clustering methods in both accuracy and efficiency. Khanh Luong, Richi Nayak |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2021 | NOCOL - Nonnegative Orthogonal Constraint Outlier Learning
Balasubramaniam Thirunavukarasu, Wathsala Anupama Mohotti, Richi Nayak, Chau Yuen |
WISE (2) | 3 |
| 2021 | Discovering cluster evolution patterns with the Cluster Association-aware matrix factorization
Wathsala Anupama Mohotti, Richi Nayak |
Knowl. Inf. Syst. | 2 |
| 2021 | Mining discriminative itemsets in data streams using the tilted-time window model
Majid Seyfi, Richi Nayak, Yue Xu 0001, Shlomo Geva |
Knowl. Inf. Syst. | 2 |
| 2021 | Active Learning for Effectively Fine-Tuning Transfer Learning to Downstream TaskabstractLanguage model (LM) has become a common method of transfer learning in Natural Language Processing (NLP) tasks when working with small labeled datasets. An LM is pretrained using an easily available large unlabelled text corpus and is fine-tuned with the labelled data to apply to the target (i.e., downstream) task. As an LM is designed to capture the linguistic aspects of semantics, it can be biased to linguistic features. We argue that exposing an LM model during fine-tuning to instances that capture diverse semantic aspects (e.g., topical, linguistic, semantic relations) present in the dataset will improve its performance on the underlying task. We propose a Mixed Aspect Sampling (MAS) framework to sample instances that capture different semantic aspects of the dataset and use the ensemble classifier to improve the classification performance. Experimental results show that MAS performs better than random sampling as well as the state-of-the-art active learning models to abuse detection tasks where it is hard to collect the labelled data for building an accurate classifier. Md. Abul Bashar, Richi Nayak |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2021 | Column-Wise Element Selection for Computationally Efficient Nonnegative Coupled Matrix Tensor FactorizationabstractCoupled Matrix Tensor Factorization (CMTF) facilitates the integration and analysis of multiple data sources and helps discover meaningful information. Nonnegative CMTF (N-CMTF) has been employed in many applications for identifying latent patterns, prediction, and recommendation. However, due to the added complexity with coupling between tensor and matrix data, existing N-CMTF algorithms exhibit poor computation efficiency. In this paper, a computationally efficient N-CMTF factorization algorithm is presented based on the column-wise element selection, preventing frequent gradient updates. Theoretical and empirical analyses show that the proposed N-CMTF factorization algorithm is not only more accurate but also more computationally efficient than existing algorithms in approximating the tensor as well as in identifying the underlying nature of factors. Balasubramaniam Thirunavukarasu, Richi Nayak, Chau Yuen, Yu-Chu Tian |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2020 | A Novel Approach to Learning Consensus and Complementary Information for Multi-View Data ClusteringabstractEffective methods are required to be developed that can deal with the multi-faceted nature of the multi-view data. We design a factorization-based loss function-based method to simultaneously learn two components encoding the consensus and complementary information present in multi-view data by using the Coupled Matrix Factorization (CMF) and Non-negative Matrix Factorization (NMF). We propose a novel optimal manifold for multi-view data which is the most consensed manifold embedded in the high-dimensional multi-view data. A new complementary enhancing term is added in the loss function to enhance the complementary information inherent in each view. An extensive experiment with diverse datasets, benchmarking the state-of-the-art multi-view clustering methods, has demonstrated the effectiveness of the proposed method in obtaining accurate clustering solution. Khanh Luong, Richi Nayak |
ICDE | 2 |
| 2020 | Regularising LSTM classifier by transfer learning for detecting misogynistic tweets with small training set
Md. Abul Bashar, Richi Nayak, Nicolas Suzor |
Knowl. Inf. Syst. | 2 |
| 2020 | Efficient Outlier Detection in Text Corpus Using Rare Frequency and RankingabstractOutlier detection in text data collections has become significant due to the need of finding anomalies in the myriad of text data sources. High feature dimensionality, together with the larger size of these document collections, presents a need for developing accurate outlier detection methods with high efficiency. Traditional outlier detection methods face several challenges including data sparseness, distance concentration, and the presence of a larger number of sub-groups when dealing with text data. In this article, we propose to address these issues by developing novel concepts such as presenting documents with the rare document frequency, finding ranking-based neighborhood for similarity computation, and identifying sub-dense local neighborhoods in high dimensions. To improve the proposed primary method based on rare document frequency, we present several novel ensemble approaches using the ranking concept to reduce the false identifications while finding the higher number of true outliers. Extensive empirical analysis shows that the proposed method and its ensemble variations improve the quality of outlier detection in document repositories as well as they are found scalable compared to the relevant benchmarking methods. Wathsala Anupama Mohotti, Richi Nayak |
ACM Trans. Knowl. Discov. Data | 2 |
| 2020 | Efficient Nonnegative Tensor Factorization via Saturating Coordinate DescentabstractWith the advancements in computing technology and web-based applications, data are increasingly generated in multi-dimensional form. These data are usually sparse due to the presence of a large number of users and fewer user interactions. To deal with this, the Nonnegative Tensor Factorization (NTF) based methods have been widely used. However existing factorization algorithms are not suitable to process in all three conditions of size, density, and rank of the tensor. Consequently, their applicability becomes limited. In this article, we propose a novel fast and efficient NTF algorithm using the element selection approach. We calculate the element importance using Lipschitz continuity and propose a saturation point-based element selection method that chooses a set of elements column-wise for updating to solve the optimization problem. Empirical analysis reveals that the proposed algorithm is scalable in terms of tensor size, density, and rank in comparison to the relevant state-of-the-art algorithms. Balasubramaniam Thirunavukarasu, Richi Nayak, Chau Yuen |
ACM Trans. Knowl. Discov. Data | 2 |
| 2019 | Transfer Learning via Feature Selection Based Nonnegative Matrix Factorization
Balasubramaniam Thirunavukarasu, Richi Nayak, Chau Yuen |
WISE | 2 |
| 2019 | Fine-grained Type Inference in Knowledge Graphs via Probabilistic and Tensor Factorization MethodsabstractKnowledge Graphs (KGs) have been proven to be incredibly useful for enriching semantic Web search results and allowing queries with a well-defined result set. In recent years much attention has been given to the task of inferring missing facts based on existing facts in a KG. Approaches have also been proposed for inferring types of entities, however these are successful in common types such as 'Person', 'Movie', or 'Actor'. There is still a large gap, however, in the inference of fine-grained types which are highly important for exploring specific lists and collections within web search. Generally there are also relatively fewer observed instances of fine-grained types present to train in KGs, and this poses challenges for the development of effective approaches. In order to address the issue, this paper proposes a new approach to the fine-grained type inference problem. This new approach is explicitly modeled for leveraging domain knowledge and utilizing additional data outside KG, that improves performance in fine-grained type inference. Further improvements in efficiency are achieved by extending the model to probabilistic inference based on entity similarity and typed class classification. We conduct extensive experiments on type triple classification and entity prediction tasks on Freebase FB15K benchmark dataset. The experiment results show that the proposed model outperforms the state-of-the-art approaches for type inference in KG, and achieves high performance results in many-to-one relation in predicting tail for KG completion task. A. B. M. Moniruzzaman, Richi Nayak, Maolin Tang, Balasubramaniam Thirunavukarasu |
WWW | 2 |
| 2019 | MH-DAGMiner: maximal hierarchical sub-DAG mining in directed weighted networks
T. M. G. Tennakoon, Richi Nayak |
Knowl. Inf. Syst. | 2 |
| 2018 | Discovering Influence Hierarchy Based on Frequent Social InteractionsabstractIn this paper, we introduce a novel problem of discovering influence hierarchy to organize influential users in a social network into different levels according to their potential of spreading influence. We present a novel approach of discovering influence hierarchy utilizing the temporal aspect and flow direction of interactions among users. The influence hierarchy has the potential to visualize the information flow of the network and identify different roles such as creators, information disseminators, emerging leaders and active followers. It is highly applicable in several domains such as sociology, marketing, political science and disaster management. T. M. G. Tennakoon, Richi Nayak |
ASONAM | 2 |
| 2018 | Learning Association Relationship and Accurate Geometric Structures for Multi-Type Relational DataabstractNon-negative Matrix Factorization (NMF) methods have been effectively used for clustering high dimensional data. Manifold learning is combined with the NMF framework to ensure the projected lower dimensional representations preserve the local geometric structure of data. In this paper, considering the context of multi-type relational data clustering, we develop a new formulation of manifold learning to be embedded in the factorization process such that the new low-dimensional space can maintain both local and global structures of original data. We also propose to include the interactions between clusters of different data types by enforcing a Normalize Cut-type constraint that leads to a comprehensive NMF-based framework. A theoretical analysis and extensive experiments are provided to validate the effectiveness of the proposed work. Khanh Luong, Richi Nayak |
ICDE | 2 |
| 2018 | An Efficient Ranking-Centered Density-Based Document Clustering Method
Wathsala Anupama Mohotti, Richi Nayak |
PAKDD (3) | 2 |
| 2018 | A Novel Technique of Using Coupled Matrix and Greedy Coordinate Descent for Multi-view Data Representation
Khanh Luong, Balasubramaniam Thirunavukarasu, Richi Nayak |
WISE (2) | 3 |
| 2017 | Efficient mining of discriminative itemsetsabstractDiscriminative itemsets can be more useful than frequent itemsets as the former identifies the frequent itemsets in one dataset with much higher frequencies than the same itemsets in other datasets. The discriminative itemsets can distinguish the target dataset from all others. The discriminative itemsets are a small subset of frequent itemsets. The efficient mining of discriminative itemsets is a challenging problem, since the Apriori property of frequent itemsets is not applicable, and the designed algorithms must deal with the exponential number of itemset combinations in more than one dataset. In this paper, a novel algorithm, called DISSparse, is proposed for efficient mining of discriminative itemsets. Two determinative heuristics are proposed for limiting the mining of discriminative itemsets to the potential discriminative itemsets. Our experiments show the efficient time and space usage of the proposed algorithm in the large and complex datasets. Majid Seyfi, Richi Nayak, Yue Xu 0001, Shlomo Geva |
WI | 2 |
| 2017 | Spatial Information Recognition in Web Documents Using a Semi-supervised Machine Learning Method
Hendi Lie, Richi Nayak, Gordon F. Wyeth |
WISE (1) | 2 |
| 2016 | Using parallel hierarchical clustering to address spatial big data challengesabstractClustering can help to make large datasets more manageable by grouping together similar objects. However, most clustering approaches are unable to scale to very large datasets (e.g. more than 10 million objects). The K-Tree is a data structure and clustering algorithm that has proven to be scalable with large streaming datasets. Here, we apply the K-Tree to spatial data (satellite images) and extend from a single threaded to a multicore environment. We show that the K-Tree is able to cluster larger dataset more efficiently than baseline approaches. Alan Woodley, Ling-Xiang Tang, Shlomo Geva, Richi Nayak, Timothy Chappell |
IEEE BigData | 4 |
| 2016 | How Relevant is the Irrelevant Data: Leveraging the Tagging Data for a Learning-to-Rank ModelabstractFor the task of tag-based item recommendations, the underlying tensor model faces several challenges such as high data sparsity and inferring latent factors effectively. To overcome the inherent sparsity issue of tensor models, we propose the graded-relevance interpretation scheme that leverages the tagging data effectively. Unlike the existing schemes, the graded-relevance scheme interprets the tagging data richly, differentiates the non-observed tagging data insightfully, and annotates each entry as one of the "relevant", "likely relevant", "irrelevant", or "indecisive" labels. To infer the latent factors of tensor models correctly to produce the high quality recommendation, we develop a novel learning-to-rank method, Go-Rank, that optimizes Graded Average Precision (GAP). Evaluating the proposed method on real-world datasets, we show that the proposed interpretation scheme produces a denser tensor model by revealing "relevant" entries from the previously assumed "irrelevant" entries. Optimizing GAP as the ranking metric, the quality of the recommendations generated by Go-Rank is found superior against the benchmarking methods. Noor Ifada, Richi Nayak |
WSDM | 2 |
| 2015 | Robust clustering of multi-type relational data via a heterogeneous manifold ensembleabstractHigh-Order Co-Clustering (HOCC) methods have attracted high attention in recent years because of their ability to cluster multiple types of objects simultaneously using all available information. During the clustering process, HOCC methods exploit object co-occurrence information, i.e., inter-type relationships amongst different types of objects as well as object affinity information, i.e., intra-type relationships amongst the same types of objects. However, it is difficult to learn accurate intra-type relationships in the presence of noise and outliers. Existing HOCC methods consider the p nearest neighbours based on Euclidean distance for the intra-type relationships, which leads to incomplete and inaccurate intra-type relationships. In this paper, we propose a novel HOCC method that incorporates multiple subspace learning with a heterogeneous manifold ensemble to learn complete and accurate intra-type relationships. Multiple subspace learning reconstructs the similarity between any pair of objects that belong to the same subspace. The heterogeneous manifold ensemble is created based on two-types of intra-type relationships learnt using p-nearest-neighbour graph and multiple subspaces learning. Moreover, in order to make sure the robustness of clustering process, we introduce a sparse error matrix into matrix decomposition and develop a novel iterative algorithm. Empirical experiments show that the proposed method achieves improved results over the state-of-art HOCC methods for FScore and NMI. Richi Nayak |
ICDE | 2 |
| 2015 | Do-Rank: DCG Optimization for Learning-to-Rank in Tag-Based Item Recommendation Systems
Noor Ifada, Richi Nayak |
PAKDD (2) | 2 |
| 2015 | FreeS: A Fast Algorithm to Discover Frequent Free Subtrees Using a Novel Canonical Form
Israt Jahan Chowdhury, Richi Nayak |
WISE (1) | 2 |
| 2015 | Semi-supervised Document Clustering via Loci
Taufik Sutanto, Richi Nayak |
WISE (2) | 2 |
| 2015 | Parallel Streaming Signature EM-tree: A Clustering Algorithm for Web Scale ApplicationsabstractThe proliferation of the web presents an unsolved problem of automatically analyzing billions of pages of natural language. We introduce a scalable algorithm that clusters hundreds of millions of web pages into hundreds of thousands of clusters. It does this on a single mid-range machine using efficient algorithms and compressed document representations. It is applied to two web-scale crawls covering tens of terabytes. ClueWeb09 and ClueWeb12 contain 500 and 733 million web pages and were clustered into 500,000 to 700,000 clusters. To the best of our knowledge, such fine grained clustering has not been previously demonstrated. Previous approaches clustered a sample that limits the maximum number of discoverable clusters. The proposed EM-tree algorithm uses the entire collection in clustering and produces several orders of magnitude more clusters than the existing algorithms. Fine grained clustering is necessary for meaningful clustering in massive collections where the number of distinct topics grows linearly with collection size. These fine-grained clusters show an improved cluster quality when assessed with two novel evaluations using ad hoc search relevance judgments and spam classifications for external validation. These evaluations solve the problem of assessing the quality of clusters where categorical labeling is unavailable and unfeasible. Christopher M. De Vries, Lance De Vine, Shlomo Geva, Richi Nayak |
WWW | 4 |
| 2014 | The Ranking Based Constrained Document Clustering Method and Its Application to Social Event Detection
Taufik Sutanto, Richi Nayak |
DASFAA (2) | 2 |
| 2014 | BOSTER: An Efficient Algorithm for Mining Frequent Unordered Induced Subtrees
Israt Jahan Chowdhury, Richi Nayak |
WISE (1) | 2 |
| 2014 | Mining Discriminative Itemsets in Data Streams
Majid Seyfi, Shlomo Geva, Richi Nayak |
WISE (1) | 3 |
| 2013 | A data mining driven risk profiling method for road asset managementabstractRoad surface skid resistance has been shown to have a strong relationship to road crash risk, however, applying the current method of using investigatory levels to identify crash prone roads is problematic as they may fail in identifying risky roads outside of the norm. The proposed method analyses a complex and formerly impenetrable volume of data from roads and crashes using data mining. This method rapidly identifies roads with elevated crash-rate, potentially due to skid resistance deficit, for investigation. A hypothetical skid resistance/crash risk curve is developed for each road segment, driven by the model deployed in a novel regression tree extrapolation method. The method potentially solves the problem of missing skid resistance values which occurs during network-wide crash analysis, and allows risk assessment of the major proportion of roads without skid resistance values. Daniel Emerson, Justin Weligamage, Richi Nayak |
KDD | 3 |
| 2013 | A Recommendation Approach Dealing with Multiple Market SegmentsabstractA new community and communication type of social networks - online dating - are gaining momentum. With many people joining in the dating network, users become overwhelmed by choices for an ideal partner. A solution to this problem is providing users with partners recommendation based on their interests and activities. Traditional recommendation methods ignore the users' needs and provide recommendations equally to all users. In this paper, we propose a recommendation approach that employs different recommendation strategies to different groups of members. A segmentation method using the Gaussian Mixture Model (GMM) is proposed to customize users' needs. Then a targeted recommendation strategy is applied to each identified segment. Empirical results show that the proposed approach outperforms several existing recommendation methods. Lin Chen 0012, Richi Nayak |
Web Intelligence | 2 |
| 2013 | A Reciprocal Collaborative Method Using Relevance Feedback and Feature ImportanceabstractIn a people-to-people matching systems, filtering is widely applied to find the most suitable matches. The results returned are either too many or only a few when the search is generic or specific respectively. The use of a sophisticated recommendation approach becomes necessary. Traditionally, the object of recommendation is the item which is inanimate. In online dating systems, reciprocal recommendation is required to suggest a partner only when the user and the recommended candidate both are satisfied. In this paper, an innovative reciprocal collaborative method is developed based on the idea of similarity and common neighbors, utilizing the information of relevance feedback and feature importance. Extensive experiments are carried out using data gathered from a real online dating service. Compared to benchmarking methods, our results show the proposed method can achieve noticeable better performance. Lin Chen 0012, Richi Nayak |
Web Intelligence | 2 |
| 2013 | A Novel Method for Finding Similarities between Unordered Trees Using Matrix Data Model
Israt Jahan Chowdhury, Richi Nayak |
WISE (1) | 2 |
| 2013 | The Heterogeneous Cluster Ensemble Method Using Hubness for Clustering Text Documents
Richi Nayak |
WISE (1) | 2 |
| 2012 | A data analytics application assessing pavement deflection factors for a road networkabstractRoad networks are a vital national property however, with the passage of time their condition deteriorates. It is critical for a road agency to know the various factors such as road conditions, traffic conditions, environmental conditions that affect the Road Pavement Deflection values. This paper proposes a data analytic application for assessing the road pavement condition. The data analytics process includes acquisition and integration of data from multiple sources, pre-processing the data and mining the useful information from the data. The generated data mining models are able to demonstrate factors that affect pavement deflection data. Outputs of this research inform the road managers about the road conditions and enable them in formulating an efficient e-government policy for budget and maintenance of this road asset. Richi Nayak, Rakesh Rawat, Justin Weligamage |
iiWAS | 1 |
| 2012 | Analyzing the Effectiveness of Graph Metrics for Anomaly Detection in Online Social Networks
Reza Hassanzadeh, Richi Nayak, Douglas Stebila |
WISE | 2 |
| 2011 | Improving Matching Process in Social Network Using Implicit and Explicit User Information
Slah Alsaleh, Richi Nayak, Yue Xu 0001, Lin Chen 0012 |
APWeb | 2 |
| 2011 | Effective Hybrid Recommendation Combining Users-Searches Correlations Using Tensors
Rakesh Rawat, Richi Nayak, Yuefeng Li 0001 |
APWeb | 2 |
| 2011 | Aggregate Distance Based Clustering Using Fibonacci Series-FIBCLUS
Rakesh Rawat, Richi Nayak, Yuefeng Li 0001, Slah Alsaleh |
APWeb | 2 |
| 2011 | Finding and Matching Communities in Social Networks Using Data MiningabstractThe rapid growth in the number of users using social networks and the information that a social network requires about their users make the traditional matching systems insufficiently adept at matching users within social networks. This paper introduces the use of clustering to form communities of users and, then, uses these communities to generate matches. Forming communities within a social network helps to reduce the number of users that the matching system needs to consider, and helps to overcome other problems from which social networks suffer, such as the absence of user activities' information about a new user. The proposed system has been evaluated on a dataset obtained from an online dating website. Empirical analysis shows that accuracy of the matching process is increased using the community information. Slah Alsaleh, Richi Nayak, Yue Xu 0001 |
ASONAM | 2 |
| 2011 | A Recommendation Method for Online Dating Networks Based on Social Relations and Demographic InformationabstractA new relationship type of social networks - online dating - are gaining popularity. With a large member base, users of a dating network are overloaded with choices about their ideal partners. Recommendation methods can be utilized to overcome this problem. However, traditional recommendation methods do not work effectively for online dating networks where the dataset is sparse and large, and a two-way matching is required. This paper applies social networking concepts to solve the problem of developing a recommendation method for online dating networks. We propose a method by using clustering, SimRank and adapted SimRank algorithms to recommend matching candidates. Empirical results show that the proposed method can achieve nearly double the performance of the traditional collaborative filtering and common neighbor methods of recommendation. Lin Chen 0012, Richi Nayak, Yue Xu 0001 |
ASONAM | 2 |
| 2011 | Assessment of Cardiovascular Disease Risk Prediction Models: Evaluation Methods
Richi Nayak, Ellen Pitt |
DASFAA (2) | 1 |
| 2011 | Road crash proneness prediction using data miningabstractDeveloping safe and sustainable road systems is a common goal in all countries. Applications to assist with road asset management and crash minimization are sought universally. This paper presents a data mining methodology using decision trees for modeling the crash proneness of road segments using available road and crash attributes. The models quantify the concept of crash proneness and demonstrate that road segments with only a few crashes have more in common with non-crash roads than roads with higher crash counts. This paper also examines ways of dealing with highly unbalanced data sets encountered in the study. Richi Nayak, Daniel Emerson, Justin Weligamage, Noppadol Piyatrapoomi |
EDBT | 1 |
| 2011 | The hidden web, XML and the Semantic Web: scientific data management perspectivesabstractThe World Wide Web no longer consists just of HTML pages. Our work sheds light on a number of trends on the Internet that go beyond simple Web pages. The hidden Web provides a wealth of data in semi-structured form, accessible through Web forms and Web services. These services, as well as numerous other applications on the Web, commonly use XML, the eXtensible Markup Language. XML has become the lingua franca of the Internet that allows customized markups to be defined for specific domains. On top of XML, the Semantic Web grows as a common structured data source. In this work, we first explain each of these developments in detail. Using real-world examples from scientific domains of great interest today, we then demonstrate how these new developments can assist the managing, harvesting, and organization of data on the Web. On the way, we also illustrate the current research avenues in these domains. We believe that this effort would help bridge multiple database tracks, thereby attracting researchers with a view to extend database technology. Fabian M. Suchanek, Aparna S. Varde, Richi Nayak, Pierre Senellart |
EDBT | 3 |
| 2011 | XML Documents Clustering Using a Tensor Space Model
Sangeetha Kutty, Richi Nayak, Yuefeng Li 0001 |
PAKDD (1) | 2 |
| 2011 | Utilizing Past Relations and User Similarities in a Social Matching System
Richi Nayak |
PAKDD (2) | 1 |
| 2010 | Personalized recommender system based on item taxonomy and folksonomyabstractItem folksonomy or tag information is popularly available on the web now. However, since tags are arbitrary words given by users, they contain a lot of noise such as tag synonyms, semantic ambiguities and personal tags. Such noise brings difficulties to improve the accuracy of item recommendations. In this paper, we propose to combine item taxonomy and folksonomy to reduce the noise of tags and make personalized item recommendations. The experiments conducted on the dataset collected from Amazon.com demonstrated the effectiveness of the proposed approaches. The results suggested that the recommendation accuracy can be further improved if we consider the viewpoints and the vocabularies of both experts and users. Huizhi Liang 0001, Yue Xu 0001, Yuefeng Li 0001, Richi Nayak |
CIKM | 4 |
| 2010 | Combining Schema and Level-Based Matching for Web Service Discovery
Alsayed Algergawy, Richi Nayak, Norbert Siegmund, Veit Köppen, Gunter Saake |
ICWE | 2 |
| 2010 | A Hybrid Approach of Personalized Web Information RetrievalabstractThis paper proposes a hybrid approach of personalized Web Information Retrieval that utilizes (1) ontology for retrieval of user's context (2) user profile that is temporarily updated according to users' browsing behavior and (3) collaborative filtering for considering recommendation of similar users. Empirical analysis reveals that Precision, Recall and F-Score of most of the queries for many users are improved with using the proposed method. Namita Mittal, Richi Nayak, Mahesh Chandra Govil, Kamal Chand Jain |
Web Intelligence | 2 |
| 2010 | Element similarity measures in XML schema matching
Alsayed Algergawy, Richi Nayak, Gunter Saake |
Inf. Sci. | 2 |
| 2009 | XCFS: an XML documents clustering approach using both the structure and the contentabstractThis paper introduces a clustering approach, XML Clustering using Frequent Substructures (XCFS) that considers both the structural and the content information of XML documents in clustering. XCFS uses frequent substructures in the form of a novel representation, Closed Frequent Embedded (CFE) subtrees to constrain the content in the clustering process. The empirical analysis ascertains that XCFS can effectively cluster even very large XML datasets and outperforms other existing methods. Sangeetha Kutty, Richi Nayak, Yuefeng Li 0001 |
CIKM | 2 |
| 2009 | Knowledge Discovery over the Deep Web, Semantic Web and XML
Aparna S. Varde, Fabian M. Suchanek, Richi Nayak, Pierre Senellart |
DASFAA | 3 |
| 2009 | HCX: an efficient hybrid clustering approach for XML documentsabstractThis paper proposes a novel Hybrid Clustering approach for XML documents (HCX) that first determines the structural similarity in the form of frequent subtrees and then uses these frequent subtrees to represent the constrained content of the XML documents in order to determine the content similarity. The empirical analysis reveals that the proposed method is scalable and accurate. Sangeetha Kutty, Richi Nayak, Yuefeng Li 0001 |
ACM Symposium on Document Engineering | 2 |
| 2009 | Thai Word Segmentation with Hidden Markov Model and Decision Tree
Poramin Bheganan, Richi Nayak, Yue Xu 0001 |
PAKDD | 2 |
| 2009 | Personalized Recommender Systems Integrating Social Tags and Item TaxonomyabstractThe social tags in web 2.0 are becoming another important information source to profile users' interests and preferences to make personalized recommendations. To solve the problem of low information sharing caused by the free-style vocabulary of tags and the long tails of the distribution of tags and items, this paper proposes an approach to integrate the social tags given by users and the item taxonomy with standard vocabulary and hierarchical structure provided by experts to make personalized recommendations. The experimental results show that the proposed approach can effectively improve the information sharing and recommendation accuracy. Huizhi Liang 0001, Yue Xu 0001, Yuefeng Li 0001, Richi Nayak, Li-Tung Weng |
Web Intelligence | 4 |
| 2008 | An Interactive Predictive Data Mining System for Informed Decision
Esther Ge, Richi Nayak |
DASFAA | 2 |
| 2008 | A User Driven Data Mining Process Model and Learning System
Esther Ge, Richi Nayak, Yue Xu 0001, Yuefeng Li 0001 |
DASFAA | 2 |
| 2008 | Expertise Analysis in a Question Answer Portal for Author RankingabstractAn online question answering (QA) portal provides users a way to socialize and help each other to solve problems. The majority of the online question answer systems use user-feedback to rank userspsila answers. This way of ranking is inefficient as it involves ongoing efforts by the users and is subjective. Currently researchers have utilized link analysis of user interactions for this task. However, this is not accurate in some circumstances. A detailed structural analysis of an online QA portal is conducted in this paper. A novel approach based on userspsila reputation reflecting the usage patterns is proposed to rank and recommend the user answers. The method is compared with a popular link topology analysis method, HITS. The result of the proposed method is promising. Lin Chen 0012, Richi Nayak |
Web Intelligence | 2 |
| 2008 | An Ontology-Based Framework for Knowledge RetrievalabstractRetrieving accurate information from the Web is a great challenge to users. The existing information retrieval systems are mostly term-based and thus need to be enhanced toward knowledge-based. User information needs need to be better captured in order to deliver personalized search results. In this paper, an ontology-based framework is proposed for capturing user information needs using a world knowledge base and the user's local instance repository. The framework aims to discover a user's background knowledge for knowledge retrieval. The evaluation result is encouraging, in which the proposed model achieved the same performance as a manual user model. Xiaohui Tao 0001, Yuefeng Li 0001, Ning Zhong 0001, Richi Nayak |
Web Intelligence | 4 |
| 2008 | Improving Web Service Discovery by Using Semantic Models
Aishwarya Bose, Richi Nayak, Peter Bruza |
WISE | 2 |
| 2008 | Fast and effective clustering of XML data using structural information
Richi Nayak |
Knowl. Inf. Syst. | 1 |
| 2007 | Ontology Mining for Semantic Interpretation of Information Needs
Xiaohui Tao 0001, Yuefeng Li 0001, Richi Nayak |
KSEM | 3 |
| 2007 | Web Service Discovery with additional Semantics and ClusteringabstractDue to the lack of semantic descriptions of the Web services, the search results returned by the service registries are effectively inadequate. This paper presents the Semantic Web services Clustering (SWSC) method that extends the semantic representation of services and groups the similar Web services in order to improve the service discovery. The empirical analysis shows the improvement in service discovery with the use of SWSC. Richi Nayak, Bryan Lee |
Web Intelligence | 1 |
| 2007 | Ontology Mining for PersonalizedWeb Information GatheringabstractIt is well accepted that ontology is useful for personalized Web information gathering. However, it is challenging to use semantic relations of "kind-of", "part-of", and "related-to" and synthesize commonsense and expert knowledge in a single computational model. In this paper, a personalized ontology model is proposed attempting to answer this challenge. A two-dimensional (Exhaustivity and Specificity) method is also presented to quantitatively analyze these semantic relations in a single framework. The proposals are successfully evaluated by applying the model to a Web information gathering system. The model is a significant contribution to personalized ontology engineering and concept-based Web information gathering in Web Intelligence. Xiaohui Tao 0001, Yuefeng Li 0001, Ning Zhong 0001, Richi Nayak |
Web Intelligence | 4 |
| 2006 | XMine: A Methodology for Mining XML Structure
Richi Nayak, Wina Iryadi |
APWeb | 1 |
| 2006 | XCLS: A Fast and Effective Clustering Algorithm for Heterogenous XML Documents
Richi Nayak, Sumei Xu |
PAKDD | 1 |
| 2006 | Investigating Semantic Measures in XML ClusteringabstractThis paper discusses the influence of semantic computation in structural data such as XML for clustering similar data. We study how the semantic similarity at individual element level influences the overall similarity of documents. The empirical results indicate that the semantic measures do not play an important role in finding clusters in structural data such as XML Richi Nayak |
Web Intelligence | 1 |
| 2006 | Automatically Acquiring Training Sets for Web Information GatheringabstractThe traditional techniques rely on human effort to acquire training sets, which is expensive and inefficient. In this paper we present an alternative method to automatically acquire training sets without heavy investment of user efforts. The proposed method tends to fill a gap for effectiveness of using Web data in Web mining, and contributes to Web information gathering. The evaluation shows that the method is adequate to yield an promising achievement. Xiaohui Tao 0001, Yuefeng Li 0001, Ning Zhong 0001, Richi Nayak |
Web Intelligence | 4 |
| 2006 | Distributed Recommender Profiling and Selection with Gittins IndicesabstractMost existing recommender systems nowadays operate in a single organizational base, and very often they do not have sufficient resources to be used in order to generate quality recommendations. Therefore, it would be beneficial if recommender systems of different organizations can cooperate together to share their resources and recommendations. In this paper, we present a distributed recommender system model that consists of multiple recommender systems from different organizations. With the hope to provide better recommendation service to users, the recommender systems can improve their performances by sharing their recommendations cooperatively. A recommender selection technique based on the Gittins indices is presented in this paper, and it makes selections based on the stability, average performance and selection frequency of the recommenders Li-Tung Weng, Yue Xu 0001, Yuefeng Li 0001, Richi Nayak |
Web Intelligence | 4 |
| 2004 | Automatic integration of Heterogenous XML-schemas
Richi Nayak, Fu Bo Xia |
iiWAS | 1 |
| 2004 | Applications of Data Mining in Web Services
Richi Nayak, Cindy Tong |
WISE | 1 |