Xiaohui Tao 0001

dblp:58/3976 · DBLP profile ↗
← Back
57ranked-venue papers in the field
11as first author
31since 2021 · last 2026
0000-0002-0020-077XORCID · conflict

Domains — venue-derived; a paper can count in several

Data Mining & Knowledge Discovery · 17 (2 first)Information Retrieval & Web Search · 16 (1 first)Database Systems & Data Management · 12 (1 first)Other / Interdisciplinary · 7 (5 first)Knowledge Engineering, Semantic Web & Information Systems · 3 (2 first)Big Data, Cloud & Distributed Data Systems · 2
YearPublicationVenuePosition
2026 Divide-and-Conquer: Cold-Start Bundle Recommendation via Mixture of Diffusion Experts
abstract
Cold-start bundle recommendation focuses on modeling new bundles with insufficient information to provide recommendations. Advanced bundle recommendation models usually learn bundle representations from multiple views (e.g., user-bundle interaction views) at both bundle and item levels. Consequently, the cold-start problem for bundles is more challenging than that for traditional items due to the dual-level multi-view complexity. For cold-start bundle recommendation, we propose a novel Mixture of Diffusion Experts (MoDiffE) framework, which employs a divide-and-conquer strategy and consists of three parts: (1) Division : The bundle cold-start problem is divided into view-specific but unified sub-problems: the poor representation of feature-missing bundles in prior-embedding models. (2) Conquest : Diffusion models uniformly solve all sub-problems by directly generating diffusion representations without depending on specific features. (3) Combination : A cold-aware hierarchical Mixture of Experts (MoE) is employed to adaptively combine results of the sub-problems into final recommendations. Additionally, MoDiffE proposes a cold-start gating augmentation method to enable gating for cold bundles. In experiments on three real-world datasets, MoDiffE significantly outperforms existing solutions in cold-start bundle recommendation. It achieves up to a 0.1027 Recall@20 improvement in cold-start scenarios and up to a 47.43% relative improvement in all-bundle scenarios.
Ming Li 0072, Lin Li 0001, Xiaohui Tao 0001, Jimmy Huang 0001
ACM Trans. Inf. Syst.3
2025 A Node-Aware Dynamic Quantization Approach for Graph Collaborative Filtering
abstract
In the realm of collaborative filtering recommendation systems, Graph Neural Networks (GNNs) have demonstrated remarkable performance but face significant challenges in deployment on resource-constrained edge devices due to their high embedding parameter requirements and computational costs. Using common quantization method directly on node embeddings may overlooks their graph based structure, causing error accumulation during message passing and degrading the quality of quantized embeddings.To address this, we propose Graph based Node-Aware Dynamic Quantization training for collaborative filtering (GNAQ), a novel quantization approach that leverages graph structural information to enhance the balance between efficiency and accuracy of GNNs for Top-K recommendation. GNAQ introduces a node-aware dynamic quantization strategy that adapts quantization scales to individual node embeddings by incorporating graph interaction relationships. Specifically, it initializes quantization intervals based on node-wise feature distributions and dynamically refines them through message passing in GNN layers. This approach mitigates information loss caused by fixed quantization scales and captures hierarchical semantic features in user-item interaction graphs. Additionally, GNAQ employs graph relation-aware gradient estimation to replace traditional straight-through estimators, ensuring more accurate gradient propagation during training. Extensive experiments on four real-world datasets demonstrate that GNAQ outperforms state-of-the-art quantization methods, including BiGeaR and N2UQ, by achieving average improvement in 27.8% Recall@10 and 17.6% NDCG@10 under 2-bit quantization. In particular, GNAQ is capable of maintaining the performance of full-precision models while reducing their model sizes by 8 to 12 times; in addition, the training time is twice as fast compared to quantization baseline methods.
Lin Li 0001, Xiaohui Tao 0001, Jianwei Zhang 0002
CIKM4
2025 SimRe: A Simulation of Memes Recreation for Memes Category Detection
Lin Li 0001, Leqi Zhong, Shaopeng Tang, Xiaohui Tao 0001
DASFAA (2)5
2025 Enhancing Multi-turn Dialogue Consistency with Localized-Generalized Persona Expansion
Yanbing Chen, Xiaohui Tao 0001, Peipei Wang 0001, Lin Li 0001
DASFAA (2)3
2025 A Unified Solution to Diverse Heterogeneities in One-Shot Federated Learning
abstract
One-Shot Federated Learning (OSFL) restricts communication between the server and clients to a single round, significantly reducing communication costs and minimizing privacy leakage risks compared to traditional Federated Learning (FL), which requires multiple rounds of communication. However, existing OSFL frameworks remain vulnerable to distributional heterogeneity, as they primarily focus on model heterogeneity while neglecting data heterogeneity. To bridge this gap, we propose FedHydra, a unified, data-free, OSFL framework designed to effectively address both model and data heterogeneity. Unlike existing OSFL approaches, FedHydra introduces a novel two-stage learning mechanism. Specifically, it incorporates model stratification and heterogeneity-aware stratified aggregation to mitigate the challenges posed by both model and data heterogeneity. By this design, the data and model heterogeneity issues are simultaneously monitored from different aspects during learning. Consequently, FedHydra can effectively mitigate both issues by minimizing their inherent conflicts. We compared FedHydra with five SOTA baselines on four benchmark datasets. Experimental results show that our method outperforms the previous OSFL methods in both homogeneous and heterogeneous settings. The code is available at https://github.com/Jun-B0518/FedHydra.
Yiliao Song, Di Wu 0050, Atul Sajjanhar, Yong Xiang 0001, Wei Zhou 0044, Xiaohui Tao 0001, Yan Li 0002, Yue Li 0017
KDD (2)7
2025 Step-wise Soft Alignment Enhanced Procedural Text Generation from Long Instructional Videos
abstract
With the rise of generative models, video-language cross-modal applications have seen significant growth. Generating procedural text from instructional videos has become a crucial task, playing a key role in both understanding visual scenes and supporting practical applications. The sequential nature of video clips is particularly important, as entities may appear across multiple clips, reflecting fine-grained intra-modal self-similarity. However, most existing training methods treat other clips in a sequence as negative samples when a target is specified, neglecting their step-wise correlations. To address this limitation, we introduce Step-wise Soft Alignment via OpTimal TrAnsport (SATA), which constructs soft positive pairs to mitigate the issue. SATA first generates a step-wise similarity matrix by leveraging visual representations and generated procedural text. It then aligns the step-wise distributions between procedural text and video clips using optimal transport. The resulting transport distance serves as a weight, treating these pairs as soft positives for contrastive learning, ultimately improving the accuracy of procedural text generation. Our experiments on the publicly available YouCookII and ActivityNet Captions datasets demonstrate the effectiveness of SATA, achieving absolute improvements of 0.5% to 1.7% and 0.6% to 14.9% in paragraph-level evaluation, respectively.
Lin Li 0001, Xian Zhong, Xiaohui Tao 0001, Jianquan Liu
ICMR4
2025 Enhancing Interpretability for Computational Personality Analysis in Education
Ahmed R. Elmahalawy, Lin Li 0001, Xiaohui Tao 0001
PAKDD (2)4
2024 Aligning Bytes with Bliss: Integrating Happiness Computing with Sociological Insight
Xiaojun Wu 0001, Lin Li 0001, Xiaohui Tao 0001, Yuefeng Li 0001
ADMA (1)3
2024 CrimeAlarm: Towards Intensive Intent Dynamics in Fine-Grained Crime Prediction
Kaixi Hu, Lin Li 0001, Qing Xie 0002, Xiaohui Tao 0001, Guandong Xu
DASFAA (7)4
2024 MealRec+: A Meal Recommendation Dataset with Meal-Course Affiliation for Personalization and Healthiness
abstract
Meal recommendation, as a typical health-related recommendation task, contains complex relationships between users, courses, and meals. Among them, meal-course affiliation associates user-meal and user-course interactions. However, an extensive literature review demonstrates that there is a lack of publicly available meal recommendation datasets including meal-course affiliation. Meal recommendation research has been constrained in exploring the impact of cooperation between two levels of interaction on personalization and healthiness. To pave the way for meal recommendation research, we introduce a new benchmark dataset called MealRec^+. Due to constraints related to user health privacy and meal scenario characteristics, the collection of data that includes both meal-course affiliation and two levels of interactions is impeded. Therefore, a simulation method is adopted to derive meal-course affiliation and user-meal interaction from the user's dining sessions simulated based on user-course interaction data. Then, two well-known nutritional standards are used to calculate the healthiness scores of meals. Moreover, we experiment with several baseline models, including separate and cooperative interaction learning methods. Our experiment demonstrates that cooperating the two levels of interaction in appropriate ways is beneficial for meal recommendations. The dataset is available on GitHub (https://github.com/WUT-IDEA/MealRecPlus).
Ming Li 0072, Lin Li 0001, Xiaohui Tao 0001, Jimmy Huang 0001
SIGIR3
2024 Erdos: A Novel Blockchain Consensus Algorithm with Equitable Node Selection and Deterministic Block Finalization
abstract
Abstract The introduction of blockchain technology has brought about significant transformation in the realm of digital transactions, providing a secure and transparent platform for peer-to-peer interactions that cannot be tampered with. The decentralised and distributed nature of blockchains guarantees the integrity and authenticity of the data, eliminating the need for intermediaries. The applications of this technology are not limited to the financial sector, but extend to various areas, such as supply chain management, identity verification, and governance. At the core of these blockchains is the consensus mechanism, which plays a crucial role in ensuring the reliability and integrity of a system. Consensus mechanisms are essential for achieving an agreement amongst network participants regarding the validity of transactions and the order in which they are recorded on the blockchain. By incorporating consensus mechanisms, blockchains ensure that all honest nodes in the network reach a consensus on whether to accept or reject a block, based on predefined rules and criteria. The aim of this study is to introduce a novel consensus mechanism named Erdos, which seeks to address the shortcomings of existing consensus algorithms, such as the Proof of Work and Proof of Stake. Erdos emphasises security, decentralisation, and fairness. One notable feature of this mechanism is its equitable node-selection algorithm, which ensures equal opportunities for all nodes to engage in block creation and validation. In addition, Erdos implements a deterministic block finalisation process that guarantees the integrity and authenticity of the blockchain. The main contribution of this research lies in its innovative approach to deterministic block finalisation, which effectively mitigates the various security risks associated with blockchain systems.
Buti Sello, Jianming Yong, Xiaohui Tao 0001
Data Sci. Eng.3
2024 Optimal Treatment Strategies for Critical Patients with Deep Reinforcement Learning
abstract
Personalized clinical decision support systems are increasingly being adopted due to the emergence of data-driven technologies, with this approach now gaining recognition in critical care. The task of incorporating diverse patient conditions and treatment procedures into critical care decision-making can be challenging due to the heterogeneous nature of medical data. Advances in Artificial Intelligence (AI), particularly Reinforcement Learning (RL) techniques, enables the development of personalized treatment strategies for severe illnesses by using a learning agent to recommend optimal policies. In this study, we propose a Deep Reinforcement Learning (DRL) model with a tailored reward function and an LSTM-GRU-derived state representation to formulate optimal treatment policies for vasopressor administration in stabilizing patient physiological states in critical care settings. Using an ICU dataset and the Medical Information Mart for Intensive Care (MIMIC-III) dataset, we focus on patients with Acute Respiratory Distress Syndrome (ARDS) that has led to Sepsis, to derive optimal policies that can prioritize patient recovery over patient survival. Both the DDQN ( RepDRL-DDQN ) and Dueling DDQN ( RepDRL-DDDQN ) versions of the DRL model surpass the baseline performance, with the proposed model’s learning agent achieving an optimal learning process across our performance measuring schemes. The robust state representation served as the foundation for enhancing the model’s performance, ultimately providing an optimal treatment policy focused on rapid patient recovery.
Simi Job, Xiaohui Tao 0001, Lin Li 0001, Haoran Xie 0001, Taotao Cai, Jianming Yong, Qing Li 0001
ACM Trans. Intell. Syst. Technol.2
2024 Boosting Healthiness Exposure in Category-Constrained Meal Recommendation Using Nutritional Standards
abstract
Food computing, a newly emerging topic, is closely linked to human life through computational methodologies. Meal recommendation, a food-related study about human health, aims to provide users a meal with courses constrained from specific categories (e.g., appetizers, main dishes) that can be enjoyed as a service. Historical interaction data, important user information, is often used by existing models to learn user preferences. However, if a user’s preferences favor less healthy meals, the model will follow that preference and make similar recommendations, potentially negatively impacting the user’s long-term health. This emphasizes the necessity for health-oriented and responsible meal recommendation systems. In this article, we propose a healthiness-aware and category-wise meal recommendation model called CateRec, which boosts healthiness exposure by using nutritional standards as knowledge to guide the model training. Two fundamental questions are raised and answered: (1) How can the healthiness of meals be evaluated? Two well-known nutritional standards from the World Health Organization and the United Kingdom Food Standards Agency are used to calculate the healthiness score of the meal. (2) How can the model training be guided in a health-oriented manner? We construct category-wise personalization partial rankings and category-wise healthiness partial rankings, and theoretically analyze that they meet the necessary properties and assumptions required to be trained by the maximum posterior estimator under Bayesian probability. The data analysis confirms the existence of user preferences leaning towards less healthy meals in two public datasets. A comprehensive experiment demonstrates that our CateRec effectively boosts healthiness exposure in terms of mean healthiness score and ranking exposure while being comparable to the state-of-the-art model in terms of recommendation accuracy.
Ming Li 0072, Lin Li 0001, Xiaohui Tao 0001, Zhongwei Xie, Qing Xie 0002, Jingling Yuan
ACM Trans. Intell. Syst. Technol.3
2024 FRAMU: Attention-Based Machine Unlearning Using Federated Reinforcement Learning
abstract
Machine Unlearning, a pivotal field addressing data privacy in machine learning, necessitates efficient methods for the removal of private or irrelevant data. In this context, significant challenges arise, particularly in maintaining privacy and ensuring model efficiency when managing outdated, private, and irrelevant data. Such data not only compromises model accuracy but also burdens computational efficiency in both learning and unlearning processes. To mitigate these challenges, we introduce a novel framework: Attention-based Machine Unlearning using Federated Reinforcement Learning (FRAMU). This framework incorporates adaptive learning mechanisms, privacy preservation techniques, and optimization strategies, making it a well-rounded solution for handling various data sources, either single-modality or multi-modality, while maintaining accuracy and privacy. FRAMU's strengths include its adaptability in fluctuating data landscapes, its ability to unlearn outdated, private, or irrelevant data, and its support for continual model evolution without compromising privacy. Our experiments, conducted on both single-modality and multi-modality datasets, revealed that FRAMU significantly outperformed baseline models. Additional assessments of convergence behavior and optimization strategies further validate the framework's utility in federated learning applications. Overall, FRAMU advances Machine Unlearning by offering a robust, privacy-preserving solution that optimizes model performance while also addressing key challenges in dynamic data environments.
Thanveer Shaik, Xiaohui Tao 0001, Lin Li 0001, Haoran Xie 0001, Taotao Cai, Xiaofeng Zhu 0001, Qing Li 0001
IEEE Trans. Knowl. Data Eng.2
2024 Decoupled Progressive Distillation for Sequential Prediction with Interaction Dynamics
abstract
Sequential prediction has great value for resource allocation due to its capability in analyzing intents for next prediction. A fundamental challenge arises from real-world interaction dynamics where similar sequences involving multiple intents may exhibit different next items. More importantly, the character of volume candidate items in sequential prediction may amplify such dynamics, making deep networks hard to capture comprehensive intents. This article presents a sequential prediction framework with Decoupled Progressive Distillation (DePoD), drawing on the progressive nature of human cognition. We redefine target and non-target item distillation according to their different effects in the decoupled formulation. This can be achieved through two aspects: (1) Regarding how to learn, our target item distillation with progressive difficulty increases the contribution of low-confidence samples in the later training phase while keeping high-confidence samples in the earlier phase. And, the non-target item distillation starts from a small subset of non-target items from which size increases according to the item frequency. (2) Regarding whom to learn from, a difference evaluator is utilized to progressively select an expert that provides informative knowledge among items from the cohort of peers. Extensive experiments on four public datasets show DePoD outperforms state-of-the-art methods in terms of accuracy-based metrics.
Kaixi Hu, Lin Li 0001, Qing Xie 0002, Jianquan Liu, Xiaohui Tao 0001, Guandong Xu
ACM Trans. Inf. Syst.5
2023 HA-CMNet: A Driver CTR Model for Vehicle-Cargo Matching in O2O Platform
Zilong Jiang, Xiang Zuo, Kaifu Yuan, Lin Li 0001, Dali Wang, Xiaohui Tao 0001
ADMA (4)6
2023 Predicting Learners' Performance Using MOOC Clickstream
Kui Xiao, Xueyan Pan, Yan Zhang 0077, Xiaohui Tao 0001, Zhifang Huang
ADMA (4)4
2023 Hyperbolic Mutual Learning for Bundle Recommendation
Haole Ke, Lin Li 0001, Peipei Wang 0001, Jingling Yuan, Xiaohui Tao 0001
DASFAA (2)5
2023 Query2Trip: Dual-Debiased Learning for Neural Trip Recommendation
Peipei Wang 0001, Lin Li 0001, Ru Wang 0001, Xiaohui Tao 0001
DASFAA (2)4
2023 L2QA: Long Legal Article Question Answering with Cascaded Key Segment Learning
Shugui Xie, Lin Li 0001, Jingling Yuan, Qing Xie 0002, Xiaohui Tao 0001
DASFAA (3)5
2023 The Impact on Employability by COVID-19 Pandemic - AI Case Studies
Venkata Bharath Bandi, Xiaohui Tao 0001, Thanveer Shaik, Jianming Yong, Ji Zhang 0001
WISE2
2023 A novel dropout mechanism with label extension schema toward text emotion classification
abstract
Researchers have been aware that emotion is not one-hot encoded in emotion-relevant classification tasks, and multiple emotions can coexist in a given sentence. Recently, several works have focused on leveraging a distribution label or a grayscale label of emotions in the classification model, which can enhance the one-hot label with additional information, such as the intensity of other emotions and the correlation between emotions. Such an approach has been proven effective in alleviating the overfitting problem and improving the model robustness by introducing a distribution learning component in the objective function. However, the effect of distribution learning cannot be fully unfolded as it can reduce the model’s discriminative ability within similar emotion categories. For example, “Sad” and “Fear” are both negative emotions. To address such a problem, we proposed a novel emotion extension scheme in the prior work (Li, Chen, Xie, Li, and Tao, 2021). The prior work incorporated fine-grained emotion concepts to build an extended label space, where a mapping function between coarse-grained emotion categories and fine-grained emotion concepts was identified. For example, sentences labeled “Joy” can convey various emotions such as enjoy, free, and leisure. The model can further benefit from the extended space by extracting dependency within fine-grained emotions when yielding predictions in the original label space. The prior work has shown that it is more apt to apply distribution learning in the extended label space than in the original space. A novel sparse connection method, i.e., Leaky Dropout, is proposed in this paper to refine the dependency-extraction step, which further improves the classification performance. In addition to the multiclass emotion classification task, we extensively experimented on sentiment analysis and multilabel emotion prediction tasks to investigate the effectiveness and generality of the label extension schema.
Zongxi Li, Xianming Li, Haoran Xie 0001, Fu Lee Wang, Mingming Leng, Qing Li 0001, Xiaohui Tao 0001
Inf. Process. Manag.7
2023 Blockchain-Enhanced Smart Contract for Cost-Effective Insurance Claims Processing
abstract
Blockchain-enabled smart contracts have revolutionized the insurance industry due to their potential to streamline backend operations, mitigate fraudulent claims, and enhance data security and transparency. Guided by the design science methodology, the authors propose two specific smart contract frameworks to enhance insurance claims processing related to vehicle damage claims and personal injury claims. These proposed frameworks can improve the overall efficiency and effectiveness of insurance claims processing by automating claims submission, review, analysis, and payment, while reducing fraud and data leakage, by merging various data sources and disintermediation. Furthermore, the authors design a smart contract template supported by eight operational algorithms to facilitate the processing of insurance claims with the help of smart contracts. This template provides practitioners with a standardized prototype for the development of secure and efficient insurance applications.
Qiping Wang 0002, Raymond Y. K. Lau, Yain-Whar Si, Haoran Xie 0001, Xiaohui Tao 0001
J. Glob. Inf. Manag.5
2022 A Deep Learning Framework for Removing Bias from Single-Photon Emission Computerized Tomography
Jia-Ching Ying, Wan-Ju Yang, Ji Zhang 0001, Yu-Ching Ni, Chia-Yu Lin, Fan-Pin Tseng, Xiaohui Tao 0001
ADMA (1)7
2022 Fast Fourier Transform and Ensemble Model to Classify Epileptic EEG Signals
abstract
The analysis of electroencephalogram (EEG) signals can provide valuable insights to the nature of many diseases such as Alzheimer, epilepsy, sleep problems and thus can improve our understating and treatment about them. One of the major EEG signal applications is related to Epilepsy. The main contribution of this work is the proposal of an effective scheme for classifying the EEG signals for the study of epilepsy based on Fast Fourier Transform (FFT) under an ensemble model. The EEG signals are decomposed into frequency bands by using Fast Fourier Transform to extract the key statistical features, which are then fed into an ensemble model to classify the epileptic patients. Three base classifiers – Naïve Bayes, Least Square Support Vector Machine and Neural Networks - are utilized to construct the ensemble framework. The final decision of classification is dependent on the aggregation of the three classifiers decisions. The experimental results demonstrated that the proposed technique is a promising tool in accurately classifying the epileptic EEG signals.
Raid Lafta, Hisham Alshaheen, Xiaohui Tao 0001, Lingling Li 0004, Ji Zhang 0001
IEEE Big Data4
2022 Contrastive and attentive graph learning for multi-view clustering
Ru Wang 0001, Lin Li 0001, Xiaohui Tao 0001, Peipei Wang 0001, Peiyu Liu 0001
Inf. Process. Manag.3
2021 Adaptive Fault Resolution for Database Replication Systems
Chee Keong Wee, Xujuan Zhou, Raj Gururajan, Xiaohui Tao 0001, Nathan Wee
ADMA4
2021 What is Next when Sequential Prediction Meets Implicitly Hard Interaction?
abstract
Hard interaction learning between source sequences and their next targets is challenging, which exists in a myriad of sequential prediction tasks. During the training process, most existing methods focus on explicitly hard interactions caused by wrong responses. However, a model might conduct correct responses by capturing a subset of learnable patterns, which results in implicitly hard interactions with some unlearned patterns. As such, its generalization performance is weakened. The problem gets more serious in sequential prediction due to the interference of substantial similar candidate targets.
Kaixi Hu, Lin Li 0001, Qing Xie 0002, Jianquan Liu, Xiaohui Tao 0001
CIKM5
2021 Event Detection in Social Media via Graph Neural Network
Wang Gao 0002, Lin Li 0001, Xiaohui Tao 0001
WISE (1)4
2021 Emerging Applications in Healthcare and Their Implications to Academia and Practice
Raj Gururajan, Xiaohui Tao 0001, Yuefeng Li 0001, Xujuan Zhou, Soman Elangovan, Srinivas Kondalsamy-Chennakesavan, Revathi Venkataraman
WISE (2)2
2021 Trio-based collaborative multi-view graph clustering with multiple constraints
Ru Wang 0001, Lin Li 0001, Xiaohui Tao 0001, Peipei Wang 0001, Peiyu Liu 0001
Inf. Process. Manag.3
2020 Context-aware Adaptive Outlier Detection in Trajectory Data
abstract
With the advent of data mining and business processes automation, outlier detection has evolved into a major problem attracting significant research in relation to several application domains. Further advances in Global Positioning system, tracking of anomalous events based on data enhances effective decision making and pro-active measures to overcome risks and avoid unwarranted outputs. Significant work has been done in trajectory outlier detection although no singular approach fits all the domains. By including position and collective outliers on the same visualizations will enhance understanding of an outlier behavior. As such, we have leveraged Hidden Markov Method for prediction-based point outlier detection and pattern mining to identify points or segments of outliers in trajectory data.
Srinivas Danda, Ji Zhang 0001, Xiaohui Tao 0001, Jerry Chun-Wei Lin, Wenbin Zhang 0002
IEEE BigData3
2020 Predicting Workplace Injuries Using Machine Learning Algorithms
abstract
Predicting workplace injury using automated techniques opens newer possibilities in evidence-based research. This paper presents our preliminary research in a PhD project in predicting workplace incidents using machine learning algorithms. The analysis on the model performance using several mainstream machine learning algorithms including random forest, k-nearest neighbor and decision tree indicated that the general performance of the decision tree model was found to be statistically higher than that of the other two algorithms.
Divya Sukumar, Ji Zhang 0001, Xiaohui Tao 0001, Xin Wang 0030, Wenbin Zhang 0002
DSAA3
2020 A Densely Connected Encoder Stack Approach for Multi-type Legal Machine Reading Comprehension
Peiran Nai, Lin Li 0001, Xiaohui Tao 0001
WISE (2)3
2019 Learning Relational Fractals for Deep Knowledge Graph Embedding in Online Social Networks
Ji Zhang 0001, Leonard Tan, Xiaohui Tao 0001, Dianwei Wang, Jia-Ching Ying, Xin Wang 0030
WISE3
2018 SLIND: Identifying Stable Links in Online Social Networks
Ji Zhang 0001, Leonard Tan, Xiaohui Tao 0001, Xiaoyao Zheng, Yonglong Luo, Jerry Chun-Wei Lin
DASFAA (2)3
2018 A Recommender System with Advanced Time Series Medical Data Analysis for Diabetes Patients in a Telehealth Environment
Raid Lafta, Ji Zhang 0001, Xiaohui Tao 0001, Jerry Chun-Wei Lin, Fulong Chen 0002, Yonglong Luo, Xiaoyao Zheng
DEXA (2)3
2018 On Link Stability Detection for Online Social Networks
Ji Zhang 0001, Xiaohui Tao 0001, Leonard Tan, Jerry Chun-Wei Lin, Hongzhou Li, Liang Chang 0003
DEXA (1)2
2017 Mining Drug Properties for Decision Support in Dental Clinics
WeePheng Goh, Xiaohui Tao 0001, Ji Zhang 0001, Jianming Yong
PAKDD (2)2
2017 A Fast Fourier Transform-Coupled Machine Learning-Based Ensemble Model for Disease Risk Prediction Using a Real-Life Dataset
Raid Lafta, Ji Zhang 0001, Xiaohui Tao 0001, Yan Li 0002, Wessam Abbas, Yonglong Luo, Fulong Chen 0002, Vincent S. Tseng
PAKDD (1)3
2017 Coupling topic modelling in opinion mining for social media analysis
abstract
Many of social media platforms such as Facebook and Twitter make it easy for everyone to share their thoughts on literally anything. Topic and opinion detection in social media facilitates the identification of emerging societal trends, analysis of public reactions to policies and business products. In this paper, we proposed a new method that combines the opining mining and context-based topic modelling to analyse public opinions on social media data. Context based topic modelling is used to categorise data in groups and discover hidden communities in data group. The unwanted data group discovered by the topic model then will be discarded. A lexicon based opinion mining method will be applied to the remaining data groups to spot out the public sentiment about the entities. A set of Tweets data on Australian Federal Election 2010 was used in our experiments. Our experimental results demonstrate that, with the help of topic modelling, our social media analysis model is accurate and effective.
Xujuan Zhou, Xiaohui Tao 0001, Md Mostafijur Rahman, Ji Zhang 0001
WI2
2016 Adopting Hybrid Descriptors to Recognise Leaf Images for Automatic Plant Specie Identification
Ali A. Al-kharaz, Xiaohui Tao 0001, Ji Zhang 0001, Raid Lafta
ADMA2
2016 IRS-HD: An Intelligent Personalized Recommender System for Heart Disease Patients in a Tele-Health Environment
Raid Lafta, Ji Zhang 0001, Xiaohui Tao 0001, Yan Li 0002, Vincent S. Tseng
ADMA3
2016 Sentiment Analysis for Depression Detection on Social Networks
Xiaohui Tao 0001, Xujuan Zhou, Ji Zhang 0001, Jianming Yong
ADMA1
2013 Minimising K-Dominating Set in Arbitrary Network Graphs
Guangyuan Wang, Hua Wang 0002, Xiaohui Tao 0001, Ji Zhang 0001
ADMA (2)3
2013 SODIT: An innovative system for outlier detection using multiple localized thresholding and interactive feedback
abstract
Outlier detection is an important long-standing research problem in data mining and has enjoyed applications in a wide range of applications in business, engineering, biology and security, etc. However, the traditional outlier detection methods inevitably need to use different parameters for detection such as those used to specify the distance or density cutoff for distinguish outliers from normal data points. Using the trial and error approach, the traditional outlier detection methods are rather tedious in parameter tuning. In this demo proposal, we introduce an innovative outlier detection system, called SODIT, that uses localized thresholding to assist the value specification of the thresholds that reflect closely the local data distribution. In addition, easy-to-use user feedback are employed to further facilitate the determination of optimal parameter values. SODIT is able to make outlier detection much easier to operate and produce more accurate, intuitive and informative results than before.
Ji Zhang 0001, Hua Wang 0002, Xiaohui Tao 0001, Lili Sun
ICDE3
2012 Unsupervised Multi-label Text Classification Using a World Knowledge Ontology
Xiaohui Tao 0001, Yuefeng Li 0001, Raymond Y. K. Lau, Hua Wang 0002
PAKDD (1)1
2012 Semantic Labelling for Document Feature Patterns Using Ontological Subjects
abstract
Finding and labelling semantic features patterns of documents in a large, spatial corpus is a challenging problem. Text documents have characteristics that make semantic labelling difficult, the rapidly increasing volume of online documents makes a bottleneck in finding meaningful textual patterns. Aiming to deal with these issues, we propose an unsupervised documnent labelling approach based on semantic content and feature patterns. A world ontology with extensive topic coverage is exploited to supply controlled, structured subjects for labelling. An algorithm is also introduced to reduce dimensionality based on the study of ontological structure. The proposed approach was promisingly evaluated by compared with typical machine learning methods including SVMs, Rocchio, and kNN.
Xiaohui Tao 0001, Yuefeng Li 0001
Web Intelligence1
2011 A Personalized Ontology Model for Web Information Gathering
abstract
As a model for knowledge description and formalization, ontologies are widely used to represent user profiles in personalized web information gathering. However, when representing user profiles, many models have utilized only knowledge from either a global knowledge base or a user local information. In this paper, a personalized ontology model is proposed for knowledge representation and reasoning over user profiles. This model learns ontological user profiles from both a world knowledge base and user local instance repositories. The ontology model is evaluated by comparing it against benchmark models in web information gathering. The results show that this ontology model is successful.
Xiaohui Tao 0001, Yuefeng Li 0001, Ning Zhong 0001
IEEE Trans. Knowl. Data Eng.1
2010 Ontology-Based Specific and Exhaustive User Profiles for Constraint Information Fusion for Multi-agents
abstract
Intelligent agents are an advanced technology utilized in Web Intelligence. When searching information from a distributed Web environment, information is retrieved by multi-agents on the client site and fused on the broker site. The current information fusion techniques rely on cooperation of agents to provide statistics. Such techniques are computationally expensive and unrealistic in the real world. In this paper, we introduce a model that uses a world ontology constructed from the Dewey Decimal Classification to acquire user profiles. By search using specific and exhaustive user profiles, information fusion techniques no longer rely on the statistics provided by agents. The model has been successfully evaluated using the large INEX data set simulating the distributed Web environment.
Xiaohui Tao 0001, Yuefeng Li 0001, Raymond Y. K. Lau, Shlomo Geva
Web Intelligence1
2009 Concept-Based, Personalized Web Information Gathering: A Survey
Xiaohui Tao 0001, Yuefeng Li 0001
KSEM1
2008 Effective pattern taxonomy mining in text documents
abstract
Many data mining techniques have been proposed for mining useful patterns in databases. However, how to effectively utilize discovered patterns is still an open research issue, especially in the domain of text mining. Most existing methods adopt term-based approaches. However, they all suffer from the problems of polysemy and synonymy. This paper presents an innovative technique, pattern taxonomy mining, to improve the effectiveness of using discovered patterns for finding useful information. Substantial experiments on RCV1 demonstrate that the proposed solution achieves encouraging performance.
Yuefeng Li 0001, Sheng-Tang Wu, Xiaohui Tao 0001
CIKM3
2008 An Ontology-Based Framework for Knowledge Retrieval
abstract
Retrieving accurate information from the Web is a great challenge to users. The existing information retrieval systems are mostly term-based and thus need to be enhanced toward knowledge-based. User information needs need to be better captured in order to deliver personalized search results. In this paper, an ontology-based framework is proposed for capturing user information needs using a world knowledge base and the user's local instance repository. The framework aims to discover a user's background knowledge for knowledge retrieval. The evaluation result is encouraging, in which the proposed model achieved the same performance as a manual user model.
Xiaohui Tao 0001, Yuefeng Li 0001, Ning Zhong 0001, Richi Nayak
Web Intelligence1
2007 Ontology Mining for Semantic Interpretation of Information Needs
Xiaohui Tao 0001, Yuefeng Li 0001, Richi Nayak
KSEM1
2007 Ontology Mining for PersonalizedWeb Information Gathering
abstract
It is well accepted that ontology is useful for personalized Web information gathering. However, it is challenging to use semantic relations of "kind-of", "part-of", and "related-to" and synthesize commonsense and expert knowledge in a single computational model. In this paper, a personalized ontology model is proposed attempting to answer this challenge. A two-dimensional (Exhaustivity and Specificity) method is also presented to quantitatively analyze these semantic relations in a single framework. The proposals are successfully evaluated by applying the model to a Web information gathering system. The model is a significant contribution to personalized ontology engineering and concept-based Web information gathering in Web Intelligence.
Xiaohui Tao 0001, Yuefeng Li 0001, Ning Zhong 0001, Richi Nayak
Web Intelligence1
2006 Automatically Acquiring Training Sets for Web Information Gathering
abstract
The traditional techniques rely on human effort to acquire training sets, which is expensive and inefficient. In this paper we present an alternative method to automatically acquire training sets without heavy investment of user efforts. The proposed method tends to fill a gap for effectiveness of using Web data in Web mining, and contributes to Web information gathering. The evaluation shows that the method is adequate to yield an promising achievement.
Xiaohui Tao 0001, Yuefeng Li 0001, Ning Zhong 0001, Richi Nayak
Web Intelligence1
2005 Information Fusion With Subject-Based Information Gathering Method for Intelligent Multi-Agent Models
Xiaohui Tao 0001, John D. King, Yuefeng Li 0001
iiWAS1