Di Wang 0004

dblp:18/5410-4 · DBLP profile ↗
← Back
47ranked-venue papers
9as first author
29since 2021 · last 2026
0000-0002-3171-4001ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 36 · 9 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 first-author · 8 since 2021Databases, data management, data science and information retrieval · 9 · 8 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Efficient Few-Step Solution Generation via Discrete Flow Matching for Combinatorial Optimization
abstract
Combinatorial optimization problems (COPs) are fundamental to many real-world applications where efficiently producing high-quality solutions is critical. Recent advances in diffusion-based non-autoregressive models have reformulated solving COPs as a generative process, achieving promising results. However, almost all of these methods still suffer from accumulated errors and high inference costs due to the multi-step stochastic denoising process. To address these issues, we propose EFLOCO, an efficient discrete flow matching method for solving COPs, learning structured and deterministic solution trajectories. EFLOCO replaces noise-driven updates with smooth and guided transitions, thereby improves inference stability and quality. Furthermore, we introduce an adaptive time-step scheduler that makes more efforts in critical transition regions, yielding strong performance under few-step constraints. Experiments on standard Traveling Salesman Problems (TSPs) and Asymmetric TSPs (ATSPs) show that our method consistently outperforms both learning-based and heuristic baselines in terms of solution quality and inference speed.
Yuanshu Li, Di Wang 0004, Wei Du 0002, Xuan Wu 0004, Peng Zhao 0018, Yubin Xiao, You Zhou 0008
AAAI2
2026 Efficient neural combinatorial optimization solver for the min-max heterogeneous capacitated vehicle routing problem
Xuan Wu 0004, Di Wang 0004, Chunguo Wu, Kaifang Qi, Chunyan Miao, Yubin Xiao, You Zhou 0008
Expert Syst. Appl.2
2026 GELD: A unified neural model for efficiently solving traveling salesman problems across different scales
Yubin Xiao, Di Wang 0004, Xuan Wu 0004, Boyang Li 0001, You Zhou 0008
Pattern Recognit.2
2025 ESGenius: Benchmarking LLMs on Environmental, Social, and Governance (ESG) and Sustainability Knowledge
abstract
Chaoyue He, Xin Zhou, Yi Wu, Xinjia Yu, Yan Zhang, Lei Zhang, Di Wang, Shengfei Lyu, Hong Xu, Wang Xiaoqiao, Wei Liu, Chunyan Miao. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Chaoyue He, Xin Zhou 0008, Xinjia Yu, Lei Zhang 0199, Di Wang 0004, Shengfei Lyu, Hong Xu 0004, Xiaoqiao Wang, Chunyan Miao
EMNLP7
2025 DGL: Dynamic Global-Local Information Aggregation for Scalable VRP Generalization with Self-Improvement Learning
abstract
The Vehicle Routing Problem (VRP) is a critical combinatorial optimization problem with wide-reaching real-world applications, particularly in logistics, transportation. While neural network-based VRP solvers have shown impressive results on test instances similar to training data, their performance often degrades when faced with varying scales and unseen distributions, limiting their practical applicability. To overcome these limitations, we introduce DGL (Dynamic Global-Local Information Aggregation), a novel model that combines global and local information to effectively solve VRPs. DGL dynamically adjusts local node selections within a localized range, capturing local invariance across problems of different scales and distributions, thereby enhancing generalization. At the same time, DGL integrates global context into the decision-making process, providing richer information for more informed decisions. Additionally, we propose a replacement-based self-improvement learning framework that leverages data augmentation and random replacement techniques, further enhancing DGL's robustness. Extensive experiments on synthetic datasets, benchmark datasets, and real-world country map instances demonstrate that DGL achieves state-of-the-art performance, particularly in generalizing to large-scale VRPs and real-world scenarios. These results showcase DGL's effectiveness in solving complex, realistic optimization challenges and highlight its potential for practical applications.
Yubin Xiao, Yuesong Wu, Di Wang 0004, Zhiguang Cao, Xuan Wu 0004, Peng Zhao 0018, Yuanshu Li, You Zhou 0008, Yuan Jiang 0007
IJCAI4
2025 Efficient Heuristics Generation for Solving Combinatorial Optimization Problems Using Large Language Models
abstract
Recent studies exploited Large Language Models (LLMs) to autonomously generate heuristics for solving Combinatorial Optimization Problems (COPs), by prompting LLMs to first provide search directions and then derive heuristics accordingly. However, the absence of task-specific knowledge in prompts often leads LLMs to provide unspecific search directions, obstructing the derivation of well-performing heuristics. Moreover, evaluating the derived heuristics remains resource-intensive, especially for those semantically equivalent ones, often requiring omissible resource expenditure. To enable LLMs to provide specific search directions, we propose the Hercules algorithm, which leverages our designed Core Abstraction Prompting (CAP) method to abstract the core components from elite heuristics and incorporate them as prior knowledge in prompts. We theoretically prove the effectiveness of CAP in reducing unspecificity and provide empirical results in this work. To reduce computing resources required for evaluating the derived heuristics, we propose few-shot Performance Prediction Prompting (PPP), a first-of-its-kind method for the Heuristic Generation (HG) task. PPP leverages LLMs to predict the fitness values of newly derived heuristics by analyzing their semantic similarity to previously evaluated ones. We further develop two tailored mechanisms for PPP to enhance predictive accuracy and determine unreliable predictions, respectively. The use of PPP makes Hercules more resource-efficient and we name this variant Hercules-P. Extensive experiments across four HG tasks, five COPs, and eight LLMs demonstrate that Hercules outperforms the state-of-the-art LLM-based HG algorithms, while Hercules-P excels at minimizing required computing resources. In addition, we illustrate the effectiveness of CAP, PPP, and the other proposed mechanisms by conducting relevant ablation studies.
Xuan Wu 0004, Di Wang 0004, Chunguo Wu, Lijie Wen 0001, Chunyan Miao, Yubin Xiao, You Zhou 0008
KDD (2)2
2025 Context Pooling: Query-specific Graph Pooling for Generic Inductive Link Prediction in Knowledge Graphs
abstract
Recent investigations on the effectiveness of Graph Neural Network (GNN)-based models for link prediction in Knowledge Graphs (KGs) show that vanilla aggregation does not significantly impact the model performance. In this paper, we introduce a novel method, named Context Pooling, to enhance GNN-based models' efficacy for link predictions in KGs. To our best of knowledge, Context Pooling is the first methodology that applies graph pooling in KGs. Additionally, Context Pooling is first-of-its-kind to enable the generation of query-specific graphs for inductive settings, where testing entities are unseen during training. Specifically, we devise two metrics, namely neighborhood precision and neighborhood recall, to assess the neighbors' logical relevance regarding the given queries, thereby enabling the subsequent comprehensive identification of only the logically relevant neighbors for link prediction. Our method is generic and assessed by being applied to two state-of-the-art (SOTA) models on three public transductive and inductive datasets, achieving SOTA performance in 42 out of 48 settings.
Zhixiang Su, Di Wang 0004, Chunyan Miao
KDD (2)2
2025 Visual-Enhanced Multimodal Framework for Flexible Job Shop Scheduling Problem
abstract
Multimodal models leverage complementary information across modalities to enrich feature representations. While visual information shows potential in representing structure for some combinatorial optimization problems (COPs), its application to complex scheduling like the Flexible Job Shop Scheduling Problem (FJSP) remains underexplored. Current learning-based FJSP solvers predominantly rely on handcrafted state features. This dependence can lead to inconsistencies and may not fully capture the problem's intricate dynamics. Crucially, these methods overlook visual modalities. Visual representations offer a distinct advantage by inherently capturing the global topological structure and complex resource interactions within the FJSP state. Unlike localized handcrafted features, this holistic, structural view provides a richer foundation for understanding scheduling complexity and making informed decisions. To overcome these limitations by leveraging visual information-known for representing topological structures and providing richer state representations-we introduce the AO-framework. This multimodal feature fusion approach enhances handcrafted state features by integrating insights from visual data. Our core contribution is a novel fusion mechanism utilizing orthogonal projection and local attention. Unlike traditional methods that often rely on simple concatenation of visual data, our method uniquely reduces redundancy by projecting global image-derived features onto local handcrafted features. This process extracts distinct information inherent to the visual modality, significantly improving the quality and complementarity of the resulting state features and enabling more informed scheduling decisions. To our knowledge, the AO-framework represents the first multimodal framework applied to scheduling problems, demonstrating the significant potential of visual information in this domain. Extensive experiments across various FJSP solvers and datasets confirm that our framework yields substantial enhancements in solution quality, decision-making capabilities, and generalization.
Peng Zhao 0018, Zhiguang Cao, Di Wang 0004, Wen Song 0004, Wei Pang 0001, You Zhou 0008, Yuan Jiang 0007
ACM Multimedia3
2025 MMESGBench: Pioneering Multimodal Understanding and Complex Reasoning Benchmark for ESG Tasks
abstract
Environmental, Social, and Governance (ESG) reports are essential for assessing sustainability, regulatory compliance, and financial transparency. However, these documents are typically long, multimodal, and structurally complex, combining dense text, tables, figures, and layout-sensitive semantics. Existing AI systems often struggle to perform reliable document-level reasoning in such settings, and no dedicated benchmark currently exists in ESG domain. To fill the gap, we introduce MMESGBench, a first-of-its-kind benchmark dataset targeted to evaluate multimodal understanding and reasoning across multi-source ESG documents. This dataset is constructed via a human-AI collaborative, multi-stage pipeline. First, a multimodal LLM generates candidate question-answer (QA) pairs by jointly interpreting textual, tabular, and visual information from layout-aware document pages. Second, an LLM verifies the semantic accuracy, completeness, and reasoning complexity of each QA pair. This automated process is followed by an expert-in-the-loop validation, where domain specialists validate and calibrate QA pairs to ensure quality, relevance, and diversity. MMESGBench comprises 933 validated QA pairs derived from 45 ESG documents, spanning across seven distinct document types and three major ESG source categories. Questions are categorized as single-page, cross-page, or unanswerable, with each accompanied by fine-grained multimodal evidence. Initial experiments validate that multimodal and retrieval-augmented models substantially outperform text-only baselines. MMESGBench is publicly available as an open-source dataset at https://github.com/Zhanglei1103/MMESGBench.
Lei Zhang 0199, Xin Zhou 0008, Chaoyue He, Di Wang 0004, Hong Xu 0004, Chunyan Miao
ACM Multimedia4
2025 Dual Operation Aggregation Graph Neural Networks for Solving Flexible Job-Shop Scheduling Problem with Reinforcement Learning
abstract
With the widespread adoption of Internet Protocol (IP) communication technology and web-based platforms, cloud manufacturing has become a significant hallmark of Industry 4.0. Integrating graph algorithms into these web-enabled environments is crucial as they facilitate the representation and analysis of complex relationships in manufacturing processes, enabling efficient decision-making and adaptability in dynamic environments. As a key scheduling problem in cloud manufacturing, the flexible job-shop scheduling problem (FJSP) finds extensive applications in real-world scenarios. However, traditional FJSP-solving methods struggle to meet the efficiency and adaptability demands of cloud manufacturing due to generalization issues and excessive computational time, while reinforcement learning-based methods fail to learn relationships between FJSP nodes, such as interactions between operations of different jobs, leading to limited interpretability and performance. To address these issues, we propose a dual operation aggregation graph neural network (GNN) for solving FJSP. Specifically, we decouple the disjunctive graph into two distinct graphs, reducing graph density and clarifying relationships between machines and operations, thus enabling more effective aggregation and understanding by neural networks. We develop two distinct graph aggregation methods to minimize the influence of non-critical machine and operation nodes on decision-making while enhancing the model's ability to account for long-term benefits. Additionally, to achieve more accurate multi-objective estimation and mitigate reward sparsity, we design a reward function that simultaneously considers machine efficiency, schedule balance, and makespan minimization. Extensive experimental results on well-known datasets demonstrate that our model outperforms state-of-the-art models and exhibits excellent generalization capabilities, effectively addressing the challenges of cloud manufacturing.
Peng Zhao 0018, You Zhou 0008, Di Wang 0004, Zhiguang Cao, Yubin Xiao, Xuan Wu 0004, Yuanshu Li, Hongjia Liu, Wei Du 0002, Yuan Jiang 0007, Liupu Wang
WWW3
2025 Improving generalization of neural Vehicle Routing Problem solvers through the lens of model architecture
Yubin Xiao, Di Wang 0004, Xuan Wu 0004, Yuesong Wu, Boyang Li 0001, Wei Du 0002, Liupu Wang, You Zhou 0008
Neural Networks2
2025 LLM-Enhanced Multi-Teacher Knowledge Distillation for Modality-Incomplete Emotion Recognition in Daily Healthcare
abstract
The critical importance of monitoring and recognizing human emotional states in healthcare has led to a surge in proposals for EEG-based multimodal emotion recognition in recent years. However, practical challenges arise in acquiring EEG signals in daily healthcare settings due to stringent data acquisition conditions, resulting in the issue of incomplete modalities. Existing studies have turned to knowledge distillation as a means to mitigate this problem by transferring knowledge from multimodal networks to unimodal ones. However, these methods are constrained by the use of a single teacher model to transfer integrated feature extraction knowledge, particularly concerning spatial and temporal features in EEG data. To address this limitation, we propose a multi-teacher knowledge distillation framework enhanced with a Large Language Model (LLM), aimed at facilitating effective feature learning in the student network by transferring knowledge of extracting integrated features. Specifically, we employ an LLM as the teacher for extracting temporal features and a graph convolutional neural network for extracting spatial features. To further enhance knowledge distillation, we introduce causal masking and a confidence indicator into the LLM to facilitate the transfer of the most discriminative features. Extensive testing on the DEAP and MAHNOB-HCI datasets demonstrates that our model outperforms existing methods in the modality-incomplete scenario. This study underscores the potential application of large models in this field.
Yuzhe Zhang 0003, Huan Liu 0012, Yang Xiao 0014, Mohammed Amoon, Dalin Zhang 0001, Di Wang 0004, Shusen Yang, Hiok Chai Quek
IEEE J. Biomed. Health Informatics6
2025 A Multi-Objective Explanation Framework for Graph Neural Networks
abstract
Graph Neural Networks (GNNs) hold promise in various application domains, but their limited explainability hinders widespread adoption, impacting customer satisfaction and loyalty. This issue intensifies when addressing diverse explanation needs of different user groups. Current GNN explanation models focus on a single objective, neglecting varied and potential conflicting user requirements, resulting in suboptimal outcomes. Moreover, existing models prioritize explanation objectives during multi-objective explanations, disrupting the intrinsic hierarchical structures and distant relationships within the graphs, further diminishing their effectiveness. To tackle these challenges, this paper introduces a novel multi-objective explanatory framework with hierarchical structure attribution for GNNs, termed HM-Explainer. This framework constructs a multi-objective explanation generation module based on Pareto theory to balance different and potentially conflicting explanatory objectives. Additionally, to embed hierarchical information into explanations, HM-Explainer designs node-level and cluster-level attribution modules to analyze the impact of input data on GNN decisions hierarchically. Furthermore, a self-attention mechanism is integrated into the node-level attribution module to account for the influence of distant neighbors. Ultimately, the efficacy of HM-Explainer is validated across multiple datasets for different GNN models through experimentation.
Yibowen Zhao, Di Wang 0004, Qingzhong Li, Li-Zhen Cui 0001
IEEE Trans. Knowl. Data Eng.3
2025 One-Shot Secure Federated K-Means Clustering Based on Density Cores
abstract
Federated clustering (FC) performs well in independent and identically distributed (IID) scenarios, but it does not perform well in non-IID scenarios. In addition, existing methods lack proof of strict privacy protection. To address the above issues, we propose a new secure federated k-means clustering framework to achieve better clustering results under privacy requirements. Specifically, for the clients, we use cluster centers (representative points) generated by k-means to represent the corresponding clusters. These representative points can effectively preserve the structure of the local data and they are encrypted by differential privacy. For the server, we propose two methods to reprocess the uploaded encrypted representative points to obtain better final cluster centers, one uses k-means, and the other considers the improved density peaks (density cores) as final centers and then sends them back to the clients. Finally, each client assigns local data to their nearest centers. Experimental results show that the proposed methods perform better than several centralized (nonfederated) classical clustering algorithms [k-means, density-based spatial clustering of applications with noise (DBSCAN), and density peak clustering (DPC)] and state-of-the-art (SOTA) centralized clustering algorithms in most cases. In particular, the proposed algorithms perform better than the SOTA FC framework k-FED (ICML2021) and MUFC (ICLR2023).
Yizhang Wang, Wei Pang 0001, Di Wang 0004, Witold Pedrycz
IEEE Trans. Neural Networks Learn. Syst.3
2025 Reinforcement Learning-Based Nonautoregressive Solver for Traveling Salesman Problems
abstract
The traveling salesman problem (TSP) is a well-known combinatorial optimization problem (COP) with broad real-world applications. Recently, neural networks (NNs) have gained popularity in this research area because as shown in the literature, they provide strong heuristic solutions to TSPs. Compared to autoregressive neural approaches, nonautoregressive (NAR) networks exploit the inference parallelism to elevate inference speed but suffer from comparatively low solution quality. In this article, we propose a novel NAR model named NAR4TSP, which incorporates a specially designed architecture and an enhanced reinforcement learning (RL) strategy. To the best of our knowledge, NAR4TSP is the first TSP solver that successfully combines RL and NAR networks. The key lies in the incorporation of NAR network output decoding into the training process. NAR4TSP efficiently represents TSP-encoded information as rewards and seamlessly integrates it into RL strategies, while maintaining consistent TSP sequence constraints during both training and testing phases. Experimental results on both synthetic and real-world TSPs demonstrate that NAR4TSP outperforms five state-of-the-art (SOTA) models in terms of solution quality, inference speed, and generalization to unseen scenarios.
Yubin Xiao, Di Wang 0004, Boyang Li 0001, Huanhuan Chen 0001, Wei Pang 0001, Xuan Wu 0004, Dong Xu 0002, Yanchun Liang 0001, You Zhou 0008
IEEE Trans. Neural Networks Learn. Syst.2
2024 Anchoring Path for Inductive Relation Prediction in Knowledge Graphs
abstract
Aiming to accurately predict missing edges representing relations between entities, which are pervasive in real-world Knowledge Graphs (KGs), relation prediction plays a critical role in enhancing the comprehensiveness and utility of KGs. Recent research focuses on path-based methods due to their inductive and explainable properties. However, these methods face a great challenge when lots of reasoning paths do not form Closed Paths (CPs) in the KG. To address this challenge, we propose Anchoring Path Sentence Transformer (APST) by introducing Anchoring Paths (APs) to alleviate the reliance of CPs. Specifically, we develop a search-based description retrieval method to enrich entity descriptions and an assessment mechanism to evaluate the rationality of APs. APST takes both APs and CPs as the inputs of a unified Sentence Transformer architecture, enabling comprehensive predictions and high-quality explanations. We evaluate APST on three public datasets and achieve state-of-the-art (SOTA) performance in 30 of 36 transductive, inductive, and few-shot experimental settings.
Zhixiang Su, Di Wang 0004, Chunyan Miao
AAAI2
2024 Distilling Autoregressive Models to Obtain High-Performance Non-autoregressive Solvers for Vehicle Routing Problems with Faster Inference Speed
abstract
Neural construction models have shown promising performance for Vehicle Routing Problems (VRPs) by adopting either the Autoregressive (AR) or Non-Autoregressive (NAR) learning approach. While AR models produce high-quality solutions, they generally have a high inference latency due to their sequential generation nature. Conversely, NAR models generate solutions in parallel with a low inference latency but generally exhibit inferior performance. In this paper, we propose a generic Guided Non-Autoregressive Knowledge Distillation (GNARKD) method to obtain high-performance NAR models having a low inference latency. GNARKD removes the constraint of sequential generation in AR models while preserving the learned pivotal components in the network architecture to obtain the corresponding NAR models through knowledge distillation. We evaluate GNARKD by applying it to three widely adopted AR models to obtain NAR VRP solvers for both synthesized and real-world instances. The experimental results demonstrate that GNARKD significantly reduces the inference time (4-5 times faster) with acceptable performance drop (2-3%). To the best of our knowledge, this study is first-of-its-kind to obtain NAR VRP solvers from AR ones through knowledge distillation.
Yubin Xiao, Di Wang 0004, Boyang Li 0001, Mingzhao Wang, Xuan Wu 0004, Changliang Zhou, You Zhou 0008
AAAI2
2024 Deep learning model for human-intuitive shoeprint reconstruction
Yan Wang 0028, Di Wang 0004, Wei Pang 0001, Daixi Li, You Zhou 0008, Dong Xu 0002, Sami Ur Rahman, Amin ur Rahman, Ahmed Ameen Fateh, Peiwu Qin
Expert Syst. Appl.3
2024 Neural Architecture Search for Text Classification With Limited Computing Resources Using Efficient Cartesian Genetic Programming
abstract
Cartesian Genetic Programming (CGP) has often been applied for Neural Architecture Search (NAS). However, the performance of CGP is less than ideal when searching for architectures with limited computing resources. To better facilitate NAS with limited computing resources, this paper proposes a crossover operator, a light-weighted age mechanism, and two adaptive mutation operators as the novel components in our Efficient Cartesian Genetic Programming (ECGP) method. To assess the performance of ECGP, we conduct extensive experiments on three text classification task datasets. The experimental results demonstrate that ECGP outperforms other NAS methods, requiring only hundreds of fitness evaluations to find architectures with competitive accuracy compared with human-designed models. Additionally, the ECGP-evolved architectures are shown as converging fast and stably, and having high-level transferability with merely a 1-2% accuracy drop. Ablation studies demonstrate the effectiveness of the proposed operators and age mechanism, and identify GRU as the most critical function in the text classification task. Finally, we summarize three design principles observed from the ECGP-evolved architectures that are in line with human-design strategies. To the best of our knowledge, this work introduces the first attention-derived NAS benchmark for the text classification task.
Xuan Wu 0004, Di Wang 0004, Huanhuan Chen 0001, Lele Yan, Yubin Xiao, Chunyan Miao, Hong-Wei Ge, Dong Xu 0002, Yanchun Liang 0001, Kangping Wang, Chunguo Wu, You Zhou 0008
IEEE Trans. Evol. Comput.2
2023 Multi-Aspect Explainable Inductive Relation Prediction by Sentence Transformer
abstract
Recent studies on knowledge graphs (KGs) show that path-based methods empowered by pre-trained language models perform well in the provision of inductive and explainable relation predictions. In this paper, we introduce the concepts of relation path coverage and relation path confidence to filter out unreliable paths prior to model training to elevate the model performance. Moreover, we propose Knowledge Reasoning Sentence Transformer (KRST) to predict inductive relations in KGs. KRST is designed to encode the extracted reliable paths in KGs, allowing us to properly cluster paths and provide multi-aspect explanations. We conduct extensive experiments on three real-world datasets. The experimental results show that compared to SOTA models, KRST achieves the best performance in most transductive and inductive test cases (4 of 6), and in 11 of 12 few-shot test cases.
Zhixiang Su, Di Wang 0004, Chunyan Miao
AAAI2
2023 VDPC: Variational density peak clustering algorithm
Yizhang Wang, Di Wang 0004, You Zhou 0008, Xiaofeng Zhang 0002, Hiok Chai Quek
Inf. Sci.2
2023 When Convolutional Network Meets Temporal Heterogeneous Graphs: An Effective Community Detection Method
abstract
Community detection has long been an important yet challenging task to analyze complex networks with a focus on detecting topological structures of graph data. Essentially, real-world graph data is generally heterogeneous which dynamically varies over time, and this invalidates most existing community detection approaches. To cope with these issues, this paper proposes the temporal-heterogeneous graph convolutional networks (THGCN) to detect communities using the learnt feature representations of a set of temporal heterogeneous graphs. Particularly, we first design a heterogeneous GCN component to represent features of heterogeneous graph at each time step. Then, a residual compressed aggregation component is proposed to learn temporal feature representations extracted from two consecutive heterogeneous graphs. These temporal features are considered to contain evolutionary patterns of underlying communities. To the best of our knowledge, this is the first attempt to detect communities from temporal heterogeneous graphs. To evaluate the model performance, extensive experiments are performed on two real-world datasets, i.e., DBLP and IMDB. The promising results have demonstrated that the proposed THGCN is superior to both benchmark and the state-of-the-art approaches, e.g., GCN, GAT, GNN, LGNN, HAN and STAR, with respect to a number of evaluation criteria.
Yaping Zheng, Xiaofeng Zhang 0002, Shiyi Chen, Xinni Zhang, Xiaofei Yang 0002, Di Wang 0004
IEEE Trans. Knowl. Data Eng.6
2022 Missing Value Imputation for Diabetes Prediction
abstract
Machine learning (ML) models have been widely used to improve the accuracy and efficiency of various types of disease diagnostic tasks. However, it is still challenging to apply ML models to perform diabetes-related prediction tasks mainly because patients' health records are sparse and have a vast amount of missing values. Missing values often break the diabetes prediction pipelines, posing challenges to existing approaches. Such problem deteriorates significantly when critical attribute values (e.g., blood test results on HbAlc, FPG and OGTT2hr) are missing. In this paper, we introduce a large-scale diabetes-related dataset named Chronic Disease Management System (CDMS) dataset, which collects the clinical records of more than 700,000 visits of over 65,000 patients across eight years. CDMS is anonymously collected and has a high percentage of missing values on several critical attributes for diabetes prediction. If not being dealt with carefully, the missing values will cause significant performance degradation of the applied ML models. In this paper, we also investigate the effectiveness of multiple data imputation methods through conducting extensive experiments using CDMS. Experimental results show that k-Nearest Neighbor Imputation (KNNI) performs better than other methods in this diabetes prediction task. Specifically, with KNNI applied, the diabetes prediction accuracy and precision are both over 0.8 using various ML predictive models.
Hangwei Qian, Di Wang 0004, Xu Guo 0002, Eng Sing Lee, Hui Hwang Teong, Ray Tian Rui Lai, Chunyan Miao
IJCNN3
2022 Efficient Reachability Query with Extreme Labeling Filter
abstract
Being a fundamental graph operator, reachability query has been widely studied by the data mining community in the past decades. In a directed acyclic graph (DAG), one vertex is reachable by another if there exists a chain of directed edges connecting the two vertexes. The state-of-the-art (SOTA) reachability query methods mostly first index all the vertexes in the underlying DAG and assign them with different labels, and then use these indexes and/or labels to efficiently filter out as many unreachable queries as possible. Thus, because a large portion of unreachable queries can be identified without evoking any tedious path-finding process, the overall time taken by a huge number of queries is much shortened with a tolerable compensation on the additional index and/or label preprocessing time and space. In this paper, we propose the Extreme Labeling Filter (ELF), which is a novel generic filter that can be applied to existing reachability query methods to additionally identify a large number of unreachable queries. Based on the analysis of the given DAG in a systematic and autonomous manner, ELF first determines whether to use predecessors or successors to label the vertexes. Based on such self-determined labels, ELF is then able to identify a large number of unreachable queries with a low time complexity of O(1). To evaluate the performance of ELF, we apply it on 4 reachability query methods (1 conventional and 3 SOTA, all designated for reachability query in DAGs) and conduct experiments on 17 datasets of different sizes. The experimental results show that by applying ELF, all methods significantly shorten the query time.
Zhixiang Su, Di Wang 0004, Xiaofeng Zhang 0002, Li-Zhen Cui 0001, Chunyan Miao
WSDM2
2022 Restorable-inpainting: A novel deep learning approach for shoeprint restoration
Yan Wang 0028, Di Wang 0004, Wei Pang 0001, Kangping Wang, Daixi Li, You Zhou 0008, Dong Xu 0002
Inf. Sci.3
2022 EEG-Video Emotion-Based Summarization: Learning With EEG Auxiliary Signals
abstract
Video summarization is the process of selecting a subset of informative keyframes to expedite storytelling with limited loss of information. In this article, we propose an EEG-Video Emotion-based Summarization (EVES) model based on a multimodal deep reinforcement learning (DRL) architecture that leverages neural signals to learn visual interestingness to produce quantitatively and qualitatively better video summaries. As such, EVES does not learn from the expensive human annotations but the multimodal signals. Furthermore, to ensure the temporal alignment and minimize the modality gap between the visual and EEG modalities, we introduce a Time Synchronization Module (TSM) that uses an attention mechanism to transform the EEG representations onto the visual representation space. We evaluate the performance of EVES on the TVSum and SumMe datasets. Based on the rank order statistics benchmarks, the experimental results show that EVES outperforms the unsupervised models and narrows the performance gap with supervised models. Furthermore, the human evaluation scores show that EVES receives a higher rating than the state-of-the-art DRL model DR-DSN by 11.4% on the coherency of the content and 7.4% on the emotion-evoking content. Thus, our work demonstrates the potential of EVES in selecting interesting content that is both coherent and emotion-evoking.
Wai-Cheong Lincoln Lew, Di Wang 0004, Kai Keng Ang, Joo-Hwee Lim, Hiok Chai Quek, Ah-Hwee Tan
IEEE Trans. Affect. Comput.2
2022 EEG-Based Emotion Recognition Using Regularized Graph Neural Networks
abstract
Electroencephalography (EEG) measures the neuronal activities in different brain regions via electrodes. Many existing studies on EEG-based emotion recognition do not fully exploit the topology of EEG channels. In this article, we propose a regularized graph neural network (RGNN) for EEG-based emotion recognition. RGNN considers the biological topology among different brain regions to capture both local and global relations among different EEG channels. Specifically, we model the inter-channel relations in EEG signals via an adjacency matrix in a graph neural network where the connection and sparseness of the adjacency matrix are inspired by neuroscience theories of human brain organization. In addition, we propose two regularizers, namely node-wise domain adversarial training (NodeDAT) and emotion-aware distribution learning (EmotionDL), to better handle cross-subject EEG variations and noisy labels, respectively. Extensive experiments on two public datasets, SEED, and SEED-IV, demonstrate the superior performance of our model than state-of-the-art models in most experimental settings. Moreover, ablation studies show that the proposed adjacency matrix and two regularizers contribute consistent and significant gain to the performance of our RGNN model. Finally, investigations on the neuronal activities reveal important brain regions and inter-channel relations for EEG-based emotion recognition.
Peixiang Zhong, Di Wang 0004, Chunyan Miao
IEEE Trans. Affect. Comput.2
2021 CARE: Commonsense-Aware Emotional Response Generation with Latent Concepts
abstract
Rationality and emotion are two fundamental elements of humans. Endowing agents with rationality and emotion has been one of the major milestones in AI. However, in the field of conversational AI, most existing models only specialize in one aspect and neglect the other, which often leads to dull or unrelated responses. In this paper, we hypothesize that combining rationality and emotion into conversational agents can improve response quality. To test the hypothesis, we focus on one fundamental aspect of rationality, i.e., commonsense, and propose CARE, a novel model for commonsense-aware emotional response generation. Specifically, we first propose a framework to learn and construct commonsense-aware emotional latent concepts of the response given an input message and a desired emotion. We then propose three methods to collaboratively incorporate the latent concepts into response generation. Experimental results on two large-scale datasets support our hypothesis and show that our model can produce more accurate and commonsense-aware emotional responses and achieve better human ratings than state-of-the-art models that only specialize in one aspect.
Peixiang Zhong, Di Wang 0004, Chen Zhang 0003, Hao Wang 0005, Chunyan Miao
AAAI2
2021 Development and validation of a practical instrument for evaluating players' familiarity with exergames
Hao Zhang 0049, Di Wang 0004, Yu Wang 0108, Ying Chi, Chunyan Miao
Int. J. Hum. Comput. Stud.2
2020 BDANN: BERT-Based Domain Adaptation Neural Network for Multi-Modal Fake News Detection
abstract
Nowadays, with the rapid growth of microblogging networks for news propagation, there are increasingly more people accessing news through such emerging social media. In the meantime, fake news now spreads at a faster pace and affects a larger population than ever before. Compared with traditional text news, the news posted on microblog often has attached images in the context. So how to correctly and autonomously detect fakes news in a multi-modal manner becomes a prominent challenge to be addressed. In this paper, we propose an end-to-end model, named BERT-based domain adaptation neural network for multi-modal fake news detection (BDANN). BDANN comprises three main modules: a multi-modal feature extractor, a domain classifier and a fake news detector. Specifically, the multi-modal feature extractor employs the pretrained BERT model to extract text features and the pretrained VGG-19 model to extract image features. The extracted features are then concatenated and fed to the detector to distinguish fake news. The role of the domain classifier is mainly to map the multi-modal features of different events to the same feature space. To assess the performance of BDANN, we conduct extensive experiments on two multimedia datasets: Twitter and Weibo. The experimental results show that BDANN outperforms the state-of-the-art models. Moreover, we further discuss the existence of noisy images in the Weibo dataset that may affect the results.
Di Wang 0004, Huanhuan Chen 0001, Wei Guo 0017, Chunyan Miao, Li-Zhen Cui 0001
IJCNN2
2020 A systematic density-based clustering method using anchor points
Yizhang Wang, Di Wang 0004, Wei Pang 0001, Chunyan Miao, Ah-Hwee Tan, You Zhou 0008
Neurocomputing2
2020 A novel hybrid deep recommendation system to differentiate user's preference and item's attractiveness
Xiaofeng Zhang 0002, Jingbin Zhong, Di Wang 0004
Inf. Sci.5
2020 McDPC: multi-center density peak clustering
Yizhang Wang, Di Wang 0004, Xiaofeng Zhang 0002, Wei Pang 0001, Chunyan Miao, Ah-Hwee Tan, You Zhou 0008
Neural Comput. Appl.2
2019 Modelling Autobiographical Memory Loss across Life Span
abstract
Neurocomputational modelling of long-term memory is a core topic in computational cognitive neuroscience, which is essential towards self-regulating brain-like AI systems. In this paper, we study how people generally lose their memories and emulate various memory loss phenomena using a neurocomputational autobiographical memory model. Specifically, based on prior neurocognitive and neuropsychology studies, we identify three neural processes, namely overload, decay and inhibition, which lead to memory loss in memory formation, storage and retrieval, respectively. For model validation, we collect a memory dataset comprising more than one thousand life events and emulate the three key memory loss processes with model parameters learnt from memory recall behavioural patterns found in human subjects of different age groups. The emulation results show high correlation with human memory recall performance across their life span, even with another population not being used for learning. To the best of our knowledge, this paper is the first research work on quantitative evaluations of autobiographical memory loss using a neurocomputational model.
Di Wang 0004, Ah-Hwee Tan, Chunyan Miao, Ahmed A. Moustafa
AAAI1
2019 An Affect-Rich Neural Conversational Model with Biased Attention and Weighted Cross-Entropy Loss
abstract
Affect conveys important implicit information in human communication. Having the capability to correctly express affect during human-machine conversations is one of the major milestones in artificial intelligence. In recent years, extensive research on open-domain neural conversational models has been conducted. However, embedding affect into such models is still under explored. In this paper, we propose an endto-end affect-rich open-domain neural conversational model that produces responses not only appropriate in syntax and semantics, but also with rich affect. Our model extends the Seq2Seq model and adopts VAD (Valence, Arousal and Dominance) affective notations to embed each word with affects. In addition, our model considers the effect of negators and intensifiers via a novel affective attention mechanism, which biases attention towards affect-rich words in input sentences. Lastly, we train our model with an affect-incorporated objective function to encourage the generation of affect-rich words in the output responses. Evaluations based on both perplexity and human evaluations show that our model outperforms the state-of-the-art baseline model of comparable size in producing natural and affect-rich responses.
Peixiang Zhong, Di Wang 0004, Chunyan Miao
AAAI2
2019 Knowledge-Enriched Transformer for Emotion Detection in Textual Conversations
abstract
Peixiang Zhong, Di Wang, Chunyan Miao. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Peixiang Zhong, Di Wang 0004, Chunyan Miao
EMNLP/IJCNLP (1)2
2019 REDPC: A residual error-based density peak clustering algorithm
Milan D. Parmar, Di Wang 0004, Xiaofeng Zhang 0002, Ah-Hwee Tan, Chunyan Miao, Jianhua Jiang, You Zhou 0008
Neurocomputing2
2019 Self-organizing neural networks for universal learning and multimodal memory encoding
Ah-Hwee Tan, Budhitama Subagdja, Di Wang 0004, Lei Meng 0001
Neural Networks3
2018 An interpretable neural fuzzy inference system for predictions of underpricing in initial public offerings
Di Wang 0004, Xiaolin Qian, Hiok Chai Quek, Ah-Hwee Tan, Chunyan Miao, Xiaofeng Zhang 0002, Geok See Ng, You Zhou 0008
Neurocomputing1
2018 Device Clustering Algorithm Based on Multimodal Data Correlation in Cognitive Internet of Things
abstract
With the development of information network, the popularity of Internet of Things (IoT) is an irreversible trend, and the intelligent demands for IoT is becoming more and more urgent. How to improve the cognitive ability of IoT is a new challenge and therefore has given rise to the emergence of cognitive IoT (CIoT). In this paper, a device-level multimodal data correlation mining model is first designed based on the canonical correlation analysis to transform the data feature into a subspace and analyze the data correlation. The correlation of the device is obtained based on the comprehensive of data correlation and the location information of the device. Then a heterogeneous clustering model (heterogeneous device clustering) is proposed by using the result of the correlation analysis to classify the device. Finally, we propose a device clustering algorithm based on multimodal data correlation for CIoT, which combines the functions of multimodal data correlation analyze with device clustering. Extensive simulations are carried out and our results show that the proposed algorithm can effectively improve the quality of data transmission and the intelligent service.
Di Wang 0004, Fuzhen Xia, Hong-Wei Ge
IEEE Internet Things J.2
2016 Self-regulated incremental clustering with focused preferences
abstract
Due to their online learning nature, incremental clustering techniques can handle a continuous stream of data. In particular, various incremental clustering techniques based on Adaptive Resonance Theory (ART) have been shown to have low computational complexity in adaptive learning and are less sensitive to noisy information. However, parameter regularization in existing ART clustering techniques is applied either on different features or on different clusters exclusively. In this paper, we introduce Interest-Focused Clustering based on Adaptive Resonance Theory (IFC-ART), which self-regulates the vigilance parameter associated with each feature and each cluster. As such, we can incorporate the domain knowledge of the data set into IFC-ART to focus on certain preferences during the self-regulated clustering process. For performance evaluation, we use a real-world data set, named American Time Use Survey (ATUS), which records nearly 160,000 telephone interviews conducted with U.S. residents from 2003 to 2014. Specifically, we conduct case studies to explore three types of interesting relationship, focusing on the wage, age, and provision of elderly care, respectively. Experimental results show that the performance of IFC-ART is highly competitive and stable when compared with two well-established clustering techniques and three ART models. In addition, we highlight the important and unexpected findings observed from the clusters discovered.
Di Wang 0004, Ah-Hwee Tan
IJCNN1
2015 Creating Autonomous Adaptive Agents in a Real-Time First-Person Shooter Computer Game
abstract
Games are good test-beds to evaluate AI methodologies. In recent years, there has been a vast amount of research dealing with real-time computer games other than the traditional board games or card games. This paper illustrates how we create agents by employing FALCON, a self-organizing neural network that performs reinforcement learning, to play a well-known first-person shooter computer game called Unreal Tournament. Rewards used for learning are either obtained from the game environment or estimated using the temporal difference learning scheme. In this way, the agents are able to acquire proper strategies and discover the effectiveness of different weapons without any guidance or intervention. The experimental results show that our agents learn effectively and appropriately from scratch while playing the game in real-time. Moreover, with the previously learned knowledge retained, our agent is able to adapt to a different opponent in a different map within a relatively short period of time.
Di Wang 0004, Ah-Hwee Tan
IEEE Trans. Comput. Intell. AI Games1
2014 Mobile humanoid agent with mood awareness for elderly care
abstract
Human, especially elderly, require frequent attention, continuous companionship, and deep understanding from the others. To provide more specific and appropriate tender care to the elderly, knowing their affective states is a great advantage. Recent work on human emotion recognition shows promising results that the expressive emotion can be successfully captured through visual, audio, and keyboard or touchpad stroke pattern signals. Furthermore, human activities are shown to be accurately recognizable with context by non-intrusive sensors within or connected to the smartphones. In this paper, we propose a computational model to characterize the affective states of the elderly based on the recognizable daily activities. Therefore, by integrating such an understanding module into a humanoid agent residing in the smartphone platform, we make the mobile agent more human-like. The initial knowledge of the activity-affect associations is taken from published work in psychology and gerontology. Based on the provided training signals, our model adapts the activity-affect knowledge accordingly. Consequently, by modeling mood awareness of the elderly, our agent can carry out more specific task and provide more appropriate tender care.
Di Wang 0004, Ah-Hwee Tan
IJCNN1
2008 A novel hybrid intelligent system: Genetic algorithm and rough set incorporated neural fuzzy inference system
abstract
This paper proposes a novel hybrid intelligent system denoted as genetic algorithm and rough set incorporated neural fuzzy inference system (GARSINFIS). Its network structure dynamically changes along with the evolving genetic algorithm based rough set clustering (GARSC) technique. When input data set is applied, only the most essential information is retained in the clustering result, as knowledge reduction is done using rough set approximations and the most optimal solution is selected by genetic algorithm. The system not only obtains promising accuracy but also possesses a great level of interpretability to meet the increasing need of understanding the inference process. In terms of TSK type of fuzzy inference system, better structural interpretability is typically manifested as employing less number of input features, less number of rules, less number of fuzzy membership functions in each feature, and less complex rules in both antecedent and consequent parts. Extensive simulations on various data sets were conducted, and the performance of GARSINFIS was benchmarked against other well established neural and neural-fuzzy systems. Experimental results have shown that GARSINFIS performs well in both accuracy and interpretability.
Di Wang 0004, Geok See Ng, Hiok Chai Quek
IEEE Congress on Evolutionary Computation1
2006 Ovarian Cancer Diagnosis Using Fuzzy Neural Networks Empowered By Evolutionary Clustering Technique
abstract
As computational power of modern computer increases exponentially, more efficient computerized solutions are possible for complex real world applications. However, the solutions are usually not interpretable to human beings such as the opaqueness of traditional neural networks. In this paper, we propose a fuzzy neural network that is empowered by genetic algorithm based rough set clustering (GARSC) technique. The system is capable to address real world problems not only with promising accuracy, but also with great interpretabilities. Ovarian cancer diagnosis is exploited to show its superior capabilities.
Di Wang 0004, Geok See Ng, Hiok Chai Quek
IEEE Congress on Evolutionary Computation1
2004 MS-TSKfnn: novel Takagi-Sugeno-Kang fuzzy neural network using ART like clustering
abstract
We propose a novel architecture of neuro-fuzzy system called modified self-organizing Takagi-Sugeno-Kang fuzzy neural network (MS-TSKfnn) that uses ART-like clustering called discrete incremental clustering (DIC). The network is able to handle online data input with significant high performance. Its ability of entirely self-organizing to form the network structure without any human supervision is the main advantage over other TSK type fuzzy rule based neuro-fuzzy systems, such as ANFIS and DENFIS. Extensive simulations were conducted using MS-TSKfnn and its performance was encouraging when benchmarked against other established neural and neuro-fuzzy system.
Di Wang 0004, Hiok Chai Quek, Geok See Ng
IJCNN1
2004 Novel Self-Organizing Takagi Sugeno Kang Fuzzy Neural Networks Based on ART-like Clustering
Di Wang 0004, Hiok Chai Quek, Geok See Ng
Neural Process. Lett.1