Junli Wang 0001

dblp:82/179-1 · also Jun-li Wang 0001 · DBLP profile ↗
← Back
47ranked-venue papers
2as first author
41since 2021 · last 2026
0000-0002-7185-9731ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 26 · 1 first-author · 24 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 7 since 2021Human-computer interaction and ubiquitous computing · 6 · 3 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Computer networks · 3 · 3 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Relation Prototype Driven Multimodal Knowledge Graph Completion
Shuyue Zhu, ChunGang Yan, Bingxuan Hou, Mengna Gao, Qingqing Hong, Dapeng Yin, Junli Wang 0001
PAKDD (2)7
2026 Layer-Adaptive-Augmentation-Based Graph Contrastive Learning With Feature Decorrelation
abstract
Graph Contrastive Learning (GCL) methods typically leverage augmentation techniques to generate different graph views for comparison, thereby learning corresponding representations for graph-related tasks in label-scarce scenarios. However, existing GCL methods suffer from two primary limitations: 1) they use predefined or one-time perturbations for augmentation, ignoring adaptive noise injection during forward propagation and thus leading to suboptimal model robustness; 2) their contrast mechanisms mainly focus on the agreement of inter-graph representations while neglecting the dimensional feature redundancy within intra-graph representations. To solve these issues, we propose Layer-adaptive-augmentation-based Graph Contrastive Learning with feature Decorrelation (LGCLD). First, the designed layer-wise adaptive augmentation method performs dynamic perturbations while maintaining the semantic similarity between augmented and original graphs, which can improve model robustness. Second, we introduce an Agreement-Decorrelation loss (AD loss) that simultaneously optimizes the agreement between graph-level representations and the feature correlation among different dimensions within each graph-level representation, promoting the model to learn informative and non-redundant graph-level representations. Furthermore, we analyze the reasonableness of AD loss through the graph information bottleneck principle. Experiments on various-domain graph datasets demonstrate that LGCLD achieves better or competitive performance compared with a series of state-of-the-art baselines.
Yuhua Xu 0005, Junli Wang 0001, Rui Duan 0003, Changjun Jiang 0002
IEEE Trans. Pattern Anal. Mach. Intell.2
2026 Intra-Class Amplitude Mixup for Data Augmentation in Image Classification
abstract
Mixup-based data augmentation improves deep neural network performance by enhancing training diversity, while amplitude mixing further enriches image style variations, benefiting tasks such as domain adaptation. However, the random selection of images for mixing in current amplitude-mixing approaches may introduce class-irrelevant information, potentially reducing their effectiveness in image classification. To address this, we propose Intra-Class Amplitude Mixup (ICAMix), a plug-and-play data augmentation method. To prevent random amplitude information from disrupting discriminative features in the mixed images, we restrict the mixing process to intra-class images, thereby enhancing the class discriminability of the generated images, making them more effective for classification tasks. Furthermore, since amplitude mixing preserves the structural integrity of images, our method effectively integrates multiple intra-class images, further enriching the diversity of features in the mixed images. Extensive experiments across various image classification tasks demonstrate that our method outperforms several baselines with minimal computational overheads. Visualizations of the generated images further validate its effectiveness. Our code is available at: https://github.com/rqfzpy/ICAMix.
Junli Wang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2026 Ensemble Graph Neural Networks With Individual Decision Feedback for Graph Classification
Mingjian Guang, Zhong Li 0006, Rui Zhang 0003, Junli Wang 0001, Dawei Cheng
IEEE Trans. Knowl. Data Eng.4
2025 Disentangle to Decay: Linear Attention with Trainable Decay Factor
abstract
Linear attention enhances inference efficiency of Transformer and has attracted research interests as an efficient backbone of language models. Existing linear attention based models usually exploit decay factor based positional encoding (PE), where attention scores decay exponentially with increasing relative distance. However, most work manually designs a non-trainable decay factor of exponential calculation, which limits further optimization. Our analysis reveals directly training decay factor is unstable because of large gradients. To address this, we propose a novel PE for linear attention named Disentangle to Decay (D2D). D2D disentangles decay factor into two parts to achieve further optimization and stable training. Moreover, D2D can be transformed into recurrent form for efficient inference. Experiments demonstrate that D2D achieves stable training of decay factor, and enhances performance of linear attention in both normal context length and length extrapolation scenarios.
Haibo Tong, Chenyang Zhang 0004, Jiayi Lin 0008, Bingxuan Hou, Qingqing Hong, Junli Wang 0001
COLING6
2025 Ontology-based Graph and Large Language Model Fusion Method for Relational Triple Extraction
abstract
In recent years, Relational Triple Extraction (RTE) methods leveraging Large Language Models (LLMs) have garnered lots of attention. Several studies attempted to harness ontology, the foundational template of knowledge graph, to assist LLMs in comprehending the structure of knowledge graph. However, the broadness of ontology within the prompts leads to unsatisfactory outcomes for RTE tasks. To address this, we proposed an Ontology-based Graph and LLM Fusion method for Relational Triple Extraction (OGLRTE). In this framework, we pioneeringly designed a two-stage RTE method centered around ontology, comprising of a relation filter and a text generator. To fully utilize the information in knowledge graph, we proposed an innovative combination of the ontology and co-occurrence graph within relation filter to repersent co-occurrence of relations in knowledge graph. Additionally, we employed established fine-tuning techniques to optimize the LLM within our triple generator, thereby enhancing its capability to extract triples. Our method surpasses traditional methods and LLM-based methods in extracting information, as evidenced by its superior performance on three public datasets: DocRED, NYT10 and CoNLL04.
Yuxiang Yao, Junli Wang 0001, ChunGang Yan
SMC2
2025 FedFAA: knowledge filtering for adaptive model aggregation in federated learning
Zihao Lu, Junli Wang 0001, Mingjian Guang
Appl. Intell.2
2025 Domain-wise knowledge decoupling for personalized federated learning via Radon transform
Zihao Lu, Junli Wang 0001, Changjun Jiang 0002
Neurocomputing2
2025 DoA-ViT: Dual-objective Affine Vision Transformer for Data Insufficiency
Junli Wang 0001
Neurocomputing2
2025 Irrelevant Patch-Masked Autoencoders for Enhancing Vision Transformers under Limited Data
Junli Wang 0001
Knowl. Based Syst.2
2025 Hierarchical neighbor-enhanced graph contrastive learning for recommendation
Hongjie Wei, Junli Wang 0001, Mingjian Guang, ChunGang Yan
Knowl. Based Syst.2
2025 Disentangled Active Learning on Graphs
Haoran Yang 0003, Junli Wang 0001, Rui Duan 0003, Changwei Wang 0001, ChunGang Yan
Neural Networks2
2025 Data-Free Knowledge Filtering and Distillation in Federated Learning
abstract
In federated learning (FL), multiple parties collaborate to train a global model by aggregating their local models while keeping private training sets isolated. One problem hindering effective model aggregation is data heterogeneity. Federated ensemble distillation tackles this problem by using fused local-model knowledge to train the global model rather than directly averaging model parameters. However, most existing methods fuse all knowledge indiscriminately, which makes the global model inherit some data-heterogeneity-caused flaws from local models. While knowledge filtering is a potential coping method, its implementation in FL is challenging due to the lack of public data for knowledge validation. To address this issue, we propose a novel data-free approach (FedKFD) that synthesizes credible labeled data to support knowledge filtering and distillation. Specifically, we construct a prediction capability description to characterize the samples where a local model makes correct predictions. FedKFD explores the intersection of local-model-input space and prediction capability descriptions with a conditional generator to synthesize consensus-labeled proxy data. With these labeled data, we filter for relevant local-model knowledge and further train a robust global model through distillation. The theoretical analysis and extensive experiments demonstrate that our approach achieves improved generalization, superior performance, and compatibility with other FL efforts.
Zihao Lu, Junli Wang 0001, Changjun Jiang 0002
IEEE Trans. Big Data2
2025 Multi-Temporal Partitioned Graph Attention Networks for Financial Fraud Detection
Mingjian Guang, Zhong Li 0006, ChunGang Yan, Yuhua Xu 0005, Junli Wang 0001, Dawei Cheng, Changjun Jiang 0002
IEEE Trans. Inf. Forensics Secur.5
2024 Unifying Homophily and Heterophily for Spectral Graph Neural Networks via Triple Filter Ensembles
abstract
Polynomial-based learnable spectral graph neural networks (GNNs) utilize polynomial to approximate graph convolutions and have achieved impressive performance on graphs. Nevertheless, there are three progressive problems to be solved. Some models use polynomials with better approximation for approximating filters, yet perform worse on real-world graphs. Carefully crafted graph learning methods, sophisticated polynomial approximations, and refined coefficient constraints leaded to overfitting, which diminishes the generalization of the models. How to design a model that retains the ability of polynomial-based spectral GNNs to approximate filters while it possesses higher generalization and performance? In this paper, we propose a spectral GNN with triple filter ensemble (TFE-GNN), which extracts homophily and heterophily from graphs with different levels of homophily adaptively while utilizing the initial features. Specifically, the first and second ensembles are combinations of a set of base low-pass and high-pass filters, respectively, after which the third ensemble combines them with two learnable coefficients and yield a graph convolution (TFE-Conv). Theoretical analysis shows that the approximation ability of TFE-GNN is consistent with that of ChebNet under certain conditions, namely it can learn arbitrary filters. TFE-GNN can be viewed as a reasonable combination of two unfolded and integrated excellent spectral GNNs, which motivates it to perform well. Experiments show that TFE-GNN achieves high generalization and new state-of-the-art performance on various real-world datasets.
Rui Duan 0003, Mingjian Guang, Junli Wang 0001, ChunGang Yan, Hongda Qi, Wenkang Su 0001, Can Tian, Haoran Yang 0003
NeurIPS3
2024 A Novel ICD Coding Method Based on Associated and Hierarchical Code Description Distillation
Junli Wang 0001
NLPCC (3)2
2024 Adaptive Knowledge Recomposition for Personalized Federated Learning via Discrete Wavelet Transform
abstract
Heterogeneous data silos hinder the application of deep learning in the Internet of Things. As a dealing scheme, personalized federated learning (pFL) distributedly customizes multiple local models for these silos. Most pFL methods directly use a global model to assist the local model optimization ignoring the performance drop caused by irrelevant or misleading global-model knowledge. To address this, we propose an adaptive knowledge recomposition approach (FedAKR), which refines relevant and correctly leading knowledge from the global model and recomposes it into the local model to promote personalization. Specifically, FedAKR provides a discrete wavelet transform-based method to recompose different kinds of knowledge in the same representation space. Facilitated by this common space, we introduce an enriched local optimization objective to establish a causal relationship between the refined global-model knowledge and recomposed local-model knowledge. The relationship guides effective and efficient knowledge refinement, thereby promoting personalization. Besides, we provide the theoretical proof of convergence for our novel pFL approach. Extensive experiments demonstrate that FedAKR achieves interpretable improvements, higher performance over 12 state-of-the-art methods, and the potential to further integrate pretrained large models.
Zihao Lu, Junli Wang 0001, Changjun Jiang 0002
IEEE Internet Things J.2
2024 Graph contrastive learning with min-max mutual information
Yuhua Xu 0005, Junli Wang 0001, Mingjian Guang, ChunGang Yan, Changjun Jiang 0002
Inf. Sci.2
2024 Graph Convolutional Networks With Adaptive Neighborhood Awareness
abstract
Graph convolutional networks (GCNs) can quickly and accurately learn graph representations and have shown powerful performance in many graph learning domains. Despite their effectiveness, neighborhood awareness remains essential and challenging for GCNs. Existing methods usually perform neighborhood-aware steps only from the node or hop level, which leads to a lack of capability to learn the neighborhood information of nodes from both global and local perspectives. Moreover, most methods learn the nodes' neighborhood information from a single view, ignoring the importance of multiple views. To address the above issues, we propose a multi-view adaptive neighborhood-aware approach to learn graph representations efficiently. Specifically, we propose three random feature masking variants to perturb some neighbors' information to promote the robustness of graph convolution operators at node-level neighborhood awareness and exploit the attention mechanism to select important neighbors from the hop level adaptively. We also utilize the multi-channel technique and introduce a proposed multi-view loss to perceive neighborhood information from multiple perspectives. Extensive experiments show that our method can better obtain graph representation and has high accuracy.
Mingjian Guang, ChunGang Yan, Yuhua Xu 0005, Junli Wang 0001, Changjun Jiang 0002
IEEE Trans. Pattern Anal. Mach. Intell.4
2024 Graph Multi-Convolution and Attention Pooling for Graph Classification
abstract
Many studies have achieved excellent performance in analyzing graph-structured data. However, learning graph-level representations for graph classification is still a challenging task. Existing graph classification methods usually pay less attention to the fusion of node features and ignore the effects of different-hop neighborhoods on nodes in the graph convolution process. Moreover, they discard some nodes directly during the graph pooling process, resulting in the loss of graph information. To tackle these issues, we propose a new Graph Multi-Convolution and Attention Pooling based graph classification method (GMCAP). Specifically, the designed Graph Multi-Convolution (GMConv) layer explicitly fuses node features learned from different perspectives. The proposed weight-based aggregation module combines the outputs of all GMConv layers, for adaptively exploiting the information over different-hop neighborhoods to generate informative node representations. Furthermore, the designed Local information and Global Attention based Pooling (LGAPool) utilizes the local information of a graph to select several important nodes and aggregates the information of unselected nodes to the selected ones by a global attention mechanism when reconstructing a pooled graph, thus effectively reducing the loss of graph information. Extensive experiments show that GMCAP outperforms the state-of-the-art methods on graph classification tasks, demonstrating that GMCAP can learn graph-level representations effectively.
Yuhua Xu 0005, Junli Wang 0001, Mingjian Guang, Changjun Jiang 0002
IEEE Trans. Pattern Anal. Mach. Intell.2
2024 Probabilistic Reachability Prediction of Unbounded Petri Nets: A Machine Learning Method
abstract
Unbounded Petri nets (UPNs) can describe and analyze discrete event systems with infinite states (DESIS). Due to the infinite state space and the combination explosion problem, the reachability analysis of UPNs is an NP-Hard problem. The existing reachability analysis methods cannot achieve an accurate result at reasonable costs (computational time and space) due to the finite reachability tree with$\omega$-numbers. Based on the idea of approximating infinite space with finite states, given some limited reachable markings of a UPN, we propose a method that can quantitatively solve the UPN’s reachability problem with machine learning. Firstly, we define the probabilistic reachability of markings and transform the UPN’s reachability problem into the prediction problem of markings. The proposed method based on positive and unlabeled learning (PUL) and bagging trains a classifier to predict the probabilistic reachability of unknown markings. Finally, to predict the markings outside the positive sample set and unlabeled sample set, an iterative strategy is designed to update the classifier. Based on seven general UPNs, the results of the experiments show that the proposed method has a good performance in the accuracy and time consumption for the UPN’s reachability problem.Note to Practitioners—In discrete event systems, the reachability problem mainly studies reachable states of the system and the relationship between states, which is the basis of the system’s states, behaviors, attributes and performance analysis. For discrete event systems with infinite states, it is hard to analyze the reachable relationship between states within a finite time due to the infinite state space and the combination explosion problem. The main motivation of the paper is to propose a method that can predict the reachable relationship between the states with a probability value within a finite time. By machine learning algorithms, the method learns the feature information of the known reachable states. The reachability of unknown states in the infinite state space can be predicted approximately. The proposed approximation method can be applied to analyze the reachability properties of general discrete event systems with infinite states, such as checking whether a fault occurs in operating systems, whether a message is delivered in communication and so on.
Hongda Qi, Mingjian Guang, Junli Wang 0001, ChunGang Yan, Changjun Jiang 0002
IEEE Trans Autom. Sci. Eng.3
2024 PHAED: A Speaker-Aware Parallel Hierarchical Attentive Encoder-Decoder Model for Multi-Turn Dialogue Generation
abstract
This paper presents a novel open-domain dialogue generation model emphasizing the differentiation of speakers in multi-turn conversations. Differing from prior work that treats the conversation history as a long text, we argue that capturing relative social relations among utterances (i.e., generated by either the same speaker or different persons) benefits the machine capturing fine-grained context information from a conversation history to improve context coherence in the generated response. Given that, we propose a Parallel Hierarchical Attentive Encoder-Decoder (PHAED) model that can effectively leverage conversation history by modeling each utterance with the awareness of its speaker and contextual associations with the same speaker's previous messages. Specifically, to distinguish the speaker roles over a multi-turn conversation (involving two speakers), we regard the utterances from one speaker as responses and those from the other as queries. After understanding queries via hierarchical encoder with inner-query and inter-query encodings, transformer-xl style decoder reuses the hidden states of previously generated responses to generate a new response. Our empirical results with three large-scale benchmarks show that PHAED significantly outperforms baseline models on both automatic and human evaluations. Furthermore, our ablation study shows that dialogue models with speaker tokens can generally decrease the possibility of generating non-coherent responses.
Ming Jiang 0018, Junli Wang 0001
IEEE Trans. Big Data3
2024 DRL-Based VNF Cooperative Scheduling Framework With Priority-Weighted Delay
abstract
Effective Service Function Chains (SFCs) mapping and Virtual Network Functions (VNFs) scheduling are crucial to ensure high-quality service provision for Internet of Things (IoT) tasks. Meeting the varying demands of multiple SFCs poses a significant challenge, particularly when working with the limited resources available in edge computing networks. Most existing working focuses on uniformly mapping and scheduling service requests in a batch processing manner within a given time period, without taking the diversity and priority of VNFs into account. When there is a sudden surge in demand, the issues of VNF queueing waiting and resources imbalance become prominent. To address the mentioned issues, this paper proposes a Deep Reinforcement Learning (DRL)-based VNF cooperative scheduling framework with priority-weighted delay. In light of the urgency of VNFs with higher priorities and the limitations of available resources, we begin by modeling an average queuing delay with priority weight based on the shortest remaining time priority technique. We then formulate a mathematical optimization problem to minimize the modeled delay in VNF scheduling process while providing suitable multidimensional resources in the edge network. Finally, a DRL method with experience replay and target Q-network is designed to effectively obtain the optimal solutions of the optimization problem from experience. The experimental results show that our proposed method outperforms its peers in terms of SFC request acceptance, delay, load balance, and resource utilization.
Junli Wang 0001, Cheng Wang 0001, ChunGang Yan
IEEE Trans. Mob. Comput.2
2024 A Multichannel Convolutional Decoding Network for Graph Classification
abstract
Graph convolutional networks (GCNs) have shown superior performance on graph classification tasks, and their structure can be considered as an encoder-decoder pair. However, most existing methods lack the comprehensive consideration of global and local in decoding, resulting in the loss of global information or ignoring some local information of large graphs. And the commonly used cross-entropy loss is essentially an encoder-decoder global loss, which cannot supervise the training states of the two local components (encoder and decoder). We propose a multichannel convolutional decoding network (MCCD) to solve the above-mentioned problems. MCCD first adopts a multichannel GCN encoder, which has better generalization than a single-channel GCN encoder since different channels can extract graph information from different perspectives. Then, we propose a novel decoder with a global-to-local learning pattern to decode graph information, and this decoder can better extract global and local information. We also introduce a balanced regularization loss to supervise the training states of the encoder and decoder so that they are sufficiently trained. Experiments on standard datasets demonstrate the effectiveness of our MCCD in terms of accuracy, runtime, and computational complexity.
Mingjian Guang, ChunGang Yan, Yuhua Xu 0005, Junli Wang 0001, Changjun Jiang 0002
IEEE Trans. Neural Networks Learn. Syst.4
2024 Stable QoE-Aware Multi-SFCs Cooperative Routing Mechanism Based on Deep Reinforcement Learning
abstract
The explosive development of the Internet of Things (IOT) has stimulated the sudden and dynamic demand pattern of potential Service Function Chain (SFC) traffic. The virtual network built based on the requirements of Multiple SFCs (Multi-SFCs), sharing the resource, should provide better Quality of Experience (QoE) for Multi-SFCs according to the real-time network status. However, simultaneous Multi-SFCs in the routing process compete for the VNF resources of the same node and preempt the same link. The instability of QoE caused by node invalidation and link faults is highlighted. Therefore, to provide a more efficient and stable service to Multi-SFCs, quantitative QoE utility and stability models and cooperative Multi-SFCs path allocation are needed. Considering the limited resources, this paper designs a stable QoE-aware Multi-SFCs Cooperative Routing Mechanism (CRM) based on Deep Reinforcement Learning (DRL). Firstly, we model the problem of Multi-SFCs cooperative routing aiming at optimizing diverse QoEs of utility, stability and delay, and formulate it as a multi-objective optimization problem. Moreover, the QoE stable queueing model based on node and link availability is formulated by Lyapunov drift, which can potentially mitigate the QoE fluctuations. Then we associate triple Dueling Deep Q-Networks (DDQNs) and propose a cooperative computing multi-task DRL to obtain the optimal path allocation policy. Experiments are conducted to verify the effectiveness and the results show that our method outperforms its peers on QoE, delay, throughput and stability.
ChunGang Yan, Junli Wang 0001, Changjun Jiang 0002
IEEE Trans. Netw. Serv. Manag.3
2024 The Probabilistic Liveness Decision Method of Unbounded Petri Nets Based on Machine Learning
abstract
The liveness of Petri nets (PNs) means that every event can occur in any state, establishing a close relationship with the deadlock-free property of existing systems. Due to the problem of state space explosion and the infinite state space of unbounded PNs (UPNs), the time complexity and space complexity of the liveness decision are difficult to give accurate measures; at least, they are both NP-hard. Except for some particular subclasses of UPNs, there has not been an accurate method to decide the liveness of generalized UPNs. Thus, a liveness decision method from a machine learning perspective is proposed to predict probability values about UPNs’ liveness within a finite time. The method aims to learn the feature information on UPNs by deep neural networks and establish the mapping relationship between the UPNs and the liveness. First, the concept of approximating infinite space with finite states is applied to generate reachability graphs at different moments, following the firing rules of a UPN. Then, the graph convolutional network (GCN)-based reachability graph feature representation module and the gated recurrent unit (GRU)-based UPN feature representation module are designed to map the reachability graphs at different moments into the low-dimensional feature space. And the feature vector that can characterize the UPN’s liveness is obtained to decide the liveness probabilistically. Finally, three datasets, including 50000 samples, are constructed. Based on these datasets and some case studies, the experimental results validate the method’s ability to make liveness decisions for UPNs, demonstrating its strong performance in terms of effectiveness and generalization.
Hongda Qi, Junli Wang 0001, ChunGang Yan, Changjun Jiang 0002
IEEE Trans. Syst. Man Cybern. Syst.2
2023 Dynamic and Static Feature-Aware Microservices Decomposition via Graph Neural Networks
Mingjian Guang, Junli Wang 0001, ChunGang Yan
KSEM (1)3
2023 Feature-wise attention based boosting ensemble method for fraud detection
Ruihao Cao, Junli Wang 0001, Mingze Mao, Guanjun Liu, Changjun Jiang 0002
Eng. Appl. Artif. Intell.2
2023 Class-homophilic-based data augmentation for improving graph neural networks
Rui Duan 0003, ChunGang Yan, Junli Wang 0001, Changjun Jiang 0002
Knowl. Based Syst.3
2023 DCOM-GNN: A Deep Clustering Optimization Method for Graph Neural Networks
Haoran Yang 0003, Junli Wang 0001, Rui Duan 0003, ChunGang Yan
Knowl. Based Syst.2
2023 Multistructure Graph Classification Method With Attention-Based Pooling
abstract
Graph neural networks (GNNs) have achieved effective performance in many graph-related tasks involving recommendation systems, social networks, and bioinformatics. Recent studies have proposed several graph pooling operators to obtain graph-level representations from node representations. Nevertheless, they usually adopt a single strategy to evaluate the importance of nodes, which may generate node rankings with weak robustness. Also, they cannot capture the different substructures of a graph since they shrink the graph layer by layer. To solve the above problems, this article proposes a Multistructure graph classification method with Attention mechanism and Convolutional neural network (CNN), called MAC. In particular, we propose a novel pooling operator, which adopts multiple strategies to evaluate the importance of nodes and updates node representations through an attention mechanism. Also, we design a hierarchical architecture for MAC to capture multiple different substructures of a graph. To further reduce the loss of graph information, we utilize 2-D CNN to generate a graph-level representation. Comparative experiments are performed on public benchmark datasets deriving from social systems, and the experimental results indicate that our method outperforms a range of state-of-the-art graph classification methods.
Yuhua Xu 0005, Junli Wang 0001, Mingjian Guang, ChunGang Yan, Changjun Jiang 0002
IEEE Trans. Comput. Soc. Syst.2
2022 MUSH: Multi-scale Hierarchical Feature Extraction for Semantic Image Synthesis
Zicong Wang, Junli Wang 0001, ChunGang Yan, Changjun Jiang 0002
ACCV (7)3
2022 LAMPT: LAbel Mask-Predicted Transformer for Extreme Multi-label Text Classification
abstract
Extreme Multi-label text Classification ( XMC) is a task of recalling the most relevant labels for each given text from an extremely large-scale label set. It is emphasized that XMC is a more complex classification task because there are two main problems: large space of labels and the labels in XMC tasks tend to be correlated. Existing methods attempt to model label correlations by viewing the XMC task as a sequence generation problem however they still suffer from (1) using the slow serial decoding strategy where labels are predicted one-by-one. (2) needing to compare a mass of label ordering strategies in the decoding stage to achieve satisfied accuracy. In this work, we propose LAbel Mask-Predicted Transformer (LAMPT) to address the both issues, which is a novel non-autoregressive generation model that (1) enriches the input raw text representation with the additional label features by fully exploiting the label dependencies, (2) allows for efficient parallel decoding thanks to its non-autoregressive decoding formulation and mask-prediced training strategy. Experimental results demonstrate that single model performance is substantially enhanced by LAMPT. On a Wiki dataset with thirty-one thousand labels, LAMPT-XLNet accuracy has gained 1.6% relative improvement on P@3 over the LightXML-XLNet. Also, the P@1 of ensemble LAMPT is 90.00%, a significant i mprovement over the state-of-the-art ensemble LightXML (transformer-based) and AttentionXML (LSTM-based), which achieve 89.45% and 87.47%, respectively.
Junli Wang 0001, ChunGang Yan
IEEE Big Data3
2022 Word Order is Considerable: Contextual Position-aware Graph Neural Network for Text Classification
abstract
Text Classification is a crucial and classical research branch in Natural Language Processing. Recently, Graph Neural Networks (GNNs) have been explored for this task since GNNs are proficient in handling non-Euclidean data with more complex relations. Despite the success, these graph-based methods have some limitations to model the word order information in the document. In this paper, to address this issue, we propose the Contextual Position-aware Graph Neural Network for text classification, which includes the Position-aware Graph Attention module and the Contextual Fusion module. The former can respectively capture left-side, self-side, and right-side word order information of each word. The latter enables the Context-fused Graph LSTM to learn word contextual representations by controlling three contextual information flows. In addition, experimental results have demonstrated that our proposed method can effectively enhance text representations, thus improving the classification performance.
Yingying Tan, Junli Wang 0001
IJCNN2
2022 Unified Multimodal Model with Unlikelihood Training for Visual Dialog
abstract
The task of visual dialog requires a multimodal chatbot to answer sequential questions from humans about image content. Prior work performs the standard likelihood training for answer generation on the positive instances (involving correct answers). However, the likelihood objective often leads to frequent and dull outputs and fails to exploit the useful knowledge from negative instances (involving incorrect answers). In this paper, we propose a Unified Multimodal Model with UnLikelihood Training, named UniMM-UL, to tackle this problem. First, to improve visual dialog understanding and generation by multi-task learning, our model extends ViLBERT from only supporting answer discrimination to holding both answer discrimination and answer generation seamlessly by different attention masks. Specifically, in order to make the original discriminative model compatible with answer generation, we design novel generative attention masks to implement the autoregressive Masked Language Modeling (autoregressive MLM) task. And to attenuate the adverse effects of the likelihood objective, we exploit unlikelihood training on negative instances to make the model less likely to generate incorrect answers. Then, to utilize dense annotations, we adopt different fine-tuning methods for both generating and discriminating answers, rather than just for discriminating answers as in the prior work. Finally, on the VisDial dataset, our model achieves the best generative results (69.23 NDCG score). And our model also yields comparable discriminative results with the state-of-the-art in both single-model and ensemble settings (75.92 and 76.17 NDCG scores).
Junli Wang 0001, Changjun Jiang 0002
ACM Multimedia2
2022 Non-Autoregressive Neural Machine Translation with Consistency Regularization Optimized Variational Framework
abstract
Variational Autoencoder (VAE) is an effective framework to model the interdependency for non-autoregressive neural machine translation (NAT).One of the prominent VAE-based NAT frameworks, LaNMT, achieves great improvements to vanilla models, but still suffers from two main issues which lower down the translation quality: (1) mismatch between training and inference circumstances and (2) inadequacy of latent representations.In this work, we target on addressing these issues by proposing posterior consistency regularization.Specifically, we first perform stochastic data augmentation on the input samples to better adapt the model for inference circumstance, and then conduct consistency training on posterior latent variables to construct a more robust latent representations without any expansion on latent size.Experiments on En<->De and En<->Ro benchmarks confirm the effectiveness of our methods with about 1.5/0.7 and 0.8/0.3BLEU points improvement to the baseline model with about 12.6× faster than autoregressive Transformer.
Junli Wang 0001, ChunGang Yan
NAACL-HLT2
2022 GNN-EA: Graph Neural Network with Evolutionary Algorithm
abstract
Recently, Graph Neural Networks (GNNs) have shown great promise in addressing various tasks with non-Euclidean data. Encouraged by the successful application on discovering convolutional and recurrent neural networks, Neural Architecture Search (NAS) is extended to alleviate the complexity of designing appropriate task-specific GNNs. Unfortunately, existing graph NAS methods are usually susceptible to unscalable depth, redundant computation, constrained search space and some other limitations. In this paper, we present an evolutionary graph neural network architecture search strategy, involving inheritance, crossover and mutation operators based on fine-grained atomic operations. Specifically, we design two novel crossover operators at different granularity levels, GNNCross and LayerCross. Experiments on three different graph learning tasks indicate that the neural architectures generated by our method exhibit comparable performance to the handcrafted and automated baseline GNN models.
Junli Wang 0001
SMC3
2022 Path-aware multi-hop graph towards improving graph learning
Rui Duan 0003, ChunGang Yan, Junli Wang 0001, Changjun Jiang 0002
Neurocomputing3
2022 Net Learning
abstract
Graph neural networks, which generalize deep learning to graph-structured data, have achieved significant improvements in numerous graph-related tasks. Petri nets (PNs), on the other hand, are mainly used for the modeling and analysis of various event-driven systems from the perspective of prior knowledge, mechanisms, and tasks. Compared with graph data, net data can simulate the dynamic behavioral features of systems and are more suitable for representing real-world problems. However, the problem of large-scale data analysis has been puzzling the PN field for decades, and thus, limited its universal applicability. In this article, a framework of net learning (NL) is proposed. NL contains the advantages of PN modeling and analysis with the advantages of graph learning computation. Then, two kinds of NL algorithms are designed for performance analysis of stochastic PNs, and more specifically, the hidden feature information of the PN is obtained by mapping net information to the low-dimensional feature space. Experiments demonstrate the effectiveness of the proposed model and algorithms on the performance analysis of stochastic PNs.
Junli Wang 0001, Hongda Qi, Mingjian Guang, ChunGang Yan, Changjun Jiang 0002
IEEE Trans. Neural Networks Learn. Syst.1
2021 Benchmark Datasets for Stochastic Petri Net Learning
abstract
The existing Stochastic Petri Net (SPN) analysis methods are based on a series of steps, including generating the reachable graph, solving the state equation, etc. Unfortunately, these methods cannot perform performance analysis when the state equation has no unique solution. The end-to-end deep learning methods can build a mapping relationship from SPN to performance indicators, which avoids solving the state equation. However, there is a lack of benchmark datasets for SPN learning and training. This paper proposes an automatic generation method of SPN datasets, including SPN random generation, data labeling, data enhancement, and filtering. To relieve the local aggregation problem of random-based organization, a grid-based data organization method is proposed to ensure the diversity of the datasets. In the experimental section, the generated datasets are trained and tested on three types of neural networks. The results show that the generated benchmark datasets are useful, and the increase of dataset size will significantly improve the learning performance.
Mingjian Guang, ChunGang Yan, Junli Wang 0001, Hongda Qi, Changjun Jiang 0002
IJCNN3
2021 Face Image Inpainting With Evolutionary Generators
abstract
Recently, deep learning has become a mainstream method of image inpainting. It can not only restore the image texture, obtain high-level abstract features of images, but also restore semantic images such as human face images. Among these methods, generative adversarial networks (GANs) using autoencoder as the generator have become the promising model for image inpainting. These models implement the end-to-end image inpainting and also generate visually reasonable and clear image structures and textures. However, GANs often have problems with gradient vanishing and model collapse during training, so we propose a Generative Adversarial Network with Evolutionary Generators (EG-GAN) and apply it in face image inpainting. To stabilize the model training process, EG-GAN trains the generator network by evolution, combines two mutation functions as a training objective to update the parameter of generator networks, and produces offspring generators through crossover, using the matcher assists the discriminator to criticize the generated image. Experiments on various face image datasets such as CelebA-HQ and CelebA show that EG-GAN successfully overcomes the gradient vanishing problem, achieves stable and efficient training, and generates visually reasonable images.
Chong Han 0004, Junli Wang 0001
IEEE Signal Process. Lett.2
2014 News Topic Evolution Tracking by Incorporating Temporal Information
Junli Wang 0001
NLPCC3
2007 AI Planning for Web Service Automatic Composition Using Petri Nets
abstract
This paper presents an AI planning method for Web service automatic composition using Predicate/Transition (Pr/T) net model. First, based on Web service description of inputs, outputs, preconditions and effects in OWL-S specification, a Pr/T net model is constructed for a service composition plan. Then, this plan is effectively solved by a reachability algorithm of the Pr/T net and a corresponding reachability graph is obtained. Moreover, a regular language of Pr/T net is generated for representing all plan paths of a service composition. Finally, a process model of composite service is extracted after normalization of the plan paths. This method provides a complete solution from modeling AI plan to constructing process model for service composition, so it is helpful for realizing automatic composition and dynamical integration of Web service.
Zhijun Ding, Junli Wang 0001
CSCWD2
2006 GAOM: Genetic Algorithm Based Ontology Matching
abstract
In this paper a genetic algorithm-based optimization procedure for ontology matching problem is presented as a feature-matching process. First, from a global view, we model the problem of ontology matching as an optimization problem of a mapping between two compared ontologies, and every ontology has its associated feature sets. Second, as a powerful heuristic search strategy, genetic algorithm is employed for the ontology matching problem. Given a certain mapping as optimizing object for GA, fitness function is defined as a global similarity measure function between two ontologies based on feature sets. Finally, a set of experiments are conducted to analysis and evaluate the performance of GA in solving ontology matching problem
Junli Wang 0001, Zhijun Ding, Changjun Jiang 0002
APSCC1
2001 The Influence of the Probability Density Function on Similartaxis in MEC
abstract
Mind evolutionary computation (MEC) is a new approach of evolutionary computation (EC). It is proved that MEC has much higher computing efficiency and convergence ability than genetic algorithms (GAs). This is because of using operation similartaxis and dissimilation rather than crossover and mutation operators in GA. The paper analyzes the influence of type of the probability density function on similartaxis in MEC. We get theoretically the relation among similartaxis calculated amount, the parameters of probability density function of scattering individuals, the size of group, the precision of solution and the distance between initial searching position and local optimum. The experiment shows that the analysis method proposed in the paper is reasonable. The analysis and experiment also shows that using different types of probability density functions doesn't make much change on similartaxis searching performance.
Chengyi Sun, Jianqing Zhang, Junli Wang 0001
FUZZ-IEEE3
2001 MEC dissimilation strategy by rejected regions
abstract
Mind evolutionary computation (MEC) is a new approach to evolutionary computation (EC). This paper presents a new dissimilation strategy using rejected regions, which can avoid searching repeatedly, so that the capability of MEC to search globally in dissimilation is enhanced. Experimental results show that basic MEC has improved considerately compared with a genetic algorithm (GA), and that the MEC dissimilation strategy using rejected regions has also advanced a lot. The reason for this is that, in the modified MEC, the regions searched in similartaxis are recorded, so that, in dissimilation, the scope of scattered individuals is reduced to the whole solution space, excluding the rejected regions. Therefore, the regions explored in dissimilation have never been searched before, and the search scope is diminished accordingly, while the capability of MEC to search globally in dissimilation is enhanced and repeated searching is avoided. It is the memory mechanism of MEC that makes the dissimilation strategy of rejected regions possible, so the probability that the individuals are scattered in the region of the global optimum has greatly increased, the calculated amount and the average evaluation time are decreased, and population convergence can be implemented in fewer generations.
Chengyi Sun, Junli Wang 0001, Jianqing Zhang
SMC2
2001 Dissimilation strategy of avoiding searching the same peak
abstract
Mind evolutionary computation (MEC) is a new approach to evolutionary computation (EC). It is proved that MEC has much higher computational efficiency and convergence ability than genetic algorithms (GAs). This is because MEC uses the operations of similartaxis and dissimilation rather than the crossover and mutation operators used in GAs, and also because of the different performance mechanisms from GAs, the memory mechanism, the evolutionary directional mechanism and the harmonizing mechanism between exploitation and exploration. This paper presents a new dissimilation strategy using rejected regions, which can avoid searching repeatedly, so that the capability of MEC to search globally in dissimilation is enhanced. It is the memory mechanism of MEC that makes the dissimilation strategy of rejected regions possible.
Jianqing Zhang, Chengyi Sun, Junli Wang 0001
SMC3