EDBT 2026 Demo / reviewers in the wild / expert
Zhenghua Xu 0001
dblp:80/11498-1
· DBLP profile ↗
37ranked-venue papers
12as first author
26since 2021 · last 2026
0000-0002-6719-7333ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 7 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 2 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 4 first-author · 8 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Modeling Uncertainty Trends for Timely Retrieval in Dynamic RAGabstractDynamic retrieval-augmented generation (RAG) allows large language models (LLMs) to fetch external knowledge on demand, offering greater adaptability than static RAG. A central challenge in this setting lies in determining the optimal timing for retrieval. Existing methods often trigger retrieval based on low token-level confidence, which may lead to delayed intervention after errors have already propagated. We introduce Entropy-Trend Constraint (ETC), a training-free method that determines optimal retrieval timing by modeling the dynamics of token-level uncertainty. Specifically, ETC utilizes first- and second-order differences of the entropy sequence to detect emerging uncertainty trends, enabling earlier and more precise retrieval. Experiments on six QA benchmarks with three LLM backbones demonstrate that ETC consistently outperforms strong baselines while reducing retrieval frequency. ETC is particularly effective in domain-specific scenarios, exhibiting robust generalization capabilities. Ablation studies and qualitative analyses further confirm that trend-aware uncertainty modeling yields more effective retrieval timing. The method is plug-and-play, model-agnostic, and readily integrable into existing decoding pipelines. Implementation code is included in the supplementary materials. Bo Li 0099, Zhenghua Xu 0001, Shikun Zhang, Wei Ye 0004 |
AAAI | 3 |
| 2026 | Language Drift in Multilingual Retrieval-Augmented Generation: Characterization and Decoding-Time MitigationabstractMultilingual Retrieval-Augmented Generation (RAG) enables large language models (LLMs) to perform knowledge-intensive tasks in multilingual settings by leveraging retrieved documents as external evidence. However, when the retrieved evidence differs in language from the user query and in-context exemplars, the model often exhibits language drift by generating responses in an unintended language. This phenomenon is especially pronounced during reasoning-intensive decoding, such as Chain-of-Thought (CoT) generation, where intermediate steps introduce further language instability. In this paper, we systematically study output language drift in multilingual RAG across multiple datasets, languages, and LLM backbones. Our controlled experiments reveal that the drift results not from comprehension failure but from decoder-level collapse, where dominant token distributions and high-frequency English patterns dominate the intended generation language. We further observe that English serves as a semantic attractor under cross-lingual conditions, emerging as both the strongest interference source and the most frequent fallback language. To mitigate this, we propose Soft Constrained Decoding (SCD), a lightweight, training-free decoding strategy that gently steers generation toward the target language by penalizing non-target-language tokens. SCD is model-agnostic and can be applied to any generation algorithm without modifying the architecture or requiring additional data. Experiments across three multilingual datasets and multiple typologically diverse languages show that SCD consistently improves language alignment and task performance, providing an effective and generalizable solution in multilingual RAG. Bo Li 0099, Zhenghua Xu 0001, Rui Xie 0003 |
AAAI | 2 |
| 2026 | AttCL-GAN: Attentional contrastive learning-based generative adversarial network for modality completion of medical images
Zhenghua Xu 0001, Jiaqi Tang 0013, Thomas Lukasiewicz |
Knowl. Based Syst. | 1 |
| 2026 | Semi-supervised medical image lesion detection based on multi-head feature fusion
Zhenghua Xu 0001, Hexiang Zhang, Runhe Yang, Weipeng Liu, Thomas Lukasiewicz |
Knowl. Based Syst. | 1 |
| 2026 | Advancing federated semi-supervised medical image segmentation: A duo of interactive denoising pseudo-labels and convolutional contrastive learning
Zhenghua Xu 0001, Bo Li 0099, Gaoxi Zhou, Xianglin Lu, Thomas Lukasiewicz |
Medical Image Anal. | 1 |
| 2026 | You Need Glimpse Before Segmentation: Stochastic Detector-Actor-Critic for Medical Image SegmentationabstractMedical images often contain more redundant background areas than natural images, potentially introducing noise and degrading image segmentation performance. Inspired by doctors' diagnostic processes, where they identify the lesion area before conducting a detailed analysis, we introduce a novel Stochastic Detector-Actor-Critic (SDAC) framework to tackle this challenge. SDAC initially glimpses the entire image using a detector network and policy gradient algorithms to filter out irrelevant background regions and focus on crucial, smaller areas for segmentation. The Actor-Critic algorithm then dynamically creates segmentation masks pixel by pixel without user intervention or coarse masks, forming a robust segmentation module. Both processes are trained jointly to reduce error propagation and ensure stability and ease of implementation. Our experiments on two commonly used medical image segmentation datasets demonstrate that SDAC achieves competitive results comparable to state-of-the-art methods while using 10x fewer parameters than the best-performing baseline in terms of DICE and IoU metrics. We also conduct detailed ablation studies to enhance understanding and facilitate practical use. Furthermore, SDAC performs well in low-resource settings (i.e., 50-shot or 100-shot), making it ideal for real-world scenarios. Its lightweight design make SDAC an excellent baseline for medical image segmentation tasks. Zhenghua Xu 0001, Bo Li 0099, Weipeng Liu, Thomas Lukasiewicz |
IEEE J. Biomed. Health Informatics | 1 |
| 2026 | AMLP: Adjustable Masking Lesion Patches for Self-Supervised Medical Image SegmentationabstractSelf-supervised masked image modeling (MIM) methods have shown promising performances on analyzing natural images. However, directly applying such methods to medical image segmentation tasks still cannot achieve satisfactory results. The challenges arise from the facts that (i) medical images are inherently more complex compared to natural images, and the subjects in medical images often exhibit more distinct contour features; (ii) moreover, the conventional high and fixed masking ratio in MIM is likely to mask the background, limiting the scope of learnable information. To address these problems, we propose a new self-supervised medical image segmentation framework, called Adjustable Masking Lesion Patches (AMLP), which employs Masked Patch Selection (MPS) strategy to identify patches with high probabilities of containing lesions to help model achieve precise lesion reconstruction. To improve the categorization of patches in MPS, we further introduce Relative Reconstruction Loss (RRL) to better learn hard-to-reconstruct lesion patches. Then, Category Consistency Loss (CCL) is proposed to refine patch categorization based on reconstruction difficulty, enhancing difference between lesions and backgrounds. Moreover, an Adjustable Masking Ratio (AMR) strategy is proposed to gradually increase the masking ratio over training to expand the scope of learnable mutual information. Extensive experiments on two medical segmentation datasets demonstrate the superior performances of the proposed AMLP w.r.t. the SOTA self-supervised methods; the results prove that AMLP effectively addresses the challenges of applying masked modeling to medical images and capturing accurate lesion details that are crucial for segmentation tasks. Xiangtao Wang, Thomas Lukasiewicz, Zhenghua Xu 0001 |
IEEE Trans. Medical Imaging | 4 |
| 2025 | Aggregated Mutual Learning between CNN and Transformer for semi-supervised medical image segmentation
Zhenghua Xu 0001, Hening Wang, Runhe Yang, Weipeng Liu, Thomas Lukasiewicz |
Knowl. Based Syst. | 1 |
| 2025 | CGraphNet: Contrastive Graph Context Prediction for Sparse Unlabeled Short Text Representation Learning on Social MediaabstractUnlabeled text representation learning (UTRL), encompassing static word embeddings such as Word2Vec and contextualized word embeddings such as bidirectional encoder representations from transformer (BERT), aims to capture semantic word relationships in a low-dimensional space without the need for manual labeling. These word embeddings are invaluable for downstream tasks such as document classification and clustering. However, the surge of short texts generated daily on social media platforms results in sparse word cooccurrences, compromising UTRL outcomes. Contextualized models such as recurrent neural network (RNN) and BERT, while impressive, often struggle with predicting the next word due to sparse word sequences in short texts. To address this, we introduce CGraphNet, a contrastive graph context prediction model designed for UTRL. This approach converts short texts into graphs, establishing links between sequentially occurring words. Information from the next word and its neighbors informs the target prediction, a process referred to as graph context prediction, mitigating sparse word cooccurrence issues in brief sentences. To minimize noise, an attention mechanism assigns importance to neighbors, while a contrastive objective encourages more distinctive representations by comparing the target word with its neighbors. Our experiments demonstrate CGraphNet's superior performance over other baselines, particularly in classification and clustering tasks on real-world datasets. Junyang Chen 0001, Jingcai Guo, Xueliang Li 0002, Huan Wang 0005, Zhenghua Xu 0001, Zhiguo Gong, Liang-Jie Zhang, Victor C. M. Leung |
IEEE Trans. Comput. Soc. Syst. | 5 |
| 2025 | Self-Supervised Medical Image Segmentation Using Deep Reinforced Adaptive MaskingabstractSelf-supervised learning aims to learn transferable representations from unlabeled data for downstream tasks. Inspired by masked language modeling in natural language processing, masked image modeling (MIM) has achieved certain success in the field of computer vision, but its effectiveness in medical images remains unsatisfactory. This is mainly due to the high redundancy and small discriminative regions in medical images compared to natural images. Therefore, this paper proposes an adaptive hard masking (AHM) approach based on deep reinforcement learning to expand the application of MIM in medical images. Unlike predefined random masks, AHM uses an asynchronous advantage actor-critic (A3C) model to predict reconstruction loss for each patch, enabling the model to learn where masking is valuable. By optimizing the non-differentiable sampling process using reinforcement learning, AHM enhances the understanding of key regions, thereby improving downstream task performance. Experimental results on two medical image datasets demonstrate that AHM outperforms state-of-the-art methods. Additional experiments under various settings validate the effectiveness of AHM in constructing masked images. Zhenghua Xu 0001, Thomas Lukasiewicz |
IEEE Trans. Medical Imaging | 1 |
| 2025 | Hybrid Reinforced Medical Report Generation With M-Linear Attention and Repetition PenaltyabstractTo reduce doctors' workload, deep-learning-based automatic medical report generation has recently attracted more and more research efforts, where deep convolutional neural networks (CNNs) are employed to encode the input images, and recurrent neural networks (RNNs) are used to decode the visual features into medical reports automatically. However, these state-of-the-art methods mainly suffer from three shortcomings: 1) incomprehensive optimization; 2) low-order and unidimensional attention; and 3) repeated generation. In this article, we propose a hybrid reinforced medical report generation method with m-linear attention and repetition penalty mechanism (HReMRG-MR) to overcome these problems. Specifically, a hybrid reward with different weights is employed to remedy the limitations of single-metric-based rewards, and a local optimal weight search algorithm is proposed to significantly reduce the complexity of searching the weights of the rewards from exponential to linear. Furthermore, we use m-linear attention modules to learn multidimensional high-order feature interactions and to achieve multimodal reasoning, while a new repetition penalty is proposed to apply penalties to repeated terms adaptively during the model's training process. Extensive experimental studies on two public benchmark datasets show that HReMRG-MR greatly outperforms the state-of-the-art baselines in terms of all metrics. The effectiveness and necessity of all components in HReMRG-MR are also proved by ablation studies. Additional experiments are further conducted and the results demonstrate that our proposed local optimal weight search algorithm can significantly reduce the search time while maintaining superior medical report generation performances. Zhenghua Xu 0001, Wenting Xu, Junyang Chen 0001, Chang Qi, Thomas Lukasiewicz |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | A Stable, Fast, and Fully Automatic Learning Algorithm for Predictive Coding NetworksabstractPredictive coding networks are neuroscience-inspired models with roots in both Bayesian statistics and neuroscience. Training such models, however, is quite inefficient and unstable. In this work, we show how by simply changing the temporal scheduling of the update rule for the synaptic weights leads to an algorithm that is much more efficient and stable than the original one, and has theoretical guarantees in terms of convergence. The proposed algorithm, that we call incremental predictive coding (iPC) is also more biologically plausible than the original one, as it it fully automatic. In an extensive set of experiments, we show that iPC constantly performs better than the original formulation on a large number of benchmarks for image classification, as well as for the training of both conditional and masked language models, in terms of test accuracy, efficiency, and convergence with respect to a large set of hyperparameters. Tommaso Salvatori, Yuhang Song 0001, Yordan Yordanov, Beren Millidge, Lei Sha, Cornelius Emde, Zhenghua Xu 0001, Rafal Bogacz, Thomas Lukasiewicz |
ICLR | 7 |
| 2024 | Multi-ConDoS: Multimodal Contrastive Domain Sharing Generative Adversarial Networks for Self-Supervised Medical Image SegmentationabstractExisting self-supervised medical image segmentation usually encounters the domain shift problem (i.e., the input distribution of pre-training is different from that of fine-tuning) and/or the multimodality problem (i.e., it is based on single-modal data only and cannot utilize the fruitful multimodal information of medical images). To solve these problems, in this work, we propose multimodal contrastive domain sharing (Multi-ConDoS) generative adversarial networks to achieve effective multimodal contrastive self-supervised medical image segmentation. Compared to the existing self-supervised approaches, Multi-ConDoS has the following three advantages: (i) it utilizes multimodal medical images to learn more comprehensive object features via multimodal contrastive learning; (ii) domain translation is achieved by integrating the cyclic learning strategy of CycleGAN and the cross-domain translation loss of Pix2Pix; (iii) novel domain sharing layers are introduced to learn not only domain-specific but also domain-sharing information from the multimodal medical images. Extensive experiments on two publicly multimodal medical image segmentation datasets show that, with only 5% (resp., 10%) of labeled data, Multi-ConDoS not only greatly outperforms the state-of-the-art self-supervised and semi-supervised medical image segmentation baselines with the same ratio of labeled data, but also achieves similar (sometimes even better) performances as fully supervised segmentation methods with 50% (resp., 100%) of labeled data, which thus proves that our work can achieve superior segmentation performances with very low labeling workload. Furthermore, ablation studies prove that the above three improvements are all effective and essential for Multi-ConDoS to achieve this very superior performance. Shuo Zhang 0017, Xiaoqian Shen, Thomas Lukasiewicz, Zhenghua Xu 0001 |
IEEE Trans. Medical Imaging | 5 |
| 2023 | MPS-AMS: Masked Patches Selection and Adaptive Masking Strategy Based Self-Supervised Medical Image SegmentationabstractExisting self-supervised learning methods based on contrastive learning and masked image modeling have demonstrated impressive performances. However, current masked image modeling methods are mainly utilized in natural images, and their applications in medical images are relatively lacking. Besides, their fixed high masking strategy limits the upper bound of conditional mutual information, and the gradient noise is considerable, making less the learned representation information. Motivated by these limitations, in this paper, we propose masked patches selection and adaptive masking strategy based self-supervised medical image segmentation method, named MPS-AMS. We leverage the masked patches selection strategy to choose masked patches with lesions to obtain more lesion representation information, and the adaptive masking strategy is utilized to help learn more mutual information and improve performance further. Extensive experiments on three public medical image segmentation datasets (BUSI, Hecktor, and Brats2018) show that our proposed method greatly outperforms the state-of-the-art self-supervised baselines. Xiangtao Wang, Shuo Zhang 0017, Junyang Chen 0001, Thomas Lukasiewicz, Zhenghua Xu 0001 |
ICASSP | 8 |
| 2023 | MvCo-DoT: Multi-View Contrastive Domain Transfer Network for Medical Report GenerationabstractIn clinical scenarios, multiple medical images with different views are usually generated at the same time, and they have high semantic consistency. However, the existing medical report generation methods cannot exploit the rich multi-view mutual information of medical images. Therefore, in this work, we propose the first multi-view medical report generation model, called MvCo-DoT. Specifically, MvCo-DoT first propose a multi-view contrastive learning (MvCo) strategy to help the deep reinforcement learning based model utilize the consistency of multi-view inputs for better model learning. Then, to close the performance gaps of using multi-view and single-view inputs, a domain transfer network is further proposed to ensure MvCo-DoT achieve almost the same performance as multi-view inputs using only single-view inputs. Extensive experiments on the IU X-Ray public dataset show that MvCo-DoT outperforms the SOTA medical report generation baselines in all metrics. Xiangtao Wang, Zhenghua Xu 0001, Wenting Xu, Junyang Chen 0001, Thomas Lukasiewicz |
ICASSP | 3 |
| 2023 | Multi-Head Feature Pyramid Networks for Breast Mass DetectionabstractAnalysis of X-ray images is one of the main tools to diagnose breast cancer. The ability to quickly and accurately detect the location of masses from the huge amount of image data is the key to reducing the morbidity and mortality of breast cancer. Currently, the main factor limiting the accuracy of breast mass detection is the unequal focus on the mass boxes, leading the network to focus too much on larger masses at the expense of smaller ones. In the paper, we propose the multi-head feature pyramid module (MHFPN) to solve the problem of unbalanced focus of target boxes during feature map fusion and design a multi-head breast mass detection network (MBMDnet). Experimental studies show that, comparing to the SOTA detection baselines, our method improves by 6.58% (in AP@50) and 5.4% (in TPR@50) on the commonly used IN-breast dataset, while about 6-8% improvements (in AP@20) are also observed on the public MIAS and BCS-DBT datasets. Hexiang Zhang, Zhenghua Xu 0001, Shuo Zhang 0017, Junyang Chen 0001, Thomas Lukasiewicz |
ICASSP | 2 |
| 2023 | Adaptive-Masking Policy with Deep Reinforcement Learning for Self-Supervised Medical Image SegmentationabstractAlthough self-supervised learning methods based on masked image modeling have achieved some success in improving the performance of deep learning models, these methods have difficulty in ensuring that the masked region is the most appropriate for each image, resulting in segmentation networks that do not get the best weights in pre-training. Therefore, we propose a new adaptive-masking policy self-supervised learning method. Specifically, we model the process of masking images as a reinforcement learning problem and use the results of the reconstruction model as a feedback signal to guide the agent to learn the masking policy to select a more appropriate mask position and size for each image, helping the reconstruction network to learn more fine-grained image representation information and thus improve the downstream segmentation model performance. We conduct extensive experiments on two datasets, Cardiac and TCIA, and the results show that our approach outperforms current state-of-the-art self-supervised learning methods. Shengxin Wang, Thomas Lukasiewicz, Zhenghua Xu 0001 |
ICME | 4 |
| 2023 | A Neural Inference of User Social Interest for Item RecommendationabstractAbstract User-generated content is daily produced in social media, as such user interest summarization is critical to distill salient information from massive information for recommendation tasks. While the interested messages (e.g., tags or posts) from a single user are usually sparse becoming a bottleneck for existing methods, we propose a neural inference method (NIGraphNet) by mining user social interest for item recommendation. It can unearth user latent topics combined with user relation learning. Specifically, we exploit a neural variational inference approach to learn the distributions between user interests and hidden topics. (We denote it as interest-topic distributions in the following.) Then, we adopt a unified graph-based training loss that jointly learns the hidden topics and user relations for item recommendation. Experiments on two datasets collected from well-known social media platforms demonstrate the superior performance of our model in the tasks of user interest summarization and item recommendation. Further discussions also show that exploiting the latent topic representations and user relations is conducive to the user’s automatic language understanding. Junyang Chen 0001, Mengzhu Wang, Ge Fan, Guo Zhong, Ou Liu, Wenfeng Du, Zhenghua Xu 0001, Zhiguo Gong |
Data Sci. Eng. | 8 |
| 2023 | Multi-modal contrastive mutual learning and pseudo-label re-learning for semi-supervised medical image segmentation
Shuo Zhang 0017, Thomas Lukasiewicz, Zhenghua Xu 0001 |
Medical Image Anal. | 5 |
| 2023 | IRLM: Inductive Representation Learning Model for Personalized POI RecommendationabstractWith the rapid development of the Internet of Things technology, the concept of smart cities that aims to help residents improve their quality of life has raised much attention in several application areas. In the context of smart cities, the provision of point of interest (POI) recommendations become an important requirement because a wide range of POIs are available for urban dwellers. Location-based social networks (LBSNs) such as Foursquare and Gowalla provide a massive volume of user check-in records that can assist users in choosing new POIs. However, user trajectories are mostly sparse in the real world. For example, users only check in a few POIs, and this makes it difficult to provide recommendations based on limited history trajectories. Though some attempts have adopted auxiliary geographical information to enhance POI recommendation, they still encounter the following problems: 1) the geographical trajectories of users are usually sparse in real-world datasets; 2) users may be more interested in the remote POIs; and 3) the previous models inherently perform transductive learning that cannot handle well the recommendation of unseen users and POIs. To address these problems, we propose an inductive representation learning model (IRLM) for location recommendation. IRLM contains two parts, namely geographic feature extraction and inductive representation learning. IRLM first captures global geographical influences among POIs through a standard Gaussian mixture model (GMM). Then IRLM adopts an attention neural network for the recommendation. Experimental results indicate that our proposed model can achieve superior performance over state-of-the-art models. Junyang Chen 0001, Mengzhu Wang, Zhenghua Xu 0001, Xueliang Li 0002, Zhiguo Gong, Kaishun Wu, Victor C. M. Leung |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2023 | Toward Knowledge as a Service (KaaS): Predicting Popularity of Knowledge Services Leveraging Graph Neural NetworksabstractKnowledge services are becoming a rising star in the family of XaaS (Everything as a Service). In recent years, people are more willing to search for answers and share their knowledge directly over the Internet, which makes the knowledge service ecosystem prosperous. In this paper, we aim to predict the popularity of knowledge services, which will benefit the downstream industries. Toward such a task, the spatial interactions (e.g., hyperlinks in Wikipedia) and temporal observations (e.g., page views) provide crucial information. However, it is difficult to utilize this information due to: (i) complicated and different usage observations, (ii) intricate and evolutionary spatial interactions, and (iii) small world trait of the network. To tackle such issues, we propose evolutionary graph convolutional recurrent neural networks (E-GCRNNs) to simultaneously model both temporal and spatial dependencies of knowledge services from their evolving networks. Additionally, a localized mini-batch training scheme is developed, which allows the E-GCRNNs to work on large-scale knowledge services network and reduce the prediction bias caused by the small world trait. Extensive experiments on real-world datasets have demonstrated that the proposed E-GCRNNs outperform baselines in terms of prediction accuracy, especially with the prediction range being longer, while remaining computationally efficient. Haozhe Lin, Yushun Fan, Jia Zhang 0001, Zhenghua Xu 0001, Thomas Lukasiewicz |
IEEE Trans. Serv. Comput. | 5 |
| 2022 | Reverse Differentiation via Predictive CodingabstractDeep learning has redefined AI thanks to the rise of artificial neural networks, which are inspired by neuronal networks in the brain. Through the years, these interactions between AI and neuroscience have brought immense benefits to both fields, allowing neural networks to be used in a plethora of applications. Neural networks use an efficient implementation of reverse differentiation, called backpropagation (BP). This algorithm, however, is often criticized for its biological implausibility (e.g., lack of local update rules for the parameters). Therefore, biologically plausible learning methods that rely on predictive coding (PC), a framework for describing information processing in the brain, are increasingly studied. Recent works prove that these methods can approximate BP up to a certain margin on multilayer perceptrons (MLPs), and asymptotically on any other complex model, and that zerodivergence inference learning (Z-IL), a variant of PC, is able to exactly implement BP on MLPs. However, the recent literature shows also that there is no biologically plausible method yet that can exactly replicate the weight update of BP on complex models. To fill this gap, in this paper, we generalize (PC and) Z-IL by directly defining it on computational graphs, and show that it can perform exact reverse differentiation. What results is the first PC (and so biologically plausible) algorithm that is equivalent to BP in the way of updating parameters on any neural network, providing a bridge between the interdisciplinary research of neuroscience and deep learning. Furthermore, the above results in particular also immediately provide a novel local and parallel implementation of BP. Tommaso Salvatori, Yuhang Song 0001, Zhenghua Xu 0001, Thomas Lukasiewicz, Rafal Bogacz |
AAAI | 3 |
| 2022 | ω-net: Dual supervised medical image segmentation with multi-dimensional self-attention and diversely-connected multi-scale convolution
Zhenghua Xu 0001, Junyang Chen 0001, Thomas Lukasiewicz, Zhigang Fu |
Neurocomputing | 1 |
| 2022 | Adversarial Caching Training: Unsupervised Inductive Network Representation Learning on Large-Scale GraphsabstractNetwork representation learning (NRL) has far-reaching effects on data mining research, showing its importance in many real-world applications. NRL, also known as network embedding, aims at preserving graph structures in a low-dimensional space. These learned representations can be used for subsequent machine learning tasks, such as vertex classification, link prediction, and data visualization. Recently, graph convolutional network (GCN)-based models, e.g., GraphSAGE, have drawn a lot of attention for their success in inductive NRL. When conducting unsupervised learning on large-scale graphs, some of these models employ negative sampling (NS) for optimization, which encourages a target vertex to be close to its neighbors while being far from its negative samples. However, NS draws negative vertices through a random pattern or based on the degrees of vertices. Thus, the generated samples could be either highly relevant or completely unrelated to the target vertex. Moreover, as the training goes, the gradient of NS objective calculated with the inner product of the unrelated negative samples and the target vertex may become zero, which will lead to learning inferior representations. To address these problems, we propose an adversarial training method tailored for unsupervised inductive NRL on large networks. For efficiently keeping track of high-quality negative samples, we design a caching scheme with sampling and updating strategies that has a wide exploration of vertex proximity while considering training costs. Besides, the proposed method is adaptive to various existing GCN-based models without significantly complicating their optimization process. Extensive experiments show that our proposed method can achieve better performance compared with the state-of-the-art models. Junyang Chen 0001, Zhiguo Gong, Wei Wang 0077, Cong Wang 0018, Zhenghua Xu 0001, Jianming Lv, Xueliang Li 0002, Kaishun Wu, Weiwen Liu |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2021 | RSG: A Simple but Effective Module for Learning Imbalanced DatasetsabstractImbalanced datasets widely exist in practice and are a great challenge for training deep neural models with a good generalization on infrequent classes. In this work, we propose a new rare-class sample generator (RSG) to solve this problem. RSG aims to generate some new samples for rare classes during training, and it has in particular the following advantages: (1) it is convenient to use and highly versatile, because it can be easily integrated into any kind of convolutional neural network, and it works well when combined with different loss functions, and (2) it is only used during the training phase, and therefore, no additional burden is imposed on deep neural networks during the testing phase. In extensive experimental evaluations, we verify the effectiveness of RSG. Furthermore, by leveraging RSG, we obtain competitive results on Imbalanced CIFAR and new state-of-the-art results on Places-LT, ImageNet-LT, and iNaturalist 2018. The source code is available at https://github.com/Jianf-Wang/RSG. Thomas Lukasiewicz, Xiaolin Hu 0001, Jianfei Cai 0001, Zhenghua Xu 0001 |
CVPR | 5 |
| 2021 | Associative Memories via Predictive CodingabstractAssociative memories in the brain receive and store patterns of activity registered by the sensory neurons, and are able to retrieve them when necessary. Due to their importance in human intelligence, computational models of associative memories have been developed for several decades now. In this paper, we present a novel neural model for realizing associative memories, which is based on a hierarchical generative network that receives external stimuli via sensory neurons. It is trained using predictive coding, an error-based learning algorithm inspired by information processing in the cortex. To test the model's capabilities, we perform multiple retrieval experiments from both corrupted and incomplete data points. In an extensive comparison, we show that this new model outperforms in retrieval accuracy and robustness popular associative memory models, such as autoencoders trained via backpropagation, and modern Hopfield networks. In particular, in completing partial data points, our model achieves remarkable results on natural image datasets, such as ImageNet, with a surprisingly high accuracy, even when only a tiny fraction of pixels of the original images is presented. Our model provides a plausible framework to study learning and retrieval of memories in the brain, as it closely mimics the behavior of the hippocampus as a memory index and generative model. Tommaso Salvatori, Yuhang Song 0001, Yujian Hong, Lei Sha, Simon Frieder, Zhenghua Xu 0001, Rafal Bogacz, Thomas Lukasiewicz |
NeurIPS | 6 |
| 2020 | Arena: A General Evaluation Platform and Building Toolkit for Multi-Agent IntelligenceabstractLearning agents that are not only capable of taking tests, but also innovating is becoming a hot topic in AI. One of the most promising paths towards this vision is multi-agent learning, where agents act as the environment for each other, and improving each agent means proposing new problems for others. However, existing evaluation platforms are either not compatible with multi-agent settings, or limited to a specific game. That is, there is not yet a general evaluation platform for research on multi-agent intelligence. To this end, we introduce Arena, a general evaluation platform for multi-agent intelligence with 35 games of diverse logics and representations. Furthermore, multi-agent intelligence is still at the stage where many problems remain unexplored. Therefore, we provide a building toolkit for researchers to easily invent and build novel multi-agent problems from the provided game set based on a GUI-configurable social tree and five basic multi-agent reward schemes. Finally, we provide Python implementations of five state-of-the-art deep multi-agent reinforcement learning baselines. Along with the baseline implementations, we release a set of 100 best agents/teams that we can train with different training schemes for each game, as the base for evaluating agents with population performance. As such, the research community can perform comparisons under a stable and uniform standard. All the implementations and accompanied tutorials have been open-sourced for the community at https://sites.google.com/view/arena-unity/. Yuhang Song 0001, Andrzej Wojcicki, Thomas Lukasiewicz, Jianyi Wang, Abi Aryan, Zhenghua Xu 0001, Mai Xu, Lianlong Wu |
AAAI | 6 |
| 2020 | Mega-Reward: Achieving Human-Level Play without Extrinsic RewardsabstractIntrinsic rewards were introduced to simulate how human intelligence works; they are usually evaluated by intrinsically-motivated play, i.e., playing games without extrinsic rewards but evaluated with extrinsic rewards. However, none of the existing intrinsic reward approaches can achieve human-level performance under this very challenging setting of intrinsically-motivated play. In this work, we propose a novel megalomania-driven intrinsic reward (called mega-reward), which, to our knowledge, is the first approach that achieves human-level performance in intrinsically-motivated play. Intuitively, mega-reward comes from the observation that infants' intelligence develops when they try to gain more control on entities in an environment; therefore, mega-reward aims to maximize the control capabilities of agents on given entities in a given environment. To formalize mega-reward, a relational transition model is proposed to bridge the gaps between direct and latent control. Experimental studies show that mega-reward (i) can greatly outperform all state-of-the-art intrinsic reward approaches, (ii) generally achieves the same level of performance as Ex-PPO and professional human-level scores, and (iii) has also a superior performance when it is incorporated with extrinsic rewards. Yuhang Song 0001, Jianyi Wang, Thomas Lukasiewicz, Zhenghua Xu 0001, Shangtong Zhang, Andrzej Wojcicki, Mai Xu |
AAAI | 4 |
| 2020 | Hybrid Deep-Semantic Matrix Factorization for Tag-Aware Personalized RecommendationabstractMatrix factorization has now become a dominant solution for personalized recommendation on the Social Web. To alleviate the cold start problem, previous approaches have incorporated various additional sources of information into traditional matrix factorization models. These upgraded models, however, achieve only "marginal" enhancements on the performance of personalized recommendation. Therefore, inspired by the recent development of deep-semantic modeling, we propose a hybrid deep-semantic matrix factorization (HDMF) model to further improve the performance of tag-aware personalized recommendation by integrating the techniques of deep-semantic modeling, hybrid learning, and matrix factorization. Experimental results show that HDMF significantly outperforms the state-of-the-art baselines in tag-aware personalized recommendation, in terms of all evaluation metrics. Zhenghua Xu 0001, Thomas Lukasiewicz, Cheng Chen 0002, Yishu Miao, Guizhi Xu |
ICASSP | 1 |
| 2020 | Can the Brain Do Backpropagation? - Exact Implementation of Backpropagation in Predictive Coding NetworksabstractBackpropagation (BP) has been the most successful algorithm used to train artificial neural networks. However, there are several gaps between BP and learning in biologically plausible neuronal networks of the brain (learning in the brain, or simply BL, for short), in particular, (1) it has been unclear to date, if BP can be implemented exactly via BL, (2) there is a lack of local plasticity in BP, i.e., weight updates require information that is not locally available, while BL utilizes only locally available information, and (3)~there is a lack of autonomy in BP, i.e., some external control over the neural network is required (e.g., switching between prediction and learning stages requires changes to dynamics and synaptic plasticity rules), while BL works fully autonomously. Bridging such gaps, i.e., understanding how BP can be approximated by BL, has been of major interest in both neuroscience and machine learning. Despite tremendous efforts, however, no previous model has bridged the gaps at a degree of demonstrating an equivalence to BP, instead, only approximations to BP have been shown. Here, we present for the first time a framework within BL that bridges the above crucial gaps. We propose a BL model that (1) produces \emph{exactly the same} updates of the neural weights as~BP, while (2)~employing local plasticity, i.e., all neurons perform only local computations, done simultaneously. We then modify it to an alternative BL model that (3) also works fully autonomously. Overall, our work provides important evidence for the debate on the long-disputed question whether the brain can perform~BP. Yuhang Song 0001, Thomas Lukasiewicz, Zhenghua Xu 0001, Rafal Bogacz |
NeurIPS | 3 |
| 2019 | Diversity-Driven Extensible Hierarchical Reinforcement LearningabstractHierarchical reinforcement learning (HRL) has recently shown promising advances on speeding up learning, improving the exploration, and discovering intertask transferable skills. Most recent works focus on HRL with two levels, i.e., a master policy manipulates subpolicies, which in turn manipulate primitive actions. However, HRL with multiple levels is usually needed in many real-world scenarios, whose ultimate goals are highly abstract, while their actions are very primitive. Therefore, in this paper, we propose a diversitydriven extensible HRL (DEHRL), where an extensible and scalable framework is built and learned levelwise to realize HRL with multiple levels. DEHRL follows a popular assumption: diverse subpolicies are useful, i.e., subpolicies are believed to be more useful if they are more diverse. However, existing implementations of this diversity assumption usually have their own drawbacks, which makes them inapplicable to HRL with multiple levels. Consequently, we further propose a novel diversity-driven solution to achieve this assumption in DEHRL. Experimental studies evaluate DEHRL with nine baselines from four perspectives in two domains; the results show that DEHRL outperforms the state-of-the-art baselines in all four aspects. Yuhang Song 0001, Jianyi Wang, Thomas Lukasiewicz, Zhenghua Xu 0001, Mai Xu |
AAAI | 4 |
| 2019 | Tumor Segmentation Based on Deeply Supervised Multi-Scale U-NetabstractAlthough deep learning has achieved great success in the field of medical image processing, the existing deep learning based medical image segmentation solutions still cannot obtain satisfactory performances for abdominal small organs and lesions due to their small object size and shape-variability. In this work, a Deeply Supervised Multi-Scale U-Net (DSMS U-Net) is proposed for more accurate segmentation performances on abdominal small organs images. DSMS U-Net integrate the existing U-Net model with a restoration decoder module and some multi-scale convolution modules. Our experiment results demonstrate that the proposed DSMS U-Net approach has much better segmentation performances than the state-of-the-art baselines. Bo Wang 0072, Zhenghua Xu 0001 |
BIBM | 3 |
| 2019 | Semi-Supervised Attention-Guided CycleGAN for Data Augmentation on Medical ImagesabstractRecently, deep learning methods, in particular, convolutional neural networks (CNNs), have made a massive breakthrough in computer vision. And a big amount of annotated data is the essential cornerstone to reach this success. However, in the medical domain, it is usually difficult (and sometimes even impossible) to get sufficient data for some specific learning tasks. Consequently, in this work, a novel data augmentation solution, called semi-supervised attention-guided CycleGAN (SSA-CycleGAN) is proposed to resolve this problem. Specifically, a cycle-consistency GANs-based model is first proposed to generate synthetic tumor (resp., normal) images from normal (resp., tumor) images. Then, a semi-supervised attention module is further proposed to enhance the model's capability in learning the important details of the training images, which in turns help the generated synthetic images become more realistic. To verify its effectiveness, experimental studies are conducted on three medical image datasets with limited amounts of MRI images, and the proposed SSA-CycleGAN is applied to generate synthetic tumor and normal MRI images for data augmentation. Experimental results show that i) SSA-CycleGAN can add (resp., remove) tumor lesions on (resp., from) the original normal (resp., tumor) images and generate very realistic synthetic tumor (resp. normal) images; and ii) in the ResNet18-based MRI image classification tasks based on these datasets, data augmentation using SSA-CycleGAN achieves much better classification performances than the classic data augmentation methods. Zhenghua Xu 0001, Chang Qi, Guizhi Xu |
BIBM | 1 |
| 2019 | Long Text Analysis Using Sliced Recurrent Neural Networks with Breaking Point Information EnrichmentabstractSliced recurrent neural networks (SRNNs) are the state-of-the-art efficient solution for long text analysis tasks; however, their slicing operations inevitably result in long-term dependency loss in lower-level networks and thus limit their accuracy. Therefore, we propose a breaking point information enrichment mechanism to strengthen dependencies between sliced subsequences without hindering parallelization. Then, the resulting BPIE-SRNN model is further extended to a bidirectional model, BPIE-BiSRNN, to utilize the dependency information in not only the previous but also the following contexts. Experiments on four large public real-world datasets demonstrate that the BPIE-SRNN and BPIE-BiSRNN models always achieve a much better accuracy than SRNNs and BiSRNNs, while maintaining a superior training efficiency. Bo Li 0099, Zehua Cheng, Zhenghua Xu 0001, Wei Ye 0004, Thomas Lukasiewicz, Shikun Zhang |
ICASSP | 3 |
| 2017 | Location-Aware News Recommendation Using Deep Localized Semantic Analysis
Cheng Chen 0002, Thomas Lukasiewicz, Xiangwu Meng, Zhenghua Xu 0001 |
DASFAA (1) | 4 |
| 2017 | Tag-Aware Personalized Recommendation Using a Hybrid Deep ModelabstractRecently, many efforts have been put into tag-aware personalized recommendation. However, due to uncontrolled vocabularies, social tags are usually redundant, sparse, and ambiguous. In this paper, we propose a deep neural network approach to solve this problem by mapping the tag-based user and item profiles to an abstract deep feature space, where the deep-semantic similarities between users and their target items (resp., irrelevant items) are maximized (resp., minimized). To ensure the scalability in practice, we further propose to improve this model's training efficiency by using hybrid deep learning and negative sampling. Experimental results show that our approach can significantly outperform the state-of-the-art baselines in tag-aware personalized recommendation (3.8 times better than the best baseline), and that using hybrid deep learning and negative sampling can dramatically enhance the model's training efficiency (hundreds of times quicker), while maintaining similar (and sometimes even better) training quality and recommendation performance. Zhenghua Xu 0001, Thomas Lukasiewicz, Cheng Chen 0002, Yishu Miao, Xiangwu Meng |
IJCAI | 1 |
| 2016 | Tag-Aware Personalized Recommendation Using a Deep-Semantic Similarity Model with Negative SamplingabstractWith the rapid growth of social tagging systems, many efforts have been put on tag-aware personalized recommendation. However, due to uncontrolled vocabularies, social tags are usually redundant, sparse, and ambiguous. In this paper, we propose a deep neural network approach to solve this problem by mapping both the tag-based user and item profiles to an abstract deep feature space, where the deep-semantic similarities between users and their target items (resp., irrelevant items) are maximized (resp., minimized). Due to huge numbers of online items, the training of this model is usually computationally expensive in the real-world context. Therefore, we introduce negative sampling, which significantly increases the model's training efficiency (109.6 times quicker) and ensures the scalability in practice. Experimental results show that our model can significantly outperform the state-of-the-art baselines in tag-aware personalized recommendation: e.g., its mean reciprocal rank is between 5.7 and 16.5 times better than the baselines. Zhenghua Xu 0001, Cheng Chen 0002, Thomas Lukasiewicz, Yishu Miao, Xiangwu Meng |
CIKM | 1 |