VLDB 2026 Research / reviewers in the wild / expert
Pan Xie
dblp:78/6247
· DBLP profile ↗
17ranked-venue papers
6as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Handling Network Faults in Distributed AI Training: Failover is Now an OptionabstractDistributed AI training often suffers from network faults. Network faults, especially at the last hop between a switch and a host, result in loss of connectivity, resulting in training job stalls and eventual failure. This is typically managed through a fail-stop mechanism, followed by a restart, incurring significant inefficiencies. We present ReCCL, the first network fault-tolerant collective communication library (CCL) that allows training progress to be preserved by seamlessly failing over to alternate paths when a network fault occurs. During failover, ReCCL keeps communication states synchronized while using dynamic channel load balancing and intra-host GPU routing to improve communication performance. Our evaluations demonstrate that ReCCL can perform failover seamlessly with minimal performance losses. Additionally, our simulations also demonstrate that failover can be effectively used to achieve significant savings in GPU hours for large-scale distributed AI training workloads. Xin Zhe Khooi, Zhuo Jiang, Pan Xie, Zhigang Cui, Meng Wang 0018, Yuze Jin, Pengfei Huo, Lulu Chen, Liaoyuan Feng, Qinlong Wang, Yongcan Wang, Jinshuai Sun, Yingkai Zhao, Haiquan Chen 0002, Yi Li 0098, Jianxi Ye, Mun Choon Chan |
EuroSys | 3 |
| 2026 | Dual Nonlinear Sparse Feature Selection Method
Pan Xie, Cong Lei, Shanwen Zhang, Shichao Zhang 0001 |
PAKDD (1) | 1 |
| 2025 | ResAdapter: Domain Consistent Resolution Adapter for Diffusion ModelsabstractRecent advancement in text-to-image models and corresponding personalized technologies enables individuals to generate high-quality and imaginative images. However, they often suffer from limitations when generating images with resolutions outside of their trained domain. To overcome this limitation, we present the resolution adapter \textbf{(ResAdapter)}, a domain-consistent adapter designed for diffusion models to generate images with unrestricted resolutions and aspect ratios. Unlike other multi-resolution generation methods that process images of static resolution with complex post-process operations, ResAdapter directly generates images with the dynamical resolution. Especially, after learning a deep understanding of pure resolution priors, ResAdapter trained on the general dataset, generates resolution-free images with personalized diffusion models while preserving their original style domain. Comprehensive experiments demonstrate that ResAdapter with only 0.5M can process images with flexible resolutions for arbitrary diffusion models. More extended experiments demonstrate that ResAdapter is compatible with other modules for image generation across a broad range of resolutions, and can be integrated into other multi-resolution model for efficiently generating higher-resolution images. Jiaxiang Cheng, Pan Xie, Xin Xia 0005, Jiashi Li, Yuxi Ren, Huixia Li, Xuefeng Xiao 0001, Shilei Wen, Lean Fu |
AAAI | 2 |
| 2025 | CNRel: Candidate Prompt Enhancement and Noise Filtering Relational Triple Extraction Framework Based on Large Language ModelsabstractRelational Triple Extraction (RTE) focuses on extracting triples from sentences, a crucial task in the automatic construction of knowledge graphs. Large Language Models (LLMs) have the ability to automatically extract triples from text through appropriate instructions or fine-tuning. However, due to the bias between LLMs training data and inference data, the previous LLM-based triple extraction method ignores many potentially valuable knowledge and lacks noise filtering, which greatly limits the capability of RTE model. To address these challenges, we propose Candidate Prompt Enhancement and Noise Filtering Relational Triple Extraction Framework Based on Large Language Models (CNRel), which combines small pre-trained language model and LLMs. Specifically, we first utilize a candidate entity pair extraction and filtering block, based on a small pre-trained language model, to extract and refine all possible entity pairs in the text, ensuring the capture of as much valuable information as possible Then, a fine-tuned LLMs such as LLaMA is then used to predict the relationship between the candidate entity pairs and extract as many triples as possible. Finally, Noise Filter block filter the extracted triples through LLMs, and remove the wrong triples, which greatly improve the precision of the RTE model. Experiments on several public datasets show that CNRel achieves state-of-the-art among all previous mainstream relational triple extraction methods, and we conduct a widely ablation experiments to reveal the contribution of each component to the overall performance. Pan Xie, Chenbin Zhao, Liangxiong Li, Jingguo Ge |
SMC | 2 |
| 2024 | G2P-DDM: Generating Sign Pose Sequence from Gloss Sequence with Discrete Diffusion ModelabstractThe Sign Language Production (SLP) project aims to automatically translate spoken languages into sign sequences. Our approach focuses on the transformation of sign gloss sequences into their corresponding sign pose sequences (G2P). In this paper, we present a novel solution for this task by converting the continuous pose space generation problem into a discrete sequence generation problem. We introduce the Pose-VQVAE framework, which combines Variational Autoencoders (VAEs) with vector quantization to produce a discrete latent representation for continuous pose sequences. Additionally, we propose the G2P-DDM model, a discrete denoising diffusion architecture for length-varied discrete sequence data, to model the latent prior. To further enhance the quality of pose sequence generation in the discrete space, we present the CodeUnet model to leverage spatial-temporal information. Lastly, we develop a heuristic sequential clustering method to predict variable lengths of pose sequences for corresponding gloss sequences. Our results show that our model outperforms state-of-the-art G2P models on the public SLP evaluation benchmark. For more generated results, please visit our project page: https://slpdiffusier.github.io/g2p-ddm. Pan Xie, Qipeng Zhang, Taiying Peng, Hao Tang 0005, Zexian Li |
AAAI | 1 |
| 2024 | ByteEdit: Boost, Comply and Accelerate Generative Image Editing
Yuxi Ren, Jie Wu 0030, Yanzuo Lu, Huafeng Kuang, Xin Xia 0005, Xionghui Wang, Yixing Zhu, Pan Xie, Shiyin Wang, Xuefeng Xiao 0001, Lean Fu |
ECCV (3) | 9 |
| 2024 | Hyper-SD: Trajectory Segmented Consistency Model for Efficient Image SynthesisabstractRecently, a series of diffusion-aware distillation algorithms have emerged to alleviate the computational overhead associated with the multi-step inference process of Diffusion Models (DMs). Current distillation techniques often dichotomize into two distinct aspects: i) ODE Trajectory Preservation; and ii) ODE Trajectory Reformulation. However, these approaches suffer from severe performance degradation or domain shifts. To address these limitations, we propose Hyper-SD, a novel framework that synergistically amalgamates the advantages of ODE Trajectory Preservation and Reformulation, while maintaining near-lossless performance during step compression. Firstly, we introduce Trajectory Segmented Consistency Distillation to progressively perform consistent distillation within pre-defined time-step segments, which facilitates the preservation of the original ODE trajectory from a higher-order perspective. Secondly, we incorporate human feedback learning to boost the performance of the model in a low-step regime and mitigate the performance loss incurred by the distillation process. Thirdly, we integrate score distillation to further improve the low-step generation capability of the model and offer the first attempt to leverage a unified LoRA to support the inference process at all steps. Extensive experiments and user studies demonstrate that Hyper-SD achieves SOTA performance from 1 to 8 inference steps for both SDXL and SD1.5. For example, Hyper-SDXL surpasses SDXL-Lightning by +0.68 in CLIP Score and +0.51 in Aes Score in the 1-step inference. Yuxi Ren, Xin Xia 0005, Yanzuo Lu, Jie Wu 0032, Pan Xie, Xuefeng Xiao 0001 |
NeurIPS | 6 |
| 2024 | UniFL: Improve Latent Diffusion Model via Unified Feedback LearningabstractLatent diffusion models (LDM) have revolutionized text-to-image generation, leading to the proliferation of various advanced models and diverse downstream applications. However, despite these significant advancements, current diffusion models still suffer from several limitations, including inferior visual quality, inadequate aesthetic appeal, and inefficient inference, without a comprehensive solution in sight. To address these challenges, we present **UniFL**, a unified framework that leverages feedback learning to enhance diffusion models comprehensively. UniFL stands out as a universal, effective, and generalizable solution applicable to various diffusion models, such as SD1.5 and SDXL.
Notably, UniFL consists of three key components: perceptual feedback learning, which enhances visual quality; decoupled feedback learning, which improves aesthetic appeal; and adversarial feedback learning, which accelerates inference.
In-depth experiments and extensive user studies validate the superior performance of our method in enhancing generation quality and inference acceleration. For instance, UniFL surpasses ImageReward by 17\% user preference in terms of generation quality and outperforms LCM and SDXL Turbo by 57\% and 20\% general preference with 4-step inference. Jie Wu 0030, Yuxi Ren, Xin Xia 0005, Huafeng Kuang, Pan Xie, Jiashi Li, Xuefeng Xiao 0001, Shilei Wen, Lean Fu, Guanbin Li |
NeurIPS | 6 |
| 2024 | Sign Language Production with Latent Motion TransformerabstractSign Language Production (SLP) is the tough task of turning sign language into sign videos. The main goal of SLP is to create these videos using a sign gloss. In this research, we’ve developed a new method to make high-quality sign videos without using human poses as a middle step. Our model works in two main parts: first, it learns from a generator and the video’s hidden features, and next, it uses another model to understand the order of these hidden features. To make this method even better for sign videos, we make several significant improvements. (i) In the first stage, we take an improved 3D VQ-GAN to learn downsampled latent representations. (ii) In the second stage, we introduce sequence-to-sequence attention to better leverage conditional information. (iii) The separated two-stage training discards the realistic visual semantic of the latent codes in the second stage. To endow the latent sequences semantic information, we extend the token-level autoregressive latent codes learning with perceptual loss and reconstruction loss for the prior model with visual perception. Compared with previous state-of-the-art approaches, our model performs consistently better on two word-level sign language datasets, i.e., WLASL and NMFs-CSL. Pan Xie, Taiying Peng, Qipeng Zhang |
WACV | 1 |
| 2024 | Modeling Balanced Explicit and Implicit Relations with Contrastive Learning for Knowledge Concept Recommendation in MOOCsabstractThe knowledge concept recommendation in Massive Open Online Courses (MOOCs) is a significant issue that has garnered widespread attention. Existing methods primarily rely on the explicit relations between users and knowledge concepts on the MOOC platforms for recommendation. However, there are numerous implicit relations (e.g., shared interests or same knowledge levels between users) generated within the users' learning activities on the MOOC platforms. Existing methods fail to consider these implicit relations, and these relations themselves are difficult to learn and represent, causing poor performance in knowledge concept recommendation and an inability to meet users' personalized needs. To address this issue, we propose a novel framework based on contrastive learning, which can represent and balance the explicit and implicit relations for knowledge concept recommendation in MOOCs (CL-KCRec). Specifically, we first construct a MOOCs heterogeneous information network (HIN) by modeling the data from the MOOC platforms. Then, we utilize a relation-updated graph convolutional network and stacked multi-channel graph neural network to represent the explicit and implicit relations in the HIN, respectively. Considering that the quantity of explicit relations is relatively fewer compared to implicit relations in MOOCs, we propose a contrastive learning with prototypical graph to enhance the representations of both relations to capture their fruitful inherent relational knowledge, which can guide the propagation of students' preferences within the HIN. Based on these enhanced representations, to ensure the balanced contribution of both towards the final recommendation, we propose a dual-head attention mechanism for balanced fusion. Experimental results demonstrate that CL-KCRec outperforms several state-of-the-art baselines on real-world datasets in terms of HR, NDCG and MRR. Hengnian Gu, Zhiyi Duan, Pan Xie, Dongdai Zhou |
WWW | 3 |
| 2023 | Assessing the Effectiveness of Deception-Based Cyber Defense with CyberBattleSim
Quan Hong, Xizhong Guo, Pan Xie, Lidong Zhai |
ICDF2C (2) | 4 |
| 2023 | Multi-scale local-temporal similarity fusion for continuous sign language recognition
Pan Xie, Zhi Cui, Mengyi Zhao, Jianwei Cui 0002, Bin Wang 0004 |
Pattern Recognit. | 1 |
| 2023 | Bidirectional Transformer GAN for Long-term Human Motion PredictionabstractThe mainstream motion prediction methods usually focus on short-term prediction, and their predicted long-term motions often fall into an average pose, i.e., the freezing forecasting problem [ 27 ]. To mitigate this problem, we propose a novel Bidirectional Transformer-based Generative Adversarial Network (BiTGAN) for long-term human motion prediction. The bidirectional setup leads to consistent and smooth generation in both forward and backward directions. Besides, to make full use of the history motions, we split them into two parts. The first part is fed to the Transformer encoder in our BiTGAN while the second part is used as the decoder input. This strategy can alleviate the exposure problem [ 37 ]. Additionally, to better maintain both the local (i.e., frame-level pose) and global (i.e., video-level semantic) similarities between the predicted motion sequence and the real one, the soft dynamic time warping (Soft-DTW) loss is introduced into the generator. Finally, we utilize a dual-discriminator to distinguish the predicted sequence at both frame and sequence levels. Extensive experiments on the public Human3.6M dataset demonstrate that our proposed BiTGAN achieves state-of-the-art performance on long-term (4 s ) human motion prediction, and reduces the average error of all actions by 4%. Mengyi Zhao, Hao Tang 0005, Pan Xie, Shuling Dai, Nicu Sebe, Wei Wang 0108 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2022 | Full transformer network with masking future for word-level sign language recognition
Pan Xie, Mingye Wang |
Neurocomputing | 2 |
| 2022 | PB-GCN: Progressive binary graph convolutional networks for skeleton-based action recognition
Mengyi Zhao, Shuling Dai, Yanjun Zhu, Hao Tang 0005, Pan Xie, Chunlei Liu 0001, Baochang Zhang 0001 |
Neurocomputing | 5 |
| 2022 | PiSLTRc: Position-Informed Sign Language Transformer With Content-Aware ConvolutionabstractSince the superiority of Transformer in learning long-term dependency, the sign language Transformer model achieves remarkable progress in Sign Language Recognition (SLR) and Translation (SLT). However, there are several issues with the Transformer that prevent it from better sign language understanding. The first issue is that the self-attention mechanism learns sign video representation in a frame-wise manner, neglecting the temporal semantic structure of sign gestures. Secondly, the attention mechanism with absolute position encoding is direction and distance unaware, thus limiting its ability. To address these issues, we propose a new model architecture, namely PiSLTRc, with two distinctive characteristics: (i) content-aware and position-aware convolution layers. Specifically, we explicitly select relevant features using a novel content-aware neighborhood gathering method. Then we aggregate these features with position-informed temporal convolution layers, thus generating robust neighborhood-enhanced sign representation. (ii) injecting the relative position information to the attention mechanism in the encoder, decoder, and even encoder-decoder cross attention. Compared with the vanilla Transformer model, our model performs consistently better on three large-scale sign language benchmarks: PHOENIX-2014, PHOENIX-2014-T and CSL. Furthermore, extensive experiments demonstrate that the proposed method achieves state-of-the-art performance on translation quality with$+1.6$BLEU improvements. Pan Xie, Mengyi Zhao |
IEEE Trans. Multim. | 1 |
| 2020 | Infusing Sequential Information into Conditional Masked Translation Model with Self-Review MechanismabstractNon-autoregressive models generate target words in a parallel way, which achieve a faster decoding speed but at the sacrifice of translation accuracy.To remedy a flawed translation by non-autoregressive models, a promising approach is to train a conditional masked translation model (CMTM), and refine the generated results within several iterations.Unfortunately, such approach hardly considers the sequential dependency among target words, which inevitably results in a translation degradation.Hence, instead of solely training a Transformer-based CMTM, we propose a Self-Review Mechanism to infuse sequential information into it.Concretely, we insert a left-to-right mask to the same decoder of CMTM, and then induce it to autoregressively review whether each generated word from CMTM is supposed to be replaced or kept.The experimental results (WMT14 En↔De and WMT16 En↔Ro) demonstrate that our model uses dramatically less training computations than the typical CMTM, as well as outperforms several state-of-the-art non-autoregressive models by over 1 BLEU.Through knowledge distillation, our model even surpasses a typical left-to-right Transformer model, while significantly speeding up decoding. Pan Xie, Zhi Cui, Xiuying Chen, Jianwei Cui 0002, Bin Wang 0004 |
COLING | 1 |