VLDB 2026 Research / reviewers in the wild / expert
Xiaolu Chen
dblp:30/3918
· DBLP profile ↗
17ranked-venue papers
5as first author
15since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Systems, architecture and hardware · 4 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | WinMamba: Multi-Scale Shifted Windows in State Space Model for 3D Object Detectionabstract3D object detection is critical for autonomous driving, yet it remains fundamentally challenging to simultaneously maximize computational efficiency and capture long-range spatial dependencies.We observed that Mamba-based models, with their linear state-space design, capture long-range dependencies at lower cost, offering a promising balance between efficiency and accuracy.However, existing methods rely on axis-aligned scanning within a fixed window, inevitably discarding spatial information. To address this problem, we propose WinMamba, a novel Mamba-based 3D feature-encoding backbone composed of stacked WinMamba blocks. To enhance the backbone with robust multi-scale representation, the WinMamba block incorporates a window-scale-adaptive module that compensates voxel features across varying resolutions during sampling. Meanwhile, to obtain rich contextual cues within the linear state space, we equip the WinMamba layer with a learnable positional encoding and a window-shift strategy.Extensive experiments on the KITTI and Waymo datasets demonstrate that WinMamba significantly outperforms the baseline. Ablation studies further validate the individual contributions of the WSF and AWF modules in improving detection accuracy. The code will be made publicly available. Longhui Zheng, Qiming Xia, Xiaolu Chen, Zhaoliang Liu, Chenglu Wen |
AAAI | 3 |
| 2026 | Exploiting point-language models with dual-prompts for 3D anomaly detection
Haote Xu, Xiaolu Chen, Haodi Xu, Yue Huang 0001, Xinghao Ding, Xiaotong Tu |
Expert Syst. Appl. | 3 |
| 2025 | Feature Disentangling Dual-stream Network for User Bias Alleviation in Social Media PredictionabstractSocial media popularity prediction is increasingly crucial for optimizing user engagement and guiding content recommendation systems. However, existing methods suffer from an excessive reliance on user information, which disproportionately influences predictions and leads to the neglect of content diversity. This oversight results in user bias, which adversely impacts the accuracy of predictions. In this paper, an approach named Feature Disentangling Dual-Stream Network (FDDN) is introduced to address this gap. In FDDN, we introduce the Multimodal Extraction Module to extract content features from different modalities. Additionally, the User Popularity Extraction Module helps to analyze the impact of user features on popularity. The Disentangled Adaptation Module distinguishes the impact of user features from content features, thus alleviating user bias and ensuring a more comprehensive prediction. Extensive experiments on a large public dataset demonstrate the robustness and effectiveness of our approach, indicating superior performance compared to existing methods. Weilong Chen, Weimin Yuan, Xiaolu Chen, Yanru Zhang, Zhu Han 0001 |
ICASSP | 4 |
| 2025 | LoadGuard: An Adaptive Deep Learning Model for Smart Meter Electricity Theft DetectionabstractModern electricity theft poses severe risks to power grid stability, particularly as cyber-attacks targeting Advanced Metering Infrastructure (AMI) become increasingly covert. This paper proposes a deep learning framework for Electricity Theft Detection (ETD) based on a Transformer encoder and a Dynamic-Weight Multi-Head Classifier (DW-MHC). The model extracts temporal load features via self-attention and employs specialized heads to capture diverse theft patterns, such as abrupt anomalies and gradual deviations. A soft-attention fusion mechanism adaptively integrates the head outputs for robust prediction. The framework also accommodates variability across residential and industrial users, whose load profiles may differ significantly in scale and regularity. Experimental results demonstrate the model's superior performance in detecting varied theft behaviors among consumers, achieving improved performance over traditional single-head and static classifiers. Xiaolu Chen, Yanru Zhang, Hao Wang 0016 |
INDIN | 1 |
| 2025 | Tri-Modal Transformers With Mixture-of-Modality-Experts for Social Media PredictionabstractWith billions of users worldwide, accurately predicting social media popularity is crucial for assessing user behavior, forecasting trends, and enhancing social interactions and business strategies. However, this task presents significant challenges. Firstly, the extraction of valuable insights is complicated by the presence of tri-modal data (visual, text, structured) and pervasive noise. Secondly, the applicability of knowledge acquired during the pre-training phase is often limited due to discrepancies with downstream prediction tasks during the fine-tuning phase. Existing methods for Social Media Popularity Prediction (SMPP), including traditional models and Visual-and-language Models (VLMs), struggle to overcome these challenges, thereby failing to achieve satisfactory accuracy. To tackle these challenges, we propose a novel approach named Tri-Modal Transformers with Mixture-of-Modality-Experts (TTME) for SMPP. TTME integrates Artificial Intelligence Generated Content to mitigate data noise and incorporate a mix of Modality Experts in pre-training phases to effectively utilize tri-modal data. Moreover, to address training disparity, we explore strategies for downstream task adaptation including the integration of diverse pre-training experts and the implementation of DistillSoftmax. Through empirical evaluation, we demonstrate that the TTME significantly improves the accuracy of social media popularity predictions, effectively utilizes tri-modal data with noise, and enhances transferring knowledge from pre-training to downstream tasks. Weilong Chen, Xiaolu Chen, Weimin Yuan, Yan Wang 0083, Yanru Zhang, Zhu Han 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | FedCoSR: Personalized Federated Learning With Contrastive Shareable Representations for Label Heterogeneity in Non-IID DataabstractHeterogeneity arising from label distribution skew and data scarcity can cause inaccuracy and unfairness in intelligent communication applications that heavily rely on distributed computing. To deal with it, this article proposes a novel personalized federated learning algorithm, named federated contrastive shareable representations (FedCoSRs), to facilitate knowledge sharing among clients while maintaining data privacy. Specifically, the parameters of local models' shallow layers and typical local representations are both considered as shareable information for the server and are aggregated globally. To address performance degradation caused by label distribution skew among clients, contrastive learning is adopted between local and global representations to enrich local knowledge. Additionally, to ensure fairness for clients with scarce data, FedCoSR introduces adaptive local aggregation to coordinate the global model involvement in each client. Our simulations demonstrate FedCoSR's effectiveness in mitigating label heterogeneity by achieving accuracy and fairness improvements over existing methods on datasets with varying degrees of label heterogeneity. Xiaolu Chen, Yanru Zhang, Hao Wang 0016 |
IEEE Trans. Cybern. | 2 |
| 2024 | Implicit Foreground-Guided Network for Anomaly Detection and LocalizationabstractAnomaly detection plays an essential role in large-scale industrial manufacturing. However, reconstruction-based anomaly detection methods, as one of the mainstream methods, are prone to incorrectly detecting background noise as anomalous regions. Therefore, inspired by multi-task learning, we propose an Implicit Foreground-guided Network (IFgNet), which consists of a Multi-Task Attention Shared (MTAS) sub-network and a discriminative sub-network. Specifically, the MTAS sub-network implements the foreground detection and reconstruction tasks within the shared network, while the discriminative sub-network performs the final anomaly detection. In the MTAS sub-network, multiple task-specific attention blocks are applied to learn task-specific features while allowing features to be shared between different tasks. Consequently, the features that contain both semantic and edge structure information are learned through the foreground detection task, which also facilitates the reconstruction task. Furthermore, the outputs of foreground detection can be utilized to refine the anomaly detection results. In this way, IFgNet effectively mitigates the influence of background noise and achieves competitive performance on the VisA and BTAD datasets with existing methods. Xiaolu Chen, Haote Xu, Chenghao Deng, Xiaotong Tu, Xinghao Ding, Yue Huang 0001 |
ICASSP | 1 |
| 2024 | SimCLIP: Refining Image-Text Alignment with Simple Prompts for Zero-/Few-shot Anomaly DetectionabstractRecently, large pre-trained vision-language models, such as CLIP, have demonstrated significant potential in zero-/few-shot anomaly detection tasks. However, existing methods not only rely on expert knowledge to manually craft extensive text prompts but also suffer from a misalignment of high-level language features with fine-level vision features in anomaly segmentation tasks. In this paper, we propose a method, named SimCLIP, which focuses on refining the aforementioned misalignment problem through bidirectional adaptation of both Multi-Hierarchy Vision Adapter (MHVA) and Implicit Prompt Tuning (IPT). In this way, our approach requires only a simple binary prompt to efficiently accomplish anomaly classification and segmentation tasks in zero-shot scenarios. Furthermore, we introduce its few-shot extension, SimCLIP+, integrating the relational information among vision embeddings and skillfully merging the cross-modal synergy information between vision and language to address downstream anomaly detection tasks. Extensive experiments on two challenging datasets prove the more remarkable generalization capacity of our method compared to the current SOTA approaches. Our code is available at https://github.com/CH-ORGI/SimCLIP. Chenghao Deng, Haote Xu, Xiaolu Chen, Haodi Xu, Xiaotong Tu, Xinghao Ding, Yue Huang 0001 |
ACM Multimedia | 3 |
| 2024 | AFSC: Adaptive Fourier Space Compression for Anomaly DetectionabstractThe primary challenge faced by reconstruction-based anomaly detection (AD) methods is that neural networks exhibit strong generalization, resulting in a high probability and accuracy of anomaly reconstruction. Several existing methods attempt to alleviate this problem by randomly masking partial image regions and reconstructing the image from partial inpaintings. However, local masking in spatial space is not guaranteed to remove anomalous regions during the testing phase and poses the risk of normal regions being inaccurately reconstructed. Hence, we explore an approach to compress the global information of the image while ensuring the loss of partial anomaly information renders it difficult to reconstruct. Inspired by the fact that each Fourier coefficient contains global information of the image, we propose an adaptive Fourier space compression (AFSC) method. Specifically, the Fourier coefficients of the input image are sparsely sampled by binary masks obtained from the AFSC module (AFSCm). In AFSCm, the masks are jointly optimized with the reconstruction network subject to sparsity constraint. The learned masks are forced to selectively retain part of the global information that is favourable to recovering normal images. In addition, we introduce an efficient Fourier convolution module that enables the network to accurately reconstruct normal regions under conditions of losing partial information. Experimental results on three benchmarks of industrial scenarios demonstrate our method (without external prior) achieves competitive results compared with recent methods. Haote Xu, Xiaolu Chen, Changxing Jing, Liyan Sun, Yue Huang 0001, Xinghao Ding |
IEEE Trans. Ind. Informatics | 3 |
| 2023 | Hierarchical Privacy-Preserved Knowledge GraphabstractThe knowledge graphs have found widespread use in numerous areas. However, as its applications expand, privacy concerns have been raised due to its ability to reveal links between entities in the real world. This brief work introduces an innovative approach to address the concern by considering the hierarchical structure of the knowledge graph to safeguard the privacy of information presented in knowledge graphs. Xiaolu Chen, Peng Yuan Zhou, Yong Liao 0003 |
ICDCS | 1 |
| 2023 | Cross-Domain Data Extraction and Knowledge Graph Construction for Dispute AnalysisabstractThis study aims to establish a comprehensive knowledge graph that spans domains and networks, with a specific focus on legal cases and their applications. The proposed methodology enables efficient collection and storage of large volumes of structured, semi-structured, and unstructured data related to cases from various sources including organizations, the government, and the internet. To analyze the relationships between roles in cases, a multimodal model is proposed to process and collect data for domain-specific knowledge graphs. Furthermore, to support social governance and public safety, a knowledge-driven intelligent recommendation algorithm is proposed in the form of question-answering, providing multiple strategies such as causal analysis, similar case matching and pre-disaster response. This work contributes to the field of artificial intelligence and natural language processing, with potential applications in legal and governmental domains, as well as in disaster response and prevention. Qinglang Guo, Xiaolu Chen, Peng Yuan Zhou, Yong Liao 0003 |
ICDCS | 2 |
| 2023 | Double-Fine-Tuning Multi-Objective Vision-and-Language Transformer for Social Media Popularity PredictionabstractSocial media popularity prediction aims to predict future interaction or attractiveness of new posts. However, in most existing works, there is a notable deficiency in the effective treatment of numerical features. Despite their significant potential to provide ample information, these features are often inadequately processed, leading to insufficiency of information acquirement. In this paper, we introduce a method, named Double-Fine-Tuning Multi-Objective Vision-and-Language Transformer (DFT-MOVLT). To supplement the information in vision-and-language pre-training (VLP), we propose compound text, which is concatenated by numerical data and text. Furthermore, during VLP, a transformer is trained using 3 objectives to ensure thorough feature extraction. Finally, for more generalized prediction, we fine-tune 2 models using different training ways and ensemble them. To evaluate the effectiveness of each mechanism adopted in the proposed method, we conduct an array of ablation experiments. Our team achieve the 3rd place in Social Media Prediction (SMP) Challenge 2023. Xiaolu Chen, Weilong Chen, Zhongjian Zhang, Lixin Duan, Yanru Zhang |
ACM Multimedia | 1 |
| 2023 | Synchronization of machine learning oscillators in complex networks
Tongfeng Weng, Xiaolu Chen, Zhuoming Ren, Huijie Yang, Jie Zhang 0012, Michael Small |
Inf. Sci. | 2 |
| 2022 | Title-and-Tag Contrastive Vision-and-Language Transformer for Social Media Popularity PredictionabstractSocial media is an indispensable part of modern life, and social media popularity prediction (SMPP) plays a vital role in practice. In current work, the inconsistency of words in labels and titles, user feature transformation, etc have not been well noticed. In this paper, we propose a novel approach named Title-and-Tag Contrastive Vision-and-Language Transformer (TTC-VLT), combining two pre-trained vision and language transformers and other two dense feature parts for this prediction task. On one hand, in order to learn the differences between titles and tags, we design title-tag contrastive learning for title-visual and tag-visual, which separately extracts multimodal information from two types of text. On the other hand, user identification features are transformed to embedding vectors to capture user attribute details. From the extensive experiments, our approach outperforms the other methods on the social media prediction dataset. Our team achieve the 2nd place on the leader board of the Social Media Prediction Challenge 2022. Weilong Chen, Weimin Yuan, Xiaolu Chen, Xinran Zhang 0006, Yanru Zhang |
ACM Multimedia | 4 |
| 2021 | Complex System Monitoring Based on Distributed Least Squares MethodabstractThe distributed monitoring framework is undoubtedly more suitable for large-scale complex industrial systems. However, most existing distributed monitoring methods ignored the information interaction between the local system and its neighbors. In this article, an improved distributed fault detection framework that considering the communication between subsystems is present. The system decomposition is optimized based on the monitoring performance with mechanism knowledge as constraints. The integration of mechanism and data is helpful to find the appropriate common variables between subsystems. The distributed partial least squares (DPLSs) algorithm is proposed to address the local monitoring challenges caused by the propagation of a common variable. The local monitoring model takes full advantage of the information from neighbors to reduce the uncertainty of the local system. Bayesian fusion performance metrics strategy is implemented to detect system status. The simulation results of the Tennessee Eastman process verify the effectiveness of the proposed scheme.Note to Practitioners—This article attempted to tackle an issue derived from distributed process monitoring of industrial processes. Even in an era of big industrial data, the fusion idea of process data and mechanism knowledge also provides a solution to the process decomposition monitoring strategy. It reduces the computational complexity, corrects the misdirection caused by the false information hidden in the measurements, and further increases the monitoring accuracy. Considering the information flowing and spreading along with the process equipment, common variables are used to describe the interaction between different subsystems. Then, the pretrained monitoring model and the online monitoring strategy are given to promote automatic implementation. The operability and monitoring accuracy of the proposed method is verified. It is suitable for process monitoring of large-scale complex industrial systems. Xiaolu Chen, Jing Wang 0016, Steven X. Ding |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2013 | Two-stage charge sensitive amplifier with self-biased MOS transistor as continuous reset systemabstractA two-stage charge sensitive amplifier architecture suitable for semiconductor radiation detector with large capacitance is proposed. The integration capacitor of the first stage can be made large to reduce gain sensibility to detector capacitance without any stability problem. Each stage uses a self-biased MOS transistor to discharge the integration capacitor. The self-bias circuit tracks process, temperature and supply voltage variations to make a relatively constant feedback resistor. This improves the gain linearity of the charge sensitive amplifier and the uniformity of multi-channel front-end electronics. The feasibility of the proposed circuit is verified by comparing the simulation results with the conventional one-stage charge sensitive amplifier and a two-stage structure with fixed-gate-voltage reset transistor. A prototype of 16-channel front-end circuit for electron collection designed in a 0.35μm CMOS technology has been measured. The area is 2.5×1.54mm2with 42 pads and the power dissipation is 60mW with power supplies of 2V and 5V. Yacong Zhang, Xiaolu Chen, Zhongjian Chen, Wengao Lu |
ISCAS | 2 |
| 2007 | Weighted Active Appearance Models
Shuchang Wang, Yangsheng Wang, Xiaolu Chen |
ICIC (1) | 3 |