EDBT 2026 Demo / reviewers in the wild / expert
Wei Liu 0183
dblp:49/3283-183
· DBLP profile ↗
19ranked-venue papers
2as first author
19since 2021 · last 2026
0000-0002-9897-1213ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 1 first-author · 15 since 2021Artificial intelligence and machine learning · 7 · 7 since 2021Computer networks · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MambaOVSR: Multiscale Fusion with Global Motion Modeling for Chinese Opera Video Super-ResolutionabstractChinese opera is celebrated for preserving classical art. However, early filming equipment limitations have degraded videos of last-century performances by renowned artists (e.g., low frame rates and resolution), hindering archival efforts. Although space-time video super-resolution (STVSR) has advanced significantly, applying it directly to opera videos remains challenging. The scarcity of datasets impedes the recovery of high-frequency details, and existing STVSR methods lack global modeling capabilities—compromising visual quality when handling opera’s characteristic large motions. To address these challenges, we pioneer a large-scale Chinese Opera Video Clip (COVC) dataset and propose the Mamba-based multiscale fusion network for space-time Opera Video Super-Resolution (MambaOVSR). Specifically, MambaOVSR involves three novel components: the Global Fusion Module (GFM) for motion modeling through a multiscale alternating scanning mechanism, and the Multiscale Synergistic Mamba Module (MSMM) for alignment across different sequence lengths. Additionally, our MambaVR block resolves feature artifacts and positional information loss during alignment. Experimental results on the COVC dataset show that MambaOVSR significantly outperforms the SOTA STVSR method by an average of 1.86 dB in terms of PSNR. Hua Chang, Xin Xu 0007, Wei Liu 0183, Wei Wang 0170, Xin Yuan 0009, Kui Jiang |
AAAI | 3 |
| 2026 | ICLR: Inter-Chrominance and Luminance Interaction for Natural Color Restoration in Low-Light Image EnhancementabstractLow-Light Image Enhancement (LLIE) task aims at improving contrast while restoring details and textures for images captured in low-light conditions. HVI color space has made significant progress in this task by enabling precise decoupling of chrominance and luminance. However, for the interaction of chrominance and luminance branches, substantial distributional differences between the two branches prevalent in natural images limit complementary feature extraction, and luminance errors are propagated to chrominance channels through the nonlinear parameter. Furthermore, for interaction between different chrominance branches, images with large homogeneous-color regions usually exhibit weak correlation between chrominance branches due to concentrated distributions. Traditional pixel-wise losses exploit strong inter-branch correlations for co-optimization, causing gradient conflicts in weakly correlated regions. Therefore, we propose an Inter-Chrominance and Luminance Interaction (ICLR) framework including a Dual-stream Interaction Enhancement Module (DIEM) and a Covariance Correction Loss (CCL). The DIEM improves the extraction of complementary information from two dimensions, fusion and enhancement, respectively. The CCL utilizes luminance residual statistics to penalize chrominance errors and balances gradient conflicts by constraining chrominance branches covariance. Experimental results on multiple datasets show that the proposed ICLR framework outperforms state-of-the-art methods. Xin Xu 0007, Wei Liu 0183, Wei Wang 0170, Kui Jiang |
AAAI | 3 |
| 2026 | FedARKS: Federated Aggregation via Robust and Discriminative Knowledge Selection and Integration for Person Re-identification
Xin Xu 0007, Binchang Ma, Zhixi Yu, Wei Liu 0183 |
AAAI | 4 |
| 2026 | Domain-Aware Suppression and Aggregation for Federated DG ReIDabstractFederated domain generalization in person re-identification (FedDG-ReID) aims to learn a privacy-preserving server model from decentralized client source domains that generalizes to unseen domains. Existing approaches enhance the generalizability of the server model by increasing the diversity of client person data. However, these methods overlook that ReID model parameters are easily biased by client-specific data distributions, leading to the capture of excessive domain-specific identity information. Such identity information (e.g., clothing style) struggles with identity information in unseen domains, thereby hindering the generalization ability of the server model. To address this, we propose a novel FedDG-ReID framework, which mainly consists of Domain-aware Parameter Suppression (DPS) and Domain-invariant Weighted Aggregation (DWA), called FedSupWA. Specifically, DPS adaptively attenuates the update magnitude of the parameters based on the fit of the parameters to the client's domain, encouraging the model to focus on more generalized domain-independent identity information, such as pedestrian contours, and other consistent information across domains. DWA enhances the server model’s generalization by evaluating the effectiveness of the client model in maintaining the consistency of pedestrian identities to measure the importance of the learned domain-independent identity information and assigning greater aggregation weights to clients that contribute more generalized information. Extensive experiments demonstrate the effectiveness of FedSupWA, showing that it achieves state-of-the-art performance. Zhixi Yu, Wei Liu 0183, Wenke Huang 0003, Bin Yang 0026, Qian Bie, Guancheng Wan, Xin Xu 0007 |
AAAI | 2 |
| 2026 | DSFMamba: Dual-state fusion mamba for long-short term motion-adaptive streaming perception in autonomous driving
Qinghua Qi, Xiaowei Xu 0003, Mingxing Deng, Wei Liu 0183, Xin Xu 0007 |
Knowl. Based Syst. | 4 |
| 2026 | Discrepancy-Consistency Bi-Knowledge Fusion for unsupervised video anomaly detection
Awei Yin, Wei Liu 0183, Xiao Wang 0029, Ao Huang, Xin Xu 0007 |
Knowl. Based Syst. | 2 |
| 2025 | VAGeo: View-specific Attention for Cross-View Object Geo-LocalizationabstractCross-view object geo-localization (CVOGL) aims to locate an object of interest in a captured ground- or drone-view image within the satellite image. However, existing works treat ground-view and drone-view query images equivalently, overlooking their inherent viewpoint discrepancies and the spatial correlation between the query image and the satellite-view reference image. To this end, this paper proposes a novel View-specific Attention Geo-localization method (VAGeo) for accurate CVOGL. Specifically, VAGeo contains two key modules: view-specific positional encoding (VSPE) module and channel-spatial hybrid attention (CSHA) module. In object-level, according to the characteristics of different viewpoints of ground and drone query images, viewpoint-specific positional codings are designed to more accurately identify the click-point object of the query image in the VSPE module. In feature-level, a hybrid attention in the CSHA module is introduced by combining channel attention and spatial attention mechanisms simultaneously for learning discriminative features. Extensive experimental results demonstrate that the proposed VAGeo gains a significant performance improvement, i.e., improving [email protected]/[email protected] on the CVOGL dataset from 45.43%/42.24% to 48.21%/45.22% for ground-view, and from 61.97%/57.66% to 66.19%/61.87% for drone-view. Xin Yuan 0009, Wei Liu 0183, Xin Xu 0007 |
ICASSP | 3 |
| 2025 | Event-based Video Person Re-identification via Cross-Modality and Temporal CollaborationabstractVideo-based person re-identification (ReID) has become increasingly important due to its applications in video surveillance applications. By employing events in video-based person ReID, more motion information can be provided between continuous frames to improve recognition accuracy. Previous approaches have assisted by introducing event data into the video person ReID task, but they still cannot avoid the privacy leakage problem caused by RGB images. In order to avoid privacy attacks and to take advantage of the benefits of event data, we consider using only event data. To make full use of the information in the event stream, we propose a Cross-Modality and Temporal Collaboration (CMTC) network for event-based video person ReID. First, we design an event transform network to obtain corresponding auxiliary information from the input of raw events. Additionally, we propose a differential modality collaboration module to balance the roles of events and auxiliaries to achieve complementary effects. Furthermore, we introduce a temporal collaboration module to exploit motion information and appearance cues. Experimental results demonstrate that our method outperforms others in the task of event-based video person ReID. Renkai Li, Xin Yuan 0009, Wei Liu 0183, Xin Xu 0007 |
ICASSP | 3 |
| 2025 | Toward Comprehensive Semantic Prompt for Region Contrastive Learning Underwater Image EnhancementabstractUnderwater image enhancement (UIE) focuses on mitigating image quality degradation due to light absorption and scattering. However, most existing methods enhance images via a global and uniform manner, neglecting the inherent semantic information in different regions, which may cause the network to easily deviate from the region’s original color. Moreover, these methods typically rely on clear images to guide network convergence, a process constrained by the limited availability of real-world datasets, making it extremely challenging to train enhancement models for various degradations. To address these challenges, this paper introduces a semantic guidance and region contrastive constraints network (SRCNet). Initially, we propose a semantic-aware RWKV (Receptance Weighted Key Value) block and a semantic prompt regularization module. These components leverage intra-target semantic correlations to preserve image details and colors within a global perceptual framework, while employing focal loss to emphasize the restoration of severely degraded regions. Subsequently, we introduce a region contrastive learning method that effectively utilizes negative samples to precisely capture features sensitive to degradation factors, thereby fostering robust feature distributions. Finally, experimental results demonstrate that our method outperforms existing state-of-the-art (SOTA) approaches. Xiao Wang 0029, Yongsheng Fu, Wei Wang 0170, Wei Liu 0183 |
ICASSP | 4 |
| 2025 | RPUDet: Learning Relational Prior and Uncertainty for Robust Aerial Object DetectionabstractAerial object detection remains a challenging task in the computer vision community. While general object detectors perform well on natural images, they struggle with aerial images due to missed detections from small, low-resolution objects and misdetections caused by classification uncertainty between semantically similar objects. To address these issues, we propose a Relational Prior and Uncertainty detector (RPUDet). RPUDet consists of two core modules: 1) Relation Aware Reasoning Module (RARM), which leverages a relational prior graph to help detect correlated objects and reduce missed detections of low-resolution objects; and 2) Uncertainty Guided Awareness Module (UGAM), which computes a uncertainty map to identify low-confidence regions, dynamically adjusts feature weights, and refines areas with high semantic ambiguity to mitigate misdetections. Additionally, to advance aerial object detection in industrial applications, we introduce the Low-Voltage Distribution Insulator Dataset (LVDID), focusing on static scenes, in contrast to existing dynamic datasets. This enables us to evaluate RPUDet's performance across diverse real-world scenarios. We evaluate RPUDet on VisDrone2019, VEDAI, and LVDID, demonstrating superior performance compared to existing methods. The code and dataset will be available at https://github.com/Godk02/RPUDet. Wei Liu 0183, Minshi Chen, Xiao Wang 0029, Xin Yuan 0009 |
ICMR | 2 |
| 2025 | Positive Style Accumulation: A Style Screening and Continuous Utilization Framework for Federated DG-ReIDabstractThe Federated Domain Generalization for Person re-identification (FedDG-ReID) aims to learn a global server model that can be effectively generalized to source and target domains through distributed source domain data. Existing methods mainly improve the diversity of samples through style transformation, which to some extent enhances the generalization performance of the model. However, we discover that not all styles contribute to the generalization performance. Therefore, we define styles that are beneficial/harmful to the model's generalization performance as positive/negative styles. Based on this, new issues arise: How to effectively screen and continuously utilize the positive styles. To solve these problems, we propose a Style Screening and Continuous Utilization (SSCU) framework. Firstly, we design a Generalization Gain-guided Dynamic Style Memory (GGDSM) for each client model to screen and accumulate generated positive styles. Specifically, the memory maintains a prototype initialized from raw data for each category, then screens positive styles that enhance the global model during training, and updates these positive styles into the memory using a momentum-based approach. Meanwhile, we propose a style memory recognition loss to fully leverage the positive styles memorized by GGDSM. Furthermore, we propose a Collaborative Style Training (CST) strategy to make full use of positive styles. Unlike traditional learning strategies, our approach leverages both newly generated styles and the accumulated positive styles stored in memory to train client models on two distinct branches. This training strategy is designed to effectively promote the rapid acquisition of new styles by the client models, ensuring that they can quickly adapt to and integrate novel stylistic variations. Simultaneously, this strategy guarantees the continuous and thorough utilization of positive styles, which is highly beneficial for the model's generalization performance. Extensive experimental results demonstrate that our method outperforms existing methods in both the source domain and the target domain. Xin Xu 0007, Chaoyue Ren, Wei Liu 0183, Wenke Huang 0003, Bin Yang 0026, Zhixi Yu, Kui Jiang |
ACM Multimedia | 3 |
| 2025 | Retrieving and Reasoning: Multivariate Feature and Attribute Cooperation for Video Anomaly DetectionabstractVideo anomaly detection (VAD), which detects abnormal patterns in video sequence, is based on several kinds of features or attributes in the existing methods. This ignores the interconnections between different features and attributes, and the initiation of an anomalous result is brought about by multiple factors. If several individual neural networks are used to perceive various types of anomalies, the system would lose awareness of the association among features and attributes, which limits the system's ability to perceive complex anomalies. In this work, we propose a dual-branch framework for VAD task, which includes deep feature retrieving and semantic attribute reasoning branch. In the former branch, three high-dimensional deep features are extracted and modeled, then the anomaly scores are obtained based on the vector retrieval database. In the latter branch, three low-dimensional semantic-level attributes are extracted for composing the attribute triplets, then use theAssociation-ruleMiningModule (AMM) to perceive potential connections among these triplets. The coefficients computed by the latter branch calibrate the anomaly scores obtained by the former while providing high-level anomaly causes. Extensive experiments show that our approach achieves state-of-the-art performance with 87.9$\%$on ShanghaiTech and 94.6$\%$on Avenue. Xingshuo Han, Xiao Wang 0029, Wei Liu 0183, Liping Ye, Xin Xu 0007 |
IEEE Signal Process. Lett. | 3 |
| 2025 | Diversity-Representativeness Replay and Knowledge Alignment for Lifelong Vehicle Re-identificationabstractLifelong Vehicle Re-Identification (LVReID) aims to match a target vehicle across multiple cameras, considering non-stationary and continuous data streams, which fits the needs of the practical application better than traditional vehicle re-identification. Nonetheless, this area has received relatively little attention. Recently, methods for Lifelong Person Re-Identification (LPReID) have been emerging, with replay-based methods achieving the best results by storing a small number of instances from previous tasks for retraining, thus effectively reducing catastrophic forgetting. However, these methods cannot be directly applied to LVReID because they fail to simultaneously consider the diversity and representativeness of replayed data, resulting in biases between the subset stored in the memory buffer and the original data. They randomly sample classes, which may not adequately represent the distribution of the original data. Additionally, these methods fail to consider the rich variation in instances of the same vehicle class due to factors such as vehicle orientation and lighting conditions. Therefore, preserving more informative classes and instances for replay helps maintain information from previous tasks and may mitigate the model's forgetting of old knowledge. In view of this, we propose a novel Diversity-Representativeness Dual-Stage Sampling Replay (DDSR) strategy for LVReID that constructs an effective memory buffer through two stages, i.e. , Cluster-Centric Class Selection and Diverse Instance Mining. Specifically, we first perform class-level sampling based on density in the clustered class-centered feature space and then further mine the diverse, high-quality instances within the selected classes. In addition, we introduce Maximum Mean Discrepancy loss to align the feature distribution between replay data and the new arrivals and apply L2 regularization in the parameter space to facilitate knowledge transfer, thus enhancing the model's generalization ability to new tasks. Extensive experiments demonstrate effective improvements of our method compared to current state-of-the-art lifelong ReID methods on the VeRi-776, VehicleID, and VERI-Wild datasets. Zhijing Wan, Xiao Wang 0029, Wei Liu 0183, Wei Wang 0170, Zheng Wang 0059, Xin Xu 0007 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2025 | Mix-Modality Person Re-Identification: A New and Practical ParadigmabstractCurrent visible-infrared cross-modality person re-identification research has only focused on exploring the bi-modality mutual retrieval paradigm, and we propose a new and more practical mix-modality retrieval paradigm. Existing Visible-Infrared Person Re-Identification (VI-ReID) methods have achieved some results in the bi-modality mutual retrieval paradigm by learning the correspondence between visible and infrared modalities. However, significant performance degradation occurs due to the modality confusion problem when these methods are applied to the new mix-modality paradigm. Therefore, this article proposes a Mix-Modality Person Re-Identification (MM-ReID) task, explores the influence of modality mixing ratio on performance, and constructs mix-modality test sets for existing datasets according to the new mix-modality testing paradigm. To solve the modality confusion problem in MM-ReID, we propose a Cross-Identity Discrimination Harmonization Loss (CIDHL) adjusting the distribution of samples in the hyperspherical feature space, pulling the centers of samples with the same identity closer, and pushing away the centers of samples with different identities while aggregating samples with the same modality and the same identity. Furthermore, we propose a Modality Bridge Similarity Optimization Strategy (MBSOS) to optimize the cross-modality similarity between the query and queried samples with the help of the similar bridge sample in the gallery. Extensive experiments demonstrate that compared to the original performance of existing cross-modality methods on MM-ReID, the addition of our CIDHL and MBSOS demonstrates a general improvement. Wei Liu 0183, Xin Xu 0007, Hua Chang, Xin Yuan 0009, Zheng Wang 0007 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2024 | Mutuality Attribute Makes Better Video Anomaly DetectionabstractVideo anomaly detection (VAD) is an essential but challenging task. Existing prevalent methods focus on analyzing the reconstruction or prediction difference between normal and abnormal patterns through multiple deep features, e.g., optic flow. However, these approaches independently use deep features to characterize attributes, ignore the mutuality among multiple deep features. Therefore, the constructed representation is limited to indirectly representing the anomaly from isolated attributes, and makes the network difficult to capture the high-level causes of anomaly. In this paper, we proposed a novel Mutuality Attribute-based Representation framework (MAR-VAD) for the VAD task, which absorbs the mutuality among deep features to characterize the mutuality attribute. Specifically, the mutuality attribute encapsulates high-level semantic information, such as the specific abnormal object or action, which mutually utilizes information from multiple deep features. In this way, the system is able to directly capture the high-level causes of anomaly, thus providing a more comprehensive perspective to accurately detect anomaly events. Following a process-transparent density estimation, we produce the final anomaly scores. Experiments show that MAR-VAD achieves state-of-the-art performance on ShanghaiTech and Avenue. Xingshuo Han, Xiao Wang 0029, Kui Jiang, Wei Liu 0183, Ruimin Hu, Xuefeng Pan, Xin Xu 0007 |
ICASSP | 4 |
| 2024 | A comprehensive survey of visible infrared person re-identification from an application perspective
Hua Chang, Xin Xu 0007, Wei Liu 0183, Lingyi Lu, Weigang Li 0004 |
Multim. Tools Appl. | 3 |
| 2024 | Contour Counts: Restricting Deformation for Accurate Animation InterpolationabstractAnimated videos with low frame rates commonly degrade the visual experience with choppy motion. Recently, video frame interpolation has been growing rapidly which can increase frame rates. However, the sparse texture information and complex motion scenes of animated videos make the objects in the frames generated by existing video frame interpolation methods appear significantly deformation, causing distortion in the content of generated frames. To address this issue, the Restricting Deformation by Contour Network (RDC-Net) is proposed to repair and fill the content by optimizing object contours and leveraging context features for high-quality animation interpolation. Specifically, the RDC-Net was proposed to learn optical flow maps utilized to effectively capture the spatial shifts in the motion subject's contours across time intervals. Furthermore, contour information is used to refine the structure of the object in optical flow estimation and moderate the object deformation scale in the generated frame. In addition, the context information is explored to characterize the motion between adjacent frames via bidirectional optical flow learning, enabling the filling of distortion in the content of generated frames by feature-filling technology. Experiments on commonly used benchmarks show our state-of-the-art performance. Xin Xu 0007, Kui Jiang, Wei Liu 0183, Zheng Wang 0007 |
IEEE Signal Process. Lett. | 4 |
| 2022 | Towards generalizable person re-identification with a bi-stream generative model
Xin Xu 0007, Wei Liu 0183, Zheng Wang 0007, Ruimin Hu |
Pattern Recognit. | 2 |
| 2021 | Deep Camera-Aware Metric Learning for Person ReidentificationabstractPerson reidentification (re‐id) suffers from a challenging issue due to the significant inconsistency of the camera network, including position, view, and brands. In this paper, we propose a deep camera‐aware metric learning (DCAML) model, where images on the identity‐level spaces are further projected into different camera‐level subspaces, which can explore the inherent relationship between identity and camera. Furthermore, we exploit dynamic training strategy to jointly multiple metrics for identity‐camera relationship learning and thus consumedly elevating the retrieval accuracy. Extensive experiments on the three public datasets demonstrated that our method performs competitive results compared to the state‐of‐the‐art person re‐id methods. Wei Liu 0183, Zhiqiang Hao, Xin Xu 0007 |
Wirel. Commun. Mob. Comput. | 1 |