EDBT 2026 Demo / reviewers in the wild / expert
Xiao Wang 0029
dblp:49/67-29
· DBLP profile ↗
39ranked-venue papers
10as first author
33since 2021 · last 2026
0000-0003-0770-9891ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 33 · 10 first-author · 27 since 2021Artificial intelligence and machine learning · 8 · 3 first-author · 7 since 2021Computer networks · 4 · 1 first-author · 3 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Robust Pedestrian Detection with Uncertain ModalityabstractExisting cross-modal pedestrian detection (CMPD) employs complementary information from RGB and thermal-infrared (TIR) modalities to detect pedestrians in 24h-surveillance systems. RGB captures rich pedestrian details under daylight, while TIR excels at night. However, TIR focuses primarily on the person's silhouette, neglecting critical texture details essential for detection. While the near-infrared (NIR) captures texture under low-light conditions, which effectively alleviates performance issues of RGB and detail loss in TIR, thereby reducing missed detections. To this end, we construct a new Triplet RGB–NIR–TIR (TRNT) dataset, comprising 8,281 pixel-aligned image triplets, establishing a comprehensive foundation for algorithmic research. However, due to the variable nature of real-world scenarios, imaging devices may not always capture all three modalities simultaneously. This results in input data with unpredictable combinations of modal types, which challenge existing CMPD methods that fail to extract robust pedestrian information under arbitrary input combinations, leading to significant performance degradation. To address these challenges, we propose the Adaptive Uncertainty-aware Network (AUNet) for accurately discriminating modal availability and fully utilizing the available information under uncertain inputs. Specifically, we introduce Unified Modality Validation Refinement (UMVR), which includes an uncertainty-aware router to validate modal availability and a semantic refinement to ensure the reliability of information within the modality. Furthermore, we design a Modality-Aware Interaction (MAI) module to adaptively activate or deactivate its internal interaction mechanisms per UMVR output, enabling effective complementary information fusion from available modalities. AUNet enables accurate modality validation and robust inference without fixed modality pairings, facilitating the effective fusion of RGB, NIR, and TIR information across diverse inputs. Qian Bie, Xiao Wang 0029, Bin Yang 0026, Zhixi Yu, Jun Chen 0001, Xin Xu 0007 |
AAAI | 2 |
| 2026 | Discrepancy-Consistency Bi-Knowledge Fusion for unsupervised video anomaly detection
Awei Yin, Wei Liu 0183, Xiao Wang 0029, Ao Huang, Xin Xu 0007 |
Knowl. Based Syst. | 3 |
| 2026 | FD-HDRMamba: Frequency-Decoupled Mamba for Multi-Exposure HDR Reconstruction
Zhehan Gong, Wei Wang 0170, Xiao Wang 0029, Xin Yuan 0009 |
IEEE Signal Process. Lett. | 3 |
| 2026 | Mining Cross-Modality Implicit Semantic Association for Unsupervised Visible-Infrared Person Re-IdentificationabstractUnsupervised visible-infrared person reidentification (US-VI-ReID) seeks to learn a cross-modality retrieval model without relying on manual annotations, thereby reducing the high cost associated with labeling. Recent large-scale vision-language pre-training models, such as CLIP, have shown significant potential in enhancing pure-vision-based person re-identification. However, existing CLIP-based US-VI-ReID methods focus on independently learning semantic information within the visible and infrared modalities. These methods overlook the mismatch between the pre-training data of CLIP and the downstream cross-modality data, resulting in substantial cross-modal semantic differences. Such inconsistent semantic information, which exhibits modality discrepancies, cannot ensure the accuracy of cross-modality associations and thus hampers the performance of cross-modality learning. To address these challenges and further explore the generalizable semantic representation across modalities in CLIP, we propose a novel framework named Mining Cross-Modality Implicit Semantic Association (MCSA), which focuses on learning a modality-invariant implicit semantic space to enhance cross-modality associations and feature learning. The proposed method comprises two key modules: Modality-invariant Prompt Learning and GCNs-Driven Collaboration Alignment. Specifically, to enable CLIP to learn modality-invariant semantics, we integrate a random color augmentation branch into the visible stream for joint contrastive learning for mining generalizable semantic representations. This ensures the color generalization of the constructed implicit semantic prompts. Moreover, within the cross-modal invariant implicit semantic space, we utilize Graph Convolutional Networks (GCNs) to uncover more reliable cross-modal associations. By integrating information from images and semantic graphs, we jointly refine cross-modal correspondences, enabling the model to perform precise cross-modal feature learning. Extensive experiments conducted on the SYSU-MM01 and RegDB datasets demonstrate the effectiveness of the proposed MCSA. The source code will be released. Bin Yang 0026, Lekai Liu, Wenke Huang 0003, Xiao Wang 0029, Bo Du 0001, Mang Ye |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2025 | TokenMatcher: Diverse Tokens Matching for Unsupervised Visible-Infrared Person Re-IdentificationabstractUnsupervised visible-infrared person re-identification (US-VI-ReID) seeks to match infrared and visible images of the same individual without the use of annotations. Current methods typically derive cross-modal correspondences through a single global feature matching process for generating pseudo labels and learning modality-invariant features. However, this matching approach is hindered by both intra-modality and inter-modality discrepancies, which result in imprecise measurements. As a consequence, the clustering of individuals with single global feature is often incomplete and unreliable, leading to suboptimal performance in cross-modal clustering tasks. To address these challenges and to extract cross-modality discriminative identity information, we propose a TokenMatcher, which encompasses three key components: Diverse Tokens Matching (DTM), Diverse Tokens Neighbor Learning (DTNL), and the Homogeneous Fusion (HF) Module. DTM utilizes multiple class tokens within the visual transformer framework to capture diverse embedding representations, thereby facilitating the integration of fine-grained information essential for reliable cross-modality correspondences. DTNL enhances the intra-modality and inter-modality consistency among diverse tokens by refining neighborhood sets with insights from neighboring tokens and camera information, promoting robust neighborhood learning and fostering discriminative identity information. Additionally, the HF module consolidates clusters of the same identity while effectively separating those of different identities. Extensive experiments conducted on the publicly available SYSU-MM01 and RegDB datasets demonstrate the efficacy of the proposed method. Xiao Wang 0029, Lekai Liu, Bin Yang 0026, Mang Ye, Zheng Wang 0007, Xin Xu 0007 |
AAAI | 1 |
| 2025 | Toward Comprehensive Semantic Prompt for Region Contrastive Learning Underwater Image EnhancementabstractUnderwater image enhancement (UIE) focuses on mitigating image quality degradation due to light absorption and scattering. However, most existing methods enhance images via a global and uniform manner, neglecting the inherent semantic information in different regions, which may cause the network to easily deviate from the region’s original color. Moreover, these methods typically rely on clear images to guide network convergence, a process constrained by the limited availability of real-world datasets, making it extremely challenging to train enhancement models for various degradations. To address these challenges, this paper introduces a semantic guidance and region contrastive constraints network (SRCNet). Initially, we propose a semantic-aware RWKV (Receptance Weighted Key Value) block and a semantic prompt regularization module. These components leverage intra-target semantic correlations to preserve image details and colors within a global perceptual framework, while employing focal loss to emphasize the restoration of severely degraded regions. Subsequently, we introduce a region contrastive learning method that effectively utilizes negative samples to precisely capture features sensitive to degradation factors, thereby fostering robust feature distributions. Finally, experimental results demonstrate that our method outperforms existing state-of-the-art (SOTA) approaches. Xiao Wang 0029, Yongsheng Fu, Wei Wang 0170, Wei Liu 0183 |
ICASSP | 1 |
| 2025 | RPUDet: Learning Relational Prior and Uncertainty for Robust Aerial Object DetectionabstractAerial object detection remains a challenging task in the computer vision community. While general object detectors perform well on natural images, they struggle with aerial images due to missed detections from small, low-resolution objects and misdetections caused by classification uncertainty between semantically similar objects. To address these issues, we propose a Relational Prior and Uncertainty detector (RPUDet). RPUDet consists of two core modules: 1) Relation Aware Reasoning Module (RARM), which leverages a relational prior graph to help detect correlated objects and reduce missed detections of low-resolution objects; and 2) Uncertainty Guided Awareness Module (UGAM), which computes a uncertainty map to identify low-confidence regions, dynamically adjusts feature weights, and refines areas with high semantic ambiguity to mitigate misdetections. Additionally, to advance aerial object detection in industrial applications, we introduce the Low-Voltage Distribution Insulator Dataset (LVDID), focusing on static scenes, in contrast to existing dynamic datasets. This enables us to evaluate RPUDet's performance across diverse real-world scenarios. We evaluate RPUDet on VisDrone2019, VEDAI, and LVDID, demonstrating superior performance compared to existing methods. The code and dataset will be available at https://github.com/Godk02/RPUDet. Wei Liu 0183, Minshi Chen, Xiao Wang 0029, Xin Yuan 0009 |
ICMR | 4 |
| 2025 | Style Separation and Content Recovery for Generalizable Sketch Re-identification and a New Benchmark
Lingyi Lu, Xin Xu 0007, Xiao Wang 0029 |
MMM (4) | 3 |
| 2025 | Retrieving and Reasoning: Multivariate Feature and Attribute Cooperation for Video Anomaly DetectionabstractVideo anomaly detection (VAD), which detects abnormal patterns in video sequence, is based on several kinds of features or attributes in the existing methods. This ignores the interconnections between different features and attributes, and the initiation of an anomalous result is brought about by multiple factors. If several individual neural networks are used to perceive various types of anomalies, the system would lose awareness of the association among features and attributes, which limits the system's ability to perceive complex anomalies. In this work, we propose a dual-branch framework for VAD task, which includes deep feature retrieving and semantic attribute reasoning branch. In the former branch, three high-dimensional deep features are extracted and modeled, then the anomaly scores are obtained based on the vector retrieval database. In the latter branch, three low-dimensional semantic-level attributes are extracted for composing the attribute triplets, then use theAssociation-ruleMiningModule (AMM) to perceive potential connections among these triplets. The coefficients computed by the latter branch calibrate the anomaly scores obtained by the former while providing high-level anomaly causes. Extensive experiments show that our approach achieves state-of-the-art performance with 87.9$\%$on ShanghaiTech and 94.6$\%$on Avenue. Xingshuo Han, Xiao Wang 0029, Wei Liu 0183, Liping Ye, Xin Xu 0007 |
IEEE Signal Process. Lett. | 2 |
| 2025 | Spatial Bi-Exploration for Robust Camouflaged Object DetectionabstractCamouflaged Object Detection (COD) aims to segment camouflaged objects hidden within their environment. Existing COD models, aside from image features, mostly focus on a single coarse-grained spatial structure, such as depth information, texture information, or edge information. However, when faced with complex scenes where the target and background textures are similar and overlapping, or when subjected to noise interference, this design often leads to insufficient detection accuracy and robustness. To address these issues, we proposed a strategy for multiple spatial explorations and designedSpatial Bi-Exploration Network (SPNet). SPNet conducts a comprehensive analysis of complex camouflage scenarios by jointly exploring depth spatial, contour spatial, and image feature information, thereby enhancing detection performance and maintaining robustness. Unlike existing methods, SPNet leverages dual exploration of depth and contour spaces to mitigate the vulnerability of coarse structures to noise. Depth spatial information aids the model in recognizing the deep relationships between objects and the background, reducing the impact of noise on object boundaries, while contour spatial information improves edge detection accuracy. This dual approach significantly enhances robustness, especially in the face of adversarial attacks. Extensive experiments on benchmark datasets demonstrate that our model not only outperforms existing methods in detection performance but also exhibits superior robustness against adversarial attacks. Xiao Wang 0029, Xin Yuan 0009, Nan Mu, Zheng Wang 0007 |
IEEE Signal Process. Lett. | 2 |
| 2025 | MonOri: Orientation-Guided PnP for Monocular 3-D Object DetectionabstractMonocular 3-D object detection is a challenging task in the field of autonomous driving and has made great progress. However, current monocular image methods tend to incorporate additional information such as pseudolabels to improve algorithm performance while overlooking the geometric relationship between the object's keypoints, resulting in low performance for occluded object detection. To address this issue, we find that introducing the orientation information of objects in the 3-D detection pipeline can help improve the detection performance of occluded objects. An orientation-guided perspective-n-point (PnP) for monocular 3-D object detection method named MonOri is presented in this article, which uses object's orientation to guide keypoints' optimization. Considering the existence of different deformation objects in the scene, we design the feature aggregation detection module (FADM), which consists of the feature focus fusion module (FFFM) and CondConv detection module (CCDM). First, FFFM can highlight signals from irregularly occluded objects, effectively modeling features of elongated and small-sized objects. This module enhances the model's ability to recognize elongated and small-sized objects in complex scenes. Then, the CCDM is designed to improve the network's ability to estimate object keypoints' location regression under occlusion conditions and minimize the network computational overhead. Finally, considering that the unoccluded portions of occluded objects are closely related to the orientation of the objects, an orientation-guided keypoints' selection module (OGKSM) is proposed to enhance the accuracy of objected optimization for keypoint positions and spatial location inference of the object. Experimental results indicate that the MonOri method achieves competitive results; it is also demonstrated that the orientation information is introduced in the PnP algorithm to estimate the object's spatial position that can mitigate the impact of occlusion on object detection, thus improving the recognition rate of occluded objects. Our code is available at https://github.com/DL-YHD/MonOri. Hongdou Yao, Jun Chen 0001, Zheng Wang 0007, Yansheng Qiu, Xiao Wang 0029, Yimin wang, Xiaoyu Chai, Chenglong Cao |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2025 | Diversity-Representativeness Replay and Knowledge Alignment for Lifelong Vehicle Re-identificationabstractLifelong Vehicle Re-Identification (LVReID) aims to match a target vehicle across multiple cameras, considering non-stationary and continuous data streams, which fits the needs of the practical application better than traditional vehicle re-identification. Nonetheless, this area has received relatively little attention. Recently, methods for Lifelong Person Re-Identification (LPReID) have been emerging, with replay-based methods achieving the best results by storing a small number of instances from previous tasks for retraining, thus effectively reducing catastrophic forgetting. However, these methods cannot be directly applied to LVReID because they fail to simultaneously consider the diversity and representativeness of replayed data, resulting in biases between the subset stored in the memory buffer and the original data. They randomly sample classes, which may not adequately represent the distribution of the original data. Additionally, these methods fail to consider the rich variation in instances of the same vehicle class due to factors such as vehicle orientation and lighting conditions. Therefore, preserving more informative classes and instances for replay helps maintain information from previous tasks and may mitigate the model's forgetting of old knowledge. In view of this, we propose a novel Diversity-Representativeness Dual-Stage Sampling Replay (DDSR) strategy for LVReID that constructs an effective memory buffer through two stages, i.e. , Cluster-Centric Class Selection and Diverse Instance Mining. Specifically, we first perform class-level sampling based on density in the clustered class-centered feature space and then further mine the diverse, high-quality instances within the selected classes. In addition, we introduce Maximum Mean Discrepancy loss to align the feature distribution between replay data and the new arrivals and apply L2 regularization in the parameter space to facilitate knowledge transfer, thus enhancing the model's generalization ability to new tasks. Extensive experiments demonstrate effective improvements of our method compared to current state-of-the-art lifelong ReID methods on the VeRi-776, VehicleID, and VERI-Wild datasets. Zhijing Wan, Xiao Wang 0029, Wei Liu 0183, Wei Wang 0170, Zheng Wang 0059, Xin Xu 0007 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2025 | Glare countering and exploiting via dual stream network for nighttime vehicle detection
Pengshu Du, Xiao Wang 0029, WeiGang Li, Xin Xu 0007 |
Vis. Comput. | 2 |
| 2024 | Mutuality Attribute Makes Better Video Anomaly DetectionabstractVideo anomaly detection (VAD) is an essential but challenging task. Existing prevalent methods focus on analyzing the reconstruction or prediction difference between normal and abnormal patterns through multiple deep features, e.g., optic flow. However, these approaches independently use deep features to characterize attributes, ignore the mutuality among multiple deep features. Therefore, the constructed representation is limited to indirectly representing the anomaly from isolated attributes, and makes the network difficult to capture the high-level causes of anomaly. In this paper, we proposed a novel Mutuality Attribute-based Representation framework (MAR-VAD) for the VAD task, which absorbs the mutuality among deep features to characterize the mutuality attribute. Specifically, the mutuality attribute encapsulates high-level semantic information, such as the specific abnormal object or action, which mutually utilizes information from multiple deep features. In this way, the system is able to directly capture the high-level causes of anomaly, thus providing a more comprehensive perspective to accurately detect anomaly events. Following a process-transparent density estimation, we produce the final anomaly scores. Experiments show that MAR-VAD achieves state-of-the-art performance on ShanghaiTech and Avenue. Xingshuo Han, Xiao Wang 0029, Kui Jiang, Wei Liu 0183, Ruimin Hu, Xuefeng Pan, Xin Xu 0007 |
ICASSP | 2 |
| 2024 | Occlusion-Aware Plane-Constraints for Monocular 3D Object DetectionabstractThe task of 3D object detection poses a significant challenge for 3D scene understanding and is primarily employed in the fields of robot control and autonomous driving. Monocular-based 3D detection methods are more cost-effective and practical than stereo-based or LiDAR-based methods. Monocular image 3D detection methods have garnered considerable attention from researchers. However, the impact of occlusion scenarios of the objects on the keypoints prediction is often overlooked. To address this issue, the present paper proposes a novel 3D monocular object detection method named MonOAPC, which is equipped with occlusion-aware plane-constraints. This method can adaptively utilize partial keypoints to infer the plane location of the object based on the level of occlusion and highlights that the introduction of plane constraints is advantageous for the 3D detection task. First, the plane information of the object in 3D space is beneficial to optimize the keypoints regression, and considering that each plane holds different significance, an adaptive plane location inference module is proposed to enhance the keypoints location regression. Second, a novel co-depth estimation module is proposed to jointly estimate the object’s spatial location through various depth inference methods, thereby improving the generalization of the object depth estimation. Furthermore, the paper demonstrates that the accuracy of 3D object detection can be indirectly improved by introducing plane information to promote keypoints regression, and that plane information is effective for monocular 3D detection. The experimental outcomes show that the MonOAPC method can attain competitive results. Hongdou Yao, Jun Chen 0001, Zheng Wang 0007, Xiao Wang 0029, Xiaoyu Chai, Yansheng Qiu |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2024 | Blind 3D Video Stabilization with Spatio-Temporally Varying Motion BlurabstractVideo stabilization is a challenging task that attempts to compensate for the overall frame shake during video acquisition. Existing three-dimensional video stabilization methods aim at modeling camera perspective projection through either data-driven training or explicit motion estimation. However, the above methods are difficult to effectively solve the issue of shaky videos with abrupt object movements, resulting in local motion blur in the direction of the movement. This phenomenon is prevalent in real-world scenarios featuring foreground blind motion scenes. Unfortunately, directly combining stabilization and deblurring methods poses challenges when dealing with this situation. In the video, the intensity of motion blur undergoes continuous changes, and the direct combination method inadequately utilizes spatiotemporal information, providing insufficient clues for cross-frame compensation. To alleviate this problem, the Cross-frame-temporal Module framework is proposed to address blind motion blur induced by various conditions, which utilizes cross-frame temporal features to estimate depth maps and camera motion. In this framework, a Blur Transform Network (BTNet) is designed to adapt to spatially varying motion blur, which transforms local regions according to the impact of blur intensities to adapt to the effects of non-uniform motion blur; furthermore, our Temporal-Aware Network (TANet) further suppresses motion blur by leveraging cross-frame temporal features. In addition, the limited availability of pair-training video data containing motion blur limits the application of this approach in practice. The Cross-frame-temporal Module framework adopts an un-pretrained in-test training strategy. Extensive experimental results have demonstrated that our method outperforms state-of-the-art methods. Hengwei Li, Wei Wang 0170, Xiao Wang 0029, Xin Yuan 0009, Xin Xu 0007 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2023 | Dual-focus: person search from Coarse-Grained Focus to Fine-Grained Focus
Wenyi Hu, Xiao Wang 0029, Zheng Wang 0007, Xin Xu 0007, Ruimin Hu |
Multim. Syst. | 2 |
| 2023 | Vertex points are not enough: Monocular 3D object detection via intra- and inter-plane constraints
Hongdou Yao, Jun Chen 0001, Zheng Wang 0007, Xiao Wang 0029, Xiaoyu Chai, Yansheng Qiu |
Neural Networks | 4 |
| 2023 | Beyond the Parts: Learning Coarse-to-Fine Adaptive Alignment Representation for Person SearchabstractPerson search is a time-consuming computer vision task that entails locating and recognizing query people in scenic pictures. Body components are commonly mismatched during matching due to position variation, occlusions, and partially absent body parts, resulting in unsatisfactory person search results. Existing approaches for extracting local characteristics of the human body using keypoint information are unable to handle the search job when distinct body parts are misaligned, ignoring to exploit multiple granularities, which is crucial in the person search process. Moreover, the alignment learning methods learn body part features with fixed and equal weights, ignoring the beneficial contextual information, e.g., the umbrella carried by the pedestrian, which supplements compelling clues for identifying the person. In this paper, we propose a Coarse-to-Fine Adaptive Alignment Representation (CFA 2 R) network for learning multiple granular features in misaligned person search in the coarse-to-fine perspective. To exploit more beneficial body parts and related context of the cropped pedestrians, we design a Part-Attentional Progressive Module (PAPM) to guide the network to focus on informative body parts and positive accessorial regions. Besides, we propose a Re-weighting Alignment Module (RAM) shedding light on more contributive parts instead of treating them equally. Specifically, adaptive re-weighted but not fixed part features are reconstructed by Re-weighting Reconstruction module, considering that different parts serve unequally during image matching. Extensive experiments conducted on CUHK-SYSU and PRW datasets demonstrate competitive performance of our proposed method. Wenxin Huang, Xuemei Jia, Xian Zhong, Xiao Wang 0029, Kui Jiang, Zheng Wang 0007 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2023 | Low-light image enhancement with joint illumination and noise data distribution transformation
Wei Wang 0170, Xiao Wang 0029, Xin Xu 0007 |
Vis. Comput. | 3 |
| 2022 | Graph-Based Structural Attributes for Vehicle Re-IdentificationabstractVehicle re-identification (Re-ID), which aims to identify the same vehicle across different surveillance cameras, is a significant application in urban operation and security. Although the existing methods have noticed the importance of local features, near-duplicated cases are still hard to be handled. The reason lies that the attribute features and personalized structure of vehicles are often ignored. In this paper, we propose a graph-based structural attribute network (GSAN), which contains an attribute feature extraction module (AFEM) and a dual-grained structural relation module (DSRM). The AFEM aims to obtain attribute features of vehicles with structural information between attributes, and the DSRM aims to make the attribute features able to represent structural relation information between parts and attributes. The result on representative datasets shows that GSAN achieves competitive improvements over the state-of-the-art methods. We also collect a dataset of vehicle images with attribute annotations. Our dataset and code are released at https://github.com/HappyBoBo0331/GSAN. Rongbo Zhang, Xian Zhong, Xiao Wang 0029, Wenxin Huang, Wenxuan Liu 0008 |
ICME | 3 |
| 2022 | UoLMM'22: 2nd International Workshop on Robust Understanding of Low-quality Multimedia Data: Unitive Enhancement, Analysis and EvaluationabstractLow-quality multimedia data (including low resolution, low illumination, defects, blurriness, etc.) often pose a challenge for content understanding, as algorithms are typically developed under ideal conditions (high resolution and good visibility). To alleviate this problem, data enhancement techniques (e.g., super-resolution, low-light enhancement, derain, and inpainting) have been proposed to restore low-quality multimedia data. Efforts are also being made to develop robust content understanding algorithms in adverse weather and lighting conditions. Some quality assessment techniques aiming at evaluating the analytical quality of data have also emerged. Even though these topics are mostly studied independently, they are tightly related in terms of ensuring a robust understanding of multimedia content. For example, enhancement should maintain the semantic consistency of the analysis, while quality assessment should consider the comprehensibility of the multimedia data. The purpose of this workshop is to bring together individuals in three areas: enhancement, analysis, and evaluation, for sharing ideas and discussion on current developments and future directions. Yang Wu 0001, Xiao Wang 0029, Jing Xiao 0004 |
ACM Multimedia | 4 |
| 2022 | Towards Causality Inference for Very Important Person LocalizationabstractVery Important Person Localization (VIPLoc) aims at detecting certain individuals in a given image, who are more attractive than others in the image. Existing uncontrolled VIPLoc benchmark assumes that the image has one single VIP, which is not suitable for actual application scenarios when multiple VIPs or no VIPs appear in the image. In this paper, we re-built a complex uncontrolled conditions (CUC) dataset to make the VIPLoc closer to the actual situation, containing no, single, and multiple VIPs. Existing methods use the hand-designed and deep learning strategies to extract the features of persons and analyze the differences between VIPs and other persons from the perspective of statistics. They are not explainable as to why the VIP located this output for that input. Thus, there exist the severe performance degradation when we use these models in real-world VIPLoc. Specifically, we establish a causal inference framework that unpacks the causes of previous methods and derives a new principled solution for VIPLoc. It treats the scene as confounding factor, allowing the ever-elusive confounding effects to be eliminated and the essential determinants to be uncovered. Through extensive experiments, our method outperforms the state-of-the-art methods on public VIPLoc datasets and the re-built CUC dataset. Xiao Wang 0029, Zheng Wang 0007, Wu Liu 0005, Xin Xu 0007, Qijun Zhao, Shin'ichi Satoh 0001 |
ACM Multimedia | 1 |
| 2022 | SAM: Self Attention Mechanism for Scene Text Recognition Based on Swin Transformer
Xiang Shuai, Xiao Wang 0029, Wei Wang 0170, Xin Yuan 0009, Xin Xu 0007 |
MMM (1) | 2 |
| 2022 | Continuous and Unified Person Re-IdentificationabstractPerson re-identification (ReID) aims to match pedestrian images across disjoint cameras. Mainstream Re-ID tasks focus on training ReID models once using all the data, which become limited in some real-world scenarios where training data tends to arrive in stages. To match scenarios where training data is incrementally available, some works began to explore ReID task that can make efficient use of piecemeal new data. However, due to the limitations of the training and testing setups, these efforts are still preliminary explorations. In this paper, we explore a novel yet harder Continuous and Unified ReID (CUReID), which not only enables to continuously learn discrimination knowledge from data streams with style differences, but also to be uniformly evaluated discriminatory capability on all the data (seen and unseen). Furthermore, we propose a novel Generalized Feature Decoupled Learning (GFDL) framework for CUReID, which characterizes by introducing alternate training with extra images to solve the problem of optimization divergence between regularisation (learning new knowledge) and generalization (anti-forgetting old knowledge) tasks. In our newly proposed benchmark setup, GFDL achieves the state-of-the-art performance. Zhu Mao, Xiao Wang 0029, Xin Xu 0007, Zheng Wang 0007, Chia-Wen Lin |
IEEE Signal Process. Lett. | 2 |
| 2021 | Very Important Person Localization in Unconstrained Conditions: A New BenchmarkabstractThis paper presents a new high-quality dataset for Very Important Person Localization (VIPLoc), named Unconstrained-7k. Generally, current datasets: 1) are limited in scale; 2) built under simple and constrained conditions, where the number of disturbing non-VIPs is not large, the scene is relatively simple, and the face of VIP is always in frontal view and salient. To tackle these problems, the proposed Unconstrained-7k dataset is featured in two aspects. First, it contains over 7,000 annotated images, making it the largest VIPLoc dataset under unconstrained conditions to date. Second, our dataset is collected freely on the Internet, including multiple scenes, where images are in unconstrained conditions. VIPs in the new dataset are in different settings, e.g., large view variation, varying sizes, occluded, and complex scenes. Meanwhile, each image has more persons (> 20), making the dataset more challenging. As a minor contribution, motivated by the observation that VIPs are highly related to not only neighbors but also iconic objects, this paper proposes a Joint Social Relation and Individual Interaction Graph Neural Networks (JSRII-GNN) for VIPLoc. Experiments show that the JSRII-GNN yields competitive accuracy on NCAA (National Collegiate Athletic Association), MS (Multi-scene), and Unconstrained-7k datasets. https://github.com/xiaowang1516/VIPLoc. Xiao Wang 0029, Zheng Wang 0007, Toshihiko Yamasaki, Wenjun Zeng 0001 |
AAAI | 1 |
| 2021 | Part-Aligned Network with Background for Misaligned Person SearchabstractPerson search is a significant computer vision task that requires addressing person detection and re-identification simultaneously. Body parts are frequently misaligned due to variation poses, occlusions, and partial missing, leading to the unsatisfied results of person search. Existing methods usually extract local features from the human body by the key point information, that cannot tackle the recognition task between a pair of persons with different body parts due to misalignment. Moreover, these methods overlook background information (e.g. the carries and the background reference object) which can also supplement effective features for representing the person. In this paper, we propose a part-aligned network with background (PANB) to address this misalignment issue. To learn local fine-grained features of different body parts, we fine-tune a parsing network to divide the body region into seven parts. In particular, our proposed method considers extracting the background features as the eighth part features to extract more robust representations, which is more rational and efficient. Furthermore, we design a reconstruction method to align the parts existing in both the query image and the cropped gallery image. Extensive experiments show that our proposed method achieves competitive performance on CUHK-SYSU and PRW datasets. Xian Zhong, Wenxin Huang, Xiao Wang 0029, Jingling Yuan |
ICASSP | 4 |
| 2021 | Consistency-Constancy Bi-Knowledge Learning for Pedestrian Detection in Night SurveillanceabstractPedestrian detection in the night surveillance is a challenging yet not largely explored task. As the success of the detector in the daytime surveillance and the convenient acquisition of all-weather data, we learn knowledge from these data to benefit pedestrian detection in night surveillance. We find two key properties of surveillance: distribution cross-time consistency and background cross-frame constancy. This paper proposes a consistency-constancy bi-knowledge learning (CCBL) for pedestrian detection in night surveillance, which is able to simultaneously achieve the night pedestrian detection's useful knowledge, coming from day and night surveillance. Firstly, based on the robustness of the existing detector in day surveillance, we obtain pedestrians' distribution in the daytime scene using the detector's detection results in the daytime scene. Based on the consistency of pedestrians' distribution during the day and night in the same scene, the pedestrian distribution from daytime is used as the consistency-knowledge for pedestrian detection in night surveillance. Secondly, the background as a constant knowledge of the surveillance scene is extractable and contributes to the division of the foreground, which contains most of the pedestrian regions and helps in pedestrian detection for night surveillance. Finally, we add bi-knowledge representation to promote each other and merge them together as the final pedestrian representation. Through extensive experiments, our CCBL significantly outperforms the state-of-the-art methods on public pedestrian detection datasets. In the NightSurveillance dataset, CCBL reduced the average missed detection rate by 3.04% compared to the existing best method. Xiao Wang 0029, Zheng Wang 0007, Wu Liu 0005, Xin Xu 0007, Jing Chen 0003, Chia-Wen Lin |
ACM Multimedia | 1 |
| 2021 | Unsupervised Vehicle Search in the Wild: A New BenchmarkabstractIn urban surveillance systems, finding a specific vehicle in video frames efficiently and accurately has always been an essential part of traffic supervision and criminal investigation. Existing studies focus on vehicle re-identification (re-ID), but vehicle search is still underexploited. These methods depend on the locations of many vehicles (bounding boxes) that are not available in most real-world applications. Therefore, the unsupervised joint study of vehicle location and identification for the observed scene is a pressing need. Inspired by person search, we conduct a study on the vehicle search while considering four main discrepancies among them, summarized as: 1) It is challenging to select the candidate regions for the observed vehicle due to the perspective differences (front or side); 2) The sides of the same type of vehicles are almost the same, resulting in smaller inter-class; 3) Lacking satisfied dataset for vehicle search to meet the practical scenarios; 4) Supervised search publishing methods rely on datasets with expensive annotations. To address these issues, we have established a new vehicle search dataset. We design an unsupervised framework on this benchmark dataset to generate pseudo labels for further training existing vehicle re-ID or person search models. Experimental results reveal that these methods turn less effective on vehicle search tasks. Therefore, the vehicle search task needs to be further developed, and this dataset can advance the research of vehicle search. Https://github.com/zsl1997/VSW. Xian Zhong, Xiao Wang 0029, Kui Jiang, Wenxuan Liu 0008, Wenxin Huang, Zheng Wang 0007 |
ACM Multimedia | 3 |
| 2021 | Illuminate Low-Light Image via Coarse-to-fine Multi-level Network
Yansheng Qiu, Jun Chen 0001, Xiao Wang 0029, Kui Jang |
MMM (1) | 3 |
| 2021 | Occluded suspect search via channel-guided mechanism
Wenxin Huang, Ruimin Hu, Xiao Wang 0029, Chao Liang 0001, Jun Chen 0001 |
Neural Comput. Appl. | 3 |
| 2021 | Spatio-Spectral Feature Fusion for Low-Light Image EnhancementabstractLow-light image enhancement aims to improve an image's visual quality, which is essential for many downstream computer vision and multimedia tasks. Existing spatial-domain low-light enhancement methods barely focus on the regions containing object boundaries, which take the most informative characteristics. However, solely focusing on enhancing high-frequency details not only causes over-sharpening of an image but also leads to color distortion. In this paper, we propose a novel spatio-spectral feature fusion network (S2F2N), that involves a frequency-feature representation branch (FRB) and a spatial-feature representation branch (SRB) to learn the domain-specific representation individually. Moreover, a spatial-channel mixed attention block (MAB) is introduced to learn the joint representation of spatio-spectral features for final image relighting. Extensive experiments on several benchmark datasets demonstrate that our method can produce high fidelity results for low-light images. Yansheng Qiu, Jun Chen 0001, Zheng Wang 0007, Xiao Wang 0029, Chia-Wen Lin |
IEEE Signal Process. Lett. | 4 |
| 2021 | Rain-Free and Residue Hand-in-Hand: A Progressive Coupled Network for Real-Time Image DerainingabstractRainy weather is a challenge for many vision-oriented tasks (e.g., object detection and segmentation), which causes performance degradation. Image deraining is an effective solution to avoid performance drop of downstream vision tasks. However, most existing deraining methods either fail to produce satisfactory restoration results or cost too much computation. In this work, considering both effectiveness and efficiency of image deraining, we propose a progressive coupled network (PCNet) to well separate rain streaks while preserving rain-free details. To this end, we investigate the blending correlations between them and particularly devise a novel coupled representation module (CRM) to learn the joint features and the blending correlations. By cascading multiple CRMs, PCNet extracts the hierarchical features of multi-scale rain streaks, and separates the rain-free content and rain streaks progressively. To promote computation efficiency, we employ depth-wise separable convolutions and a U-shaped structure, and construct CRM in an asymmetric architecture to reduce model parameters and memory footprint. Extensive experiments are conducted to evaluate the efficacy of the proposed PCNet in two aspects: (1) image deraining on several synthetic and real-world rain datasets and (2) joint image deraining and downstream vision tasks (e.g., object detection and segmentation). Furthermore, we show that the proposed CRM can be easily adopted to similar image restoration tasks including image dehazing and low-light enhancement with competitive performance. The source code is available at https://github.com/kuijiang0802/PCNet. Kui Jiang, Zhongyuan Wang 0001, Peng Yi 0002, Chen Chen 0001, Zheng Wang 0007, Xiao Wang 0029, Junjun Jiang, Chia-Wen Lin |
IEEE Trans. Image Process. | 6 |
| 2020 | When Pedestrian Detection Meets Nighttime Surveillance: A New BenchmarkabstractPedestrian detection at nighttime is a crucial and frontier problem in surveillance, but has not been well explored by the computer vision and artificial intelligence communities. Most of existing methods detect pedestrians under favorable lighting conditions (e.g. daytime) and achieve promising performances. In contrast, they often fail under unstable lighting conditions (e.g. nighttime). Night is a critical time for criminal suspects to act in the field of security. The existing nighttime pedestrian detection dataset is captured by a car camera, specially designed for autonomous driving scenarios. The dataset for nighttime surveillance scenario is still vacant. There are vast differences between autonomous driving and surveillance, including viewpoint and illumination. In this paper, we build a novel pedestrian detection dataset from the nighttime surveillance aspect: NightSurveillance1. As a benchmark dataset for pedestrian detection at nighttime, we compare the performances of state-of-the-art pedestrian detectors and the results reveal that the methods cannot solve all the challenging problems of NightSurveillance. We believe that NightSurveillance can further advance the research of pedestrian detection, especially in the field of surveillance security at nighttime. Xiao Wang 0029, Jun Chen 0001, Zheng Wang 0007, Wu Liu 0005, Shin'ichi Satoh 0001, Chao Liang 0001, Chia-Wen Lin |
IJCAI | 1 |
| 2020 | S3D: Scalable Pedestrian Detection via Score Scale Surface DiscriminationabstractPedestrian detection has remained an important research topic in both the computer vision and multimedia communities because of its importance in practical applications, such as driving assistance and video surveillance. Existing methods compare the response score with a fixed threshold to determine whether a candidate region contains pedestrians and produce dissatisfactory results that contain either missed detections or false detections, which are difficult to balance. This situation has a serious impact under the condition of variable scale. This paper investigates the functional relationship between the scores and scales of pedestrians. By designing experiments with multiple scales, we have found a discriminant surface in the score scale space. Pedestrians can be distinguished at various scale levels according to their locations on the discriminant surface. The proposed approach is evaluated using four challenging pedestrian detection datasets, including Caltech, INRIA, ETH, and KITTI, and the superior experimental results are achieved when compared with baseline methods. Xiao Wang 0029, Chao Liang 0001, Chen Chen 0001, Jun Chen 0001, Zheng Wang 0007, Zhen Han 0002, Chunxia Xiao |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2020 | Listen, Look, and Find the One: Robust Person Search with Multimodality IndexabstractPerson search with one portrait, which attempts to search the targets in arbitrary scenes using one portrait image at a time, is an essential yet unexplored problem in the multimedia field. Existing approaches, which predominantly depend on the visual information of persons, cannot solve problems when there are variations in the person’s appearance caused by complex environments and changes in pose, makeup, and clothing. In contrast to existing methods, in this article, we propose an associative multimodality index for person search with face, body, and voice information. In the offline stage, an associative network is proposed to learn the relationships among face, body, and voice information. It can adaptively estimate the weights of each embedding to construct an appropriate representation. The multimodality index can be built by using these representations, which exploit the face and voice as long-term keys and the body appearance as a short-term connection. In the online stage, through the multimodality association in the index, we can retrieve all targets depending only on the facial features of the query portrait. Furthermore, to evaluate our multimodality search framework and facilitate related research, we construct the Cast Search in Movies with Voice (CSM-V) dataset, a large-scale benchmark that contains 127K annotated voices corresponding to tracklets from 192 movies. According to extensive experiments on the CSM-V dataset, the proposed multimodality person search framework outperforms the state-of-the-art methods. Xiao Wang 0029, Wu Liu 0005, Jun Chen 0001, Xiaobo Wang 0001, Chenggang Yan 0001, Tao Mei 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2017 | Low-resolution pedestrian detection via a novel resolution-score discriminative surfaceabstractPedestrian detection, as an important task in video surveillance and forensics applications, has been widely studied. However, its performance is unsatisfactory especially in the low resolution conditions. In realistic scenarios, the size of pedestrians in the images is often small, and detection can be challenging. To solve this problem, this paper proposes a novel resolution-score discriminative surface method to investigate the variation behaviors of detection scores under different pedestrian and non-pedestrian image resolutions. The discriminative surface consists of a series of positive and negative resolution-score lines, and each of them is a connected line to depict the variation relationship between pedestrian's detection scores under various image resolutions. On this basis, the resolution-score discriminative surface can classify a resolution-score line as a pedestrian or not according to whether it lies in the positive or the negative region. Experimental results on two public datasets and one campus surveillance dataset demonstrate the effectiveness of the proposed method. Xiao Wang 0029, Jun Chen 0001, Chao Liang 0001, Chen Chen 0001, Zheng Wang 0007, Ruimin Hu |
ICME | 1 |
| 2015 | Object Detection in Low-Resolution Image via Sparse Representation
Wenhua Fang, Jun Chen 0001, Chao Liang 0001, Xiao Wang 0029, Yuanyuan Nan, Ruimin Hu |
MMM (1) | 4 |
| 2014 | Pedestrian detection from salient regionsabstractClassic algorithms of pedestrian detection usually locate the latent position via sliding window techniques, which resize the matching window and/or original images at different scales and scan the image. However, this method has two main drawbacks. First, resizing at a fix rate cannot search through the whole scale space, resulting in the failure of accurate object location. Second, resizing and scanning at various scales is usually time-consuming, which is improper for practical applications. To conquer the above difficulties, a novel pedestrian detection method with salient information is proposed. In this paper, the salient detection model and the traditional covariance matrix descriptor are combined in a Bayesian framework to detect pedestrians in the still image. Finally, the efficiency of our approach compared with state-of-the-art results is demonstrated on the public INRIA dataset. Xiao Wang 0029, Jun Chen 0001, Wenhua Fang, Chao Liang 0001, Chunjie Zhang 0001, Ruimin Hu |
ICIP | 1 |