VLDB 2026 Research / reviewers in the wild / expert
Wei Li 0110
dblp:64/6025-110
· DBLP profile ↗
31ranked-venue papers
8as first author
24since 2021 · last 2026
0000-0002-8278-1765ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 29 · 8 first-author · 22 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Monocular Vehicle Pose and Shape Reconstruction via Dynamic Context Adaptation and Progressive Geometry RefinementabstractAccurate reconstruction of 3D vehicle pose and shape from monocular images is challenging, particularly for distant objects in autonomous driving. Existing methods often suffer from geometric ambiguity in depth estimation and structural hollowness in shape recovery, primarily due to inadequate multi-scale feature aggregation and unflexible prior modeling. To overcome these limitations, MonoVPR is proposed, a novel framework integrating dynamic context adaptation and progressive geometry refinement. Specifically, a Hierarchical Dual-Context Attention (HDCA) module is introduced to resolve scale-dependent degradation through gated cross-attention across multi-resolution feature maps, dynamically fusing object-centric geometric cues with scene-centric semantics. For shape refinement, the Bounded Iterative Mesh Refiner (BIMR) progressively optimizes template-guided deformations via multi-head attention and a tanh-bounded correction loop, ensuring physically plausible reconstructions.Extensive experiments on the ApolloCar3D benchmark demonstrate MonoVPR achieves state-of-the-art performance, showing exceptional capability in reconstructing geometrically consistent shapes and precise poses for challenging long-range scenarios. Wei Li 0110, Long Ji, Xiao Wu 0001, Zhaoquan Yuan, Penglin Dai |
AAAI | 1 |
| 2026 | InfoCom: Kilobyte-Scale Communication-Efficient Collaborative Perception with Information BottleneckabstractPrecise environmental perception is critical for the reliability of autonomous driving systems. While collaborative perception mitigates the limitations of single-agent perception through information sharing, it encounters a fundamental communication-performance trade-off. Existing communication-efficient approaches typically assume MB-level data transmission per collaboration, which may fail due to practical network constraints. To address these issues, we propose InfoCom, an information-aware framework establishing the pioneering theoretical foundation for communication-efficient collaborative perception via extended Information Bottleneck principles. Departing from mainstream feature manipulation, InfoCom introduces a novel information purification paradigm that theoretically optimizes the extraction of minimal sufficient task-critical information under Information Bottleneck constraints. Its core innovations include: i) An Information-Aware Encoding condensing features into minimal messages while preserving perception-relevant information; ii) A Sparse Mask Generation identifying spatial cues with negligible communication cost; and iii) A Multi-Scale Decoding that progressively recovers perceptual information through mask-guided mechanisms rather than simple feature reconstruction. Comprehensive experiments across multiple datasets demonstrate that InfoCom achieves near-lossless perception while reducing communication overhead from megabyte to kilobyte-scale, representing 440-fold and 90-fold reductions per agent compared to Where2comm and ERMVP, respectively. Quanmin Wei, Penglin Dai, Wei Li 0110, Bingyi Liu, Xiao Wu 0001 |
AAAI | 3 |
| 2026 | I2CD: An Invertible Causal Framework for Compositional Zero-Shot Learning via Disentangle-Compose-DisentangleabstractCompositional Zero-Shot Learning (CZSL) addresses the challenge of recognizing unseen attribute-object compositions in images, representing a fundamental challenge in artificial intelligence. Current approaches, which primarily focus on semantic alignment or distribution independence of primitives, have not achieved effective state-object decoupling and causal interventional invariance, limiting their performance on unseen compositions. To tackle this challenge, this study introduces I2CD (Invertible Causal framework via Disentangle-Compose-Disentangle), a novel framework that integrates invertible neural networks with causal intervention techniques to achieve state-object disentanglement. The framework employs a disentangle-compose-disentangle mechanism for counterfactual generation within the disentangled representation space, ensuring that modifications to one primitive (attribute or object) maintain independence from the other, thus enabling robust causal disentanglement. Representational consistency is maintained through semantic alignment between initial disentangled representations and their recomposed-then-disentangled counterparts with corresponding textual concepts. Comprehensive evaluations on three benchmark datasets—MIT-States, UT-Zappos, and C-GQA—demonstrate the framework's effectiveness in achieving both disentanglement and compositional generalization in CZSL tasks. Zhaoquan Yuan, Yuankang Pan, Ao Luo, Wei Li 0110, Xiao Wu 0001, Changsheng Xu |
AAAI | 5 |
| 2026 | Prohibited Item Detection in X-ray images based on refined surface perception
Wei Li 0110, Zhaoquan Yuan |
J. Vis. Commun. Image Represent. | 3 |
| 2026 | Energy-based causal disentanglement for compositional zero-shot learning
Yuankang Pan, Zhaoquan Yuan, Wei Li 0110 |
Multim. Syst. | 4 |
| 2026 | Learning Unknowns Without Forgetting Knowns: Compositional and Bidirectional Low-Rank Adaptive Open-World Detection TransformerabstractOpen-World Object Detection (OWOD) aims to detect unseen objects as “unknown” while incrementally learning them without catastrophic forgetting. This problem presents two major challenges: (1) the lack of annotations for unknown objects during training, and (2) the risk of catastrophic forgetting during model updates. To address these issues, we propose the COmpositional and Bidirectional low-Rank Adaptive open-world detection transformer (COBRA)-a novel framework built upon a pre-trained Deformable DETR model. Specifically, COBRA first employs an attentional filtering mechanism that prunes previously known (P-Known) and currently known (C-Known) objects, yielding a purified set of candidateunknowns. To system-atically pseudo-label theseunknowns, we introduce a Primitive Composition Recognition (PCR) module, which evaluates set-level similarity between candidate objects and learned primitives, enabling accurate labeling ofpseudo-unknowns. To mitigate catastrophic forgetting during incremental updates, COBRA leverages Bidirectional Low-Rank Adaptation (Bi-LoRA)-a parameter-efficient mechanism that supports forward knowledge transfer and stable backward integration. Together, these components form a synergistic pipeline for continual object discovery and knowledge consolidation. Extensive experiments on MS COCO and PASCAL VOC demonstrate that our rehearsal-free COBRA framework outperforms SAM-powered methods in unknown recall while achieving lower forgetting compared to rehearsal-based competitors. Yuankang Pan, Zhaoquan Yuan, Xiao Wu 0001, Wei Li 0110, Ao Luo, Changsheng Xu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | PSGNet: Pure Smoke Image Generation With Gradient and Style LearningabstractThe realistic and controllable generation of pure smoke is critical for smoke image editing, smoke visual special effects generation, and smoke data synthesizing within security scenarios. It is a relatively underexplored topic and continues to present significant challenges. Existing methods face challenges in the generation of smoke with intricate details and the regulation of various smoke styles. In this paper, a Pure Smoke image Generation Network (PSGNet) is proposed with a gradient and style learning approach to generate realistic and controllable smoke images. To achieve flexibility in control across the spatial dimension, the smoke shape mask is used to encode spatial details, such as the location and contour of the smoke, along with other related properties. To enhance the physical realism of synthesized smoke, a novel gradient-based learning framework is proposed to generate smoke gradient features, highlighting a special focus on explicitly encoding and exploiting gradient information. This framework uses a smoke gradient learning architecture that captures the subtle structures and patterns characteristic of real smoke, enabling the generation of highly realistic smoke with rich, fine-scale detail. In addition, a spatially aware style learning strategy is proposed to provide fine-grained control over smoke attributes such as density, color, and overall look. It is able to effectively model style features across both channel and spatial dimensions, thereby enabling spatially aware style manipulation. By combining the gradient module with this style learning framework, the method produces smoke that exhibits rich visual details and customizable image styles. Experiments conducted on six benchmark datasets demonstrate that the proposed PSGNet significantly outperforms the state-of-the-art approaches. Jian-Jun Qiao, Xiao Wu 0001, Zhi-Qi Cheng, Wei Li 0110, Zhaoquan Yuan |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2025 | CoPEFT: Fast Adaptation Framework for Multi-Agent Collaborative Perception with Parameter-Efficient Fine-TuningabstractMulti-agent collaborative perception is expected to significantly improve perception performance by overcoming the limitations of single-agent perception through exchanging complementary information. However, training a robust collaborative perception model requires collecting sufficient training data that covers all possible collaboration scenarios, which is impractical due to intolerable deployment costs. Hence, the trained model is not robust against new traffic scenarios with inconsistent data distribution and fundamentally restricts its real-world applicability. Further, existing methods, such as domain adaptation, have mitigated this issue by exposing the deployment data during the training stage but incur a high training cost, which is infeasible for resource-constrained agents. In this paper, we propose a Parameter-Efficient Fine-Tuning-based lightweight framework, CoPEFT, for fast adapting a trained collaborative perception model to new deployment environments under low-cost conditions. CoPEFT develops a Collaboration Adapter and Agent Prompt to perform macro-level and micro-level adaptations separately. Specifically, the Collaboration Adapter utilizes the inherent knowledge from training data and limited deployment data to adapt the feature map to new data distribution. The Agent Prompt further enhances the Collaboration Adapter by inserting fine-grained contextual information about the environment. Extensive experiments demonstrate that our CoPEFT surpasses existing methods with less than 1\% trainable parameters, proving the effectiveness and efficiency of our proposed method. Quanmin Wei, Penglin Dai, Wei Li 0110, Bingyi Liu, Xiao Wu 0001 |
AAAI | 3 |
| 2025 | Dynamic Object Queries for Transformer-based Incremental Object DetectionabstractIncremental object detection (IOD) aims to sequentially learn new classes, while maintaining the capability to locate and identify old ones. Prior methodologies mainly tackle catastrophic forgetting through knowledge distillation and exemplar replay, ignoring the conflict between limited model capacity and increasing knowledge. In this paper, we propose the Dynamic object Query-based DEtection TRansformer (DyQ-DETR), which incrementally expands the model representation ability to achieve stability-plasticity tradeoff. First, a new set of learnable object queries are fed into the decoder to represent new classes. Second, we propose the isolated bipartite matching for object queries in different phases, based on disentangled self-attention. Thanks to the separate supervision and computation over object queries, we further present the risk-balanced partial calibration for effective exemplar replay. Extensive experiments demonstrate that DyQ-DETR significantly surpasses the state-of-the-art methods, with limited parameter overhead. The code is available at https://github.com/THUzhangjic/DyQ-DETR. Jichuan Zhang, Wei Li 0110, Shuang Cheng, Yali Li 0001, Shengjin Wang |
ICASSP | 2 |
| 2025 | HOPNet: Learning Hand-Object-Person Interaction Network for Hand Contact State DetectionabstractThe detection of hand contact states, which involves identifying interactions between hands and objects or other entities, is essential for the development of human-computer interaction systems and the comprehension of social dynamics. Previous approaches have made progress in modeling hand-object interactions. Nonetheless, they neglect critical cues between their hands and bodies, as well as those of others, thus constraining their ability to accurately detect interpersonal contact. The task remains challenging due to frequent occlusions, especially in crowded multi-person scenarios with complex contexts. In this paper, a novel hand-object-person interaction network, called HOPNet, is proposed to model contextual information between hands and objects, as well as between hands and bodies. Specifically, HOPNet consists of two components: (i) the Hand-Object Relation (HOR) module analyzes interaction patterns between hands and objects, capturing spatial and semantic relationships; (ii) the Contrastive Spatial Refinement (CSR) module learns hand-body interactions through contrastive geometric embedding and relative spatial enhancement, improving interpersonal contact recognition in crowded scenarios. Experiments on ContactHands and 100DOH datasets demonstrate that HOPNet outperforms state-of-the-art methods. Wei Li 0110, Yizhao Wan, Xiao Wu 0001, Jianshuai Wang, Penglin Dai, Zhaoquan Yuan |
ACM Multimedia | 1 |
| 2025 | DualEnhance: External Multimodal Foundation Models Guidance and Internal Fast-Slow Teacher RegulationabstractSource-Free Domain Adaptive Object Detection addresses cross-domain detection on an unlabeled target domain without accessing source data. Existing methods implement self-training with Mean Teacher but are bottlenecked by error accumulation from noisy pseudo-labels generated via recursive teacher-student updates. This issue is handled through the proposed dual enhancements: (1) External Guidance via Multimodal Foundation Models (FMs); (2) Internal Regulation through Fast-Slow Teacher. First, despite FMs' multimodal comprehension, their semantic misalignment with a specific task introduces noise during adaptation. Bidirectional Distillation mitigates this by calibrating the FM using task-specific knowledge transferred from the source detector. The aligned cross-modal knowledge then propagates through high-quality pseudo-label generation. Second, the conventional Mean Teacher suffers from plasticity-stability dilemma, where rapid adaptation corrupts historical knowledge. Fast-Slow Teacher introduces dual-velocity knowledge consolidation: The Fast Teacher dynamically captures emerging domain features, while the Slow Teacher preserves stable historical knowledge and periodically resets the Fast Teacher, establishing an error-correcting dynamic equilibrium. Experiments show our method achieves significant improvements over SOTA. Qi He 0007, Xiao Wu 0001, Jun-Yan He, Wei Li 0110, Zhaoquan Yuan |
ACM Multimedia | 4 |
| 2025 | HighlightNet: Learning Highlight-Guided Attention Network for Nighttime Vehicle DetectionabstractVehicle detection at night is a crucial task in Intelligent Transportation Systems. Due to the complex lighting environment, vehicle detection at night remains a challenging task. Headlights and taillights are essential cues to identify vehicles at night. However, existing methods struggle to effectively utilize the light information of the vehicle. This paper proposes a novel highlight-guided framework to identify vehicles, named HighlightNet, by utilizing both the illumination data from the vehicle lights and the reflective properties of vehicles. The framework combines vehicle detection and highlight area recognition via dual-branch joint learning. To ensure that both branches focus on the highlighted regions, Feature Similarity Awareness Attention (FSAA) is introduced to capture the common attention regions of different branches. Highlight Region Perception (HRP) is proposed to exclude streetlights and other reflective illuminations from the FSAA output, which generates a mask map capable of differentiating the foreground from the background of highlighted areas. It improves the allocation of feature weights and adaptively modifies the distribution within the dual-branch configuration. Furthermore, to address the severe pixel imbalance between the highlighted area and the background, Adaptive Spatial Balance (ASB) loss is introduced to allocate the attention towards prospective vehicle regions while diminishing the emphasis on background regions. Extensive experiments conducted on the BDD100K-Night dataset and a newly acquired dataset specifically designed for nighttime surveillance, called the NightVehicle dataset, demonstrate that HighlightNet outperforms the state-of-the-art methods for nighttime vehicle detection. Yu-Pei Song, Xiao Wu 0001, Wei Li 0110, Tingquan He, Dongfeng Hu, Qiang Peng |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2024 | CAPNet: Cartoon Animal Parsing with Spatial Learning and Structural ModelingabstractCartoon animal parsing aims to segment the body parts such as heads, arms, legs and tails of cartoon animals. Different from previous parsing tasks, cartoon animal parsing faces new challenges, including irregular body structures, abstract drawing styles and diverse animal categories. Existing methods have difficulties when addressing these challenges caused by the spatial and structural properties of cartoon animals. To address these challenges, a novel spatial learning and structural modeling network, named CAPNet, is proposed for cartoon animal parsing. It aims to address the critical problems of spatial perception, structure modeling and spatial-structural consistency learning. A spatial-aware learning module integrates deformable convolutions to learn spatial features of diverse cartoon animals. The multi-task edge and center point prediction mechanism is incorporated to capture the intricate spatial patterns. A structural modeling method is proposed to model the complex structural representations of cartoon animals, which integrates a graph neural network with a shape-aware relation learning module. To mitigate the significant differences among animals, a spatial and structural consistency learning strategy is proposed to capture and learn feature correlations across different animal species. Extensive experiments conducted on benchmark datasets demonstrate the effectiveness of the proposed approach, which outperforms the state-of-the-art methods. Jian-Jun Qiao, Meng-Yu Duan, Xiao Wu 0001, Wei Li 0110 |
ACM Multimedia | 4 |
| 2023 | CPNet: Cartoon Parsing with Pixel and Part CorrelationabstractCartoon parsing, the task of segmenting constituent parts such as heads, arms, and legs of cartoon characters, holds substantial significance for applications in the animation industry and emerging metaverse. Nonetheless, this domain presents considerable challenges stemming from complex visual appearances, irregular structures, abstract drawing styles, among other factors. In this paper, a novel Cartoon Parsing Network (CPNet) is introduced to address these challenges. CPNet skillfully leverages the spatial and semantic correlations of pixels to discern intricate and visually akin appearances. Furthermore, it employs both local and global correlations of constituent parts to differentiate irregular and abstract body sections. Specifically, the pixels of the cartoon image are interconnected by capitalizing on the spatial and semantic correlations. To this end, a center point predictor, working in tandem with a pixel-aware attention, facilitates the exploration of pixel-level correlation learning. Additionally, the various constituent parts are meticulously organized to resonate with the intrinsic physiological structure of a cartoon character. The character's graph structure is assembled and analyzed by an edge-aware graph neural network, thereby linking adjacent parts and assimilating local correlations. A part-guided non-local attention mechanism is fashioned to correlate individual parts with the entire body, thereby modeling global connections. In addition, a new dataset named CartoonSet is curated and annotated explicitly for cartoon parsing. Experiments carried out on both cartoon parsing and human parsing datasets yield compelling results, thereby attesting to the efficacy and innovativeness of the proposed method. Jian-Jun Qiao, Jie Zhang 0179, Xiao Wu 0001, Yu-Pei Song, Wei Li 0110 |
ACM Multimedia | 5 |
| 2023 | Improving Anomaly Segmentation with Multi-Granularity Cross-Domain AlignmentabstractAnomaly segmentation plays a crucial role in identifying anomalous objects within images, which facilitates the detection of road anomalies for autonomous driving. Although existing methods have shown impressive results in anomaly segmentation using synthetic training data, the domain discrepancies between synthetic training data and real test data are often neglected. To address this issue, Multi-Granularity Cross-Domain Alignment (MGCDA) framework is proposed for anomaly segmentation in complex driving environments. It uniquely combines a new Multi-source Domain Adversarial Training (MDAT) module and a novel Cross-domain Anomaly-aware Contrastive Learning (CACL) method to boost the generality of the model, seamlessly integrating multi-domain data at both scene and sample levels. Multi-source domain adversarial loss and a dynamic label smoothing strategy are integrated into MDAT module to facilitate the acquisition of domain-invariant features at the scene level, through adversarial training across multiple stages. CACL aligns sample-level representations with contrastive loss on cross-domain data, which utilizes an anomaly-aware sampling strategy to efficiently sample hard samples and anchors. The proposed framework has decent properties of parameter-free during the inference stage and is compatible with other anomaly segmentation networks. Experimental conducted on Fishyscapes and RoadAnomaly datasets demonstrate that the proposed framework achieves the state-of-the-art performance. Ji Zhang 0027, Xiao Wu 0001, Zhi-Qi Cheng, Qi He 0007, Wei Li 0110 |
ACM Multimedia | 5 |
| 2023 | Learning Surface-awareness Network for X-Ray Prohibited Item DetectionabstractX-ray image security detection is a crucial method used to identify various types of prohibited items in luggage. However, the unique characteristics of X-ray imaging can result in the loss of intricate surface details, leading to subpar detection of prohibited items within X-ray images. In this paper, a Surface-aware Prohibited Item X-ray Detection Network (SPIXDet) is proposed to address this issue, which incorporates two key components: the Boundary Aggregation Module (BAM) and the Global Cross-Feature Downsampling layer (GCFD). The BAM module effectively mines image edge information while minimizing the number of parameters involved. Meanwhile, the GCFD module is introduced to mitigate chaotic interference caused by undifferentiated boundary boosting. The surface-aware capability of the model can be enhanced through the BAM and GCFD module. Furthermore, the Focal-SIoU loss function is introduced to increase positioning accuracy and optimize the model training process. To validate the effectiveness of our model, extensive experiments are conducted on the SIXray100 dataset, and the results demonstrate the advantages of SPIXDet compared to other X-ray prohibited item detection methods. Wei Li 0110, Zhaoquan Yuan, Xiao Wu 0001 |
MMAsia | 2 |
| 2022 | Learning Selective Assignment Network for Scene-Aware Vehicle DetectionabstractDeep learning has shown remarkable success in data-driven vehicle detection, relying on collected training samples from known scenes. A challenging problem arises when these detectors handle agnostic scenes, while keeping the performance of previous ones. To address this issue, a feasible remedy is to learn a set of domain-adaptive detectors by aligning the features from one scene to another. However, the improvement obtained in this way is inflexible despite the progress in object detection. An important reason is that the memory sizes grow massively with deliberately saving all scenes-independent detectors, while ignoring the relationship among different scenes. In this paper, we aim to bridge the gap between scene diversification and object consistency for scene-aware vehicle detection. Specifically, a novel structured network is proposed to integrate selective assignment of scene-specific parameters into the vehicle detection framework. Extensive experiments conducted on different scenes including BDD, Cityscapes-car, CARPK, etc, demonstrate that the proposed method achieves impressive performance, while keeping the performance of previous scenes as the scene changes. Zhenting Wang, Wei Li 0110, Xiao Wu 0001, Luhan Sheng |
ICIP | 2 |
| 2022 | Learning Graph-based Residual Aggregation Network for Group Activity RecognitionabstractGroup activity recognition aims to understand the overall behavior performed by a group of people. Recently, some graph-based methods have made progress by learning the relation graphs among multiple persons. However, the differences between an individual and others play an important role in identifying confusable group activities, which have not been elaborately explored by previous methods. In this paper, a novel Graph-based Residual AggregatIon Network (GRAIN) is proposed to model the differences among all persons of the whole group, which is end-to-end trainable. Specifically, a new local residual relation module is explicitly proposed to capture the local spatiotemporal differences of relevant persons, which is further combined with the multi-graph relation networks. Moreover, a weighted aggregation strategy is devised to adaptively select multi-level spatiotemporal features from the appearance-level information to high level relations. Finally, our model is capable of extracting a comprehensive representation and inferring the group activity in an end-to-end manner. The experimental results on two popular benchmarks for group activity recognition clearly demonstrate the superior performance of our method in comparison with the state-of-the-art methods. Wei Li 0110, Tianzhao Yang, Xiao Wu 0001, Zhaoquan Yuan |
IJCAI | 1 |
| 2022 | Learning Action-guided Spatio-temporal Transformer for Group Activity RecognitionabstractLearning spatial and temporal relations among people plays an important role in recognizing group activity. Recently, transformer-based methods have become popular solutions due to the proposal of self-attention mechanism. However, the person-level features are fed directly into the self-attention module without any refinement. Moreover, group activity in a clip often involves unbalanced spatio-temporal interactions, where only a few persons with special actions are critical to identifying different activities. It is difficult to learn the spatio-temporal interactions due to the lack of elaborately modeling the action dependencies among all people. In this paper, a novel Action-guided Spatio-Temporal transFormer (ASTFormer) is proposed to capture the interaction relations for group activity recognition by learning action-centric aggregation and modeling spatio-temporal action dependencies. Specifically, ASTFormer starts with assigning all persons in each frame to the latent actions, while an action-centric aggregation strategy is performed by weighting the sum of residuals for each latent action under the supervision of global action information. Then, a dual-branch transformer is proposed to refine the inter- and intra-frame action-level features, where two encoders with the self-attention mechanism are employed to select important tokens. Next, a semantic action graph is explicitly devised to model the dynamic action-wise dependencies. Finally, our model is capable of boosting group activity recognition by fusing these important cues, while only requiring video-level action labels. Extensive experiments on two popular benchmarks (Volleyball and Collective Activity) demonstrate the superior performance of our method in comparison with the state-of-the-art methods using only raw RGB frames as input. Wei Li 0110, Tianzhao Yang, Xiao Wu 0001, Xian-Jun Du, Jian-Jun Qiao |
ACM Multimedia | 1 |
| 2022 | Real-time Semantic Segmentation with Parallel Multiple Views Feature AugmentationabstractReal-time semantic segmentation is essential for many practical applications, which utilizes attention-based feature aggregation into lightweight structures to improve accuracy and efficiency. However, existing attention-based methods ignore 1) high-level and low-level feature augmentation guided by spatial information, and 2) low-level feature augmentation guided by semantic context, so that feature gaps between multi-level features and noise of low-level spatial details still exist. To address these problems, a new real-time semantic segmentation network, called MvFSeg, is proposed. In MvFSeg, parallel convolution with multiple depths is designed as a context head to generate and integrate multi-view features with larger receptive fields. Moreover, MvFSeg designs multiple views feature augmentation strategies that exploit spatial and semantic guidance for shallow and deep feature augmentation in an inter-layer and intra-layer manner. These strategies eliminate feature gaps between multi-level features, filter out the noise of spatial details, and provide spatial and semantic guidance for multi-level features. By combining multi-view features and augmented features from the lightweight networks with progressive dense aggregation structures, MvFSeg effectively captures invariance at various scales and generates high-quality segmentation results. Experiments conducted on Cityscapes and CamVid benchmark show that MvFSeg outperforms existing state-of-the-art methods. Jian-Jun Qiao, Zhi-Qi Cheng, Xiao Wu 0001, Wei Li 0110, Ji Zhang 0027 |
ACM Multimedia | 4 |
| 2022 | CrossNet: Boosting Crowd Counting with LocalizationabstractGenerating high-quality density maps is a crucial step in crowd counting. It is obvious that exploiting the head location of the people can naturally highlight the crowded area and eliminate the interference of background noise. However, existing crowd counting methods are still tricky to reasonably use location in density generation. In this paper, a novel location-guided framework named CrossNet is proposed for crowd counting, which integrates location supervision into density maps through dual-branch joint training. First, a new branching network is proposed to localize the potential positions of pedestrians. With the help of supervision induced from the localization branch, Location Enhancement (LE) module is designed to obtain high-quality density maps by positioning foreground regions. Second, Adaptive Density Awareness Attention (ADAA) module is engaged to enhance localization accuracy, which can efficiently use the density of the counting branch to adaptively capture the error-prone dense areas of the location maps. Finally, Density Awareness Localization (DAL) loss is offered to allocate attention to the crowd density levels, which delivers more focus on regions with high densities and less concentration on areas with low densities. Extensive experiments conducted on four benchmark datasets demonstrate that the proposed method outperforms the state-of-the-art approaches both in crowd counting and crowd localization. Ji Zhang 0027, Zhi-Qi Cheng, Xiao Wu 0001, Wei Li 0110, Jian-Jun Qiao |
ACM Multimedia | 4 |
| 2022 | SWNet: A Deep Learning Based Approach for Splashed Water Detection on RoadabstractAdverse weather conditions seriously threaten the traffic safety, especially for rainy days with the ponding water on the road surface, which potentially result in vehicle crashes, person injuries and crash fatalities. Automatic splashed water detection based on surveillance videos is an attractive way to effectively prevent the traffic accidents. However, surveillance videos exhibit great variations with lighting changes, illumination conditions and complex backgrounds, which pose great difficulties in automatic recognition. In this paper, a novel deep learning based approach is proposed to detect the splashed water. To the best of our knowledge, this is the first work on this topic based on deep learning. An effective semantic segmentation network, called SWNet, is novelly proposed to extract the potential splashed water regions. An encoder-decoder structure is designed to capture the visual characteristics of splashed water. SWNet achieves high efficiency by reusing pooling indices and adopting the light-weight decoder. With the multi-scale feature fusion structure, SWNet integrates the coarse semantic information and detailed appearance information, which significantly boosts the accuracy and refines the edge segmentation. A weighted cross entropy loss for splashed water is adopted to cope with the unbalanced distribution between splashed water and backgrounds. Moreover, a splashed water attention module is designed to focus on the salient regions of moving vehicles and splashed water, by performing attention mechanism to integrate global contextual information in semantic segmentation. Experiments conducted on a newly collected splashed water dataset demonstrate the effectiveness and efficiency of the proposed approach, which outperforms the state-of-the-art methods. Jian-Jun Qiao, Xiao Wu 0001, Jun-Yan He, Wei Li 0110, Qiang Peng |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2021 | Vehicle Counting Network with Attention-based Mask Refinement and Spatial-awareness Block LossabstractVehicle counting aims to calculate the number of vehicles in congested traffic scenes. Although object detection and crowd counting have made tremendous progress with the development of deep learning, vehicle counting remains a challenging task, due to scale variations, viewpoint changes, inconsistent location distributions, diverse visual appearances and severe occlusions. In this paper, a well-designed Vehicle Counting Network (VCNet) is novelly proposed to alleviate the problem of scale variation and inconsistent spatial distribution in congested traffic scenes. Specifically, VCNet is composed of two major components: (i) To capture multi-scale vehicles across different types and camera viewpoints, an effective multi-scale density map estimation structure is designed by building an attention-based mask refinement module. The multi-branch structure with hybrid dilated convolution blocks is proposed to assign receptive fields to generate multi-scale density maps. To efficiently aggregate multi-scale density maps, the attention-based mask refinement is well-designed to highlight the vehicle regions, which enables each branch to suppress the scale interference from other branches. (ii) In order to capture the inconsistent spatial distributions, a spatial-awareness block loss (SBL) based on the region-weighted reward strategy is proposed to calculate the loss of different spatial regions including sparse, congested and occluded regions independently by dividing the density map into different regions. Extensive experiments conducted on three benchmark datasets, TRANCOS, VisDrone2019 Vehicle and CVCSet demonstrate that the proposed VCNet outperforms the state-of-the-art approaches in vehicle counting. Moreover, the proposed idea can be applicable for crowd counting, which produces competitive results on ShanghaiTech crowd counting dataset. Ji Zhang 0027, Jian-Jun Qiao, Xiao Wu 0001, Wei Li 0110 |
ACM Multimedia | 4 |
| 2021 | Semantic-aware visual attributes learning for zero-shot recognition
Yurui Xie, Tiecheng Song, Wei Li 0110 |
J. Vis. Commun. Image Represent. | 3 |
| 2020 | CODAN: Counting-driven Attention Network for Vehicle Detection in Congested ScenesabstractAlthough recent object detectors have shown excellent performance for vehicle detection, they are incompetent for scenarios with a relatively large number of vehicles. In this paper, we explore the dense vehicle detection given the number of vehicles. Existing crowd counting methods cannot directly applied for dense vehicle detection due to insufficient description of density map, and the lack of effective constraint for mining the spatial awareness of dense vehicles. Inspired by these observations, a conceptually simple yet efficient framework, called CODAN, is proposed for dense vehicle detection. The proposed approach is composed of three major components: (i) an efficient strategy for generating multi-scale density maps (MDM) is designed to represent the vehicle counting, which can capture the global semantics and spatial information of dense vehicles, (ii) a multi-branch attention module (MAM) is proposed to bridging the gap between object counting and vehicle detection framework, (iii) with the well-designed density maps as explicit supervision, an effective counting-awareness loss (C-Loss) is employed to guide the attention learning by building the pixel-level constrain. Extensive experiments conducted on four benchmark datasets demonstrate that the proposed method outperforms the state-of-the-art methods. The impressive results indicate that vehicle detection and counting can be mutually supportive, which is an important and meaningful finding. Wei Li 0110, Zhenting Wang, Xiao Wu 0001, Ji Zhang 0027, Qiang Peng, Hongliang Li 0001 |
ACM Multimedia | 1 |
| 2020 | HeadNet: An End-to-End Adaptive Relational Network for Head DetectionabstractHead detection plays an important role in localizing and identifying persons from visual data. Most existing methods treat head detection as a specific form of object detection. Head detection is nontrivial due to the considerable difficulty in building the local and global information under conditions of unconstrained pose and orientation. To address these issues, this paper presents an effective adaptive relational network to capture context information, which is greatly helpful to suppress missed detection. We show that the fundamental contextual properties, such as the global shape priors from different heads and the local adjacent relationship between the head and shoulders, can be systematically quantified by visual operators. Specifically, we propose a two-step search algorithm to quantify the global intergroup conflict with adaptive scale, pose and viewpoint. Meanwhile, a structured feature module is introduced to capture the local relation of intraindividual stability. Finally, the global priors and local relation are integrated seamlessly into a single-stage head detector that is end-to-end trainable. An extensive ablation analysis demonstrates the effectiveness of our approach. We achieve state-of-the-art results on two challenging datasets, i.e., HollywoodHeads and Brainwash. Wei Li 0110, Hongliang Li 0001, Qingbo Wu 0001, Fanman Meng, Linfeng Xu 0001, King Ngi Ngan |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2018 | Adaptive Multi-Scale Information Flow for Object Detection
Wei Li 0110, Qingbo Wu 0001, Fanman Meng |
BMVC | 2 |
| 2017 | Store classification using Text-Exemplar-Similarity and Hypotheses-Weighted-CNN
Chao Huang 0003, Hongliang Li 0001, Wei Li 0110, Qingbo Wu 0001, Linfeng Xu 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2017 | Improving object proposals with top-down cues
Wei Li 0110, Hongliang Li 0001, Bing Luo 0003, Hengcan Shi, Qingbo Wu 0001, King Ngi Ngan |
Signal Process. Image Commun. | 1 |
| 2017 | Blind Image Quality Assessment Based on Rank-Order Regularized RegressionabstractBlind image quality assessment (BIQA) aims to estimate the subjective quality of a query image without access to the reference image. Existing learning-based methods typically train a regression function by minimizing the average error between subjective opinion scores and model predictions. However, minimizing average error does not necessarily lead to correct quality rank-orders between the test images, which is a highly desirable property of image quality models. In this paper, we propose a novel rank-order regularized regression model to address this problem. The key idea is to introduce a pairwise rank-order constraint into the maximum margin regression framework, aiming to better preserve the correct perceptual preference. To the best of our knowledge, this is the first attempt to incorporate rank-order constraints into margin-based quality regression model. By combing with a new local spatial structure feature, we achieve highly consistent quality prediction with human perception. Experimental results show that the proposed method outperforms many state-of-the-art BIQA metrics on popular publicly available IQA databases (i.e., LIVE-II, TID2013, VCL@FER, LIVEMD, and ChallengeDB). Qingbo Wu 0001, Hongliang Li 0001, Zhou Wang 0001, Fanman Meng, Bing Luo 0003, Wei Li 0110, King Ngi Ngan |
IEEE Trans. Multim. | 6 |
| 2016 | Person re-identification based on multi-region-set ensembles
Wei Li 0110, Chao Huang 0003, Bing Luo 0003, Fanman Meng, Tiecheng Song, Hengcan Shi |
J. Vis. Commun. Image Represent. | 1 |