VLDB 2026 Research / reviewers in the wild / expert
Wentong Li 0001
dblp:86/5922-1
· DBLP profile ↗
28ranked-venue papers
7as first author
25since 2021 · last 2026
0000-0002-2715-0995ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 6 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 3 first-author · 19 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Text-guided Controllable Diffusion for Realistic Camouflage Images GenerationabstractCamouflage Images Generation (CIG) is an emerging research area that focuses on synthesizing images in which objects are harmoniously blended and exhibit high visual consistency with their surroundings. Existing methods perform CIG by either fusing objects into specific backgrounds or outpainting the surroundings via foreground object-guided diffusion. However, they often fail to obtain natural results because they overlook the logical relationship between camouflaged objects and background environments. To address this issue, we propose CT-CIG, a Controllable Text-guided Camouflage Images Generation method that produces realistic and logically plausible camouflage images. Leveraging Large Visual Language Models (VLM), we design a Camouflage-Revealing Dialogue Mechanism (CRDM) to annotate existing camouflage datasets with high-quality text prompts. Subsequently, the constructed image-prompt pairs are utilized to finetune Stable Diffusion, incorporating a lightweight controller to guide the location and shape of camouflaged objects for enhanced camouflage scene fitness. Moreover, we design a Frequency Interaction Refinement Module (FIRM) to capture high-frequency texture features, facilitating the learning of complex camouflage patterns. Extensive experiments, including CLIPScore evaluation and camouflage effectiveness assessment, demonstrate the semantic alignment of our generated text prompts and CT-CIG's ability to produce photorealistic camouflage images. Yuhang Qian, Haiyan Chen 0001, Wentong Li 0001, Ningzhong Liu, Jie Qin 0004 |
AAAI | 3 |
| 2025 | Uncertainty-Instructed Structure Injection for Generalizable HD Map ConstructionabstractReliable high-definition (HD) map construction is crucial for the driving safety of autonomous vehicles. Although recent studies demonstrate improved performance, their generalization capability across unfamiliar driving scenes remains unexplored. To tackle this issue, we propose UIGenMap, an uncertainty-instructed structure injection approach for generalizable HD map vectorization, which concerns the uncertainty resampling in statistical distribution and employs explicit instance features to reduce excessive reliance on training data. Specifically, we introduce the perspective-view (PV) detection branch to obtain explicit structural features, in which the uncertainty-aware decoder is designed to dynamically sample probability distributions considering the difference in scenes. With probabilistic embedding and selection, UI2DPrompt is proposed to construct PV-learnable prompts. These PV prompts are integrated into the map decoder by designed hybrid injection to compensate for neglected instance structures. To ensure real-time inference, a lightweight Mimic Query Distillation is designed to learn from PV prompts, which can serve as an efficient alternative to the flow of PV branches. Extensive experiments on challenging geographically disjoint (geo-based) data splits demonstrate that our UIGen-Map achieves superior performance, with +5.7 mAP improvement on the nuScenes dataset. Source code is available at https://github.com/xiaolul2/UIGenMap. Ruizi Yang, Song Wang 0019, Wentong Li 0001, Junbo Chen, Jianke Zhu |
CVPR | 4 |
| 2025 | PointLoRA: Low-Rank Adaptation with Token Selection for Point Cloud LearningabstractSelf-supervised representation learning for point cloud has demonstrated effectiveness in improving pre-trained model performance across diverse tasks. However, as pre-trained models grow in complexity, fully fine-tuning them for downstream applications demands substantial computational and storage resources. Parameter-efficient fine-tuning (PEFT) methods offer a promising solution to mitigate these resource requirements, yet most current approaches rely on complex adapter and prompt mechanisms that increase tunable parameters. In this paper, we propose PointLoRA, a simple yet effective method that combines low-rank adaptation (LoRA) with multi-scale token selection to efficiently fine-tune point cloud models. Our approach embeds LoRA layers within the most parameter-intensive components of point cloud transformers, reducing the need for tunable parameters while enhancing global feature capture. Additionally, multi-scale token selection extracts critical local information to serve as prompts for downstream fine-tuning, effectively complementing the global context captured by LoRA. The experimental results across various pre-trained models and three challenging public datasets demonstrate that our approach achieves competitive performance with only 3.43% of the trainable parameters, making it highly effective for resource-constrained applications. Source code is available at: https://github.com/songw-zju/PointLoRA. Song Wang 0019, Lingdong Kong, Jianyun Xu, Chunyong Hu, Gongfan Fang, Wentong Li 0001, Jianke Zhu, Xinchao Wang |
CVPR | 7 |
| 2025 | Inst3D-LMM: Instance-Aware 3D Scene Understanding with Multi-modal Instruction TuningabstractDespite encouraging progress in 3D scene understanding, it remains challenging to develop an effective Large Multimodal Model (LMM) that is capable of understanding and reasoning in complex 3D environments. Most previous methods typically encode 3D point and 2D image features separately, neglecting interactions between 2D semantics and 3D object properties, as well as the spatial relationships within the 3D environment. This limitation not only hinders comprehensive representations of 3D scene, but also compromises training and inference efficiency. To address these challenges, we propose a unified Instance-aware 3DLarge Multi-modal Model (Inst3D-LMM) to deal with multiple 3D scene understanding tasks simultaneously. To obtain the fine-grained instance-level visual tokens, we first introduce a novel Multi-view Cross-Modal Fusion (MCMF) module to inject the multi-view 2D semantics into their corresponding 3D geometric features. For scene-level relation-aware tokens, we further present a 3D Instance Spatial Relation (3D-ISR) module to capture the intricate pairwise spatial relationships among objects. Additionally, we perform end-to-end multi-task instruction tuning simultaneously without the subsequent task-specific fine-tuning. Extensive experiments demonstrate that our approach outperforms the state-of-the-art methods across 3D scene understanding, reasoning and grounding tasks. Source code is available at: https://github.com/hanxunyu/Inst3D-LMM. Hanxun Yu, Wentong Li 0001, Song Wang 0019, Junbo Chen, Jianke Zhu |
CVPR | 2 |
| 2025 | VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLMabstractVideo Large Language Models (Video LLMs) have recently exhibited remarkable capabilities in general video understanding. However, they mainly focus on holistic comprehension and struggle with capturing fine-grained spatial and temporal details. Besides, the lack of high-quality object-level video instruction data and a comprehensive benchmark further hinders their advancements. To tackle these challenges, we introduce the VideoRefer Suite to empower Video LLM for finer-level spatial-temporal video understanding, i.e., enabling perception and reasoning on any objects throughout the video. Specially, we thoroughly develop VideoRefer Suite across three essential aspects: dataset, model, and benchmark. Firstly, we introduce a multi-agent data engine to meticulously curate a largescale, high-quality object-level video instruction dataset, termed VideoRefer-700K. Next, we present the VideoRefer model, which equips a versatile spatial-temporal object encoder to capture precise regional and sequential representations. Finally, we meticulously create a VideoRefer-Bench to comprehensively assess the spatial-temporal understanding capability of a Video LLM, evaluating it across various aspects. Extensive experiments and analyses demonstrate that our VideoRefer model not only achieves promising performance on video referring benchmarks but also facilitates general video understanding capabilities. Yuqian Yuan, Wentong Li 0001, Zesen Cheng, Boqiang Zhang, Xin Li 0056, Deli Zhao, Wenqiao Zhang, Yueting Zhuang, Jianke Zhu, Lidong Bing |
CVPR | 3 |
| 2025 | BEVFog: Enhancing Vision-Based Roadside 3D Object Detection Robustness Under Foggy ConditionsabstractRoadside cameras play a crucial role in extending the perception range of intelligent transportation infrastructures, offering a cost-effective means for long-range 3D object detection. However, under foggy conditions, the visibility degradation severely diminishes the performance of vision-based roadside perception systems, as reduced contrast and distorted monocular cues impair bird’s-eye-view (BEV) depth estimation. To address this critical challenge, we propose BEVFog, a novel end-to-end framework that simultaneously restores fog-degraded semantic features and leverages fog density as an auxiliary depth prior to enhance detection robustness. BEVFog integrates two key components: a lightweight, region-aware defogging module with CNN-predicted, differentiable filters, and a depth adjustment module that uses Fog-Cue Conditional Kernels (FCK) to modulate BEV features based on local scattering, correcting depth distortions and enhancing geometric cues. To facilitate rigorous evaluation, we construct Foggy DAIR-V2X-I, a synthetic benchmark that introduces ten fog levels into the large-scale DAIR-V2X-I roadside dataset using an atmospheric scattering model. Extensive experiments demonstrate that BEVFog narrows the performance gap between heavy-fog and clear-weather conditions from over 10% AP to less than 2% AP for vehicles, while achieving up to 8% AP improvements for pedestrians and cyclists compared to the best dehazing-plus-detection baselines. This work paves the way for safer, all-weather autonomous infrastructure perception, and highlights the importance of integrating low-level restoration with high-level 3D reasoning in adverse conditions. Gaoyuan Miao, Wentong Li 0001, Rong Quan, Jie Qin 0004 |
ECAI | 2 |
| 2025 | Efficient Semi-DETR: Real Time End-to-End Semi-Supervised Object DetectionabstractAlthough DETR-based methods have achieved considerable success in semi-supervised object detection (SSOD), several challenges remain unresolved: (1) To obtain higher-quality pseudo-labels for training, teacher models often adopt larger-scale architectures, which severely impacts inference speed. In contrast, lightweight detectors may compromise the effectiveness of existing SSOD methods. (2) Bipartite matching utilizing the Hungarian algorithm does not fully leverage potentially valuable pseudo-labels. (3) Current methods alleviate the negative impact of low-quality pseudo-labels through one-to-many assignment, yet this often leads to issues such as duplicate detections. To tackle these challenges, we propose a lightweight end-to-end semi-supervised object detection framework called Efficient Semi-DETR. Specifically, to improve accuracy while reducing inference latency, we introduce a heterogeneous teacher-student framework, which leverages a Collaborative Auxiliary Head to better mine potentially valuable pseudo-labels. To tackle the issue of duplicate detections, we propose Efficient Query Matching to improve training efficiency and enhance the detection of small objects. Notably, Efficient Semi-DETR achieves 44.92 mAP with only 10% of the annotated MS-COCO data, surpassing state-of-the-art methods, while its inference latency is only a quarter of existing methods. Zihao Xin, Wentong Li 0001, Jie Qin 0004, Shengjun Huang |
ECAI | 2 |
| 2025 | Reliable and Calibrated Semantic Occupancy Prediction by Hybrid Uncertainty LearningabstractVision-centric semantic occupancy prediction plays a crucial role in autonomous driving, which requires accurate and reliable predictions from low-cost sensors. Although having notably narrowed the accuracy gap with LiDAR, there is still few research effort to explore the reliability and calibration in predicting semantic occupancy from camera. In this paper, we conduct a comprehensive evaluation of existing semantic occupancy prediction models from a reliability perspective for the first time. Despite the gradual alignment of camera-based models with LiDAR in terms of accuracy, a significant reliability gap still persists. To address this concern, we propose ReliOcc, a method designed to enhance the reliability of camera-based occupancy networks. ReliOcc provides a plug-and-play scheme for existing models, which integrates hybrid uncertainty from individual voxels with sampling-based noise and relative voxels through mix-up learning. Besides, an uncertainty-aware calibration strategy is devised to further improve model reliability in offline mode. Extensive experiments under various settings demonstrate that ReliOcc significantly enhances the reliability of learned model while maintaining the accuracy for both geometric and semantic predictions. Notably, our proposed approach exhibits robustness to sensor failures and out of domain noises during inference. Song Wang 0019, Zhongdao Wang, Wentong Li 0001, Bailan Feng, Junbo Chen, Jianke Zhu |
IJCAI | 4 |
| 2025 | DCN: Decoupled-Coupled Network for Text-based Person SearchabstractText-based person search aims to identify a person based on textual descriptions, by simultaneously addressing person detection and cross-modal alignment between text queries and person images. Existing approaches often struggle with conflicts in exploiting proposals across these two sub-tasks. Specifically, cross-modal alignment requires highly precise proposals, while person detection can tolerate a certain degree of proposal inaccuracy but always needs a large number of proposals. In this paper, we propose the Decoupled-Coupled Network (DCN) to tackle the above conflicts. We first attempt to resolve the above conflicts by proposing a Decoupled Proposal Selection (DPS) strategy, inspired by the divide-and-conquer principle. DPS adaptively selects the most suitable proposals for each sub-task, ensuring their distinct requirements are adequately met. We further present a Coupled Cascade Refinement (CCR) module to jointly optimize both sub-tasks in a multi-stage manner, progressively improving detection accuracy and fostering cross-modal alignment between text and person image. In addition, we introduce two types of objective functions to optimize the inherently multi-positive contrastive learning challenge. Extensive experiments conducted on two benchmarks demonstrate the effectiveness and superiority of our DCN over existing competitors. Rong Quan, Liangxu Su, Wentong Li 0001, Yichao Yan, Jie Qin 0004 |
MMAsia | 5 |
| 2025 | MUVR: A Multi-Modal Untrimmed Video Retrieval Benchmark with Multi-Level Visual CorrespondenceabstractWe propose the Multi-modal Untrimmed Video Retrieval task, along with a new benchmark (MUVR) to advance video retrieval for long-video platforms. MUVR aims to retrieve untrimmed videos containing relevant segments using multi-modal queries. It has the following features: 1) Practical retrieval paradigm: MUVR supports video-centric multi-modal queries, expressing fine-grained retrieval needs through long text descriptions, video tag prompts, and mask prompts. It adopts a one-to-many retrieval paradigm and focuses on untrimmed videos, tailored for long-video platform applications. 2) Multi-level visual correspondence: To cover common video categories (e.g., news, travel, dance) and precisely define retrieval matching criteria, we construct multi-level visual correspondence based on core video content (e.g., news events, travel locations, dance moves) which users are interested in and want to retrieve. It covers six levels: copy, event, scene, instance, action, and others. 3) Comprehensive evaluation criteria: We develop 3 versions of MUVR (i.e., Base, Filter, QA). MUVR-Base/Filter evaluates retrieval models, while MUVR-QA assesses MLLMs in a question-answering format. We also propose a Reranking Score to evaluate the reranking ability of MLLMs. MUVR consists of 53K untrimmed videos from the video platform Bilibili, with 1,050 multi-modal queries and 84K matches. Extensive evaluations of 3 state-of-the-art video retrieval models, 6 image-based VLMs, and 10 MLLMs are conducted. MUVR reveals the limitations of retrieval methods in processing untrimmed videos and multi-modal queries, as well as MLLMs in multi-video understanding and reranking. Our code and benchmark is available at https://github.com/debby-0527/MUVR. Qijia Lu, Jiawei Niu, Qingzhi He, Shiping Ge, Ethan Q. Chen, Wentong Li 0001, Limin Wang 0002, Jie Qin 0004 |
NeurIPS | 12 |
| 2025 | EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World?abstractThe emergence of multimodal large language models (MLLMs) has driven breakthroughs in egocentric vision applications. These applications necessitate persistent, context-aware understanding of objects, as users interact with tools in dynamic and cluttered environments. However, existing embodied benchmarks primarily focus on static scene exploration, emphasizing object's appearance and spatial attributes while neglecting the assessment of dynamic changes arising from users' interactions.capabilities in object-level spatiotemporal reasoning required for real-world interactions.To address this gap, we introduce EOC-Bench, an innovative benchmark designed to systematically evaluate object-centric embodied cognition in dynamic egocentric scenarios.Specially, EOC-Bench features 3,277 meticulously annotated QA pairs categorized into three temporal categories: Past, Present, and Future, covering 11 fine-grained evaluation dimensions and 3 visual object referencing types.To ensure thorough assessment, we develop a mixed-format human-in-the-loop annotation frameworkBased on EOC-Bench, we conduct comprehensive evaluations of various proprietary, open-source, and object-level MLLMs. EOC-Bench serves as a crucial tool for advancing the embodied object cognitive capabilities of MLLMs, establishing a robust foundation for developing reliable core models for embodied systems. Yuqian Yuan, Ronghao Dang, Wentong Li 0001, Xin Li 0056, Deli Zhao, Fan Wang 0019, Wenqiao Zhang, Jun Xiao 0001, Yueting Zhuang |
NeurIPS | 4 |
| 2025 | Large Models are Good Annotators for Zero-Shot LearningabstractHuman-annotated attributes serve as effective semantic label embeddings for zero-shot learning (ZSL); however, their annotation is labor-intensive and difficult to scale. Recent studies have explored weakly supervised semantic label embeddings to reduce human effort, but these methods often fail to capture visual similarity and underperform compared to human-annotated semantics. In this work, we propose a minimally supervised yet effective approach: GPT- and CLIP-powered attributes (GCAtt). Specifically, we introduce a three-step interaction process with ChatGPT-comprising preliminary design, hierarchical refinement, and specific value determination-to generate attributes that are both category-shared and discriminative for classification. Additionally, we develop a method that encodes attributes and their values as potential text pairings, leveraging CLIP's retrieval capabilities for annotation. Experimental results on four widely used benchmarks demonstrate that GCAtt consistently outperforms human-annotated semantics. Code and data are available at https://github.com/RowenaHe/GCAtt. Qingzhi He, Wentong Li 0001, Shengcai Liao, Rong Quan, Tong Cui, Jie Qin 0004 |
SIGIR | 3 |
| 2025 | TokenPacker: Efficient Visual Projector for Multimodal LLM
Wentong Li 0001, Yuqian Yuan, Jian Liu 0012, Dongqi Tang, Song Wang 0019, Jie Qin 0004, Jianke Zhu, Lei Zhang 0006 |
Int. J. Comput. Vis. | 1 |
| 2024 | Fine-Grained Multi-View Hand Reconstruction Using Inverse RenderingabstractReconstructing high-fidelity hand models with intricate textures plays a crucial role in enhancing human-object interaction and advancing real-world applications. Despite the state-of-the-art methods excelling in texture generation and image rendering, they often face challenges in accurately capturing geometric details. Learning-based approaches usually offer better robustness and faster inference, which tend to produce smoother results and require substantial amounts of training data. To address these issues, we present a novel fine-grained multi-view hand mesh reconstruction method that leverages inverse rendering to restore hand poses and intricate details. Firstly, our approach predicts a parametric hand mesh model through Graph Convolutional Networks (GCN) based method from multi-view images. We further introduce a novel Hand Albedo and Mesh (HAM) optimization module to refine both the hand mesh and textures, which is capable of preserving the mesh topology. In addition, we suggest an effective mesh-based neural rendering scheme to simultaneously generate photo-realistic image and optimize mesh geometry by fusing the pre-trained rendering network with vertex features. We conduct the comprehensive experiments on InterHand2.6M, DeepHandMesh and dataset collected by ourself, whose promising results show that our proposed approach outperforms the state-of-the-art methods on both reconstruction accuracy and rendering quality. Code and dataset are publicly available at https://github.com/agnJason/FMHR. Qijun Gan, Wentong Li 0001, Jinwei Ren, Jianke Zhu |
AAAI | 2 |
| 2024 | MGMap: Mask-Guided Learning for Online Vectorized HD Map ConstructionabstractCurrently, high-definition (HD) map construction leans towards a lightweight online generation tendency, which aims to preserve timely and reliable road scene information. However, map elements contain strong shape priors. Subtle and sparse annotations make current detection-based frameworks ambiguous in locating relevant feature scopes and cause the loss of detailed structures in prediction. To alleviate these problems, we propose MGMap, a mask-guided approach that effectively highlights the informative regions and achieves precise map element localization by introducing the learned masks. Specifically, MGMap employs learned masks based on the enhanced multi-scale BEV features from two perspectives. At the instance level, we propose the Mask-activated instance (MAI) decoder, which incorporates global instance and structural information into instance queries by the activation of instance masks. At the point level, a novel position-guided mask patch refinement (PG-MPR) module is designed to refine point locations from a finer-grained perspective, enabling the extraction of point-specific patch information. Compared to the baselines, our proposed MGMap achieves a notable improvement of around 10 mAP for different input modalities. Extensive experiments also demonstrate that our approach showcases strong robustness and generalization capabilities. Our code can be found at https://github.com/xiaolul2/MGMap. Song Wang 0019, Wentong Li 0001, Ruizi Yang, Junbo Chen, Jianke Zhu |
CVPR | 3 |
| 2024 | Not All Voxels are Equal: Hardness-Aware Semantic Scene Completion with Self-DistillationabstractSemantic scene completion, also known as semantic oc-cupancy prediction, can provide dense geometric and semantic information for autonomous vehicles, which attracts the increasing attention of both academia and industry. Un-fortunately, existing methods usually formulate this task as a voxel-wise classification problem and treat each voxel equally in 3D space during training. As the hard voxels have not been paid enough attention, the performance in some challenging regions is limited. The 3D dense space typically contains a large number of empty voxels, which are easy to learn but require amounts of computation due to handling all the voxels uniformly for the existing models. Further-more, the voxels in the boundary region are more challenging to differentiate than those in the interior. In this paper, we propose HASSC approach to train the semantic scene completion model with hardness-aware design. The global hardness from the network optimization process is defined for dynamical hard voxel selection. Then, the local hard-ness with geometric anisotropy is adopted for voxel- wise refinement. Besides, self-distillation strategy is introduced to make training process stable and consistent. Extensive experiments show that our HASSC scheme can effectively promote the accuracy of the baseline model without incur-ring the extra inference cost. Source code is available at: https://github.com/songw-zju/HASSC. Song Wang 0019, Wentong Li 0001, Wenyu Liu 0005, Junbo Chen, Jianke Zhu |
CVPR | 3 |
| 2024 | Osprey: Pixel Understanding with Visual Instruction TuningabstractMultimodal large language models (MLLMs) have recently achieved impressive general-purpose vision-language capabilities through visual instruction tuning. However, current MLLMs primarily focus on image-level or box-level understanding, falling short in achieving fine-grained vision-language alignment at pixel level. Besides, the lack of mask-based instruction data limits their ad-vancements. In this paper, we propose Osprey, a mask-text instruction tuning approach, to extend MLLMs by incor-porating fine-grained mask regions into language instruction, aiming at achieving pixel-wise visual understanding. To achieve this goal, we first meticulously curate a mask-based region-text dataset with 724K samples, and then design a vision-language model by injecting pixel-level representation into LLM. Specifically, Osprey adopts a convolutional CLIP backbone as the vision encoder and employs a mask-aware visual extractor to extract precise visual mask features from high resolution input. Experimen-tal results demonstrate Osprey's superiority in various region understanding tasks, showcasing its new capability for pixel-level instruction tuning. In particular, Osprey can be integrated with Segment Anything Model (SAM) seamlessly to obtain multi-granularity semantics. The source code, dataset and demo can be found at https://github.com/CircleRadon/Osprey. Yuqian Yuan, Wentong Li 0001, Jian Liu 0012, Dongqi Tang, Xinjie Luo, Chi Qin, Lei Zhang 0006, Jianke Zhu |
CVPR | 2 |
| 2024 | Label-efficient Semantic Scene Completion with Scribble Annotations
Song Wang 0019, Wentong Li 0001, Hao Shi 0004, Kailun Yang 0001, Junbo Chen, Jianke Zhu |
IJCAI | 3 |
| 2024 | Box2Mask: Box-Supervised Instance Segmentation via Level-Set EvolutionabstractIn contrast to fully supervised methods using pixel-wise mask labels, box-supervised instance segmentation takes advantage of simple box annotations, which has recently attracted increasing research attention. This paper presents a novel single-shot instance segmentation approach, namely Box2Mask, which integrates the classical level-set evolution model into deep neural network learning to achieve accurate mask prediction with only bounding box supervision. Specifically, both the input image and its deep features are employed to evolve the level-set curves implicitly, and a local consistency module based on a pixel affinity kernel is used to mine the local context and spatial relations. Two types of single-stage frameworks, i.e., CNN-based and transformer-based frameworks, are developed to empower the level-set evolution for box-supervised instance segmentation, and each framework consists of three essential components: instance-aware decoder, box-level matching assignment and level-set evolution. By minimizing the level-set energy function, the mask map of each instance can be iteratively optimized within its bounding box annotation. The experimental results on five challenging testbeds, covering general scenes, remote sensing, medical and scene text images, demonstrate the outstanding performance of our proposed Box2Mask approach for box-supervised instance segmentation. In particular, with the Swin-Transformer large backbone, our Box2Mask obtains 42.4% mask AP on COCO, which is on par with the recently developed fully mask-supervised methods. Wentong Li 0001, Wenyu Liu 0005, Jianke Zhu, Miaomiao Cui, Risheng Yu, Xian-Sheng Hua 0001, Lei Zhang 0006 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | LiDAR2Map: In Defense of LiDAR-Based Semantic Map Construction Using Online Camera DistillationabstractSemantic map construction under bird's-eye view (BEV) plays an essential role in autonomous driving. In contrast to camera image, LiDAR provides the accurate 3D observations to project the captured 3D features onto BEV space inherently. However, the vanilla LiDAR-based BEV feature often contains many indefinite noises, where the spatial features have little texture and semantic cues. In this paper, we propose an effective LiDAR-based method to build semantic map. Specifically, we introduce a BEV pyramid feature decoder that learns the robust multi-scale BEV features for semantic map construction, which greatly boosts the accuracy of the LiDAR-based method. To mitigate the defects caused by lacking semantic cues in LiDAR data, we present an online Camera-to-LiDAR distillation scheme to facilitate the semantic learning from image to point cloud. Our distillation scheme consists of feature-level and log it-level distillation to absorb the semantic information from camera in BEV. The experimental results on challenging nuScenes dataset demonstrate the efficacy of our proposed LiDAR2Map on semantic map construction, which significantly outperforms the previous LiDAR-based methods over 27.9% mIoU and even performs better than the state-of-the-art camera-based approaches. Source code is available at: https://github.com/songw-zjuILiDAR2Map. Song Wang 0019, Wentong Li 0001, Wenyu Liu 0005, Jianke Zhu |
CVPR | 2 |
| 2023 | Point2Mask: Point-supervised Panoptic Segmentation via Optimal TransportabstractWeakly-supervised image segmentation has recently attracted increasing research attentions, aiming to avoid the expensive pixel-wise labeling. In this paper, we present an effective method, namely Point2Mask, to achieve high-quality panoptic prediction using only a single random point annotation per target for training. Specifically, we formulate the panoptic pseudo-mask generation as an Optimal Transport (OT) problem, where each ground-truth (gt) point label and pixel sample are defined as the label supplier and consumer, respectively. The transportation cost is calculated by the introduced task-oriented maps, which focus on the category-wise and instance-wise differences among the various thing and stuff targets. Furthermore, a centroid-based scheme is proposed to set the accurate unit number for each gt point supplier. Hence, the pseudo-mask generation is converted into finding the optimal transport plan at a globally minimal transportation cost, which can be solved via the Sinkhorn-Knopp Iteration. Experimental results on Pascal VOC and COCO demonstrate the promising performance of our proposed Point2Mask approach to point-supervised panoptic segmentation. Source code is available at: https://github.com/LiWentomng/Point2Mask. Wentong Li 0001, Yuqian Yuan, Song Wang 0019, Jianke Zhu, Jianshu Li, Jian Liu 0012, Lei Zhang 0006 |
ICCV | 1 |
| 2023 | Label-efficient Segmentation via Affinity PropagationabstractWeakly-supervised segmentation with label-efficient sparse annotations has attracted increasing research attention to reduce the cost of laborious pixel-wise labeling process, while the pairwise affinity modeling techniques play an essential role in this task. Most of the existing approaches focus on using the local appearance kernel to model the neighboring pairwise potentials. However, such a local operation fails to capture the long-range dependencies and ignores the topology of objects. In this work, we formulate the affinity modeling as an affinity propagation process, and propose a local and a global pairwise affinity terms to generate accurate soft pseudo labels. An efficient algorithm is also developed to reduce significantly the computational cost. The proposed approach can be conveniently plugged into existing segmentation networks. Experiments on three typical label-efficient segmentation tasks, i.e. box-supervised instance segmentation, point/scribble-supervised semantic segmentation and CLIP-guided semantic segmentation, demonstrate the superior performance of the proposed approach. Wentong Li 0001, Yuqian Yuan, Song Wang 0019, Wenyu Liu 0005, Dongqi Tang, Jian Liu 0012, Jianke Zhu, Lei Zhang 0006 |
NeurIPS | 1 |
| 2023 | Improving Nighttime Driving-Scene Segmentation via Dual Image-Adaptive Learnable FiltersabstractSemantic segmentation on driving-scene images is vital for autonomous driving. Although encouraging performance has been achieved on daytime images, the performance on nighttime images are less satisfactory due to the insufficient exposure and the lack of labeled data. To address these issues, we present an add-on module called dual image-adaptive learnable filters (DIAL-Filters) to improve the semantic segmentation in nighttime driving conditions, aiming at exploiting the intrinsic features of driving-scene images under different illuminations. DIAL-Filters consist of two parts, including an image-adaptive processing module (IAPM) and a learnable guided filter (LGF). With DIAL-Filters, we design both unsupervised and supervised frameworks for nighttime driving-scene segmentation, which can be trained in an end-to-end manner. Specifically, the IAPM module consists of a small convolutional neural network with a set of differentiable image filters, where each image can be adaptively enhanced for better segmentation with respect to the different illuminations. The LGF is employed to enhance the output of segmentation network to get the final segmentation result. The DIAL-Filters are light-weight and efficient and they can be readily applied for both daytime and nighttime images. Our experiments show that DAIL-Filters can significantly improve the supervised segmentation performance on ACDC_Night and NightCity datasets, while it demonstrates the state-of-the-art performance on unsupervised nighttime semantic segmentation on Dark Zurich and Nighttime Driving testbeds. Codes and models are available athttps://github.com/wenyyu/IA-Seg. Wenyu Liu 0005, Wentong Li 0001, Jianke Zhu, Miaomiao Cui, Xuansong Xie, Lei Zhang 0006 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Oriented RepPoints for Aerial Object DetectionabstractIn contrast to the generic object, aerial targets are often non-axis aligned with arbitrary orientations having the cluttered surroundings. Unlike the mainstreamed approaches regressing the bounding box orientations, this paper proposes an effective adaptive points learning approach to aerial object detection by taking advantage of the adaptive points representation, which is able to capture the geometric information of the arbitrary-oriented instances. To this end, three oriented conversion functions are presented to facilitate the classification and localization with accurate orientation. Moreover, we propose an effective quality assessment and sample assignment scheme for adaptive points learning toward choosing the representative oriented reppoints samples during training, which is able to capture the non-axis aligned features from adjacent objects or background noises. A spatial constraint is introduced to penalize the outlier points for roust adaptive learning. Experimental results on four challenging aerial datasets including DOTA, HRSC2016, UCAS-AOD and DIOR-R, demonstrate the efficacy of our proposed approach. The source code is availabel at: https://github.com/LiWentomng/OrientedRepPoints. Wentong Li 0001, Kaixuan Hu, Jianke Zhu |
CVPR | 1 |
| 2022 | Box-Supervised Instance Segmentation with Level Set Evolution
Wentong Li 0001, Wenyu Liu 0001, Jianke Zhu, Miaomiao Cui, Xian-Sheng Hua 0001, Lei Zhang 0006 |
ECCV (29) | 1 |
| 2019 | S3OD: Single Stage Small Object Detector from Scratch for Remote Sensing Images
Feng Yang 0001, Wentong Li 0001, Wanyi Li 0002, Peng Wang 0024 |
ICIG (3) | 2 |
| 2019 | Multi-Scale Object Detection in Satellite Imagery Based On YOLTabstractMulti-scale object detection (MOD) is one of the remaining challenges for satellite imagery. To improve the performance of MOD task, YOLT (You Only Look Twice) has achieved a good accuracy in high resolution remote sensing images. Motivated by the state-of-art object detection method for satellite imagery, we explored and achieved the state-of-the-art accuracy based on the standard YOLT for MOD task by providing a novel method with enough experimental results and model comparison on the typical multi-scale satellite imagery dataset. First, we divide objects into three categories according to the scale of objects. Then, different training strategies are used to train the classifier and detector for different scale objects. Finally, multi-scale detection chips are stitched and fused to get more accurate localization and classification as the final predicted results for MOD in satellite imagery. Experiments have been conducted over dataset from the second stage of AIIA1Cup Competition of Typical Object Recognition for Satellite Imagery in Small Samples compared with the standard YOLT and Faster R-CNN, which demonstrates the effectiveness and the comparable detection performance of our proposed pipeline. Wentong Li 0001, Wanyi Li 0002, Feng Yang 0001, Peng Wang 0024 |
IGARSS | 1 |
| 2018 | Unscented Particle Double Layer FilterabstractThe Particle filter (PF) provides a general numerical tool to deal with the non-Gaussian filtering problems, but it has the particle depletion problem and so on. The unscented particle filter (UPF) can solve the problem of particle depletion, but it has the computationally intensive problem and so on. To overcome these problems, the unscented particle double layer filter (UPDLF) is proposed. The proposed algorithm uses the PF algorithm to replace the state transition density function in the UKF algorithm, and updates the weights of each deterministic sampling point based on the new measurements. Finally, the state estimation at each time is obtained. The numerical simulation with two examples shows that the proposed filter outperforms the PF algorithm and the UPF algorithm. Feng Yang 0001, Litao Zheng, Wentong Li 0001, Yongting Wang, Pengxiang Wang 0001 |
FUSION | 3 |