Zhigang Wang 0002

dblp:35/1989-2 · DBLP profile ↗
← Back
36ranked-venue papers
2as first author
31since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 30 · 1 first-author · 28 since 2021Graphics, computer vision, multimedia, augmented reality and games · 26 · 2 first-author · 23 since 2021Systems, architecture and hardware · 4 · 4 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 EvoMoE: Expert Evolution in Mixture of Experts for Multimodal Large Language Models
abstract
Recent advancements have shown that the Mixture of Experts (MoE) approach significantly enhances the capacity of large language models (LLMs) and improves performance on downstream tasks. Building on these promising results, multi-modal large language models (MLLMs) have increasingly adopted MoE techniques. However, existing multi-modal MoE tuning methods typically face two key challenges: expert uniformity and router rigidity. Expert uniformity occurs because MoE experts are often initialized by simply replicating the FFN parameters from LLMs, leading to homogenized expert functions and weakening the intended diversification of the MoE architecture. Meanwhile, router rigidity stems from the prevalent use of static linear routers for expert selection, which fail to distinguish between visual and textual tokens, resulting in similar expert distributions for image and text. To address these limitations, we propose EvoMoE, an innovative MoE tuning framework. EvoMoE introduces a meticulously designed expert initialization strategy that progressively evolves multiple robust experts from a single trainable expert, a process termed expert evolution that specifically targets severe expert homogenization. Furthermore, we introduce the Dynamic Token-aware Router (DTR), a novel routing mechanism that allocates input tokens to appropriate experts based on their modality and intrinsic token values. This dynamic routing is facilitated by hypernetworks, which dynamically generate routing weights tailored for each individual token. Extensive experiments demonstrate that EvoMoE significantly outperforms other sparse MLLMs across a variety of multi-modal benchmarks, including MME, MMBench, TextVQA, and POPE. Our results highlight the effectiveness of EvoMoE in enhancing the performance of MLLMs by addressing the critical issues of expert uniformity and router rigidity.
Linglin Jing, Zhigang Wang 0002, Wang Lan, Weiyun Wang, Wenhai Wang, Qingpei Guo
AAAI3
2025 Learning 2D Invariant Affordance Knowledge for 3D Affordance Grounding
abstract
3D Object Affordance Grounding aims to predict the functional regions on a 3D object and has laid the foundation for a wide range of applications in robotics. Recent advances tackle this problem via learning a mapping between 3D regions and a single human-object interaction image. However, the geometric structure of the 3D object and the object in the human-object interaction image are not always consistent, leading to poor generalization. To address this issue, we propose to learn generalizable invariant affordance knowledge from multiple human-object interaction images within the same affordance category. Specifically, we introduce the Multi-Image Guided Invariant-Feature-Aware 3D Affordance Grounding (MIFAG) framework. It grounds 3D object affordance regions by identifying common interaction patterns across multiple human-object interaction images. First, the Invariant Affordance Knowledge Extraction Module (IAM) utilizes an iterative updating strategy to gradually extract aligned affordance knowledge from multiple images and integrate it into an affordance dictionary. Then, the Affordance Dictionary Adaptive Fusion Module (ADM) learns comprehensive point cloud representations that consider all affordance candidates in multiple images. Besides, the Multi-Image and Point Affordance (MIPA) benchmark is constructed and our method outperforms existing state-of-the-art methods on various experimental comparisons.
Xianqiang Gao 0001, Pingrui Zhang, Delin Qu, Dong Wang 0028, Zhigang Wang 0002, Yan Ding 0002, Bin Zhao 0001
AAAI5
2025 Efficient Diffusion as Low Light Enhancer
abstract
The computational burden of the iterative sampling process remains a major challenge in diffusion-based LowLight Image Enhancement (LLIE). Current acceleration methods, whether training-based or training-free, often lead to significant performance degradation, highlighting the trade-off between performance and efficiency. In this paper, we identify two primary factors contributing to performance degradation: fitting errors and the inference gap. Our key insight is that fitting errors can be mitigated by linearly extrapolating the incorrect score functions, while the inference gap can be reduced by shifting the Gaussian flow to a reflectance-aware residual space. Based on the above insights, we design Reflectance-Aware Trajectory Refinement (RATR) module, a simple yet effective module to refine the teacher trajectory using the reflectance component of images. Following this, we introduce Reflectance-aware Diffusion with Distilled Trajectory (ReDDiT), an efficient and flexible distillation framework tailored for LLIE. Our framework achieves comparable performance to previous diffusion-based methods with redundant steps in just 2 steps while establishing new state-of-the-art (SOTA) results with 8 or 4 steps. Comprehensive experimental evaluations on 10 benchmark datasets validate the effectiveness of our method, consistently outperforming existing SOTA methods. Code is available at project page.
Guanzhou Lan, Qianli Ma 0008, Zhigang Wang 0002, Dong Wang 0028, Xuelong Li 0001, Bin Zhao 0001
CVPR4
2025 Think Small, Act Big: Primitive Prompt Learning for Lifelong Robot Manipulation
abstract
Building a lifelong robot that can effectively leverage prior knowledge for continuous skill acquisition remains significantly challenging. Despite the success of experience replay and parameter-efficient methods in alleviating catastrophic forgetting problem, naively applying these methods causes a failure to leverage the shared primitives between skills. To tackle these issues, we propose Primitive Prompt Learning (PPL), to achieve lifelong robot manipulation via reusable and extensible primitives. Within our two stage learning scheme, we first learn a set of primitive prompts to represent shared primitives through multi-skills pre-training stage, where motion-aware prompts are learned to capture semantic and motion shared primitives across different skills. Secondly, when acquiring new skills in lifelong span, new prompts are concatenated and optimized with frozen pretrained prompts, boosting the learning via knowledge transfer from old skills to new ones. For evaluation, we construct a large-scale skill dataset and conduct extensive experiments in both simulation and real-world tasks, demonstrating PPL’s superior performance over state-of-the-art methods.
Yuanqi Yao, Siao Liu, Haoming Song, Delin Qu, Yan Ding 0002, Bin Zhao 0001, Zhigang Wang 0002, Xuelong Li 0001, Dong Wang 0028
CVPR8
2025 AerialVG: A Challenging Benchmark for Aerial Visual Grounding by Exploring Positional Relations
abstract
Visual grounding (VG) aims to localize target objects in an image based on natural language descriptions. In this paper, we propose AerialVG, a new task focusing on visual grounding from aerial views. Compared to traditional VG, AerialVG poses new challenges, \emph{e.g.}, appearance-based grounding is insufficient to distinguish among multiple visually similar objects, and positional relations should be emphasized. Besides, existing VG models struggle when applied to aerial imagery, where high-resolution images cause significant difficulties. To address these challenges, we introduce the first AerialVG dataset, consisting of 5K real-world aerial images, 50K manually annotated descriptions, and 103K objects. Particularly, each annotation in AerialVG dataset contains multiple target objects annotated with relative spatial relations, requiring models to perform comprehensive spatial reasoning. Furthermore, we propose an innovative model especially for the AerialVG task, where a Hierarchical Cross-Attention is devised to focus on target regions, and a Relation-Aware Grounding module is designed to infer positional relations. Experimental results validate the effectiveness of our dataset and method, highlighting the importance of spatial reasoning in aerial visual grounding. The code and dataset will be released.
Junli Liu, Zhigang Wang 0002, Chi Yan, Dong Wang 0028, Xuelong Li 0001, Bin Zhao 0001
ICCV3
2025 Open-Vocabulary Octree-Graph for 3D Scene Understanding
abstract
Open-vocabulary 3D scene understanding is indispensable for embodied agents. Recent works leverage pretrained vision-language models (VLMs) for object segmentation and project them to point clouds to build 3D maps. Despite progress, a point cloud is a set of unordered coordinates that requires substantial storage space and does not directly convey occupancy information or spatial relation, making existing methods inefficient for downstream tasks, e.g., path planning and text-based object retrieval. To address these issues, we propose \textbf{Octree-Graph}, a novel scene representation for open-vocabulary 3D scene understanding. Specifically, a Chronological Group-wise Segment Merging (CGSM) strategy and an Instance Feature Aggregation (IFA) algorithm are first designed to get 3D instances and corresponding semantic features. Subsequently, an adaptive-octree structure is developed that stores semantics and depicts the occupancy of an object adjustably according to its shape. Finally, the Octree-Graph is constructed where each adaptive-octree acts as a graph node, and edges describe the spatial relations among nodes. Extensive experiments on various tasks are conducted on several widely-used datasets, demonstrating the versatility and effectiveness of our method. Code is available \href{https://github.com/yifeisu/OV-Octree-Graph}{here}.
Zhigang Wang 0002, Yifei Su, Dong Wang 0028, Xuelong Li 0001, Bin Zhao 0001
ICCV1
2025 MoMa-Kitchen: A 100K+ Benchmark for Affordance-Grounded Last-Mile Navigation in Mobile Manipulation
abstract
In mobile manipulation, navigation and manipulation are often treated as separate problems, resulting in a significant gap between merely approaching an object and engaging with it effectively. Many navigation approaches primarily define success by proximity to the target, often overlooking the necessity for optimal positioning that facilitates subsequent manipulation. To address this, we introduce MoMa-Kitchen, a benchmark dataset comprising over 100k samples that provide training data for models to learn optimal final navigation positions for seamless transition to manipulation. Our dataset includes affordance-grounded floor labels collected from diverse kitchen environments, in which robotic mobile manipulators of different models attempt to grasp target objects amidst clutter. Using a fully automated pipeline, we simulate diverse real-world scenarios and generate affordance labels for optimal manipulation positions. Visual data are collected from RGB-D inputs captured by a first-person view camera mounted on the robotic arm, ensuring consistency in viewpoint during data collection. We also develop a lightweight baseline model, NavAff, for navigation affordance grounding that demonstrates promising performance on the MoMa-Kitchen benchmark. Our approach enables models to learn affordance-based final positioning that accommodates different arm types and platform heights, thereby paving the way for more robust and generalizable integration of navigation and manipulation in embodied AI. Project page: \href{https://momakitchen.github.io/}{https://momakitchen.github.io/}.
Pingrui Zhang, Xianqiang Gao 0001, Kehui Liu, Dong Wang 0028, Zhigang Wang 0002, Bin Zhao 0001, Yan Ding 0002, Xuelong Li 0001
ICCV6
2025 Cocube: a Tabletop Modular Multi-Robot Platform for Education and Research
abstract
This paper presents CoCube a tabletop modular robotics platform designed for robotics education and multirobot algorithm research. CoCube is characterized by its low cost low floors high ceilings and wide walls offering flexibility and broad applicability across various use cases. The platform comprises four key components: CoCube robots which integrate wireless communication movement and interaction; CoModules which provide versatile external functionality; CoMaps which enable high-precision localization via microdot patterns on regular printed paper; and CoTags for interaction. CoCube operates on MicroBlocks a blocks programming language for physical computing inspired by MIT Scratch a widely-used coding language with a simple visual interface that makes programming accessible to young learners. It offers users both flexibility and ease of use with advanced API support for more complex applications. This paper details the design of the CoCube platform and demonstrates its potential in both educational and research contexts.
Songyi Zhu, Zhonghan Tang, Jialing Han, Zemin Lin, Zhongrui You, John Maloney, Bernat Romagosa, Bin Zhao 0001, Zhigang Wang 0002, Zhinan Zhang, Xuelong Li 0001
ICRA12
2025 COHERENT: Collaboration of Heterogeneous Multi-Robot System with Large Language Models
abstract
Leveraging the powerful reasoning capabilities of large language models (LLMs), recent LLM-based robot task planning methods yield promising results. However, they mainly focus on single or multiple homogeneous robots on simple tasks. Practically, complex long-horizon tasks always require collaboration among multiple heterogeneous robots especially with more complex action spaces, which makes these tasks more challenging. To this end, we propose COHERENT, a novel LLM-based task planning framework for collaboration of heterogeneous multi-robot systems including quadrotors, robotic dogs, and robotic arms. Specifically, a Proposal-Execution-Feedback-Adjustment (PEFA) mechanism is designed to decompose and assign actions for individual robots, where a centralized task assigner makes a task planning proposal to decompose the complex task into subtasks, and then assigns subtasks to robot executors. Each robot executor selects a feasible action to implement the assigned subtask and reports self-reflection feedback to the task assigner for plan adjustment. The PEFA loops until the task is completed. Moreover, we create a challenging heterogeneous multi-robot task planning benchmark encompassing 100 complex long-horizon tasks. The experimental results show that our work surpasses the previous methods by a large margin in terms of success rate and execution efficiency. The experimental videos, code, and benchmark are released at https://github.com/MrKeee/COHERENT.
Kehui Liu, Dong Wang 0028, Zhigang Wang 0002, Xuelong Li 0001, Bin Zhao 0001
ICRA4
2024 Color Event Enhanced Single-Exposure HDR Imaging
abstract
Single-exposure high dynamic range (HDR) imaging aims to reconstruct the wide-range intensities of a scene by using its single low dynamic range (LDR) image, thus providing significant efficiency. Existing methods pay high attention to restoring the luminance by inversing the tone-mapping process, while the color in the over-/under-exposed area cannot be well restored due to the information loss of the single LDR image. To address this issue, we introduce color events into the imaging pipeline, which record asynchronous pixel-wise color changes in a high dynamic range, enabling edge-like scene perception under challenging lighting conditions. Specifically, we propose a joint framework that incorporates color events and a single LDR image to restore both content and color of an HDR image, where an exposureaware transformer (EaT) module is designed to propagate the informative hints, provided by the normal-exposed LDR regions and the event streams, to the missing areas. In this module, an exposure-aware mask is estimated to suppress distractive information and strengthen the restoration of the over-/under-exposed regions. To our knowledge, we are the first to use color events to enhance single-exposure HDR imaging. We also contribute corresponding datasets, consisting of synthesized datasets and a real-world dataset collected by a DAVIS346-color camera. The datasets can be found at https://www.kaggle.com/datasets/mengyaocui/ce-hdr. Extensive experiments demonstrate the effectiveness of the proposed method.
Zhigang Wang 0002, Dong Wang 0028, Bin Zhao 0001, Xuelong Li 0001
AAAI2
2024 X4D-SceneFormer: Enhanced Scene Understanding on 4D Point Cloud Videos through Cross-Modal Knowledge Transfer
abstract
The field of 4D point cloud understanding is rapidly developing with the goal of analyzing dynamic 3D point cloud sequences. However, it remains a challenging task due to the sparsity and lack of texture in point clouds. Moreover, the irregularity of point cloud poses a difficulty in aligning temporal information within video sequences. To address these issues, we propose a novel cross-modal knowledge transfer framework, called X4D-SceneFormer. This framework enhances 4D-Scene understanding by transferring texture priors from RGB sequences using a Transformer architecture with temporal relationship mining. Specifically, the framework is designed with a dual-branch architecture, consisting of an 4D point cloud transformer and a Gradient-aware Image Transformer (GIT). The GIT combines visual texture and temporal correlation features to offer rich semantics and dynamics for better point cloud representation. During training, we employ multiple knowledge transfer techniques, including temporal consistency losses and masked self-attention, to strengthen the knowledge transfer between modalities. This leads to enhanced performance during inference using single-modal 4D point cloud inputs. Extensive experiments demonstrate the superior performance of our framework on various 4D point cloud video understanding tasks, including action recognition, action segmentation and semantic segmentation. The results achieve 1st places, i.e., 85.3% (+7.9%) accuracy and 47.3% (+5.0%) mIoU for 4D action segmentation and semantic segmentation, on the HOI4D challenge, outperforming previous state-of-the-art by a large margin. We release the code at https://github.com/jinglinglingling/X4D.
Linglin Jing, Ying Xue 0003, Xu Yan 0005, Chaoda Zheng, Dong Wang 0028, Ruimao Zhang, Zhigang Wang 0002, Hui Fang 0003, Bin Zhao 0001, Zhen Li 0026
AAAI7
2024 Point-PEFT: Parameter-Efficient Fine-Tuning for 3D Pre-trained Models
abstract
The popularity of pre-trained large models has revolutionized downstream tasks across diverse fields, such as language, vision, and multi-modality. To minimize the adaption cost for downstream tasks, many Parameter-Efficient Fine-Tuning (PEFT) techniques are proposed for language and 2D image pre-trained models. However, the specialized PEFT method for 3D pre-trained models is still under-explored. To this end, we introduce Point-PEFT, a novel framework for adapting point cloud pre-trained models with minimal learnable parameters. Specifically, for a pre-trained 3D model, we freeze most of its parameters, and only tune the newly added PEFT modules on downstream tasks, which consist of a Point-prior Prompt and a Geometry-aware Adapter. The Point-prior Prompt adopts a set of learnable prompt tokens, for which we propose to construct a memory bank with domain-specific knowledge, and utilize a parameter-free attention to enhance the prompt tokens. The Geometry-aware Adapter aims to aggregate point cloud features within spatial neighborhoods to capture fine-grained geometric information through local interactions. Extensive experiments indicate that our Point-PEFT can achieve better performance than the full fine-tuning on various downstream tasks, while using only 5% of the trainable parameters, demonstrating the efficiency and effectiveness of our approach. Code is released at https://github.com/Ivan-Tang-3D/Point-PEFT.
Ray Zhang 0002, Zoey Guo, Xianzheng Ma, Bin Zhao 0001, Zhigang Wang 0002, Dong Wang 0028, Xuelong Li 0001
AAAI6
2024 HPL-ESS: Hybrid Pseudo-Labeling for Unsupervised Event-based Semantic Segmentation
abstract
Event-based semantic segmentation has gained popularity due to its capability to deal with scenarios under high-speed motion and extreme lighting conditions, which cannot be addressed by conventional RGB cameras. Since it is hard to annotate event data, previous approaches rely on event-to-image reconstruction to obtain pseudo labels for training. However, this will inevitably introduce noise, and learning from noisy pseudo labels, especially when generated from a single source, may reinforce the errors. This drawback is also called confirmation bias in pseudo-labeling. In this paper, we propose a novel hybrid pseudo-labeling framework for unsupervised event-based semantic segmentation, HPL-ESS, to alleviate the influence of noisy pseudo labels. Specifically, we first employ a plain unsupervised domain adaptation framework as our baseline, which can generate a set of pseudo labels through self-training. Then, we incorporate offline event-to-image re-construction into the framework, and obtain another set of pseudo labels by predicting segmentation maps on the re-constructed images. A noisy label learning strategy is designed to mix the two sets of pseudo labels and enhance the quality. Moreover, we propose a soft prototypical alignment (SPA) module to further improve the consistency of target domain features. Extensive experiments show that the proposed method outperforms existing state-of-the-art methods by a large margin on benchmarks (e.g., +5.88% accuracy, +10.32% mIoU on DSEC-Semantic dataset), and even surpasses several supervised methods.
Linglin Jing, Zhigang Wang 0002, Xu Yan 0005, Dong Wang 0028, Gerald Schaefer, Hui Fang 0003, Bin Zhao 0001, Xuelong Li 0001
CVPR4
2024 GS-SLAM: Dense Visual SLAM with 3D Gaussian Splatting
abstract
In this paper, we introduce GS-SLAM that first utilizes 3D Gaussian representation in the Simultaneous Localization and Mapping (SLAM) system. It facilitates a better bal-ance between efficiency and accuracy. Compared to recent SLAM methods employing neural implicit representations, our method utilizes a real-time differentiable splatting ren-dering pipeline that offers significant speedup to map opti-mization and RGB-D rendering. Specifically, we propose an adaptive expansion strategy that adds new or deletes noisy 3D Gaussians in order to efficiently reconstruct new observed scene geometry and improve the mapping of pre-viously observed areas. This strategy is essential to ex-tend 3D Gaussian representation to reconstruct the whole scene rather than synthesize a static object in existing meth-ods. Moreover, in the pose tracking process, an effective coarse-to-fine technique is designed to select reliable 3D Gaussian representations to optimize camera pose, resulting in runtime reduction and robust estimation. Our method achieves competitive performance compared with existing state-of-the-art real-time methods on the Replica, TUM-RGBD datasets. Project page: https://gs-slam.github.io/.
Chi Yan, Delin Qu, Dan Xu 0002, Bin Zhao 0001, Zhigang Wang 0002, Dong Wang 0028, Xuelong Li 0001
CVPR5
2024 Any2Point: Empowering Any-Modality Large Models for Efficient 3D Understanding
Ray Zhang 0002, Jiaming Liu 0003, Zoey Guo, Bin Zhao 0001, Zhigang Wang 0002, Peng Gao 0007, Hongsheng Li 0001, Dong Wang 0028, Xuelong Li 0001
ECCV (36)6
2024 A Coarse-to-Fine Reconstruction Framework for Non-Lambertian Photometric Stereo
abstract
Photometric stereo aims to regress object surface normal from a set of images observed under varying illuminations. Although existing methods have achieved promising results, the irregular high-frequency detail is ignored, especially in complex and tiny surface folds. To address this problem, a coarse-to-fine reconstruction framework is proposed for non-Lambertian photometric stereo. Specifically, a coarse network is designed to roughly predict object surface normal, which learns the mapping from observed images to coarse surface normal. Then, to deal with the high-frequency information loss, we introduce a fine network to extract high-frequency information by leveraging both coarse surface normal and observation images. Meanwhile, to provide more supervision, we design a reconstruction module to reconstruct observed images from predicted surface normal and illuminations. Extensive experiments have demonstrated that the proposed method outperforms existing works and restores high-frequency detail effectively. In addition, the proposed method promotes the robustness under sparse illuminations.
Zhigang Wang 0002, Peipei Gu, Bin Zhao 0001, Xuelong Li 0001
ICME1
2024 SAM-E: Leveraging Visual Foundation Model with Sequence Imitation for Embodied Manipulation
abstract
Acquiring a multi-task imitation policy in 3D manipulation poses challenges in terms of scene understanding and action prediction. Current methods employ both 3D representation and multi-view 2D representation to predict the poses of the robot’s end-effector. However, they still require a considerable amount of high-quality robot trajectories, and suffer from limited generalization in unseen tasks and inefficient execution in long-horizon reasoning. In this paper, we propose **SAM-E**, a novel architecture for robot manipulation by leveraging a vision-foundation model for generalizable scene understanding and sequence imitation for long-term action reasoning. Specifically, we adopt Segment Anything (SAM) pre-trained on a huge number of images and promptable masks as the foundation model for extracting task-relevant features, and employ parameter-efficient fine-tuning on robot data for a better understanding of embodied scenarios. To address long-horizon reasoning, we develop a novel multi-channel heatmap that enables the prediction of the action sequence in a single pass, notably enhancing execution efficiency. Experimental results from various instruction-following tasks demonstrate that SAM-E achieves superior performance with higher execution efficiency compared to the baselines, and also significantly improves generalization in few-shot adaptation to new tasks.
Chenjia Bai, Haoran He, Zhigang Wang 0002, Bin Zhao 0001, Xiu Li 0001, Xuelong Li 0001
ICML4
2024 Kinematic-aware Prompting for Generalizable Articulated Object Manipulation with LLMs
abstract
Generalizable articulated object manipulation is essential for home-assistant robots. Recent efforts focus on imitation learning from demonstrations or reinforcement learning in simulation, however, due to the prohibitive costs of real-world data collection and precise object simulation, it still remains challenging for these works to achieve broad adaptability across diverse articulated objects. Recently, many works have tried to utilize the strong in-context learning ability of Large Language Models (LLMs) to achieve generalizable robotic manipulation, but most of these researches focus on high-level task planning, sidelining low-level robotic control. In this work, building on the idea that the kinematic structure of the object determines how we can manipulate it, we propose a kinematic-aware prompting framework that prompts LLMs with kinematic knowledge of objects to generate low-level motion trajectory waypoints, supporting various object manipulation. To effectively prompt LLMs with the kinematic structure of different objects, we design a unified kinematic knowledge parser, which represents various articulated objects as a unified textual description containing kinematic joints and contact location. Building upon this unified description, a kinematic-aware planner model is proposed to generate precise 3D manipulation waypoints via a designed kinematic-aware chain-of-thoughts prompting method. Our evaluation spanned 48 instances across 16 distinct categories, revealing that our framework not only outperforms traditional methods on 8 seen categories but also shows a powerful zero-shot capability for 8 unseen articulated object categories with only 17 demonstrations. Moreover, the real-world experiments on 7 different object categories prove our framework’s adaptability in practical scenarios. Code is released at https://github.com/GeWu-Lab/LLM_articulated_object_manipulation.
Wenke Xia, Dong Wang 0028, Xincheng Pang, Zhigang Wang 0002, Bin Zhao 0001, Di Hu 0001, Xuelong Li 0001
ICRA4
2024 Depth Helps: Improving Pre-trained RGB-based Policy with Depth Information Injection
abstract
3D perception ability is crucial for generalizable robotic manipulation. While recent foundation models have made significant strides in perception and decision-making with RGB-based input, their lack of 3D perception limits their effectiveness in fine-grained robotic manipulation tasks. To address these limitations, we propose a Depth Information Injection (DI2) framework that leverages the RGB-Depth modality for policy fine-tuning, while relying solely on RGB images for robust and efficient deployment. Concretely, we introduce the Depth Completion Module (DCM) to extract the spatial prior knowledge related to depth information and generate virtual depth information from RGB inputs to aid policy deployment. Further, we propose the Depth-Aware Codebook (DAC) to eliminate noise and reduce the cumulative error from the depth prediction. In the inference phase, this framework employs RGB inputs and accurately predicted depth data to generate the manipulation action. We conduct experiments on simulated LIBERO environments and real-world scenarios, and the experiment results prove that our method could effectively enhance the pre-trained RGB-based policy with 3D perception ability for robotic manipulation. The website is released at https://gewu-lab.github.io/DepthHelps-IROS2024.
Xincheng Pang, Wenke Xia, Zhigang Wang 0002, Bin Zhao 0001, Di Hu 0001, Dong Wang 0028, Xuelong Li 0001
IROS3
2024 LiveScene: Language Embedding Interactive Radiance Fields for Physical Scene Control and Rendering
abstract
This paper scales object-level reconstruction to complex scenes, advancing interactive scene reconstruction. We introduce two datasets, OmniSim and InterReal, featuring 28 scenes with multiple interactive objects. To tackle the challenge of inaccurate interactive motion recovery in complex scenes, we propose LiveScene, a scene-level language-embedded interactive radiance field that efficiently reconstructs and controls multiple objects. By decomposing the interactive scene into local deformable fields, LiveScene enables separate reconstruction of individual object motions, reducing memory consumption. Additionally, our interaction-aware language embedding localizes individual interactive objects, allowing for arbitrary control using natural language. Our approach demonstrates significant superiority in novel view synthesis, interactive scene control, and language grounding performance through extensive experiments. Project page: https://livescenes.github.io.
Delin Qu, Pingrui Zhang, Xianqiang Gao 0001, Bin Zhao 0001, Zhigang Wang 0002, Dong Wang 0028, Xuelong Li 0001
NeurIPS6
2024 Optics-driven drone
Xuelong Li 0001, Zhigang Wang 0002, Bin Zhao 0001
Sci. China Inf. Sci.3
2023 One-Shot High-Fidelity Talking-Head Synthesis with Deformable Neural Radiance Field
abstract
Talking head generation aims to generate faces that maintain the identity information of the source image and imitate the motion of the driving image. Most pioneering methods rely primarily on 2D representations and thus will inevitably suffer from face distortion when large head rotations are encountered. Recent works instead employ explicit 3D structural representations or implicit neural rendering to improve performance under large pose changes. Nevertheless, the fidelity of identity and expression is not so desirable, especially for novel-view synthesis. In this paper, we propose HiDe-NeRF, which achieves high-fidelity and free-view talking-head synthesis. Drawing on the recently proposed Deformable Neural Radiance Fields, HiDe-NeRF represents the 3D dynamic scene into a canonical appearance field and an implicit deformation field, where the former comprises the canonical source face and the latter models the driving pose and expression. In particular, we improve fidelity from two aspects: (i) to enhance identity expressiveness, we design a generalized appearance module that leverages multi-scale volume features to preserve face shape and details; (ii) to improve expression preciseness, we propose a lightweight deformation module that explicitly decouples the pose and expression to enable precise expression modeling. Extensive experiments demonstrate that our proposed approach can generate better results than previous works. Project page: https://www.waytron.net/hidenerf/
Weichuang Li, Longhao Zhang, Dong Wang 0028, Bin Zhao 0001, Zhigang Wang 0002, Mulin Chen, Bang Zhang, Zhongjian Wang, Liefeng Bo, Xuelong Li 0001
CVPR5
2023 Fully Self-Supervised Depth Estimation from Defocus Clue
abstract
Depth-from-defocus (DFD), modeling the relationship between depth and defocus pattern in images, has demonstrated promising performance in depth estimation. Recently, several self-supervised works try to overcome the difficulties in acquiring accurate depth ground-truth. However, they depend on the all-in-focus (AIF) images, which cannot be captured in real-world scenarios. Such limitation discourages the applications of DFD methods. To tackle this issue, we propose a completely self-supervised framework that estimates depth purely from a sparse focal stack. We show that our framework circumvents the needs for the depth and AIF image ground-truth, and receives superior predictions, thus closing the gap between the theoretical success of DFD works and their applications in the real world. In particular, we propose (i) a more realistic setting for DFD tasks, where no depth or AIF image ground-truth is available; (ii) a novel self- supervision framework that provides reliable predictions of depth and AIF image under the challenging setting. The proposed framework uses a neural model to predict the depth and AIF image, and utilizes an optical model to validate and refine the prediction. We verify our framework on three benchmark datasets with rendered focal stacks and real focal stacks. Qualitative and quantitative evaluations show that our method provides a strong baseline for self- supervised DFD tasks. The source code is publicly avail- able at https://github.com/Ehzoahis/DEReD.
Haozhe Si, Bin Zhao 0001, Dong Wang 0028, Mulin Chen, Zhigang Wang 0002, Xuelong Li 0001
CVPR6
2023 Propagate and Calibrate: Real-Time Passive Non-Line-of-Sight Tracking
abstract
Non-line-of-sight (NLOS) tracking has drawn increasing attention in recent years, due to its ability to detect object motion out of sight. Most previous works on NLOS tracking rely on active illumination, e.g., laser, and suffer from high cost and elaborate experimental conditions. Besides, these techniques are still far from practical application due to oversimplified settings. In contrast, we propose a purely passive method to track a person walking in an invisible room by only observing a relay wall, which is more in line with real application scenarios, e.g., security. To excavate imperceptible changes in videos of the relay wall, we introduce difference frames as an essential carrier of temporal-local motion messages. In addition, we propose PAC-Net, which consists of alternating propagation and calibration, making it capable of leveraging both dynamic and static messages on a frame-level granularity. To evaluate the proposed method, we build and publish the first dynamic passive NLOS tracking dataset, NLOS-Track, which fills the vacuum of realistic NLOS datasets. NLOS-Track contains thousands of NLOS video clips and corresponding trajectories. Both real-shot and synthetic data are included. Our codes and dataset are available at https://againstentropy.github.io/NLOS-Track/.
Zhigang Wang 0002, Bin Zhao 0001, Dong Wang 0028, Mulin Chen, Xuelong Li 0001
CVPR2
2023 Multi-Modal Beam Selection: A Transfer Methodology for Multi-Frequency
abstract
This paper investigates beam selection for multiple-frequency via deep learning. Existing learning-based beam selection methods are typically data-hungry to train the neural network. However, collecting sufficient data is a major challenge that hinders the generalization of networks. To address this challenge, this paper develops a frequency transfer method and proposes a multi-modal beam selection network, named FtransNet, which demonstrates superior generalization performance to different scenarios and carrier frequencies. Compared to the traditional approach of only transferring a model, we also transfer and augment high-quality samples to the new frequency, which significantly enlarges the environment features and path loss features in the dataset. Moreover, we embed the relative locations and the reflection features of the environment to assist the best beam selection, which further enhances the generalization of the network. Simulation results show that the proposed FtransNet outperforms the existing network with high beam selection accuracy.
Dong Wang 0028, Mei Tu, Xiangfeng Gao, Zhigang Wang 0002
GLOBECOM6
2023 ViewRefer: Grasp the Multi-view Knowledge for 3D Visual Grounding
abstract
Understanding 3D scenes from multi-view inputs has been proven to alleviate the view discrepancy issue in 3D visual grounding. However, existing methods normally neglect the view cues embedded in the text modality and fail to weigh the relative importance of different views. In this paper, we propose ViewRefer, a multi-view framework for 3D visual grounding exploring how to grasp the view knowledge from both text and 3D modalities. For the text branch, ViewRefer leverages the diverse linguistic knowledge of large-scale language models, e.g., GPT, to expand a single grounding text to multiple geometry-consistent descriptions. Meanwhile, in the 3D modality, a transformer fusion module with inter-view attention is introduced to boost the interaction of objects across views. On top of that, we further present a set of learnable multi-view prototypes, which memorize scene-agnostic knowledge for different views, and enhance the framework from two perspectives: a view-guided attention module for more robust text features, and a view-guided scoring strategy during the final prediction. With our designed paradigm, ViewRefer achieves superior performance on three benchmarks and surpasses the second-best by +2.8%, +1.5%, and +1.35% on Sr3D, Nr3D, and ScanRefer.
Zoey Guo, Ray Zhang 0002, Dong Wang 0028, Zhigang Wang 0002, Bin Zhao 0001, Xuelong Li 0001
ICCV5
2023 Towards Nonlinear-Motion-Aware and Occlusion-Robust Rolling Shutter Correction
abstract
This paper addresses the problem of rolling shutter correction in complex nonlinear and dynamic scenes with extreme occlusion. Existing methods suffer from two main drawbacks. Firstly, they face challenges in estimating the accurate correction field due to the uniform velocity assumption, leading to significant image correction errors under complex motion. Secondly, the drastic occlusion in dynamic scenes prevents current solutions from achieving better image quality because of the inherent difficulties in aligning and aggregating multiple frames. To tackle these challenges, we model the curvilinear trajectory of pixels analytically and propose a geometry-based Quadratic Rolling Shutter (QRS) motion solver, which precisely estimates the high-order correction field of individual pixels. Besides, to reconstruct high-quality occlusion frames in dynamic scenes, we present a 3D video architecture that effectively Aligns and Aggregates multi-frame context, namely, RSA2-Net. We evaluate our method across a broad range of cameras and video sequences, demonstrating its significant superiority. Specifically, our method surpasses the state-of-the-art by +4.98, +0.77, and +4.33 of PSNR on CarlaRS, Fastec-RS, and BS-RSC datasets, respectively. Code is available at https://github.com/DelinQu/qrsc.
Delin Qu, Yizhen Lao, Zhigang Wang 0002, Dong Wang 0028, Bin Zhao 0001, Xuelong Li 0001
ICCV3
2022 Implicit Sample Extension for Unsupervised Person Re-Identification
abstract
Most existing unsupervised person re-identification (Re-ID) methods use clustering to generate pseudo labels for model training. Unfortunately, clustering sometimes mixes different true identities together or splits the same identity into two or more sub clusters. Training on these noisy clusters substantially hampers the Re-ID accuracy. Due to the limited samples in each identity, we suppose there may lack some underlying information to well reveal the accurate clusters. To discover these information, we propose an Implicit Sample Extension (ISE) method to generate what we call support samples around the cluster boundaries. Specifically, we generate support samples from actual samples and their neighbouring clusters in the embedding space through a progressive linear interpolation (PLI) strategy. PLI controls the generation with two critical factors, i.e., 1) the direction from the actual sample towards its K-nearest clusters and 2) the degree for mixing up the context information from the K-nearest clusters. Meanwhile, given the support samples, ISE further uses a label-preserving loss to pull them towards their corresponding actual samples, so as to compact each cluster. Consequently, ISE reduces the “sub and mixed” clustering errors, thus improving the Re-ID performance. Extensive experiments demonstrate that the proposed method is effective and achieves state-of-the-art performance for unsupervised person Re-ID. Code is available at: https://github.com/PaddlePaddle/PaddleClas.
Xinyu Zhang 0015, Zhigang Wang 0002, Jian Wang 0066, Errui Ding, Qinfeng Shi, Zhaoxiang Zhang 0001, Jingdong Wang 0001
CVPR3
2022 UFO: Unified Feature Optimization
Teng Xi, Yifan Sun 0003, Deli Yu, Bi Li 0005, Nan Peng, Xinyu Zhang 0015, Zhigang Wang 0002, Jian Wang 0066, Haocheng Feng, Junyu Han, Jingtuo Liu, Errui Ding, Jingdong Wang 0001
ECCV (26)8
2022 Self-Guided Hard Negative Generation for Unsupervised Person Re-Identification
abstract
Recent unsupervised person re-identification (reID) methods mostly apply pseudo labels from clustering algorithms as supervision signals. Despite great success, this fashion is very likely to aggregate different identities with similar appearances into the same cluster. In result, the hard negative samples, playing important role in training reID models, are significantly reduced. To alleviate this problem, we propose a self-guided hard negative generation method for unsupervised person re-ID. Specifically, a joint framework is developed which incorporates a hard negative generation network (HNGN) and a re-ID network. To continuously generate harder negative samples to provide effective supervisions in the contrastive learning, the two networks are alternately trained in an adversarial manner to improve each other, where the reID network guides HNGN to generate challenging data and HNGN enforces the re-ID network to enhance discrimination ability. During inference, the performance of re-ID network is improved without introducing any extra parameters. Extensive experiments demonstrate that the proposed method significantly outperforms a strong baseline and also achieves better results than state-of-the-art methods.
Zhigang Wang 0002, Jian Wang 0066, Xinyu Zhang 0015, Errui Ding, Jingdong Wang 0001, Zhaoxiang Zhang 0001
IJCAI2
2021 Unsupervised Multi-Source Domain Adaptation for Person Re-Identification
abstract
Unsupervised domain adaptation (UDA) methods for person re-identification (re-ID) aim at transferring re-ID knowledge from labeled source data to unlabeled target data. Although achieving great success, most of them only use limited data from a single-source domain for model pre-training, making the rich labeled data insufficiently exploited. To make full use of the valuable labeled data, we introduce the multi-source concept into UDA person re-ID field, where multiple source datasets are used during training. However, because of domain gaps, simply combining different datasets only brings limited improvement. In this paper, we try to address this problem from two perspectives, i.e. domain-specific view and domain-fusion view. Two constructive modules are proposed, and they are compatible with each other. First, a rectification domain-specific batch normalization (RDSBN) module is explored to simultaneously reduce domain-specific characteristics and increase the distinctiveness of person features. Second, a graph convolutional network (GCN) based multi-domain information fusion (MDIF) module is developed, which minimizes domain distances by fusing features of different domains. The proposed method outperforms state-of-the-art UDA person re-ID methods by a large margin, and even achieves comparable performance to the supervised approaches without any post-processing techniques.
Zechen Bai, Zhigang Wang 0002, Jian Wang 0066, Di Hu 0001, Errui Ding
CVPR2
2019 Weather recognition via classification labels and weather-cue maps
Bin Zhao 0001, Lulu Hua, Xuelong Li 0001, Xiaoqiang Lu, Zhigang Wang 0002
Pattern Recognit.5
2018 A CNN-RNN architecture for multi-label weather recognition
Bin Zhao 0001, Xuelong Li 0001, Xiaoqiang Lu, Zhigang Wang 0002
Neurocomputing4
2018 Video Synopsis in Complex Situations
abstract
Video synopsis is an effective technique for surveillance video browsing and storage. However, most of the existing video synopsis approaches are not suitable for complex situations, especially crowded scenes. This is because these approaches heavily depend on the preprocessing results of foreground segmentation and multiple objects tracking, but the preprocessing techniques usually achieve poor performance in crowded scenes. To address this problem, we propose a comprehensive video synopsis approach which can be applied to scenes with drastically varying crowdedness. The proposed approach differs significantly from the existing methods, and has several appealing properties. First, we propose to detect the crowdedness of a given video, then, extract object tubes in sparse periods and extract video clips in crowded periods, respectively. Through such a solution, the poor performance of preprocessing techniques in crowded scenes can be avoided by extracting the whole video frames. Second, we propose a group-partition algorithm which can discovers the relationships among moving objects and alleviates several segmentation and tracking errors. Third, a group-based greedy optimization algorithm is proposed to automatically determine the length of a synopsis video. Besides, we present extensive experiments that demonstrate the effectiveness and efficiency of the proposed approach.
Xuelong Li 0001, Zhigang Wang 0002, Xiaoqiang Lu
IEEE Trans. Image Process.2
2017 A Multi-Task Framework for Weather Recognition
abstract
Weather recognition is important in practice, while this task has not been thoroughly explored so far. The current trend of dealing with this task is treating it as a single classification problem, i.e., determining whether a given image belongs to a certain weather category or not. However, weather recognition differs significantly from traditional image classification, since several weather features may appear simultaneously. In this case, a simple classification result is insufficient to describe the weather condition. To address this issue, we propose to provide auxiliary weather related information for comprehensive weather description. Specifically, semantic segmentation of weather-cues, such as blue sky and white clouds, is exploited as an auxiliary task in this paper. Moreover, a convolutional neural network (CNN) based multi-task framework is developed which aims to concurrently tackle weather category classification task and weather-cues segmentation task. Due to the intrinsic relationships between these two tasks, exploring auxiliary semantic segmentation of weather-cues can also help to learn discriminative features for the classification task, and thus obtain superior accuracy. To verify the effectiveness of the proposed approach, extra segmentation masks of weather-cues are generated manually on an existing weather image dataset. Experimental results have demonstrated the superior performance of our approach. The enhanced dataset, source codes and pre-trained models are available at https://github.com/wzgwzg/Multitask_Weather.
Xuelong Li 0001, Zhigang Wang 0002, Xiaoqiang Lu
ACM Multimedia2
2016 Surveillance Video Synopsis via Scaling Down Objects
abstract
Video synopsis is an effective technique to provide a compact representation of the original video by removing spatiotemporal redundancies and by preserving the essential activities. Most current approaches for video synopsis will cause collisions among objects, especially when the video is condensed much. In this paper, we present an approach for video synopsis to reduce the collisions. Our approach first shifts active objects along the time axis to compact the original video. Then, the sizes of the objects are reduced when collisions occur. Meanwhile, the geometric centroids of the objects will be kept unchanged to preserve the location information. Our contributions are threefold. First, an approach is proposed to decrease collisions in the synopsis video through reducing the sizes of the objects. Second, an optimization framework is developed to indicate the optimal time position and the appropriate reduction coefficient for each object. Finally, some metrics are proposed, and several experiments are carried out to evaluate the proposed approach. The experiments have demonstrated that the synopsis video produced by our approach has much fewer collisions while the compression ratio is high.
Xuelong Li 0001, Zhigang Wang 0002, Xiaoqiang Lu
IEEE Trans. Image Process.2