EDBT 2026 Demo / reviewers in the wild / expert
Wenzhao Zheng
dblp:230/1277
· DBLP profile ↗
54ranked-venue papers
10as first author
49since 2021 · last 2026
0000-0001-7188-3734ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 45 · 10 first-author · 41 since 2021Graphics, computer vision, multimedia, augmented reality and games · 40 · 8 first-author · 35 since 2021Systems, architecture and hardware · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ODEt(ODEl): Shortcutting the Time and the Length in Diffusion and Flow Models for Faster Sampling
Denis A. Gudovskiy, Wenzhao Zheng, Tomoyuki Okuno, Yohei Nakata, Kurt Keutzer |
WACV | 2 |
| 2026 | LiDAR-FMC: Accurate and Robust Human Capture From Point-Cloud Video
Bohao Fan, Wenzhao Zheng, Jianjiang Feng, Jie Zhou 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | OccSora: 4D Occupancy Generation Models as World Simulators for Autonomous DrivingabstractUnderstanding the evolution of 3D scenes is crucial for autonomous driving. While conventional methods describe scene development through individual instance motions, world models provide a generative framework for modeling overall scene dynamics. However, most existing approaches rely on autoregressive next-token prediction, which suffers from error accumulation and limited global spatiotemporal reasoning, leading to degraded long-term consistency. To address these issues, we propose a diffusion-based 4D occupancy generation model, OccSora, to simulate 3D world evolution for autonomous driving. A 4D scene tokenizer is introduced to obtain compact spatiotemporal representations and enable high-quality reconstruction of long occupancy sequences. We then train a diffusion transformer on these representations to generate 4D occupancy conditioned on trajectory prompts. Experiments on the nuScenes dataset with Occ3D annotations show that OccSora can generate 16s videos with authentic 3D layout and strong temporal consistency. With trajectory-aware 4D generation, OccSora has the potential to serve as a world simulator for autonomous driving decision-making. Project page: https://wzzheng.net/OccSora. Wenzhao Zheng, Yilong Ren, Han Jiang 0003, Zhiyong Cui, Haiyang Yu 0002, Jiwen Lu |
IEEE Trans. Image Process. | 2 |
| 2025 | GaussianFormer-2: Probabilistic Gaussian Superposition for Efficient 3D Occupancy Predictionabstract3D semantic occupancy prediction has garnered attention as an important task for the robustness of vision-centric autonomous driving, which predicts fine-grained geometry and semantics of the surrounding scene. Most existing methods leverage dense grid-based scene representations, overlooking the spatial sparsity of the driving scenes, which leads to computational redundancy. Although 3D semantic Gaussian serves as an object-centric sparse alternative, most of the Gaussians still describe the empty region with low efficiency. To address this, we propose a probabilistic Gaussian superposition model which interprets each Gaussian as a probability distribution of its neighborhood being occupied and conforms to probabilistic multiplication to derive the overall geometry. Furthermore, we adopt the exact Gaussian mixture model for semantics calculation to avoid unnecessary overlapping of Gaussians. To effectively initialize Gaussians in non-empty region, we design a distribution-based initialization module which learns the pixel-aligned occupancy distribution instead of the depth of surfaces. We conduct extensive experiments on nuScenes and KITTI-360 datasets and our GaussianFormer-2 achieves state-of-the-art performance with high efficiency. Yuanhui Huang 0002, Amonnut Thammatadatrakoon, Wenzhao Zheng, Dalong Du, Jiwen Lu |
CVPR | 3 |
| 2025 | Segment Any Motion in VideosabstractMoving object segmentation is a crucial task for achieving a high-level understanding of visual scenes and has numerous downstream applications. Humans can effortlessly segment moving objects in videos. Previous work has largely relied on optical flow to provide motion cues; however, this approach often results in imperfect predictions due to challenges such as partial motion, complex deformations, motion blur and background distractions. We propose a novel approach for moving object segmentation that combines long-range trajectory motion cues with DINO-based semantic features and leverages SAM2 for pixel-level mask densification through an iterative prompting strategy. Our model employs Spatio-Temporal Trajectory Attention and Motion-Semantic Decoupled Embedding to prioritize motion while integrating semantic support. Extensive testing on diverse datasets demonstrates state-of-the-art performance, excelling in challenging scenarios and fine-grained segmentation of multiple objects. Our code is available at https://motion-seg.github.io/. Wenzhao Zheng, Chenfeng Xu, Kurt Keutzer, Shanghang Zhang, Angjoo Kanazawa, Qianqian Wang 0002 |
CVPR | 2 |
| 2025 | DeSiRe-GS: 4D Street Gaussians for Static-Dynamic Decomposition and Surface Reconstruction for Urban Driving ScenesabstractWe present DeSiRe-GS, a self-supervised gaussian splatting representation, enabling effective static-dynamic decomposition and high-fidelity surface reconstruction in complex driving scenarios. Our approach employs a two-stage optimization pipeline of dynamic street Gaussians. In the first stage, we extract 2D motion masks based on the observation that 3D Gaussian Splatting inherently can reconstruct only the static regions in dynamic environments. These extracted 2D motion priors are then mapped into the Gaussian space in a differentiable manner, leveraging an efficient formulation of dynamic Gaussians in the second stage. Combined with the introduced geometric regularizations, our method are able to address the over-fitting issues caused by data sparsity in autonomous driving, reconstructing physically plausible Gaussians that align with object surfaces rather than floating in air. Furthermore, we introduce temporal cross-view consistency to ensure coherence across time and viewpoints, resulting in high-quality surface reconstruction. Comprehensive experiments demonstrate the efficiency and effectiveness of DeSiRe-GS, surpassing prior self-supervised arts and achieving accuracy comparable to methods relying on external 3D bounding box annotations. Code is available at https://github.com/chengweialan/DeSiRe-GS Chensheng Peng, Chenfeng Xu, Yichen Xie 0002, Wenzhao Zheng, Kurt Keutzer, Masayoshi Tomizuka |
CVPR | 6 |
| 2025 | GaussianWorld: Gaussian World Model for Streaming 3D Occupancy Predictionabstract3D occupancy prediction is important for autonomous driving due to its comprehensive perception of the surroundings. To incorporate sequential inputs, most existing methods fuse representations from previous frames to infer the current 3D occupancy. However, they fail to consider the continuity of driving scenarios and ignore the strong prior provided by the evolution of 3D scenes (e.g., only dynamic objects move). In this paper, we propose a world-model-based framework to exploit the scene evolution for perception. We reformulate 3D occupancy prediction as a 4D occupancy forecasting problem conditioned on the current sensor input. We decompose the scene evolution into three factors: 1) ego motion alignment of static scenes; 2) local movements of dynamic objects; and 3) completion of newly-observed scenes. We then employ a Gaussian world model (GaussianWorld) to explicitly exploit these priors and infer the scene evolution in the 3D Gaussian space considering the current RGB observation. We evaluate the effectiveness of our framework on the widely used nuScenes dataset. Our GaussianWorld improves the performance of the single-frame counterpart by over 2% in mIoU without introducing additional computations. Code: https://github.com/zuosc19/GaussianWorld. Sicheng Zuo, Wenzhao Zheng, Yuanhui Huang 0002, Jie Zhou 0001, Jiwen Lu |
CVPR | 2 |
| 2025 | SpectralAR: Spectral Autoregressive Visual GenerationabstractAutoregressive visual generation has garnered increasing attention due to its scalability and compatibility with other modalities compared with diffusion models. Most existing methods construct visual sequences as spatial patches for autoregressive generation. However, image patches are inherently parallel, contradicting the causal nature of autoregressive modeling. To address this, we propose a Spectral AutoRegressive (SpectralAR) visual generation framework, which realizes causality for visual sequences from the spectral perspective. Specifically, we first transform an image into ordered spectral tokens with Nested Spectral Tokenization, representing lower to higher frequency components. We then perform autoregressive generation in a coarse-to-fine manner with the sequences of spectral tokens. By considering different levels of detail in images, our SpectralAR achieves both sequence causality and token efficiency without bells and whistles. We conduct extensive experiments on ImageNet-1K for image reconstruction and autoregressive generation, and SpectralAR achieves 3.02 gFID with only 64 tokens and 310M parameters. Project page: https://huang-yh.github.io/spectralar/. Yuanhui Huang 0002, Wenzhao Zheng, Yueqi Duan, Jie Zhou 0001, Jiwen Lu |
ICCV | 3 |
| 2025 | Authentic 4D Driving Simulation with a Video Generation Model
Wenzhao Zheng, Dalong Du, Yilong Ren, Han Jiang 0003, Zhiyong Cui, Haiyang Yu 0002, Jie Zhou 0001, Shanghang Zhang |
ICCV | 2 |
| 2025 | EmbodiedOcc: Embodied 3D Occupancy Prediction for Vision-Based Online Scene Understandingabstract3D occupancy prediction provides a comprehensive description of the surrounding scenes and has become an essential task for 3D perception. Most existing methods focus on offline perception from one or a few views and cannot be applied to embodied agents that demand to gradually perceive the scene through progressive embodied exploration. In this paper, we formulate an embodied 3D occupancy prediction task to target this practical scenario and propose a Gaussian-based EmbodiedOcc framework to accomplish it. We initialize the global scene with uniform 3D semantic Gaussians and progressively update local regions observed by the embodied agent. For each update, we extract semantic and structural features from the observed image and efficiently incorporate them via deformable cross-attention to refine the regional Gaussians. Finally, we employ Gaussian-to-voxel splatting to obtain the global 3D occupancy from the updated 3D Gaussians. Our EmbodiedOcc assumes an unknown (i.e., uniformly distributed) environment and maintains an explicit global memory of it with 3D Gaussians. It gradually gains knowledge through the local refinement of regional Gaussians, which is consistent with how humans understand new scenes through embodied exploration. We reorganize an EmbodiedOcc-ScanNet benchmark based on local annotations to facilitate the evaluation of the embodied 3D occupancy prediction task. Our EmbodiedOcc outperforms existing methods by a large margin and accomplishes the embodied occupancy prediction with high accuracy and efficiency. Code: https://github.com/YkiWu/EmbodiedOcc. Wenzhao Zheng, Sicheng Zuo, Yuanhui Huang 0002, Jie Zhou 0001, Jiwen Lu |
ICCV | 2 |
| 2025 | D3QE: Learning Discrete Distribution Discrepancy-Aware Quantization Error for Autoregressive-Generated Image Detection
Yanran Zhang, Bingyao Yu, Yu Zheng 0015, Wenzhao Zheng, Yueqi Duan, Lei Chen 0069, Jie Zhou 0001, Jiwen Lu |
ICCV | 4 |
| 2025 | PlaneRAS: Learning Planar Primitives for 3D Plane Recovery
Wenzhao Zheng, Linqing Zhao, Zelan Zhu, Jiwen Lu, Xiuzhuang Zhou |
ICCV | 2 |
| 2025 | Learning Counterfactually Decoupled Attention for Open-World Model AttributionabstractIn this paper, we propose a Counterfactually Decoupled Attention Learning (CDAL) method for open-world model attribution. Existing methods rely on handcrafted design of region partitioning or feature space, which could be confounded by the spurious statistical correlations and struggle with novel attacks in open-world scenarios. To address this, CDAL explicitly models the causal relationships between the attentional visual traces and source model attribution, and counterfactually decouples the discriminative model-specific artifacts from confounding source biases for comparison. In this way, the resulting causal effect provides a quantification on the quality of learned attention maps, thus encouraging the network to capture essential generation patterns that generalize to unseen source models by maximizing the effect. Extensive experiments on existing open-world model attribution benchmarks show that with minimal computational overhead, our method consistently improves state-of-the-art models by large margins, particularly for unseen novel attacks. Source code: https://github.com/yzheng97/CDAL. Yu Zheng 0015, Boyang Gong, Fanye Kong, Yueqi Duan, Bingyao Yu, Wenzhao Zheng, Lei Chen 0069, Jiwen Lu, Jie Zhou 0001 |
ICCV | 6 |
| 2025 | UniDrive: Towards Universal Driving Perception Across Camera ConfigurationsabstractVision-centric autonomous driving has demonstrated excellent performance with economical sensors. As the fundamental step, 3D perception aims to infer 3D information from 2D images based on 3D-2D projection. This makes driving perception models susceptible to sensor configuration (e.g., camera intrinsics and extrinsics) variations. However, generalizing across camera configurations is important for deploying autonomous driving models on different car models. In this paper, we present UniDrive, a novel framework for vision-centric autonomous driving to achieve universal perception across camera configurations. We deploy a set of unified virtual cameras and propose a ground-aware projection method to effectively transform the original images into these unified virtual views. We further propose a virtual configuration optimization method by minimizing the expected projection error between original and virtual cameras. The proposed virtual camera projection can be applied to existing 3D perception methods as a plug-and-play module to mitigate the challenges posed by camera parameter variability, resulting in more adaptable and reliable driving perception models. To evaluate the effectiveness of our framework, we collect a dataset on CARLA by driving the same routes while only modifying the camera configurations. Experimental results demonstrate that our method trained on one specific camera configuration can generalize to varying configurations with minor performance degradation. Wenzhao Zheng, Xiaonan Huang, Kurt Keutzer |
ICLR | 2 |
| 2025 | SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model InferenceabstractIn vision-language models (VLMs), visual tokens usually consume a significant amount of computational overhead, despite their sparser information density compared to text tokens. To address this, most existing methods learn a network to prune redundant visual tokens and require additional training data. Differently, we propose an efficient training-free token optimization mechanism dubbed SparseVLM without extra parameters or fine-tuning costs. Concretely, given that visual tokens complement text tokens in VLMs for linguistic reasoning, we select visual-relevant text tokens to rate the significance of vision tokens within the self-attention matrix extracted from the VLMs. Then we progressively prune irrelevant tokens. To maximize sparsity while retaining essential information, we introduce a rank-based strategy to adaptively determine the sparsification ratio for each layer, alongside a token recycling method that compresses pruned tokens into more compact representations. Experimental results show that our SparseVLM improves the efficiency of various VLMs across a range of image and video understanding tasks. In particular, when LLaVA is equipped with SparseVLM, it achieves a 54% reduction in FLOPs, lowers CUDA time by 37%, and maintains an accuracy rate of 97%. Our code is available at https://github.com/Gumpest/SparseVLMs. Yuan Zhang 0020, Chun-Kai Fan, Junpeng Ma, Wenzhao Zheng, Tao Huang 0020, Kuan Cheng, Denis A. Gudovskiy, Tomoyuki Okuno, Yohei Nakata, Kurt Keutzer, Shanghang Zhang |
ICML | 4 |
| 2025 | Lightstereo: Channel Boost is All You Need for Efficient 2D Cost AggregationabstractWe present LightStereo, a cutting-edge stereomatching network crafted to accelerate the matching process. Departing from conventional methodologies that rely on aggregating computationally intensive 4D costs, LightStereo adopts the 3D cost volume as a lightweight alternative. While similar approaches have been explored previously, our breakthrough lies in enhancing performance through a dedicated focus on the channel dimension of the 3D cost volume, where the distribution of matching costs is encapsulated. Our exhaustive exploration has yielded plenty of strategies to amplify the capacity of the pivotal dimension, ensuring both precision and efficiency. We compare the proposed LightStereo with existing state-of-the-art methods across various benchmarks, which demonstrate its superior performance in speed, accuracy, and resource utilization. LightStereo achieves a competitive EPE metric in the SceneFlow datasets while demanding a minimum of only 22 GFLOPs and 17 ms of runtime, and ranks 1st on KITTI 2015 among real-time models. Our comprehensive analysis reveals the effect of 2 D cost aggregation for stereo matching, paving the way for realworld applications of efficient stereo systems. Code is available at https://github.com/XiandaGuo/OpenStereo. Xianda Guo, Chenming Zhang, Youmin Zhang 0008, Wenzhao Zheng, Dujun Nie, Matteo Poggi, Long Chen 0005 |
ICRA | 4 |
| 2025 | SliceOcc: Indoor 3D Semantic Occupancy Prediction with Vertical Slice Representationabstract3D semantic occupancy prediction is a crucial task in visual perception, as it requires the simultaneous comprehension of both scene geometry and semantics. It plays a crucial role in understanding 3D scenes and has great potential for various applications, such as robotic vision perception and autonomous driving. Many existing works utilize planar-based representations such as Bird's Eye View (BEV) and Tri-Perspective View (TPV). These representations aim to simplify the complexity of 3D scenes while preserving essential object information, thereby facilitating efficient scene representation. However, in dense indoor environments with prevalent occlusions, directly applying these planar-based methods often leads to difficulties in capturing global semantic occupancy, ultimately degrading model performance. In this paper, we present a new vertical slice representation that divides the scene along the vertical axis and projects spatial point features onto the nearest pair of parallel planes. To utilize these slice features, we propose SliceOcc, an RGB camera-based model specifically tailored for indoor 3D semantic occupancy prediction. SliceOcc utilizes pairs of slice queries and cross-attention mechanisms to extract planar features from input images. These local planar features are then fused to form a global scene representation, which is employed for indoor occupancy prediction. Experimental results on the EmbodiedScan dataset demonstrate that SliceOcc achieves a mIoU of 15.45 % across 81 indoor categories, setting a new state-of-the-art performance among RGB camera-based models for indoor 3D semantic occupancy prediction. Jianing Li 0001, Ming Lu 0002, Hao Wang 0073, Chenyang Gu, Wenzhao Zheng, Shanghang Zhang |
ICRA | 6 |
| 2025 | Pseudo Depth Meets Gaussian: A Feed-forward RGB SLAM Baseline
Linqing Zhao, Xiuwei Xu, Wenzhao Zheng, Yansong Tang, Haibin Yan, Jiwen Lu |
IROS | 5 |
| 2025 | EmbodiedOcc++: Boosting Embodied 3D Occupancy Prediction with Plane Regularization and Uncertainty SamplerabstractOnline 3D occupancy prediction provides a comprehensive spatial understanding of embodied environments. While the innovative EmbodiedOcc framework utilizes 3D semantic Gaussians for progressive indoor occupancy prediction, it overlooks the geometric characteristics of indoor environments, which are primarily characterized by planar structures. This paper introduces EmbodiedOcc++, enhancing the original framework with two key innovations: a Geometry-guided Refinement Module (GRM) that constrains Gaussian updates through plane regularization, along with a Semantic-aware Uncertainty Sampler (SUS) that enables more effective updates in overlapping regions between consecutive frames. GRM regularizes the position update to align with surface normals. It determines the adaptive regularization weight using curvature-based and depth-based constraints, allowing semantic Gaussians to align accurately with planar surfaces while adapting in complex regions. To effectively improve geometric consistency from different views, SUS adaptively selects proper Gaussians to update. Comprehensive experiments on the EmbodiedOcc-ScanNet benchmark demonstrate that EmbodiedOcc++ achieves state-of-the-art performance across different settings. Our method demonstrates improved edge accuracy and retains more geometric details while ensuring computational efficiency, which is essential for online embodied perception. The code will be released at: https://github.com/PKUHaoWang/EmbodiedOcc2. Hao Wang 0073, Xiaobao Wei, Xiaoan Zhang, Jianing Li 0001, Chengyu Bai, Ying Li 0128, Ming Lu 0002, Wenzhao Zheng, Shanghang Zhang |
ACM Multimedia | 8 |
| 2025 | DrivingRecon: Large 4D Gaussian Reconstruction Model For Autonomous DrivingabstractLarge reconstruction model has remarkable progress, which can directly predict 3D or 4D representations for unseen scenes and objects. However, current work has not systematically explored the potential of large reconstruction models in the field of autonomous driving. To achieve this, we introduce the Large 4D Gaussian Reconstruction Model (DrivingRecon). With an elaborate and simple framework design, it not only ensures efficient and high-quality reconstruction, but also provides potential for downstream tasks. There are two core contributions: firstly, the Prune and Dilate Block (PD-Block) is proposed to prune redundant and overlapping Gaussian points and dilate Gaussian points for complex objects. Then, dynamic and static decoupling is tailored to better learn the temporary-consistent geometry across different time. Experimental results demonstrate that DrivingRecon significantly improves scene reconstruction quality compared to existing methods. Furthermore, we explore applications of DrivingRecon in model pre-training, vehicle type adaptation, and scene editing. Our code will be available. Hao Lu 0009, Tianshuo Xu, Wenzhao Zheng, Dalong Du, Masayoshi Tomizuka, Kurt Keutzer, Ying-Cong Chen |
NeurIPS | 3 |
| 2025 | Point3R: Streaming 3D Reconstruction with Explicit Spatial Pointer MemoryabstractDense 3D scene reconstruction from an ordered sequence or unordered image collections is a critical step when bringing research in computer vision into practical scenarios. Following the paradigm introduced by DUSt3R, which unifies an image pair densely into a shared coordinate system, subsequent methods maintain an implicit memory to achieve dense 3D reconstruction from more images. However, such implicit memory is limited in capacity and may suffer from information loss of earlier frames. We propose Point3R, an online framework targeting dense streaming 3D reconstruction. To be specific, we maintain an explicit spatial pointer memory directly associated with the 3D structure of the current scene. Each pointer in this memory is assigned a specific 3D position and aggregates scene information nearby in the global coordinate system into a changing spatial feature. Information extracted from the latest frame interacts explicitly with this pointer memory, enabling dense integration of the current observation into the global coordinate system. We design a 3D hierarchical position embedding to promote this interaction and design a simple yet effective fusion mechanism to ensure that our pointer memory is uniform and efficient. Our method achieves competitive or state-of-the-art performance on various tasks with low training costs. Code: https://github.com/YkiWu/Point3R. Wenzhao Zheng, Jie Zhou 0001, Jiwen Lu |
NeurIPS | 2 |
| 2025 | QuadricFormer: Scene as Superquadrics for 3D Semantic Occupancy Predictionabstract3D occupancy prediction is crucial for robust autonomous driving systems as it enables comprehensive perception of environmental structures and semantics. Most existing methods employ dense voxel-based scene representations, ignoring the sparsity of driving scenes and resulting in inefficiency. Recent works explore object-centric representations based on sparse Gaussians, but their ellipsoidal shape prior limits the modeling of diverse structures. In real-world driving scenes, objects exhibit rich geometries (e.g., cuboids, cylinders, and irregular shapes), necessitating excessive ellipsoidal Gaussians densely packed for accurate modeling, which leads to inefficient representations. To address this, we propose to use geometrically expressive superquadrics as scene primitives, enabling efficient representation of complex structures with fewer primitives through their inherent shape diversity. We develop a probabilistic superquadric mixture model, which interprets each superquadric as an occupancy probability distribution with a corresponding geometry prior, and calculates semantics through probabilistic mixture. Building on this, we present QuadricFormer, a superquadric-based model for efficient 3D occupancy prediction, and introduce a pruning-and-splitting module to further enhance modeling efficiency by concentrating superquadrics in occupied regions. Extensive experiments on the nuScenes and KITTI-360 datasets demonstrate that QuadricFormer achieves state-of-the-art performance while maintaining superior efficiency. Code is available at https://github.com/zuosc19/QuadricFormer. Sicheng Zuo, Wenzhao Zheng, Xiaoyong Han, Longchao Yang, Jiwen Lu |
NeurIPS | 2 |
| 2025 | Probabilistic deep metric learning for hyperspectral image classification
Chengkun Wang, Wenzhao Zheng, Xian Sun 0001, Jie Zhou 0001, Jiwen Lu |
Pattern Recognit. | 2 |
| 2025 | LiDAR-HMR: 3D Human Mesh Recovery From LiDARabstractHuman mesh recovery (HMR) holds significant utility in many applications. Studying HMR involving various types of sensors is necessary, as it enables the acquisition of human meshes in diverse scenes. Unlike HMR based on RGB images, HMR based on LiDAR has received considerably less attention in previous works. The major challenge in estimating human poses and meshes from sparse point clouds lies in the sparsity, noise, and incompletion of LiDAR point clouds. To address these challenges, we propose a LiDAR-based 3D human mesh recovery algorithm, called LiDAR-HMR. This algorithm involves estimating a sparse representation of a human (3D human pose) and gradually reconstructing the body mesh. To better leverage the 3D structural information of point clouds, we propose a point-cloud-to-SMPL pipeline that uses the original point cloud features to guide the reconstruction. The experimental results on four publicly available datasets demonstrate the effectiveness of LiDAR-HMR. The codes are available athttps://github.com/soullessrobot/LiDAR-HMR. Bohao Fan, Wenzhao Zheng, Jianjiang Feng, Jie Zhou 0001 |
IEEE Trans. Multim. | 2 |
| 2024 | SelfOcc: Self-Supervised Vision-Based 3D Occupancy Predictionabstract3D occupancy prediction is an important task for the robustness of vision-centric autonomous driving, which aims to predict whether each point is occupied in the surrounding 3D space. Existing methods usually require 3D occupancy labels to produce meaningful results. However, it is very laborious to annotate the occupancy status of each voxel. In this paper, we propose SelfOcc to explore a self-supervised way to learn 3D occupancy using only video sequences. We first transform the images into the 3D space (e.g., bird's eye view) to obtain 3D representation of the scene. We directly impose constraints on the 3D representations by treating them as signed distance fields. We can then render 2D images of previous and future frames as self-supervision signals to learn the 3D representations. We propose an MVS-embedded strategy to directly optimize the SDF-induced weights with multiple depth proposals. Our SelfOcc out-performs the previous best method SceneRF by 58.7% using a single frame as input on SemanticKITTI and is the first self-supervised work that produces reasonable 3D occupancy for surround cameras on nuScenes. SelfOcc produces high-quality depth and achieves state-of-the-art results on novel depth synthesis, monocular depth estimation, and surround-view depth estimation on the SemanticKITTI, KITTI-2015, and nuScenes, respectively. Code: https://github.com/huang-yh/SelfOcc. Yuanhui Huang 0002, Wenzhao Zheng, Borui Zhang, Jie Zhou 0001, Jiwen Lu |
CVPR | 2 |
| 2024 | LowRankOcc: Tensor Decomposition and Low-Rank Recovery for Vision-Based 3D Semantic Occupancy PredictionabstractIn this paper, we present a tensor decomposition and low-rank recovery approach (LowRankOcc) for vision-based 3D semantic occupancy prediction. Conventional methods model outdoor scenes with fine-grained 3D grids, but the sparsity of non-empty voxels introduces consider-able spatial redundancy, leading to potential overfitting risks. In contrast, our approach leverages the intrinsic low-rank property of 3D occupancy data, factorizing voxel representations into low-rank components to efficiently mitigate spatial redundancy without sacrificing performance. Specifically, we present the Vertical-Horizontal (VH) de-composition block factorizes 3D tensors into vertical vectors and horizontal matrices. With our “decomposition-encoding-recovery” framework, we encode 3D contexts with only 1/2D convolutions and poolings, and subsequently recover the encoded compact yet informative context features back to voxel representations. Experimental results demonstrate that LowRankOcc achieves state-of-the-art performances in semantic scene completion on the Se-manticKITTI dataset and 3D occupancy prediction on the nuScenes dataset. Linqing Zhao, Xiuwei Xu, Ziwei Wang 0010, Borui Zhang, Wenzhao Zheng, Dalong Du, Jie Zhou 0001, Jiwen Lu |
CVPR | 6 |
| 2024 | GaussianFormer: Scene as Gaussians for Vision-Based 3D Semantic Occupancy Prediction
Yuanhui Huang 0002, Wenzhao Zheng, Jie Zhou 0001, Jiwen Lu |
ECCV (27) | 2 |
| 2024 | SpatialFormer: Towards Generalizable Vision Transformers with Explicit Spatial Understanding
Han Xiao 0010, Wenzhao Zheng, Sicheng Zuo, Peng Gao 0007, Jie Zhou 0001, Jiwen Lu |
ECCV (13) | 2 |
| 2024 | OccWorld: Learning a 3D Occupancy World Model for Autonomous Driving
Wenzhao Zheng, Yuanhui Huang 0002, Borui Zhang, Yueqi Duan, Jiwen Lu |
ECCV (13) | 1 |
| 2024 | GenAD: Generative End-to-End Autonomous Driving
Wenzhao Zheng, Xianda Guo, Chenming Zhang, Long Chen 0005 |
ECCV (65) | 1 |
| 2024 | Path Choice Matters for Clear Attributions in Path MethodsabstractRigorousness and clarity are both essential for interpretations of DNNs to engender human trust. Path methods are commonly employed to generate rigorous attributions that satisfy three axioms. However, the meaning of attributions remains ambiguous due to distinct path choices. To address the ambiguity, we introduce Concentration Principle, which centrally allocates high attributions to indispensable features, thereby endowing aesthetic and sparsity. We then present SAMP, a model-agnostic interpreter, which efficiently searches the near-optimal path from a pre-defined set of manipulation paths. Moreover, we propose the infinitesimal constraint (IC) and momentum strategy (MS) to improve the rigorousness and optimality. Visualizations show that SAMP can precisely reveal DNNs by pinpointing salient image pixels.
We also perform quantitative experiments and observe that our method significantly outperforms the counterparts. Borui Zhang, Wenzhao Zheng, Jie Zhou 0001, Jiwen Lu |
ICLR | 2 |
| 2024 | Introspective Deep Metric LearningabstractThis paper proposes an introspective deep metric learning (IDML) framework for uncertainty-aware comparisons of images. Conventional deep metric learning methods focus on learning a discriminative embedding to describe the semantic features of images, which ignore the existence of uncertainty in each image resulting from noise or semantic ambiguity. Training without awareness of these uncertainties causes the model to overfit the annotated labels during training and produce overconfident judgments during inference. Motivated by this, we argue that a good similarity model should consider the semantic discrepancies with awareness of the uncertainty to better deal with ambiguous images for more robust training. To achieve this, we propose to represent an image using not only a semantic embedding but also an accompanying uncertainty embedding, which describes the semantic characteristics and ambiguity of an image, respectively. We further propose an introspective similarity metric to make similarity judgments between images considering both their semantic differences and ambiguities. The gradient analysis of the proposed metric shows that it enables the model to learn at an adaptive and slower pace to deal with the uncertainty during training. Our framework attains state-of-the-art performance on the widely used CUB-200-2011, Cars196, and Stanford Online Products datasets for image retrieval. We further evaluate our framework for image classification on the ImageNet-1 K, CIFAR-10, and CIFAR-100 datasets, which shows that equipping existing data mixing methods with the proposed introspective metric consistently achieves better results (e.g., +0.44% for CutMix on ImageNet-1 K). Chengkun Wang, Wenzhao Zheng, Jie Zhou 0001, Jiwen Lu |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | SPTR: Structure-Preserving Transformer for Unsupervised Indoor Depth CompletionabstractRecovering a dense depth map from a pair of indoor RGB and sparse depth images in an unsupervised manner is paramount in applications such as autonomous driving and 3D reconstruction. Most existing methods leverage sparse depth maps to directly estimate the dense depth map with the pixel-wise regression constraints over the input known depth. However, such regression constraints independently compare per-pixel depth values, which ignore the important 3D structures hidden behind depth maps and result in severe structural distortion and poor robustness. In this paper, we propose a Structure-Preserving Encoding (SPE) module by reformulating depth completion as the process of 3D structure generation. The generated structure should recover the complete scene and also consist with the known partial structure, so that the learned depth features from this task are able to encode rich structural information. In addition, SPE hierarchically interpolates and propagates the 3D structures into dense structure-aware positional encodings, which further boosts the information interactions between RGB and depth features via our transformer. Extensive experiments on VOID and NYUv2 demonstrate that SPTR outperforms the state-of-the-art methods by a large margin across various densities of input depths and a strong generalization ability to other datasets. Linqing Zhao, Wenzhao Zheng, Yueqi Duan, Jie Zhou 0001, Jiwen Lu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | StructLane: Leveraging Structural Relations for Lane DetectionabstractAccurately detecting the lanes plays a significant role in various autonomous and assistant driving scenarios. It is a highly structured task as lanes in the 3D world are continuous and parallel to each other. While most existing methods focus on how to inject structural priors into the representation of each lane, we propose a StructLane method to further leverage the structural relations among lanes for more accurate and robust lane detection. To achieve this, we explicitly encode the structural relations using a set of relational templates in a learned structural space. We then employ the attention mechanism to enable interactions between templates and image features to incorporate structural relational priors. Our StructLane can be applied to existing lane detection methods as a plug-and-play module to improve their performance. Extensive experiments on the widely used CULane, TuSimple, and LLAMAS datasets demonstrate that StructLane consistently improves the performance of state-of-the-art models across all datasets and backbones. Visualization results also demonstrate the robustness of our StructLane compared with existing methods due to the leverage of structural relations. Codes will be released at https://github.com/lqzhao/StructLane. Linqing Zhao, Wenzhao Zheng, Jie Zhou 0001, Jiwen Lu |
IEEE Trans. Image Process. | 2 |
| 2024 | Hardness-Aware Scene Synthesis for Semi-Supervised 3D Object Detectionabstract3D object detection aims to recover the 3D information of concerning objects and serves as the fundamental task of autonomous driving perception. Its performance greatly depends on the scale of labeled training data, yet it is costly to obtain high-quality annotations for point cloud data. This motivates the use of semi-supervised learning which can additionally exploit unlabeled data to further boost the performance. While 2D semi-supervised learning methods focus on generating pseudo-labels for unlabeled existing samples as supplements for training, the structural nature of 3D point cloud data facilitates the composition of objects and backgrounds to synthesize realistic scenes. Motivated by this, we propose a hardness-aware scene synthesis (HASS) method to generate adaptive synthetic scenes to improve the generalization of the detection models. We obtain pseudo-labels for unlabeled objects and generate diverse scenes with different compositions of objects and backgrounds. As the scene synthesis is sensitive to the quality of pseudo-labels, we further propose a hardness-aware strategy to reduce the effect of low-quality pseudo-labels. In addition, we maintain a dynamic pseudo- database to ensure the diversity and quality of synthetic scenes. Extensive experimental results on the widely used KITTI and Waymo datasets demonstrate the superiority of the proposed HASS method, which outperforms existing semi-supervised learning methods on 3D object detection. We also conducted a series of experiments to analyze the effectiveness of our method including pseudo-label quality analysis, the effect of different filtering and thresholding strategies, and ablations of each component. Wenzhao Zheng, Jiwen Lu, Haibin Yan |
IEEE Trans. Multim. | 2 |
| 2023 | A Simple Baseline for Multi-Camera 3D Object Detectionabstract3D object detection with surrounding cameras has been a promising direction for autonomous driving. In this paper, we present SimMOD, a Simple baseline for Multi-camera Object Detection, to solve the problem. To incorporate multiview information as well as build upon previous efforts on monocular 3D object detection, the framework is built on sample-wise object proposals and designed to work in a twostage manner. First, we extract multi-scale features and generate the perspective object proposals on each monocular image. Second, the multi-view proposals are aggregated and then iteratively refined with multi-view and multi-scale visual features in the DETR3D-style. The refined proposals are endto-end decoded into the detection results. To further boost the performance, we incorporate the auxiliary branches alongside the proposal generation to enhance the feature learning. Also, we design the methods of target filtering and teacher forcing to promote the consistency of two-stage training. We conduct extensive experiments on the 3D object detection benchmark of nuScenes to demonstrate the effectiveness of SimMOD and achieve competitive performance. Code will be available at https://github.com/zhangyp15/SimMOD. Wenzhao Zheng, Guan Huang 0003, Jiwen Lu, Jie Zhou 0001 |
AAAI | 2 |
| 2023 | Tri-Perspective View for Vision-Based 3D Semantic Occupancy PredictionabstractModern methods for vision-centric autonomous driving perception widely adopt the bird's-eye-view (BEV) representation to describe a 3D scene. Despite its better efficiency than voxel representation, it has difficulty describing the fine-grained 3D structure of a scene with a single plane. To address this, we propose a tri-perspective view (TPV) representation which accompanies BEV with two additional perpendicular planes. We model each point in the 3D space by summing its projected features on the three planes. To lift image features to the 3D TPV space, we further propose a transformer-based TPV encoder (TPVFormer) to obtain the TPV features effectively. We employ the attention mechanism to aggregate the image features corresponding to each query in each TPV plane. Experiments show that our model trained with sparse supervision effectively predicts the semantic occupancy for all voxels. We demonstrate for the first time that using only camera inputs can achieve comparable performance with LiDAR-based methods on the LiDAR segmentation task on nuScenes. Code: https://github.com/wzzheng/TPVFormer. Yuanhui Huang 0002, Wenzhao Zheng, Jie Zhou 0001, Jiwen Lu |
CVPR | 2 |
| 2023 | Deep Factorized Metric LearningabstractLearning a generalizable and comprehensive similarity metric to depict the semantic discrepancies between images is the foundation of many computer vision tasks. While existing methods approach this goal by learning an ensemble of embeddings with diverse objectives, the backbone network still receives a mix of all the training signals. Differently, we propose a deep factorized metric learning (DFML) method to factorize the training signal and employ different samples to train various components of the backbone network. We factorize the network to different sub-blocks and devise a learnable router to adaptively allocate the training samples to each sub-block with the objective to capture the most information. The metric model trained by DFML capture different characteristics with different sub-blocks and constitutes a generalizable metric when using all the sub-blocks. The proposed DFML achieves state-of-the-art performance on all three benchmarks for deep metric learning including CUB-200-20ll, Cars196, and Stanford Online Products. We also generalize DFML to the image classification task on ImageNet-1K and observe consistent improvement in accuracy/computation trade-off. Specifically, we improve the performance of ViT-B on ImageNet (+0.2% accuracy) with less computation load (-24% FLOPs).11Code is available at: https://github.com/wangck20/DFML. Chengkun Wang, Wenzhao Zheng, Jie Zhou 0001, Jiwen Lu |
CVPR | 2 |
| 2023 | OPERA: Omni-Supervised Representation Learning with Hierarchical SupervisionsabstractThe pretrain-finetune paradigm in modern computer vision facilitates the success of self-supervised learning, which tends to achieve better transferability than supervised learning. However, with the availability of massive labeled data, a natural question emerges: how to train a better model with both self and full supervision signals? In this paper, we propose Omni-suPErvised Representation leArning with hierarchical supervisions (OPERA) as a solution. We provide a unified perspective of supervisions from labeled and unlabeled data and propose a unified framework of fully supervised and self-supervised learning. We extract a set of hierarchical proxy representations for each image and impose self and full supervisions on the corresponding proxy representations. Extensive experiments on both convolutional neural networks and vision transformers demonstrate the superiority of OPERA in image classification, segmentation, and object detection.1 Chengkun Wang, Wenzhao Zheng, Jie Zhou 0001, Jiwen Lu |
ICCV | 2 |
| 2023 | SurroundOcc: Multi-Camera 3D Occupancy Prediction for Autonomous Drivingabstract3D scene understanding plays a vital role in vision-based autonomous driving. While most existing methods focus on 3D object detection, they have difficulty describing real-world objects of arbitrary shapes and infinite classes. Towards a more comprehensive perception of a 3D scene, in this paper, we propose a SurroundOcc method to predict the 3D occupancy with multi-camera images. We first extract multi-scale features for each image and adopt spatial 2D-3D attention to lift them to the 3D volume space. Then we apply 3D convolutions to progressively upsample the volume features and impose supervision on multiple levels. To obtain dense occupancy prediction, we design a pipeline to generate dense occupancy ground truth without expansive occupancy annotations. Specifically, we fuse multi-frame LiDAR scans of dynamic objects and static scenes separately. Then we adopt Poisson Reconstruction to fill the holes and voxelize the mesh to get dense occupancy labels. Extensive experiments on nuScenes and SemanticKITTI datasets demonstrate the superiority of our method. Code and dataset are available at https://github.com/weiyithu/SurroundOcc. Yi Wei 0003, Linqing Zhao, Wenzhao Zheng, Jie Zhou 0001, Jiwen Lu |
ICCV | 3 |
| 2023 | Token-Label Alignment for Vision TransformersabstractData mixing strategies (e.g., CutMix) have shown the ability to greatly improve the performance of convolutional neural networks (CNNs). They mix two images as inputs for training and assign them with a mixed label with the same ratio. While they are shown effective for vision transformers (ViTs), we identify a token fluctuation phenomenon that has suppressed the potential of data mixing strategies. We empirically observe that the contributions of input tokens fluctuate as forward propagating, which might induce a different mixing ratio in the output tokens. The training target computed by the original data mixing strategy can thus be inaccurate, resulting in less effective training. To address this, we propose a token-label alignment (TL-Align) method to trace the correspondence between transformed tokens and the original tokens to maintain a label for each to-ken. We reuse the computed attention at each layer for efficient token-label alignment, introducing only negligible additional training costs. Extensive experiments demonstrate that our method improves the performance of ViTs on image classification, semantic segmentation, objective detection, and transfer learning tasks. Code is available at: https://github.com/Euphoria16/TL-Align. Han Xiao 0010, Wenzhao Zheng, Jie Zhou 0001, Jiwen Lu |
ICCV | 2 |
| 2023 | Bort: Towards Explainable Neural Networks with Bounded Orthogonal Constraint
Borui Zhang, Wenzhao Zheng, Jie Zhou 0001, Jiwen Lu |
ICLR | 2 |
| 2023 | Deep Metric Learning With Adaptively Composite Dynamic ConstraintsabstractIn this paper, we propose a deep metric learning with adaptively composite dynamic constraints (DML-DC) method for image retrieval and clustering. Most existing deep metric learning methods impose pre-defined constraints on the training samples, which might not be optimal at all stages of training. To address this, we propose a learnable constraint generator to adaptively produce dynamic constraints to train the metric towards good generalization. We formulate the objective of deep metric learning under a proxy Collection, pair Sampling, tuple Construction, and tuple Weighting (CSCW) paradigm. For proxy collection, we progressively update a set of proxies using a cross-attention mechanism to integrate information from the current batch of samples. For pair sampling, we employ a graph neural network to model the structural relations between sample-proxy pairs to produce the preservation probabilities for each pair. Having constructed a set of tuples based on the sampled pairs, we further re-weight each training tuple to adaptively adjust its effect on the metric. We formulate the learning of the constraint generator as a meta-learning problem, where we employ an episode-based training scheme and update the generator at each iteration to adapt to the current model status. We construct each episode by sampling two subsets of disjoint labels to simulate the procedure of training and testing and use the performance of the one-gradient-updated metric on the validation subset as the meta-objective of the assessor. We conducted extensive experiments on five widely used benchmarks under two evaluation protocols to demonstrate the effectiveness of the proposed framework. Wenzhao Zheng, Jiwen Lu, Jie Zhou 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Attributable Visual Similarity LearningabstractThis paper proposes an attributable visual similarity learning (AVSL) framework for a more accurate and ex-plainable similarity measure between images. Most existing similarity learning methods exacerbate the unexplain-ability by mapping each sample to a single point in the em-bedding space with a distance metric (e.g., Mahalanobis distance, Euclidean distance). Motivated by the human se-mantic similarity cognition, we propose a generalized simi-larity learning paradigm to represent the similarity between two images with a graph and then infer the overall simi-larity accordingly. Furthermore, we establish a bottom-up similarity construction and top-down similarity inference framework to infer the similarity based on semantic hier-archy consistency. We first identify unreliable higher-level similarity nodes and then correct them using the most co-herent adjacent lower-level similarity nodes, which simulta-neously preserve traces for similarity attribution. Extensive experiments on the CUB-200-2011, Cars196, and Stanford Online Products datasets demonstrate significant improve-ments over existing deep similarity learning methods and verify the interpretability of our framework.11Code: https://github.com/zbr17/AVSL. Borui Zhang, Wenzhao Zheng, Jie Zhou 0001, Jiwen Lu |
CVPR | 2 |
| 2022 | Dimension Embeddings for Monocular 3D Object DetectionabstractMost existing deep learning-based approaches for monocular 3D object detection directly regress the dimensions of objects and overlook their importance in solving the illposed problem. In this paper, we propose a general method to learn appropriate embeddings for dimension estimation in monocular 3D object detection. Specifically, we consider two intuitive clues in learning the dimension-aware embeddings with deep neural networks. First, we constrain the pair-wise distance on the embedding space to reflect the similarity of corresponding dimensions so that the model can take advantage of inter-object information to learn more discriminative embeddings for dimension estimation. Second, we propose to learn representative shape templates on the dimension-aware embedding space. Through the attention mechanism, each object can interact with the learnable templates and obtain the attentive dimensions as the initial estimation, which is further refined by the combined features from both the object and the attentive templates. Experimental results on the well-established KITTI dataset demonstrate the proposed method of dimension embeddings can bring consistent improvements with negligible computation cost overhead. We achieve new state-of-the-art performance on the KITTI 3D object detection benchmark. Wenzhao Zheng, Guan Huang 0003, Dalong Du, Jie Zhou 0001, Jiwen Lu |
CVPR | 2 |
| 2022 | Dynamic Metric Learning with Cross-Level Concept Distillation
Wenzhao Zheng, Yuan Huang 0002, Borui Zhang, Jie Zhou 0001, Jiwen Lu |
ECCV (24) | 1 |
| 2021 | Deep Compositional Metric LearningabstractIn this paper, we propose a deep compositional metric learning (DCML) framework for effective and generalizable similarity measurement between images. Conventional deep metric learning methods minimize a discriminative loss to enlarge interclass distances while suppressing intraclass variations, which might lead to inferior generalization performance since samples even from the same class may present diverse characteristics. This motivates the adoption of the ensemble technique to learn a number of sub-embeddings using different and diverse subtasks. However, most subtasks impose weaker or contradictory constraints, which essentially sacrifices the discrimination ability of each sub-embedding to improve the generalization ability of their combination. To achieve a better generalization ability without compromising, we propose to separate the sub-embeddings from direct supervisions from the subtasks and apply the losses on different composites of the sub-embeddings. We employ a set of learnable compositors to combine the sub-embeddings and use a self-reinforced loss to train the compositors, which serve as relays to distribute the diverse training signals to avoid destroying the discrimination ability. Experimental results on the CUB-200-2011, Cars196, and Stanford Online Products datasets demonstrate the superior performance of our framework.1 Wenzhao Zheng, Chengkun Wang, Jiwen Lu, Jie Zhou 0001 |
CVPR | 1 |
| 2021 | Deep Relational Metric LearningabstractThis paper presents a deep relational metric learning (DRML) framework for image clustering and retrieval. Most existing deep metric learning methods learn an embedding space with a general objective of increasing interclass distances and decreasing intraclass distances. However, the conventional losses of metric learning usually suppress intraclass variations which might be helpful to identify samples of unseen classes. To address this problem, we propose to adaptively learn an ensemble of features that characterizes an image from different aspects to model both interclass and intraclass distributions. We further employ a relational module to capture the correlations among each feature in the ensemble and construct a graph to represent an image. We then perform relational inference on the graph to integrate the ensemble and obtain a relation-aware embedding to measure the similarities. Extensive experiments on the widely-used CUB-200-2011, Cars196, and Stanford Online Products datasets demonstrate that our framework improves existing deep metric learning methods and achieves very competitive results.1 Wenzhao Zheng, Borui Zhang, Jiwen Lu, Jie Zhou 0001 |
ICCV | 1 |
| 2021 | Hardness-Aware Deep Metric LearningabstractThis paper presents a hardness-aware deep metric learning (HDML) framework for image clustering and retrieval. Most existing deep metric learning methods employ the hard negative mining strategy to alleviate the lack of informative samples for training. However, this mining strategy only utilizes a subset of training data, which may not be enough to characterize the global geometry of the embedding space comprehensively. To address this problem, we perform linear interpolation on embeddings to adaptively manipulate their hardness levels and generate corresponding label-preserving synthetics for recycled training so that information buried in all samples can be fully exploited and the metric is always challenged with proper difficulty. As a single synthetic for each sample may still not be enough to describe the unobserved distributions of the training data which is crucial for the generalization performance, we further extend HDML to generate multiple synthetics for each sample. We propose a randomly hardness-aware deep metric learning (HDML-R) method and an adaptively hardness-aware deep metric learning (HDML-A) method to sample multiple random and adaptive directions, respectively, for hardness-aware synthesis. Since the generated multiple synthetics might not all be useful and adaptive, we propose a synthetic selection method with three criteria for the selection of qualified synthetics that are beneficial to the training of the metric. Extensive experimental results on the widely used CUB-200-2011, Cars196, Stanford Online Products, In-Shop Clothes Retrieval, and VehicleID datasets demonstrate the effectiveness of the proposed framework. Wenzhao Zheng, Jiwen Lu, Jie Zhou 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2020 | Deep Metric Learning via Adaptive Learnable AssessmentabstractIn this paper, we propose a deep metric learning via adaptive learnable assessment (DML-ALA) method for image retrieval and clustering, which aims to learn a sample assessment strategy to maximize the generalization of the trained metric. Unlike existing deep metric learning methods that usually utilize a fixed sampling strategy like hard negative mining, we propose a sequence-aware learnable assessor which re-weights each training example to train the metric towards good generalization. We formulate the learning of this assessor as a meta-learning problem, where we employ an episode-based training scheme and update the assessor at each iteration to adapt to the current model status. We construct each episode by sampling two subsets of disjoint labels to simulate the procedure of training and testing and use the performance of one-gradient-updated metric on the validation subset as the meta-objective of the assessor. Experimental results on the widely used CUB-200-2011, Cars196, and Stanford Online Products datasets demonstrate the effectiveness of the proposed approach. Wenzhao Zheng, Jiwen Lu, Jie Zhou 0001 |
CVPR | 1 |
| 2020 | Structural Deep Metric Learning for Room Layout Estimation
Wenzhao Zheng, Jiwen Lu, Jie Zhou 0001 |
ECCV (18) | 1 |
| 2020 | Deep Adversarial Metric LearningabstractLearning an effective distance measurement between sample pairs plays an important role in visual analysis, where the training procedure largely relies on hard negative samples. However, hard negative samples usually account for the tiny minority in the training set, which may fail to fully describe the data distribution close to the decision boundary. In this paper, we present a deep adversarial metric learning (DAML) framework to generate synthetic hard negatives from the original negative samples, which is widely applicable to existing supervised deep metric learning algorithms. Different from existing sampling strategies which simply ignore numerous easy negatives, our DAML aim to exploit them by generating synthetic hard negatives adversarial to the learned metric as complements. We simultaneously train the feature embedding and hard negative generator in an adversarial manner, so that adequate and targeted synthetic hard negatives are created to learn more precise distance metrics. As a single transformation may not be powerful enough to describe the global input space under the attack of the hard negative generator, we further propose a deep adversarial multi-metric learning (DAMML) method by learning multiple local transformations for more complete description. We simultaneously exploit the collaborative and competitive relationships among multiple metrics, where the metrics display unity against the generator for effective distance measurement as well as compete for more training data through a metric discriminator to avoid overlapping. Extensive experimental results on five benchmark datasets show that our DAML and DAMML effectively boost the performance of existing deep metric learning approaches through adversarial learning. Yueqi Duan, Jiwen Lu, Wenzhao Zheng, Jie Zhou 0001 |
IEEE Trans. Image Process. | 3 |
| 2019 | Hardness-Aware Deep Metric LearningabstractThis paper presents a hardness-aware deep metric learning (HDML) framework. Most previous deep metric learning methods employ the hard negative mining strategy to alleviate the lack of informative samples for training. However, this mining strategy only utilizes a subset of training data, which may not be enough to characterize the global geometry of the embedding space comprehensively. To address this problem, we perform linear interpolation on embeddings to adaptively manipulate their hard levels and generate corresponding label-preserving synthetics for recycled training, so that information buried in all samples can be fully exploited and the metric is always challenged with proper difficulty. Our method achieves very competitive performance on the widely used CUB-200-2011, Cars196, and Stanford Online Products datasets. Wenzhao Zheng, Zhaodong Chen 0001, Jiwen Lu, Jie Zhou 0001 |
CVPR | 1 |
| 2018 | Deep Adversarial Metric LearningabstractLearning an effective distance metric between image pairs plays an important role in visual analysis, where the training procedure largely relies on hard negative samples. However, hard negatives in the training set usually account for the tiny minority, which may fail to fully describe the distribution of negative samples close to the margin. In this paper, we propose a deep adversarial metric learning (DAML) framework to generate synthetic hard negatives from the observed negative samples, which is widely applicable to supervised deep metric learning methods. Different from existing metric learning approaches which simply ignore numerous easy negatives, the proposed DAML exploits them to generate potential hard negatives adversarial to the learned metric as complements. We simultaneously train the hard negative generator and feature embedding in an adversarial manner, so that more precise distance metrics can be learned with adequate and targeted synthetic hard negatives. Extensive experimental results on three benchmark datasets including CUB-200-2011, Cars196 and Stanford Online Products show that DAML effectively boosts the performance of existing deep metric learning approaches through adversarial learning. Yueqi Duan, Wenzhao Zheng, Xudong Lin 0003, Jiwen Lu, Jie Zhou 0001 |
CVPR | 2 |