VLDB 2026 Research / reviewers in the wild / expert
Zhongdao Wang
dblp:211/7222
· DBLP profile ↗
35ranked-venue papers
7as first author
28since 2021 · last 2026
0000-0002-4483-8783ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 31 · 6 first-author · 25 since 2021Graphics, computer vision, multimedia, augmented reality and games · 26 · 6 first-author · 19 since 2021Systems, architecture and hardware · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Unsupervised Temporal Correspondence Learning for Unified Video Object RemovalabstractVideo object removal aims at erasing a target object in the entire video and filling holes with plausible contents, given an object mask in the first frame as input. Existing solutions mostly break down the task into (supervised) mask tracking and (self-supervised) video completion, and then separately tackle them with tailored designs. In this paper, we introduce a new setup, coined as unified video object removal, where mask tracking and completion are addressed within a unified framework. Despite introducing more challenges, the setup is promising for future practical usage. We embrace the observation that these two sub-tasks have strong inherent connections in terms of pixel-level temporal correspondence. Making full use of the connections could be beneficial considering the complexity of both algorithm and deployment. We propose a single network linking the two sub-tasks by inferring temporal correspondences across multiple frames, i.e., correspondences between valid-valid (V-V) pixel pairs for mask tracking and correspondences between valid-hole (V-H) pixel pairs for video completion. Thanks to the unified setup, the network can be learned end-to-end in a totally unsupervised fashion without any annotations. We demonstrate that our method can generate visually pleasing results and perform favorably against existing separate solutions in realistic test cases. Zhongdao Wang, Jinglu Wang, Xiao Li 0030, Yali Li 0001, Yan Lu 0001, Shengjin Wang |
IEEE Trans. Image Process. | 1 |
| 2025 | ReNeg: Learning Negative Embedding with Reward GuidanceabstractIn text-to-image (T2I) generation applications, negative embeddings have proven to be a simple yet effective approach for enhancing generation quality. Typically, these negative embeddings are derived from user-defined negative prompts, which, while being functional, are not necessarily optimal. In this paper, we introduce ReNeg, an end-to-end method designed to learn improved Negative embeddings guided by a Reward model. We employ a reward feedback learning framework and integrate classifier-free guidance (CFG) into the training process, which was previously utilized only during inference, thus enabling the effec tive learning of negative embeddings. We also propose two strategies for learning both global and per-sample negative embeddings. Extensive experiments show that the learned negative embedding significantly outperforms null-text and handcrafted counterparts, achieving substantial improvements in human preference alignment. Additionally, the negative embedding learned within the same text embedding space exhibits strong generalization capabilities. For example, using the same CLIP text encoder, the negative embedding learned on SD1.5 can be seamlessly transferred to text-to-image or even text-to-video models such as ControlNet, ZeroScope, and VideoCrafter2, resulting in consistent performance improvements across the board. Code is available at https://github.com/AMD-AIG-AIMA/ReNeg. Xiaomin Li 0001, Yixuan Liu 0004, Takashi Isobe, Xu Jia 0012, Qinpeng Cui, Dong Zhou 0003, Dong Li 0025, You He 0002, Huchuan Lu, Zhongdao Wang, Emad Barsoum |
CVPR | 10 |
| 2025 | Turbo2K: Towards Ultra-Efficient and High-Quality 2K Video Synthesis
Wenbo Li 0002, Zhongdao Wang, Haoze Sun, Bangzhen Liu, Haoyu Chen 0003, Aoxue Li, Lei Zhu 0003 |
ICCV | 3 |
| 2025 | TurboVSR: Fantastic Video Upscalers and Where to Find ThemabstractDiffusion-based generative models have demonstrated exceptional promise in the video super-resolution (VSR) task, achieving a substantial advancement in detail generation relative to prior methods. However, these approaches face significant computational efficiency challenges. For instance, current techniques may require tens of minutes to super-resolve a mere 2-second, 1080p video. In this paper, we present TurboVSR, an ultra-efficient diffusion-based video super-resolution model. Our core design comprises three key aspects: (1) We employ an autoencoder with a high compression ratio of 32$\times$32$\times$8 to reduce the number of tokens. (2) Highly compressed latents pose substantial challenges for training. We introduce factorized conditioning to mitigate the learning complexity: we first learn to super-resolve the initial frame; subsequently, we condition the super-resolution of the remaining frames on the high-resolution initial frame and the low-resolution subsequent frames. (3) We convert the pre-trained diffusion model to a shortcut model to enable fewer sampling steps, further accelerating inference. As a result, TurboVSR performs on par with state-of-the-art VSR methods, while being 100+ times faster, taking only 7 seconds to process a 2-second long 1080p video. TurboVSR also supports image resolution by considering image as a one-frame video. Our efficient design makes SR beyond 1080p possible, results on 4K (3648$\times$2048) image SR show surprising fine details. Zhongdao Wang, Guodongfang Zhao, Bailan Feng, Wenbo Li 0002 |
ICCV | 1 |
| 2025 | PVMamba: Parallelizing Vision Mamba via Dynamic State Aggregation
Zhongdao Wang, Chao Ma 0004 |
ICCV | 2 |
| 2025 | TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference OptimizationabstractRecent advancements in reinforcement learning from human feedback have shown that utilizing fine-grained token-level reward models can substantially enhance the performance of Proximal Policy Optimization (PPO) in aligning large language models. However, it is challenging to leverage such token-level reward as guidance for Direct Preference Optimization (DPO), since DPO is formulated as a sequence-level bandit problem. To address this challenge, this work decomposes the sequence-level PPO into a sequence of token-level proximal policy optimization problems and then frames the problem of token-level PPO with token-level reward guidance, from which closed-form optimal token-level policy and the corresponding token-level reward can be derived. Using the obtained reward and Bradley-Terry model, this work establishes a framework of computable loss functions with token-level reward guidance for DPO, and proposes a practical reward guidance based on the induced DPO reward. This formulation enables different tokens to exhibit varying degrees of deviation from reference policy based on their respective rewards. Experiment results demonstrate that our method achieves substantial performance improvements over DPO, with win rate gains of up to 7.5 points on MT-Bench, 6.2 points on AlpacaEval 2, and 4.3 points on Arena-Hard. Code is available at https://github.com/dvlab-research/TGDPO. Mingkang Zhu, Xi Chen 0119, Zhongdao Wang, Bei Yu 0001, Hengshuang Zhao, Jiaya Jia |
ICML | 3 |
| 2025 | Efficient 3D Perception on Multi-Sweep Point Cloud with Gumbel Spatial PruningabstractThis paper studies point cloud perception within outdoor environments. Existing methods face limitations in recognizing objects located at a distance or occluded, due to the sparse nature of outdoor point clouds. In this work, we observe a significant mitigation of this problem by accumulating multiple temporally consecutive LiDAR sweeps, resulting in a remarkable improvement in perception accuracy. However, the computation cost also increases, hindering previous approaches from utilizing a large number of LiDAR sweeps. To tackle this challenge, we find that a considerable portion of points in the accumulated point cloud is redundant, and discarding these points has minimal impact on perception accuracy. We introduce a simple yet effective Gumbel Spatial Pruning (GSP) layer that dynamically prunes points based on a learned end-toend sampling. The GSP layer is decoupled from other network components and thus can be seamlessly integrated into existing point cloud network architectures. Extensive experiments show that our pruning strategy improves several perception algorithms in multiple tasks. Xueqian Zhang, Zhongdao Wang, Bailan Feng, Ke Xu 0001 |
ICRA | 4 |
| 2025 | Bi-Stream Knowledge Transfer for Semi-Supervised 3D Point Cloud Object Detectionabstract3D point cloud object detection plays an important role in autonomous driving. However, labeling 3D object boxes is expensive and time-consuming, limiting the number of annotated point clouds used in fully-supervised training. This has led to a rise in semi-supervised 3D object detection research, which aims to improve model performance by leveraging both labeled and unlabeled point clouds. Existing methods typically rely on the Mean Teacher (MT) paradigm, which uses unlabeled instances discovered by the teacher with confidence scores higher than certain thresholds to train the student. However, this leads to a loss of information as it overlooks ambiguous instances from the teacher that could also contain valuable knowledge. To address this issue, we propose a Bi-Stream Knowledge Transfer (BiKT) framework that fully exploits and transfers knowledge from both confident and ambiguous instances to the student network. Specifically, all pseudo labels are allocated into two knowledge streams, the deterministic stream and the noisy stream, and then subsequently guide the student network through bi-level supervision. We also introduce a Dynamic Stream Switching (DSS) algorithm that sets the stream boundary tailored for the current learning status. To further improve the quality of pseudo labels in the knowledge streams, we propose a Diffusive Label Denoising (DLD) module, which is trained by explicitly generating noised instances and then learning to denoise them, as in diffusion models. Experiments show the state-of-the-art performance of our BiKT on the ONCE validation and testing sets, as well as the robust generalization capability when confronted with diverse base detectors, increased amount of unlabeled data, and distinct datasets (e.g., Waymo), unveiling the power of semi-supervised learning in 3D object detection. Jilai Zheng, Pin Tang, Xiangxuan Ren, Zhongdao Wang, Chao Ma 0004 |
ICRA | 4 |
| 2025 | Reliable and Calibrated Semantic Occupancy Prediction by Hybrid Uncertainty LearningabstractVision-centric semantic occupancy prediction plays a crucial role in autonomous driving, which requires accurate and reliable predictions from low-cost sensors. Although having notably narrowed the accuracy gap with LiDAR, there is still few research effort to explore the reliability and calibration in predicting semantic occupancy from camera. In this paper, we conduct a comprehensive evaluation of existing semantic occupancy prediction models from a reliability perspective for the first time. Despite the gradual alignment of camera-based models with LiDAR in terms of accuracy, a significant reliability gap still persists. To address this concern, we propose ReliOcc, a method designed to enhance the reliability of camera-based occupancy networks. ReliOcc provides a plug-and-play scheme for existing models, which integrates hybrid uncertainty from individual voxels with sampling-based noise and relative voxels through mix-up learning. Besides, an uncertainty-aware calibration strategy is devised to further improve model reliability in offline mode. Extensive experiments under various settings demonstrate that ReliOcc significantly enhances the reliability of learned model while maintaining the accuracy for both geometric and semantic predictions. Notably, our proposed approach exhibits robustness to sensor failures and out of domain noises during inference. Song Wang 0019, Zhongdao Wang, Wentong Li 0001, Bailan Feng, Junbo Chen, Jianke Zhu |
IJCAI | 2 |
| 2025 | JarvisArt: Liberating Human Artistic Creativity via an Intelligent Photo Retouching AgentabstractPhoto retouching has become integral to contemporary visual storytelling, enabling users to capture aesthetics and express creativity. While professional tools such as Adobe Lightroom offer powerful capabilities, they demand substantial expertise and manual effort. In contrast, existing AI-based solutions provide automation but often suffer from limited adjustability and poor generalization, failing to meet diverse and personalized editing needs. To bridge this gap, we introduce JarvisArt, a multi-modal large language model (MLLM)-driven agent that understands user intent, mimics the reasoning process of professional artists, and intelligently coordinates over 200 retouching tools within Lightroom. JarvisArt undergoes a two-stage training process: an initial Chain-of-Thought supervised fine-tuning to establish basic reasoning and tool-use skills, followed by Group Relative Policy Optimization for Retouching (GRPO-R) to further enhance its decision-making and tool proficiency. We also propose the Agent-to-Lightroom Protocol to facilitate seamless integration with Lightroom. To evaluate performance, we develop MMArt-Bench, a novel benchmark constructed from real-world user edits. JarvisArt demonstrates user-friendly interaction, superior generalization, and fine-grained control over both global and local adjustments, paving a new avenue for intelligent photo retouching. Notably, it outperforms GPT-4o with a 60\% improvement in average pixel-level metrics on MMArt-Bench for content fidelity, while maintaining comparable instruction-following capabilities. Yunlong Lin, Zixu Lin, Kunjie Lin, Jinbin Bai, Panwang Pan, Chenxin Li, Haoyu Chen 0003, Zhongdao Wang, Xinghao Ding, Wenbo Li 0002, Shuicheng Yan |
NeurIPS | 8 |
| 2024 | SparseOcc: Rethinking Sparse Latent Representation for Vision-Based Semantic Occupancy PredictionabstractVision-based perception for autonomous driving requires an explicit modeling of a 3D space, where 2D latent representations are mapped and subsequent 3D operators are applied. However, operating on dense latent spaces introduces a cubic time and space complexity, which limits scalability in terms of perception range or spatial resolution. Existing approaches compress the dense representation using projections like Bird's Eye View (BEV) or Tri-Perspective View (TPV). Although efficient, these projections result in information loss, especially for tasks like semantic occupancy prediction. To address this, we propose SparseOcc, an efficient occupancy network inspired by sparse point cloud processing. It utilizes a lossless sparse latent representation with three key innovations. Firstly, a 3D sparse diffuser performs latent completion using spatially decomposed 3D sparse convolutional kernels. Secondly, a feature pyramid and sparse interpolation enhance scales with information from others. Finally, the transformer head is redesigned as a sparse variant. SparseOcc achieves a remarkable 74.9% reduction on FLOPs over the dense baseline. Interestingly, it also improves accuracy, from 12.8% to 14.1% mIOU, which in part can be attributed to the sparse representation's ability to avoid hallucinations on empty voxels. Pin Tang, Zhongdao Wang, Jilai Zheng, Xiangxuan Ren, Bailan Feng, Chao Ma 0004 |
CVPR | 2 |
| 2024 | DiffusionTrack: Point Set Diffusion Model for Visual Object TrackingabstractExisting Siamese or transformer trackers commonly pose visual object tracking as a one-shot detection problem, i.e., locating the target object in a single forward evaluation scheme. Despite the demonstrated success, these trackers may easily drift towards distractors with similar appear-ance due to the single forward evaluation scheme lacking self-correction. To address this issue, we cast visual tracking as a point set based denoising diffusion process and propose a novel generative learning based tracker, dubbed Diffusion Track. Our DiffusionTrack possesses two appealing properties: 1) It follows a novel noise-to-target tracking paradigm that leverages multiple denoising diffusion steps to localize the target in a dynamic searching man-ner per frame. 2) It models the diffusion process using a point set representation, which can better handle appear-ance variations for more precise localization. One side benefit is that DiffusionTrack greatly simplifies the post-processing, e.g. removing window penalty scheme. Without bells and whistles, our DiffusionTrack achieves leading per-formance over the state-of-the-art trackers and runs in real-time. The code is in https://github.com/VISION-SJTU/DiffusionTrack. Zhongdao Wang, Chao Ma 0004 |
CVPR | 2 |
| 2024 | PIXART-Σ: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation
Junsong Chen, Chongjian Ge, Enze Xie, Yue Wu 0002, Lewei Yao, Xiaozhe Ren, Zhongdao Wang, Ping Luo 0002, Huchuan Lu, Zhenguo Li |
ECCV (32) | 7 |
| 2024 | Segment, Lift and Fit: Automatic 3D Shape Labeling from 2D Prompts
Zhongdao Wang, Enze Xie, Bailan Feng, Ze Yuan, Ke Xu 0001, Ping Luo 0002 |
ECCV (84) | 3 |
| 2024 | OccGen: Generative Multi-modal 3D Occupancy Prediction for Autonomous Driving
Zhongdao Wang, Pin Tang, Jilai Zheng, Xiangxuan Ren, Bailan Feng, Chao Ma 0004 |
ECCV (20) | 2 |
| 2024 | VEON: Vocabulary-Enhanced Occupancy Prediction
Jilai Zheng, Pin Tang, Zhongdao Wang, Xiangxuan Ren, Bailan Feng, Chao Ma 0004 |
ECCV (54) | 3 |
| 2024 | LogoSticker: Inserting Logos Into Diffusion Models for Customized Generation
Mingkang Zhu, Xi Chen 0119, Zhongdao Wang, Hengshuang Zhao, Jiaya Jia |
ECCV (60) | 3 |
| 2024 | PixArt-α: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis
Junsong Chen, Chongjian Ge, Lewei Yao, Enze Xie, Zhongdao Wang, James T. Kwok, Ping Luo 0002, Huchuan Lu, Zhenguo Li |
ICLR | 6 |
| 2024 | Taming Diffusion Prior for Image Super-Resolution with Domain Shift SDEsabstractDiffusion-based image super-resolution (SR) models have attracted substantial interest due to their powerful image restoration capabilities. However, prevailing diffusion models often struggle to strike an optimal balance between efficiency and performance. Typically, they either neglect to exploit the potential of existing extensive pretrained models, limiting their generative capacity, or they necessitate a dozens of forward passes starting from random noises, compromising inference efficiency. In this paper, we present DoSSR, a $\textbf{Do}$main $\textbf{S}$hift diffusion-based SR model that capitalizes on the generative powers of pretrained diffusion models while significantly enhancing efficiency by initiating the diffusion process with low-resolution (LR) images. At the core of our approach is a domain shift equation that integrates seamlessly with existing diffusion models. This integration not only improves the use of diffusion prior but also boosts inference efficiency. Moreover, we advance our method by transitioning the discrete shift process to a continuous formulation, termed as DoS-SDEs. This advancement leads to the fast and customized solvers that further enhance sampling efficiency. Empirical results demonstrate that our proposed method achieves state-of-the-art performance on synthetic and real-world datasets, while notably requiring $\textbf{\emph{only 5 sampling steps}}$. Compared to previous diffusion prior based methods, our approach achieves a remarkable speedup of 5-7 times, demonstrating its superior efficiency. Qinpeng Cui, Yixuan Liu 0004, Xinyi Zhang 0008, Qiqi Bao 0001, Qingmin Liao, liwang Amd, Zicheng Liu 0001, Zhongdao Wang, Emad Barsoum |
NeurIPS | 9 |
| 2024 | QuadMamba: Learning Quadtree-based Selective Scan for Visual State Space ModelabstractRecent advancements in State Space Models, notably Mamba, have demonstrated superior performance over the dominant Transformer models, particularly in reducing the computational complexity from quadratic to linear. Yet, difficulties in adapting Mamba from language to vision tasks arise due to the distinct characteristics of visual data, such as the spatial locality and adjacency within images and large variations in information granularity across visual tokens. Existing vision Mamba approaches either flatten tokens into sequences in a raster scan fashion, which breaks the local adjacency of images, or manually partition tokens into windows, which limits their long-range modeling and generalization capabilities. To address these limitations, we present a new vision Mamba model, coined QuadMamba, that effectively captures local dependencies of varying granularities via quadtree-based image partition and scan. Concretely, our lightweight quadtree-based scan module learns to preserve the 2D locality of spatial regions within learned window quadrants. The module estimates the locality score of each token from their features, before adaptively partitioning tokens into window quadrants. An omnidirectional window shifting scheme is also introduced to capture more intact and informative features across different local regions. To make the discretized quadtree partition end-to-end trainable, we further devise a sequence masking strategy based on Gumbel-Softmax and its straight-through gradient estimator. Extensive experiments demonstrate that QuadMamba achieves state-of-the-art performance in various vision tasks, including image classification, object detection, instance segmentation, and semantic segmentation. Our code and models will be released. Zhongdao Wang, Chao Ma 0004 |
NeurIPS | 3 |
| 2023 | Identity-Seeking Self-Supervised Representation Learning for Generalizable Person Re-identificationabstractThis paper aims to learn a domain-generalizable (DG) person re-identification (ReID) representation from large-scale videos without any annotation. Prior DG ReID methods employ limited labeled data for training due to the high cost of annotation, which restricts further advances. To overcome the barriers of data and annotation, we propose to utilize large-scale unsupervised data for training. The key issue lies in how to mine identity information. To this end, we propose an Identity-seeking Self-supervised Representation learning (ISR) method. ISR constructs positive pairs from inter-frame images by modeling the instance association as a maximum-weight bipartite matching problem. A reliability-guided contrastive loss is further presented to suppress the adverse impact of noisy positive pairs, ensuring that reliable positive pairs dominate the learning process. The training cost of ISR scales approximately linearly with the data size, making it feasible to utilize large-scale data for training. The learned representation exhibits superior generalization ability. Without human annotation and fine-tuning, ISR achieves 87.0% Rank-1 on Market-1501 and 56.4% Rank-1 on MSMT17, outperforming the best supervised domain-generalizable method by 5.0% and 19.5%, respectively. In the pre-training→fine-tuning scenario, ISR achieves state-of-the-art performance, with 88.4% Rank-1 on MSMT17. The code is at https://github.com/dcp15/ISR_ICCV2023_Oral. Zhaopeng Dou, Zhongdao Wang, Yali Li 0001, Shengjin Wang |
ICCV | 2 |
| 2023 | MetaBEV: Solving Sensor Failures for 3D Detection and Map SegmentationabstractPerception systems in modern autonomous driving vehicles typically take inputs from complementary multi-modal sensors, e.g., LiDAR and cameras. However, in real-world applications, sensor corruptions and failures lead to inferior performances, thus compromising autonomous safety. In this paper, we propose a robust framework, called MetaBEV, to address extreme real-world environments, involving overall six sensor corruptions and two extreme sensor-missing situations. In MetaBEV, signals from multiple sensors are first processed by modal-specific encoders. Subsequently, a set of dense BEV queries are initialized, termed meta-BEV. These queries are then processed iteratively by a BEV-Evolving decoder, which selectively aggregates deep features from either LiDAR, cameras, or both modalities. The updated BEV representations are further leveraged for multiple 3D prediction tasks. Additionally, we introduce a new M2oE structure to alleviate the performance drop on distinct tasks in multi-task joint learning. Finally, MetaBEV is evaluated on the nuScenes dataset with 3D object detection and BEV map segmentation tasks. Experiments show MetaBEV outperforms prior arts by a large margin on both full and corrupted modalities. For instance, when the LiDAR signal is missing, MetaBEV improves 35.5% detection NDS and 17.7% segmentation mIoU upon the vanilla BEVFusion [25] model; and when the camera signal is absent, MetaBEV still achieves 69.2% NDS and 53.7% mIoU, which is even higher than previous works that perform on full-modalities. Moreover, MetaBEV performs moderately against previous methods in both canonical perception and multi-task learning settings, refreshing state-of-the-art nuScenes BEV map segmentation with 70.4% mIoU. Chongjian Ge, Junsong Chen, Enze Xie, Zhongdao Wang, Lanqing Hong, Huchuan Lu, Zhenguo Li, Ping Luo 0002 |
ICCV | 4 |
| 2022 | Reliability-Aware Prediction via Uncertainty Learning for Person Image Retrieval
Zhaopeng Dou, Zhongdao Wang, Yali Li 0001, Shengjin Wang |
ECCV (14) | 2 |
| 2022 | How to Synthesize a Large-Scale and Trainable Micro-Expression Dataset?
Yuchi Liu, Zhongdao Wang, Tom Gedeon, Liang Zheng 0001 |
ECCV (8) | 2 |
| 2022 | Progressive-Granularity Retrieval Via Hierarchical Feature Alignment for Person Re-IdentificationabstractPerson re-identification (re-ID) aims to match pedestrian images from non-overlapping cameras. It is a challenging task because of the feature misalignment problem caused by occlusion. In this paper, inspired by the coarse-to-fine nature of human perception, we propose a novel Progressive-Granularity Retrieval (PGR) method to tackle this issue. Specifically, (i) we define instance-level, part-level and pixel-level features for an image. PGR learns these features by a single feature extractor to capture hierarchical clues in the image. (ii) These features are inherently related but different in perceptual granularity, and they can provide complementary information. For each type of feature, we propose a corresponding similarity metric to achieve hierarchical feature alignment. (iii) In training, we learn the model end-to-end. In inference, a progressive retrieval strategy is introduced to efficiently aggregate the complementary information provided by these features. Extensive experiments on three bench-marks of both occluded and holistic-body re-ID tasks show the effectiveness of the proposed method. Especially, our method significantly outperforms state-of-the-art by 4.5% Rank-1 score on the challenging Occluded-Duke dataset. Zhaopeng Dou, Zhongdao Wang, Yali Li 0001, Shengjin Wang |
ICASSP | 2 |
| 2022 | Self-Supervised Learning via Maximum Entropy CodingabstractA mainstream type of current self-supervised learning methods pursues a general-purpose representation that can be well transferred to downstream tasks, typically by optimizing on a given pretext task such as instance discrimination. In this work, we argue that existing pretext tasks inevitably introduce biases into the learned representation, which in turn leads to biased transfer performance on various downstream tasks. To cope with this issue, we propose Maximum Entropy Coding (MEC), a more principled objective that explicitly optimizes on the structure of the representation, so that the learned representation is less biased and thus generalizes better to unseen downstream tasks. Inspired by the principle of maximum entropy in information theory, we hypothesize that a generalizable representation should be the one that admits the maximum entropy among all plausible representations. To make the objective end-to-end trainable, we propose to leverage the minimal coding length in lossy data coding as a computationally tractable surrogate for the entropy, and further derive a scalable reformulation of the objective that allows fast computation. Extensive experiments demonstrate that MEC learns a more generalizable representation than previous methods based on specific pretext tasks. It achieves state-of-the-art performance consistently on various downstream tasks, including not only ImageNet linear probe, but also semi-supervised classification, object detection, instance segmentation, and object tracking. Interestingly, we show that existing batch-wise and feature-wise self-supervised objectives could be seen equivalent to low-order approximations of MEC. Code and pre-trained models are available at https://github.com/xinliu20/MEC. Zhongdao Wang, Yali Li 0001, Shengjin Wang |
NeurIPS | 2 |
| 2022 | Adaptive Affinity for Associations in Multi-Target Multi-Camera TrackingabstractData associations in multi-target multi-camera tracking (MTMCT) usually estimate affinity directly from re-identification (re-ID) feature distances. However, we argue that it might not be the best choice given the difference in matching scopes between re-ID and MTMCT problems. Re-ID systems focus on global matching, which retrieves targets from all cameras and all times. In contrast, data association in tracking is a local matching problem, since its candidates only come from neighboring locations and time frames. In this paper, we design experiments to verify such misfit between global re-ID feature distances and local matching in tracking, and propose a simple yet effective approach to adapt affinity estimations to corresponding matching scopes in MTMCT. Instead of trying to deal with all appearance changes, we tailor the affinity metric to specialize in ones that might emerge during data associations. To this end, we introduce a new data sampling scheme with temporal windows originally used for data associations in tracking. Minimizing the mismatch, the adaptive affinity module brings significant improvements over global re-ID distance, and produces competitive performance on CityFlow and DukeMTMC datasets. Yunzhong Hou, Zhongdao Wang, Shengjin Wang, Liang Zheng 0001 |
IEEE Trans. Image Process. | 2 |
| 2021 | Do Different Tracking Tasks Require Different Appearance Models?abstractTracking objects of interest in a video is one of the most popular and widely applicable problems in computer vision. However, with the years, a Cambrian explosion of use cases and benchmarks has fragmented the problem in a multitude of different experimental setups. As a consequence, the literature has fragmented too, and now novel approaches proposed by the community are usually specialised to fit only one specific setup. To understand to what extent this specialisation is necessary, in this work we present UniTrack, a solution to address five different tasks within the same framework. UniTrack consists of a single and task-agnostic appearance model, which can be learned in a supervised or self-supervised fashion, and multiple ``heads'' that address individual tasks and do not require training. We show how most tracking tasks can be solved within this framework, and that the same appearance model can be successfully used to obtain results that are competitive against specialised methods for most of the tasks considered. The framework also allows us to analyse appearance models obtained with the most recent self-supervised methods, thus extending their evaluation and comparison to a larger variety of important problems. Zhongdao Wang, Hengshuang Zhao, Yali Li 0001, Shengjin Wang, Philip Torr 0001, Luca Bertinetto |
NeurIPS | 1 |
| 2020 | Softmax Dissection: Towards Understanding Intra- and Inter-Class Objective for Embedding LearningabstractThe softmax loss and its variants are widely used as objectives for embedding learning applications like face recognition. However, the intra- and inter-class objectives in Softmax are entangled, therefore a well-optimized inter-class objective leads to relaxation on the intra-class objective, and vice versa. In this paper, we propose to dissect Softmax into independent intra- and inter-class objective (D-Softmax) with a clear understanding. It is straightforward to tune each part to the best state with D-Softmax as objective.Furthermore, we find the computation of the inter-class part is redundant and propose sampling-based variants of D-Softmax to reduce the computation cost. The face recognition experiments on regular-scale data show D-Softmax is favorably comparable to existing losses such as SphereFace and ArcFace. Experiments on massive-scale data show the fast variants significantly accelerates the training process (such as 64×) with only a minor sacrifice in performance, outperforming existing acceleration methods of Softmax in terms of both performance and efficiency. Lanqing He, Zhongdao Wang, Yali Li 0001, Shengjin Wang |
AAAI | 2 |
| 2020 | Circle Loss: A Unified Perspective of Pair Similarity OptimizationabstractThis paper provides a pair similarity optimization viewpoint on deep feature learning, aiming to maximize the within-class similarity $s_p$ and minimize the between-class similarity $s_n$. We find a majority of loss functions, including the triplet loss and the softmax cross-entropy loss, embed $s_n$ and $s_p$ into similarity pairs and seek to reduce $(s_n-s_p)$. Such an optimization manner is inflexible, because the penalty strength on every single similarity score is restricted to be equal. Our intuition is that if a similarity score deviates far from the optimum, it should be emphasized. To this end, we simply re-weight each similarity to highlight the less-optimized similarity scores. It results in a Circle loss, which is named due to its circular decision boundary. The Circle loss has a unified formula for two elemental deep feature learning paradigms, \emph {i.e.}, learning with class-level labels and pair-wise labels. Analytically, we show that the Circle loss offers a more flexible optimization approach towards a more definite convergence target, compared with the loss functions optimizing $(s_n-s_p)$. Experimentally, we demonstrate the superiority of the Circle loss on a variety of deep feature learning tasks. On face recognition, person re-identification, as well as several fine-grained image retrieval datasets, the achieved performance is on par with the state of the art. Yifan Sun 0003, Changmao Cheng, Yuhan Zhang 0004, Chi Zhang 0026, Liang Zheng 0001, Zhongdao Wang |
CVPR | 6 |
| 2020 | Towards Real-Time Multi-Object Tracking
Zhongdao Wang, Liang Zheng 0001, Yixuan Liu 0004, Yali Li 0001, Shengjin Wang |
ECCV (11) | 1 |
| 2020 | CycAs: Self-supervised Cycle Association for Learning Re-identifiable Descriptions
Zhongdao Wang, Liang Zheng 0001, Yixuan Liu 0004, Yifan Sun 0003, Yali Li 0001, Shengjin Wang |
ECCV (11) | 1 |
| 2020 | Node-Adaptive Multi-Graph Fusion Using Extreme Value TheoryabstractThis letter considers the problem of grouping data by their underlying categories with inputs from multiple sources, known as the multi-view clustering problem. One of the most fundamental challenges lies in how to benefit from the complementary information in the multi-view data, so that clustering on such data consistently achieves higher accuracy than clustering on each single-view component. In this letter, to tackle the multi-view clustering problem, we propose a novel approach to fuse multiple affinity graphs computed in each single view to a unified affinity graph, so that single-view affinity-based clustering methods can be accordingly applied on it. The edges in the unified affinity graph between a node and its neighbors are computed as weighted average over the corresponding edges from multiple single graphs, and the weights here are adaptive to each node, estimated using the Extreme Value Theory (EVT). Experiments on two challenging multi-view clustering tasks show that, combined with existing off-the-shelf single-view clustering algorithms, the proposed graph fusion method brings consistently performance gain compared with naive graph fusion baselines. Zhongdao Wang, Yali Li 0001, Shengjin Wang |
IEEE Signal Process. Lett. | 2 |
| 2019 | Linkage Based Face Clustering via Graph Convolution NetworkabstractIn this paper, we present an accurate and scalable approach to the face clustering task. We aim at grouping a set of faces by their potential identities. We formulate this task as a link prediction problem: a link exists between two faces if they are of the same identity. The key idea is that we find the local context in the feature space around an instance (face) contains rich information about the linkage relationship between this instance and its neighbors. By constructing sub-graphs around each instance as input data, which depict the local context, we utilize the graph convolution network (GCN) to perform reasoning and infer the likelihood of linkage between pairs in the sub-graphs. Experiments show that our method is more robust to the complex distribution of faces than conventional methods, yielding favorably comparable results to state-of-the-art methods on standard face clustering benchmarks, and is scalable to large datasets. Furthermore, we show that the proposed method does not need the number of clusters as prior, is aware of noises and outliers, and can be extended to a multi-view version for more accurate clustering accuracy. Zhongdao Wang, Liang Zheng 0001, Yali Li 0001, Shengjin Wang |
CVPR | 1 |
| 2017 | Orientation Invariant Feature Embedding and Spatial Temporal Regularization for Vehicle Re-identificationabstractIn this paper, we tackle the vehicle Re-identification (ReID) problem which is of great importance in urban surveillance and can be used for multiple applications. In our vehicle ReID framework, an orientation invariant feature embedding module and a spatial-temporal regularization module are proposed. With orientation invariant feature embedding, local region features of different orientations can be extracted based on 20 key point locations and can be well aligned and combined. With spatial-temporal regularization, the log-normal distribution is adopted to model the spatial-temporal constraints and the retrieval results can be refined. Experiments are conducted on public vehicle ReID datasets and our proposed method achieves state-of-the-art performance. Investigations of the proposed framework is conducted, including the landmark regressor and comparisons with attention mechanism. Both the orientation invariant feature embedding and the spatio-temporal regularization achieve considerable improvements. Zhongdao Wang, Luming Tang, Xihui Liu, Zhuliang Yao, Shuai Yi, Shengjin Wang, Hongsheng Li 0001, Xiaogang Wang 0001 |
ICCV | 1 |