Zhiwen Cao

dblp:42/10448 · DBLP profile ↗
← Back
24ranked-venue papers
8as first author
22since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 5 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 HiFi-Mesh: High-Fidelity Efficient 3D Mesh Generation via Compact Autoregressive Dependence
abstract
High-fidelity 3D meshes can be tokenized into one-dimension (1D) sequences and directly modeled using autoregressive approaches for faces and vertices. However, existing methods suffer from insufficient resource utilization, resulting in slow inference and the ability to handle only small-scale sequences, which severely constrains the expressible structural details. We introduce the Latent Autoregressive Network (LANE), which incorporates compact autoregressive dependencies in the generation process, achieving a 6× improvement in maximum generatable sequence length compared to existing methods. To further accelerate inference, we propose the Adaptive Computation Graph Reconfiguration (AdaGraph) strategy, which effectively overcomes the efficiency bottleneck of traditional serial inference through spatiotemporal decoupling in the generation process. Experimental validation demonstrates that LANE achieves superior performance across generation speed, structural detail, and geometric consistency, providing an effective solution for high-quality 3D mesh generation.
Tao Tan 0002, Qinquan Gao, Zhiwen Cao, Xiaohong Liu 0001, Yue Sun 0001
AAAI4
2025 Desmoke-VIT: A Red Bias-Resistant Approach to Unpaired Laparoscopic Smoke Removal
abstract
Robot Assisted Surgery (RAS) has seen rapid advancements, enhancing precision while reducing invasiveness and improving patient outcomes. However, smoke generated during procedures obstructs the surgical field, posing a significant challenge. Among various surgical scenarios, laparoscopic smoke removal has made some progress, demonstrating handling of smoke distribution. Nevertheless, existing approaches demonstrate limited efficacy in addressing laparoscopic images, often resulting in incomplete smoke removal, particularly in the presence of the red bias phenomenon. To address these issues, we propose Desmoke-VIT, a novel method for smoke removal that leverages attention mechanisms to capture intricate image features. Our approach builds upon the CycleGAN framework for unpaired image-toimage translation and incorporates Vision Transformer modules to enhance feature representation and extraction. Comprehensive evaluations on related datasets demonstrate that Desmoke-VIT achieves state-of-the-art performance across automated metrics. Our model attains these exceptional results while operating with the fewest parameters. By utilizing a style transfer paradigm, our model effectively removes smoke from surgical images while preserving critical visual details, enabling clearer and more reliable surgical imagery. Our implementation is available at link.
Hejing Cai, Zhiwen Cao, Cui Tang
BIBM4
2025 Probabilistic Token Alignment for Large Language Model Fusion
abstract
Training large language models (LLMs) from scratch can yield models with unique functionalities and strengths, but it is costly and often leads to redundant capabilities. A more cost-effective alternative is to fuse existing pre-trained LLMs with different architectures into a more powerful model. However, a key challenge in existing model fusion is their dependence on manually predefined vocabulary alignment, which may not generalize well across diverse contexts, leading to performance degradation in several evaluation. To solve this, we draw inspiration from distribution learning and propose the probabilistic token alignment method as a general and soft mapping for alignment, named as PTA-LLM. Our approach innovatively reformulates token alignment into a classic mathematical problem: optimal transport, seamlessly leveraging distribution-aware learning to facilitate more coherent model fusion. Apart from its inherent generality, PTA-LLM exhibits interpretability from a distributional perspective, offering insights into the essence of the token alignment. Empirical results demonstrate that probabilistic token alignment enhances the target model's performance across multiple capabilities.
Runjia Zeng, James Liang, Cheng Han 0001, Zhiwen Cao, Xiaojun Quan, Victor Y. Chen, Lifu Huang, Tong Geng, Qifan Wang 0001, Dongfang Liu
NeurIPS4
2025 A Data Perspective on Enhanced Identity Preservation for Diffusion Personalization
abstract
Large text-to-image models have revolutionized the ability to generate imagery using natural language. However, particularly unique or personal visual concepts, such as pets and furniture, will not be captured by the original model. This has led to interest in how to personalize a text-to-image model. Despite significant progress, this task remains a formidable challenge, particularly in preserving the subject's identity. Most researchers attempt to address this issue by modifying model architectures. These methods are capable of keeping the subject structure and color but fail to preserve identity details. Towards this issue, our approach takes a data-centric perspective. We introduce a novel regularization dataset generation strategy on both the text and image level. This strategy enables the model to preserve fine details of the desired subjects, such as text and logos. Our method is architecture-agnostic and can be flexibly applied on various text-to-image models. We show on established benchmarks that our data-centric approach forms the new state of the art in terms of identity preservation and text alignment.
Xingzhe He, Zhiwen Cao, Nicholas I. Kolkin, Lantao Yu, Kun Wan 0001, Helge Rhodin, Ratheesh Kalarot
WACV2
2025 Graph-Regularized Consensus Learning and Diversity Representation for unsupervised multi-view feature selection
Shengke Xu, Xijiong Xie, Zhiwen Cao
Knowl. Based Syst.3
2025 Partition-Level Tensor Learning-Based Multiview Unsupervised Feature Selection
abstract
Multiview unsupervised feature selection is an emerging direction in the machine learning community because of its ability to identify informative patterns and reduce the dimensionality of multiview data. Although numerous methods have been proposed and shown to be effective, they have some limitations: 1) most existing algorithms fail to improve the model performance along the view dimension; 2) they rarely incorporate more discriminative partition information; and 3) the negative effects of marginal samples are not considered. To solve these problems, we propose a novel method termed as partition-level tensor learning-based multiview unsupervised feature selection (PTFS). The proposed method optimizes a low-rank constrained tensor assembled by the inner product of base partition matrices. By doing so, PTFS simultaneously leverages the high-order view correlation and indirectly integrates discriminative partition information. Besides, a statistic-based adaptive self-paced strategy is introduced to ensure that confident samples are prioritized for training the model. Moreover, an effective alternating optimization method is designed to solve the resulting optimization problem. Extensive experiments on ten datasets demonstrate the effectiveness and efficiency of the proposed method compared to the state-of-the-art methods. The code is available at https://github.com/HdTgon/2023-TNNLS-PTFS.
Zhiwen Cao, Xijiong Xie
IEEE Trans. Neural Networks Learn. Syst.1
2024 ProMotion: Prototypes as Motion Learners
abstract
In this work, we introduce PRoMoTION, a unified proto-typical transformer-based framework engineered to model fundamental motion tasks. PRoMoTION offers a range of compelling attributes that set it apart from current task-specific paradigms. (1) We adopt a prototypical perspective, establishing a unified paradigm that harmonizes disparate motion learning approaches. This novel paradigm stream-lines the architectural design, enabling the simultaneous assimilation of diverse motion information. (2) We capitalize on a dual mechanism involving the feature denoiser and the prototypical learner to decipher the intricacies of motion. This approach effectively circumvents the pitfalls of ambiguity in pixel-wise feature matching, significantly bolstering the robustness of motion representation. (3)) We demon-strate a profound degree of transferability across distinct motion patterns. This inherent versatility reverberates robustly across a comprehensive spectrum of both 2D and 3D downstream tasks. Empirical results demonstrate that PRoMOTION outperforms various well-known specialized architectures, achieving 0.54 and 0.054$AbsRel$error on the Sintel and KITTI depth datasets, 1.04 and 2.01 average endpoint error on the clean and final pass of Sintel flow benchmark, and 4.30 F1-all error on the KITTI flow bench-mark. For its efficacy, we hope our work can catalyze a paradigm shift in universal models in computer vision.
Yawen Lu, Dongfang Liu, Qifan Wang 0001, Cheng Han 0001, Yiming Cui 0002, Zhiwen Cao, Xueling Zhang, Victor Y. Chen, Heng Fan 0001
CVPR6
2024 Dr.Bokeh: DiffeRentiable Occlusion-Aware Bokeh Rendering
abstract
Bokeh is widely used in photography to draw attention to the subject while effectively isolating distractions in the background. Computational methods can simulate bokeh effects without relying on a physical camera lens, but the inaccurate lens modeling in existing filtering-based meth-ods leads to artifacts that need post-processing or learning-based methods to fix. We propose Dr.Bokeh, a novel ren-dering method that addresses the issue by directly correcting the defect that violates physics in the current filtering-based bokeh rendering equation. Dr.Bokeh first preprocesses the input RGBD to obtain a layered scene representation. Dr.Bokeh then takes the layered representation and user-defined lens parameters to render photo-realistic lens blur based on the novel occlusion-aware bokeh rendering method. Experiments show that the non-learning based renderer Dr.Bokeh outperforms state-of-the-art bokeh ren-dering algorithms in terms of photo-realism. In addition, extensive quantitative and qualitative evaluations show that the more accurate lens model pushes the limit of depth-from-defocus.
Yichen Sheng, Zixun Yu, Lu Ling, Zhiwen Cao, Xuaner Cecilia Zhang, Xin Lu 0006, Ke Xian, Haiting Lin, Bedrich Benes
CVPR4
2024 Prototypical Transformer As Unified Motion Learners
abstract
In this work, we introduce the Prototypical Transformer (ProtoFormer), a general and unified framework that approaches various motion tasks from a prototype perspective. ProtoFormer seamlessly integrates prototype learning with Transformer by thoughtfully considering motion dynamics, introducing two innovative designs. First, Cross-Attention Prototyping discovers prototypes based on signature motion patterns, providing transparency in understanding motion scenes. Second, Latent Synchronization guides feature representation learning via prototypes, effectively mitigating the problem of motion uncertainty. Empirical results demonstrate that our approach achieves competitive performance on popular motion tasks such as optical flow and scene depth. Furthermore, it exhibits generality across various downstream tasks, including object tracking and video stabilization.
Cheng Han 0001, Yawen Lu, James Liang, Zhiwen Cao, Qifan Wang 0001, Qiang Guan, Sohail A. Dianat, Raghuveer M. Rao, Tong Geng, Zhiqiang Tao, Dongfang Liu
ICML5
2024 Structure learning with consensus label information for multi-view unsupervised feature selection
Zhiwen Cao, Xijiong Xie
Expert Syst. Appl.1
2024 Multi-view unsupervised feature selection with consensus partition and diverse graph
Zhiwen Cao, Xijiong Xie
Inf. Sci.1
2024 Multi-view unsupervised complementary feature selection with multi-order similarity learning
Zhiwen Cao, Xijiong Xie
Knowl. Based Syst.1
2024 Low-Complexity Subarray-Based Adaptive Detection for Multichannel Application in Inhomogeneous Clutter Environments
abstract
Multichannel adaptive detection (MAD) can achieve better performance compared with the constant false alarm rate (CFAR) methods in target detection in inhomogeneous clutter. However, its application still faces many challenges, such as the lack of sufficient training samples and huge computational costs. In this letter, a low-complexity reduced-dimension MAD (RD-MAD) scheme in an inhomogeneous clutter environment is proposed based on arbitrary subarray synthesis. By this scheme, we derive the RD generalized likelihood ratio test (GLRT). The theoretical performance of the proposed method is analyzed, including the CFAR property, RD performance, and computational complexity. Finally, with tri-channel X-band airborne radar real data, the detection performance of the proposed RD-MAD scheme is verified. Compared with the existing detectors, the proposed detector can provide better detection performance in sample-insufficient environments with much lower computational complexity.
Zhiwen Cao, Ning Cui, Kun Xing, Weijian Liu 0001, Zhongjun Yu
IEEE Geosci. Remote. Sens. Lett.1
2024 Label-Efficient Video Object Segmentation With Motion Clues
abstract
Video object segmentation (VOS) plays an important role in video analysis and understanding, which in turn facilitates a number of diverse applications, including video editing, video rendering, and augmented reality / virtual reality. However, existing deep learning-based approaches rely heavily on a large number of pixel-wise annotated video frames to achieve promising results, which is notoriously laborious and costly. To address this, in this paper, we formulate unsupervised video object detection by exploring simulated dense labels and explicit motion clues. Specifically, we first propose an effective video label generator network based on the sparsely annotated frames and the flow motion between them. It can largely alleviate our dependence and limitation on the sparse labels. Furthermore, we propose a transformer-based architecture to model the appearance and motion clues simultaneously with the cross-attention module, in order to maximally overcome non-linear motion with potential occlusions. Extensive experiments show that the proposed method outperforms recent VOS methods on four popular benchmarks (i.e., DAVIS-16, FBMS, Youtube-VOS and SegTrack-v2). Moreover, the proposed method can be further applied to a wide range of wild scenes such as wild forests and animals. Because of its effectiveness and generalization, we believe that our method could serve as a useful basis for alleviating the dependence on dense annotation in video data.
Yawen Lu, Jie Zhang 0066, Su Sun, Zhiwen Cao, Songlin Fei, Baijian Yang 0001, Victor Y. Chen
IEEE Trans. Circuits Syst. Video Technol.5
2023 E2VPT: An Effective and Efficient Approach for Visual Prompt Tuning
abstract
As the size of transformer-based, models continues to grow, fine-tuning these large-scale pretrained vision models for new tasks has become increasingly parameter-intensive. Parameter-efficient learning has been developed to reduce the number of tunable parameters during fine-tuning. Although these methods show promising results, there is still a significant performance gap compared to full fine-tuning. To address this challenge, we propose an Effective and Efficient Visual Prompt Tuning (E2VPT) approach for large-scale transformer-based model adaptation. Specifically, we introduce a set of learnable key-value prompts and visual prompts into self-attention and input layers, respectively, to improve the effectiveness of model fine-tuning. Moreover, we design a prompt pruning procedure to systematically prune low importance prompts while preserving model performance, which largely enhances the model’s efficiency. Empirical results demonstrate that our approach outperforms several state-of-the-art baselines on two benchmarks, with considerably low parameter usage (e.g., 0.32% of model parameters on VTAB-1k). Our code is available at https://github.com/ChengHan111/E2VPT.
Cheng Han 0001, Qifan Wang 0001, Yiming Cui 0002, Zhiwen Cao, Wenguan Wang, Siyuan Qi, Dongfang Liu
ICCV4
2023 Joint learning of graph and latent representation for unsupervised feature selection
Xijiong Xie, Zhiwen Cao, Feixiang Sun
Appl. Intell.2
2023 Consensus cluster structure guided multi-view unsupervised feature selection
Zhiwen Cao, Xijiong Xie, Feixiang Sun, Jiabei Qian
Knowl. Based Syst.1
2022 Towards Unbiased Label Distribution Learning for Facial Pose Estimation Using Anisotropic Spherical Gaussian
Zhiwen Cao, Dongfang Liu, Qifan Wang 0001, Victor Y. Chen
ECCV (12)1
2022 Physical Attack on Monocular Depth Estimation with Optimal Adversarial Patches
Zhiyuan Cheng 0010, James Liang, Hongjun Choi, Guanhong Tao 0001, Zhiwen Cao, Dongfang Liu, Xiangyu Zhang 0001
ECCV (38)5
2022 DG-Labeler and DGL-MOTS Dataset: Boost the Autonomous Driving Perception
abstract
Multi-object tracking and segmentation (MOTS) is a critical task for autonomous driving applications. The existing MOTS studies face two critical challenges: 1) the published datasets inadequately capture the real-world complexity for network training to address various driving settings; 2) the working pipeline annotation tool is under-studied in the literature to improve the quality of MOTS learning examples. In this work, we introduce the DG-Labeler and DGL-MOTS dataset to facilitate the training data annotation for the MOTS task and accordingly improve network training accuracy and efficiency. DG-Labeler uses the novel Depth-Granularity Module to depict the instance spatial relations and produce fine-grained instance masks. Annotated by DG-Labeler, our DGL-MOTS dataset exceeds the prior effort (i.e., KITTI MOTS and BDD100K) in data diversity, annotation quality, and temporal representations. Results on extensive cross-dataset evaluations indicate significant performance improvements for several state-of-the-art methods trained on our DGL-MOTS dataset. We believe our DGL-MOTS Dataset and DG-Labeler hold the valuable potential to boost the visual perception of future transportation. Our dataset and code are available here1.
Yiming Cui 0002, Zhiwen Cao, Chloe Yixin Xie, Xingyu Jiang 0001, Feng Tao 0002, Victor Y. Chen, Dongfang Liu
WACV2
2021 TF-Blender: Temporal Feature Blender for Video Object Detection
abstract
Video objection detection is a challenging task because isolated video frames may encounter appearance deterioration, which introduces great confusion for detection. One of the popular solutions is to exploit the temporal information and enhance per-frame representation through aggregating features from neighboring frames. Despite achieving improvements in detection, existing methods focus on the selection of higher-level video frames for aggregation rather than modeling lower-level temporal relations to increase the feature representation. To address this limitation, we propose a novel solution named TF-Blender, which includes three modules: 1) Temporal relation models the relations between the current frame and its neigh-boring frames to preserve spatial information. 2). Feature adjustment enriches the representation of every neigh-boring feature map; 3) Feature blender combines outputs from the first two modules and produces stronger features for the later detection tasks. For its simplicity, TF-Blender can be effortlessly plugged into any detection network to improve detection behavior. Extensive evaluations on ImageNet VID and YouTube-VIS benchmarks indicate the performance guarantees of using TF-Blender on recent state-of-the-art methods. Code is available at https://github.com/goodproj13/TF-Blender.
Yiming Cui 0002, Liqi Yan, Zhiwen Cao, Dongfang Liu
ICCV3
2021 A Vector-based Representation to Enhance Head Pose Estimation
abstract
This paper proposes to use the three vectors in a rotation matrix as the representation in head pose estimation and develops a new neural network based on the characteristic of such representation. We address two potential issues existed in current head pose estimation works: 1. Public datasets for head pose estimation use either Euler angles or quaternions to annotate data samples. However, both of these annotations have the issue of discontinuity and thus could result in some performance issues in neural network training. 2. Most research works report Mean Absolute Error (MAE) of Euler angles as the measurement of performance. We show that MAE may not reflect the actual behavior especially for the cases of profile views. To solve these two problems, we propose a new annotation method which uses three vectors to describe head poses and a new measurement Mean Absolute Error of Vectors (MAEV) to assess the performance. We also train a new neural network to predict the three vectors with the constraints of orthogonality. Our proposed method achieves state-of-the-art results on both AFLW2000 and BIWI datasets. Experiments show our vector-based annotation method can effectively reduce prediction errors for large pose angles.
Zhiwen Cao, Zongcheng Chu, Dongfang Liu, Victor Y. Chen
WACV1
2020 A Large-scale Simulation Dataset: Boost the Detection Accuracy for Special Weather Conditions
abstract
Object detection is a fundamental task for autonomous driving systems. One bottleneck hindering detection accuracy is a shortage of well-annotated image data. Virtual reality has provided a feasible low-cost way to facilitate computer vision related developments. In autonomous driving area, existing public datasets from real world generally have data biases and cannot represent a wide range of weather conditions, such as rainy or snowy roads. To address this challenge, we introduce a new large-scale simulation dataset which is generated by an automated pipeline from a high realism video game. Our dataset focuses on weather conditions, which can be adopted to train networks to effectively detect objects under such conditions. We use extensive experiments to evaluate our dataset by comparing it with public datasets. The experiment results show that networks trained with our dataset outperform the networks trained by other public datasets. Our work demonstrates the effectiveness of using simulation data to address real-world challenges in the practice of object detection.
Dongfang Liu, Yiming Cui 0002, Zhiwen Cao, Victor Y. Chen
IJCNN3
2020 Indoor Navigation for Mobile Agents: A Multimodal Vision Fusion Model
abstract
Indoor navigation is a challenging task for mobile agents. The latest vision-based indoor navigation methods make remarkable progress in this field but do not fully leverage visual information for policy learning and struggle to perform well in unseen scenes. To address the existing limitations, we present a multimodal vision fusion model (MVFM). We implement a joint modality of different image recognition networks for navigation policy learning. The proposed model incorporates object detection for target searching, depth estimation for distance prediction, and semantic segmentation to depict the walkable region. In design, our model provides holistic vision knowledge for navigation. Evaluation on AI2-THOR indicates that MVFM improves on the results of a strong baseline model by 3.49% for Success weighted by Path Length (SPL) and 4% for success rate respectively. In comparison with other state-of-the-art systems, MVFM performs in the lead in terms of SPL and success rate. Extensive experiments show the effectiveness of the proposed model.
Dongfang Liu, Yiming Cui 0002, Zhiwen Cao, Victor Y. Chen
IJCNN3