VLDB 2026 Research / reviewers in the wild / expert
Sanyuan Zhao
dblp:04/7667
· DBLP profile ↗
36ranked-venue papers
2as first author
21since 2021 · last 2026
0000-0001-9386-9677ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 1 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 1 first-author · 13 since 2021Computer networks · 1Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Language Interprets Vision: Adaptive Encoding and Decoding for Referring Image SegmentationabstractReferring image segmentation aims to segment the referent with natural linguistic expressions. Due to the distinct modality properties of the image and language, it is challenging to effectively align token embeddings with visual regions. Different from existing methods of coordinate linguistics for the specific visual region, we propose a novel referring image segmentation paradigm, language interprets vision (LIV), which densely fine-grained aligns the visual and linguistic modalities, and fuse the multi-modal biases effectively. LIV resorts to re-encoding visual features on compositional dimensions of, which interprets vision through linguistic expression and makes cross-modality alignment denser. More specifically, we innovatively consider the adjacency of visual regions on the channel level to promote channel semantic consistency and propagate fine-grained semantics in the whole segmentation procedure. In addition, we also theoretically analyze that LIV effectively enriches the representation space and makes the comprehensive modality-fused biases more generalized, which boosts the precision of mask prediction. Extensive experimental results on three benchmarks validate that our proposed framework significantly outperforms other methods by a remarkable margin. Qi A, Sanyuan Zhao, Xingping Dong, Jianbing Shen |
Comput. Vis. Media | 2 |
| 2026 | DriveGen: Shared Video-Condition Encoding for Autonomous Multi-View Video GenerationabstractCorner cases, such as severe weather and abnormal lighting, present significant challenges in autonomous driving. The main obstacles involve large-scale data collection and costly annotations. Leveraging generative models to expand corner-case data based on existing annotations offers a promising solution. Unlike monocular videos, multi-view videos introduce an additional "view" dimension, increasing the consistency requirements and making precise control of annotations more challenging. Existing methods decouple multi-view videos along the temporal and view-spatial axes, using separate attention mechanisms, which causes motion discrepancies and limits consistency. Additionally, current approaches employ an independent adapter or ControlNet to encode different 3D annotations, leading to high computational costs and suboptimal alignment between annotations and video latents. These issues arise from neglecting the temporal-spatial relationship and insufficient alignment between 3D annotations and video latents. To address these challenges, we propose DriveGen, which uses 4D position embeddings to encode the positional information of multi-view videos. DriveGen also designs Dual-Scale Full Attention to ensure both global and local spatiotemporal consistency. Furthermore, our Shared Video-Condition Encoding (SVCE) Mechanism converts 3D annotations into 2D masks and encodes both video and annotation sequences using a 3D VAE, requiring only 0.37 M learnable parameters to achieve pixel-level alignment and improving generation quality. Numerous experiments have proven that DriveGen has reached the state-of-the-art, capable of generating high-quality controlled autonomous driving videos. Yuhao Kang, Sanyuan Zhao, Xiameng Qin, Junyu Han, Ji Tao |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2025 | World Knowledge-Enhanced Reasoning Using Instruction-Guided Interactor in Autonomous DrivingabstractThe Multi-modal Large Language Models (MLLMs) with extensive world knowledge have revitalized autonomous driving, particularly in reasoning tasks within perceivable regions. However, when faced with perception-limited areas (dynamic or static occlusion regions), MLLMs struggle to effectively integrate perception ability with world knowledge for reasoning. These perception-limited regions can conceal crucial safety information, especially for vulnerable road users. In this paper, we propose a framework, which aims to improve autonomous driving performance under perception-limited conditions by enhancing the integration of perception capabilities and world knowledge. Specifically, we propose a plug-and-play instruction-guided interaction module that bridges modality gaps and significantly reduces the input sequence length, allowing it to adapt effectively to multi-view video inputs. Furthermore, to better integrate world knowledge with driving-related tasks, we have collected and refined a large-scale multi-modal dataset that includes 2 million natural language QA pairs, 1.7 million grounding task data. To evaluate the model’s utilization of world knowledge, we introduce an object-level risk assessment dataset comprising 200K QA pairs, where the questions necessitate multi-step reasoning leveraging world knowledge for resolution. Extensive experiments validate the effectiveness of our proposed method. Mingliang Zhai, Zengyuan Guo, Ningrui Yang, Xiameng Qin, Sanyuan Zhao, Junyu Han, Ji Tao, Yuwei Wu 0001, Yunde Jia |
AAAI | 6 |
| 2025 | A Progressive Approach to Learn Global and Local Multi-view Features for 3D Visual Grounding
Ken Yang, Sanyuan Zhao |
ICIG (2) | 2 |
| 2025 | Modality Confusion Learning: A Versatile Framework for Visible-Infrared Re-identification
Sanyuan Zhao, Mang Ye, Ruigang Yang, Jianbing Shen |
Int. J. Comput. Vis. | 2 |
| 2025 | Bilateral Cross-Modality Graph Matching Attention for Feature Fusion in Visual Question AnsweringabstractAnswering semantically complicated questions according to an image is challenging in a visual question answering (VQA) task. Although the image can be well represented by deep learning, the question is always simply embedded and cannot well indicate its meaning. Besides, the visual and textual features have a gap for different modalities, it is difficult to align and utilize the cross-modality information. In this article, we focus on these two problems and propose a graph matching attention (GMA) network. First, it not only builds graph for the image but also constructs graph for the question in terms of both syntactic and embedding information. Next, we explore the intramodality relationships by a dual-stage graph encoder and then present a bilateral cross-modality GMA to infer the relationships between the image and the question. The updated cross-modality features are then sent into the answer prediction module for final answer prediction. Experiments demonstrate that our network achieves the state-of-the-art performance on the GQA dataset and the VQA 2.0 dataset. The ablation studies verify the effectiveness of each module in our GMA network. Jianjian Cao, Xiameng Qin, Sanyuan Zhao, Jianbing Shen |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | High-Fidelity and High-Efficiency Talking Portrait Synthesis With Detail-Aware Neural Radiance FieldsabstractIn this paper, we propose a novel rendering framework based on neural radiance fields (NeRF) named HH-NeRF that can generate high-resolution audio-driven talking portrait videos with high fidelity and fast rendering. Specifically, our framework includes a detail-aware NeRF module and an efficient conditional super-resolution module. First, a detail-aware NeRF is proposed to efficiently generate a high-fidelity low-resolution talking head, by using the encoded volume density estimation and audio-eye-aware color calculation. This module can capture natural eye blinks and high-frequency details, and maintain a similar rendering time as previous fast methods. Secondly, we present an efficient conditional super-resolution module on the dynamic scene to directly generate the high-resolution portrait with our low-resolution head. Incorporated with the prior information, such as depth map and audio features, our new proposed efficient conditional super resolution module can adopt a lightweight network to efficiently generate realistic and distinct high-resolution videos. Extensive experiments demonstrate that our method can generate more distinct and fidelity talking portraits on high resolution (900 × 900) videos compared to state-of-the-art methods. Muyu Wang, Sanyuan Zhao, Xingping Dong, Jianbing Shen |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2024 | RepVF: A Unified Vector Fields Representation for Multi-task 3D Perception
Chunliang Li, Wencheng Han, Jun Yin 0003, Sanyuan Zhao, Jianbing Shen |
ECCV (32) | 4 |
| 2024 | GTMS: A Gradient-Driven Tree-Guided Mask-Free Referring Image Segmentation Method
Haoxin Lyu, Tianxiong Zhong, Sanyuan Zhao |
ECCV (66) | 3 |
| 2024 | RT-VIS: Real-Time Video Instance Segmentation with Light-Weight Decoupled Framework
Tianze Cao, Sanyuan Zhao |
PRCV (10) | 2 |
| 2024 | A simple but effective vision transformer framework for visible-infrared person re-identification
Yudong Li 0002, Sanyuan Zhao, Jianbing Shen |
Comput. Vis. Image Underst. | 2 |
| 2024 | A New Framework of Collaborative Learning for Adaptive Metric DistillationabstractThis article presents a new adaptive metric distillation approach that can significantly improve the student networks' backbone features, along with better classification results. Previous knowledge distillation (KD) methods usually focus on transferring the knowledge across the classifier logits or feature structure, ignoring the excessive sample relations in the feature space. We demonstrated that such a design greatly limits performance, especially for the retrieval task. The proposed collaborative adaptive metric distillation (CAMD) has three main advantages: 1) the optimization focuses on optimizing the relationship between key pairs by introducing the hard mining strategy into the distillation framework; 2) it provides an adaptive metric distillation that can explicitly optimize the student feature embeddings by applying the relation in the teacher embeddings as supervision; and 3) it employs a collaborative scheme for effective knowledge aggregation. Extensive experiments demonstrated that our approach sets a new state-of-the-art in both the classification and retrieval tasks, outperforming other cutting-edge distillers under various settings. Mang Ye, Yan Wang 0116, Sanyuan Zhao, Ping Li 0016, Jianbing Shen |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | Using Self-Supervised Dual Constraint Contrastive Learning for Cross-Modal RetrievalabstractIn this work, we present a self-supervised dual constraint contrastive method for efficiently fine-tuning the vision-language pre-trained (VLP) models that have achieved great success on various cross-modal tasks, since full fine-tune these pre-trained models is computationally expensive and tend to result in catastrophic forgetting restricted by the size and quality of labeled datasets. Our approach freezes the pre-trained VLP models as the fundamental, generalized, and transferable multimodal representation and incorporates lightweight parameters to learn domain and task-specific features without labeled data. We demonstrated that our self-supervised dual contrastive model performs better than previous fine-tuning methods on MS COCO and Flickr 30K datasets on the cross-modal retrieval task, with an even more pronounced improvement in zero-shot performance. Furthermore, experiments on the MOTIF dataset prove that our self-supervised approach remains effective when trained on a small, out-of-domain dataset without overfitting. As a plug-and-play method, our proposed method is agnostic to the underlying models and can be easily integrated with different VLP models, allowing for the potential incorporation of future advancements in VLP models. Xintong Wang 0001, Liang Ding 0006, Sanyuan Zhao, Chris Biemann |
ECAI | 4 |
| 2023 | Spectrum-irrelevant fine-grained representation for visible-infrared person re-identification
Jiahao Gong, Sanyuan Zhao, Kin-Man Lam 0001, Xin Gao 0001, Jianbing Shen |
Comput. Vis. Image Underst. | 2 |
| 2023 | Dual-Semantic Consistency Learning for Visible-Infrared Person Re-IdentificationabstractVisible-Infrared person Re-Identification (VI-ReID) conducts comprehensive identity analysis on non-overlapping visible and infrared camera sets for intelligent surveillance systems, which face huge instance variations derived from modality discrepancy. Existing methods employ different kinds of network structure to extract modality-invariant features. Differently, we propose a novel framework, named Dual-Semantic Consistency Learning Network (DSCNet), which attributes modality discrepancy to channel-level semantic inconsistency. DSCNet optimizes channel consistency from two aspects, fine-grained inter-channel semantics, and comprehensive inter-modality semantics. Furthermore, we propose Joint Semantics Metric Learning to simultaneously optimize the distribution of the channel-and-modality feature embeddings. It jointly exploits the correlation between channel-specific and modality-specific semantics in a fine-grained manner. We conduct a series of experiments on the SYSU-MM01 and RegDB datasets, which validates that DSCNet delivers superiority compared with current state-of-the-art methods. On the more challenging SYSU-MM01 dataset, our network can achieve 73.89% Rank-1 accuracy and 69.47% mAP value. Our code is available athttps://github.com/bitreidgroup/DSCNet. Yuhao Kang, Sanyuan Zhao, Jianbing Shen |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2022 | Modality Synergy Complement Learning with Cascaded Aggregation for Visible-Infrared Person Re-Identification
Sanyuan Zhao, Yuhao Kang, Jianbing Shen |
ECCV (14) | 2 |
| 2022 | Interaction and Alignment for Visible-Infrared Person Re-IdentificationabstractVisible-Infrared Person Re-Identification (VI-ReID) is a challenging person matching problem and is also a practical solution for intelligent surveillance systems at night. Due to the heterogeneity between visible and infrared modalities, the retrieval performance is seriously damaged. To address the issue of the discrepancy of the information between visible and infrared modalities, many works have been proposed. However, the relationship between cross-modality samples has rarely been mined. In this paper, we propose a Cross-modality Interaction and Alignment (CIA) module to solve the discrepancy problem. Through transforming the information between different modalities, the module guides the network to capture the modality-shared feature, which is beneficial to address the cross-modality discrepancy. Meanwhile, to better supervise the network, an enhanced contrastive loss is introduced. Contributed by the further optimization in the distance between intra-class samples, the network gains more effective supervision. Extensive experiments on two benchmark datasets show that our method achieves an excellent performance in VI-ReID. Jiahao Gong, Sanyuan Zhao, Kin-Man Lam 0001 |
ICPR | 2 |
| 2021 | Cross-Modality Person Re-Identification via Modality Confusion and Center AggregationabstractCross-modality person re-identification is a challenging task due to large cross-modality discrepancy and intramodality variations. Currently, most existing methods focus on learning modality-specific or modality-shareable features by using the identity supervision or modality label. Different from existing methods, this paper presents a novel Modality Confusion Learning Network (MCLNet). Its basic idea is to confuse two modalities, ensuring that the optimization is explicitly concentrated on the modality-irrelevant perspective. Specifically, MCLNet is designed to learn modality-invariant features by simultaneously minimizing inter-modality discrepancy while maximizing cross-modality similarity among instances in a single framework. Furthermore, an identity-aware marginal center aggregation strategy is introduced to extract the centralization features, while keeping diversity with a marginal constraint. Finally, we design a camera-aware learning scheme to enrich the discriminability. Extensive experiments on SYSU-MM01 and RegDB datasets show that MCLNet outperforms the state-of-the-art by a large margin. On the large-scale SYSU-MM01 dataset, our model can achieve 65.40 % and 61.98 % in terms of Rank-1 accuracy and mAP value. Xin Hao, Sanyuan Zhao, Mang Ye, Jianbing Shen |
ICCV | 2 |
| 2021 | Dual Attention Based Network with Hierarchical ConvLSTM for Video Object Segmentation
Zongji Zhao, Sanyuan Zhao |
PRCV (4) | 2 |
| 2021 | Video person re-identification with global statistic pooling and self-attention distillation
Gaojie Lin, Sanyuan Zhao, Jianbing Shen |
Neurocomputing | 2 |
| 2021 | Real-time and light-weighted unsupervised video object segmentation network
Zongji Zhao, Sanyuan Zhao, Jianbing Shen |
Pattern Recognit. | 2 |
| 2020 | Self-Learning With Rectification Strategy for Human ParsingabstractIn this paper, we solve the sample shortage problem in the human parsing task. We begin with the self-learning strategy, which generates pseudo-labels for unlabeled data to retrain the model. However, directly using noisy pseudo-labels will cause error amplification and accumulation. Considering the topology structure of human body, we propose a trainable graph reasoning method that establishes internal structural connections between graph nodes to correct two typical errors in the pseudo-labels, i.e., the global structural error and the local consistency error. For the global error, we first transform category-wise features into a high-level graph model with coarse-grained structural information, and then decouple the high-level graph to reconstruct the category features. The reconstructed features have a stronger ability to represent the topology structure of the human body. Enlarging the receptive field of features can effectively reducing the local error. We first project feature pixels into a local graph model to capture pixel-wise relations in a hierarchical graph manner, then reverse the relation information back to the pixels. With the global structural and local consistency modules, these errors are rectified and confident pseudo-labels are generated for retraining. Extensive experiments on the LIP and the ATR datasets demonstrate the effectiveness of our global and local rectification modules. Our method outperforms other state-of-the-art methods in supervised human parsing tasks. Zhiyuan Liang, Sanyuan Zhao, Jiahao Gong, Jianbing Shen |
CVPR | 3 |
| 2020 | Efficient Light Deep Network for Street Scene ParsingabstractThe semantic segmentation is a dense pixel label pre-diction task, which takes quite a lot of resources and computation cost in most of the time. In our approach, we pay attention to balance the speed and better performance which outperforms the state of the art in speed and accuracy for real-time performance. We come up with the idea of new efficient deep backbone that can extract more semantic details, reduce the computation cost and be easy to deploy at the same time. We call our new backbone as Cascaded Mobile Network, which is proved to be very useful. Our proposed model achieves 72.1 mIOU on the CityScapes val, and 69.5 on CamVid. We achieve good balance between speed and accuracy. ZheHui Wang, Sanyuan Zhao, Jianbing Shen, Zhengchao Lei |
VCIP | 2 |
| 2020 | Multiple people tracking with articulation detection and stitching strategy
Yuanpei Liu, Junbo Yin, Dajiang Yu, Sanyuan Zhao, Jianbing Shen |
Neurocomputing | 4 |
| 2020 | Video semantic segmentation via feature propagation with holistic attention
Junrong Wu, Zongzheng Wen, Sanyuan Zhao, Kele Huang |
Pattern Recognit. | 3 |
| 2019 | Learning Unsupervised Video Object Segmentation Through Visual AttentionabstractThis paper conducts a systematic study on the role of visual attention in Unsupervised Video Object Segmentation (UVOS) tasks. By elaborately annotating three popular video segmentation datasets (DAVIS, Youtube-Objects and SegTrack V2) with dynamic eye-tracking data in the UVOS setting, for the first time, we quantitatively verified the high consistency of visual attention behavior among human observers, and found strong correlation between human attention and explicit primary object judgements during dynamic, task-driven viewing. Such novel observations provide an in-depth insight into the underlying rationale behind UVOS. Inspired by these findings, we decouple UVOS into two sub-tasks: UVOS-driven Dynamic Visual Attention Prediction (DVAP) in spatiotemporal domain, and Attention-Guided Object Segmentation (AGOS) in spatial domain. Our UVOS solution enjoys three major merits: 1) modular training without using expensive video segmentation annotations, instead, using more affordable dynamic fixation data to train the initial video attention module and using existing fixation-segmentation paired static/image data to train the subsequent segmentation module; 2) comprehensive foreground understanding through multi-source learning; and 3) additional interpretability from the biologically-inspired and assessable attention. Experiments on popular benchmarks show that, even without using expensive video object mask annotations, our model achieves compelling performance in comparison with state-of-the-arts. Wenguan Wang, Hongmei Song, Shuyang Zhao, Jianbing Shen, Sanyuan Zhao, Steven C. H. Hoi, Haibin Ling |
CVPR | 5 |
| 2019 | Multi-scale Capsule Attention-Based Salient Object Detection with Multi-crossed Layer ConnectionsabstractWith the popularization of convolutional networks being used for saliency models, saliency detection performance has achieved significant improvement. However, how to integrate accurate and crucial features for modeling saliency is still underexplored. In this paper, we present CapSalNet, which includes a multi-scale Capsule attention module and multi-crossed layer connections for Salient object detection. We first propose a novel capsule attention model, which integrates multi-scale contextual information with dynamic routing. Then, our model adaptively learns to aggregate multi-level features by using multi-crossed skip-layer connections. Finally, the predicted results are efficiently fused to generate the final saliency map in a coarse-to-fine manner. Comprehensive experiments on four benchmark datasets demonstrate that our proposed algorithm outperforms existing state-of-the-art approaches. Sanyuan Zhao, Jianbing Shen, Kin-Man Lam 0001 |
ICME | 2 |
| 2019 | High-speed video salient object detection with temporal propagation using correlation filter
Sanyuan Zhao, Zhengchao Lei, Jianbing Shen, Yuanyuan Pang |
Neurocomputing | 2 |
| 2019 | A stable long-term object tracking method with re-detection strategy
Sanyuan Zhao, Qinghao Meng, Jianbing Shen |
Pattern Recognit. Lett. | 2 |
| 2018 | Pyramid Dilated Deeper ConvLSTM for Video Salient Object Detection
Hongmei Song, Wenguan Wang, Sanyuan Zhao, Jianbing Shen, Kin-Man Lam 0001 |
ECCV (11) | 3 |
| 2018 | Scene text recognition using residual convolutional recurrent neural network
Zhengchao Lei, Sanyuan Zhao, Hongmei Song, Jianbing Shen |
Mach. Vis. Appl. | 2 |
| 2017 | Co-saliency Detection Based on Siamese Network
Zhengchao Lei, Weiyan Chai, Sanyuan Zhao, Hongmei Song, Fengxia Li |
MSN | 3 |
| 2017 | Diffusion-based saliency detection with optimal seed selection scheme
Sanyuan Zhao, Zhengchao Lei, Meiling Sun, Jianbing Shen |
Neurocomputing | 1 |
| 2017 | A hierarchical visual saliency detection method by combining distinction and background probability maps
Sanyuan Zhao, Jianbing Shen, Fengxia Li |
Multim. Syst. | 1 |
| 2015 | Virtual experiments for introduction of computing: Using virtual reality technologyabstractIntroduction to Computing is a public course for the first-year non-major undergraduate students, aiming at training students for the abilities in computer science and technology with computational thinking. However, as new computer technologies emerge continuously and rapidly, it is required for this course to accommodate more and more knowledge. Therefore the teaching contents are growing enormously, which makes it very difficult to cover all of them in limited hours, and therefore sets an obstacle in understanding computing principles and building up a clear and general picture of computing, especially for non-major students. As computer science and technology are becoming more and more essential for various disciplines and majors, it is urgent for the education community to find out an effective and propagable way to solve this problem. In this regard, we employ virtual reality technology to the experiment teaching of this course, and have developed 18 virtual experiments to support the whole teaching process. For example, Turing machine is a basic model for computer science and technology. However, since it is not a real machine, it is not easy for the students to imagine the working process of Turing machine and understand the related concepts. Another example, the execution of an instruction is very important to understand the principles of computer organization. However, as the information flow is invisible, it is difficult and time-consuming for the teachers to explain how an instruction is executed inside a computer. Therefore, 3D modeling and animation techniques are used to demonstrate the invisible micro-structure of computers, and human-machine interaction and visualization techniques are used to present the internal process of information evolution, thus constructing a complete virtual experiment system of this course, including demonstration experiments, verification experiments and interaction experiments. Our virtual experiments have applied software copyrights and served more than 12,000 students from five universities of China since 2013. The evaluation demonstrates that the virtual experiments have produced excellent results in both teaching effectiveness and learning efficiency, relieved the conflicts between limited hours and vast knowledge, and helped students understand and build up the knowledge of computing. Fengxia Li, Jun Zheng 0007, Sanyuan Zhao |
FIE | 4 |
| 2012 | Totally-corrective boosting using continuous-valued weak learnersabstractThe Boosting algorithm has two main variants: the gradient Boosting and the totally-corrective column-generation Boosting. Recently, the latter has received increasing attention since it exhibits a better convergence property, thus resulting in more efficient strong learners. In this work, we point out that the totally-corrective column-generation Boosting is equivalent to the gradient-descent method for the gradient Boosting in the weak-learner selection criterion, but uses additional totally-corrective updates for the weak-learner weights. Therefore, other techniques for the gradient Boosting that produce continuous-valued weak learners, e.g. step-wise direct minimization and Newtons method, may also be used in combination with the totally-corrective procedure. In this work we take the well known AdaBoost algorithm as an example, and show that employing the continuous-valued weak learners improves the performance when used with the totally-corrective weak-learner weight update. Chensheng Sun, Sanyuan Zhao, Jiwei Hu, Kin-Man Lam 0001 |
ICASSP | 2 |