VLDB 2026 Research / reviewers in the wild / expert
Yue Zhan
dblp:181/3946
· DBLP profile ↗
11ranked-venue papers
4as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 3 since 2021Systems, architecture and hardware · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RelPose-TTA: Energy-based relative pose correction for test-time adaptation of category-level object pose estimation
Yue Zhan, Xin Wang 0135, Zhaoxiang Liu, Shiguo Lian, Tangwen Yang |
Image Vis. Comput. | 1 |
| 2025 | Data Leakage Detection in Large Vision-Language Models via Multimodal Perturbation
Xin Wang 0135, Zhaoxiang Liu, Yue Zhan, Kaikai Zhao, Kai Wang 0012, Shiguo Lian |
ICIG (1) | 3 |
| 2025 | MGNet: Mask Guided Transparent Object Depth Completion with Hybrid CNN-Transformer Network
Kailong Xie, Yue Zhan, Tangwen Yang |
PRCV (11) | 2 |
| 2025 | MambaSOD: Dual Mamba-driven cross-modal fusion network for RGB-D Salient Object Detection
Yue Zhan, Zhihong Zeng, Haijun Liu 0001, Xiaoheng Tan, Yinli Tian |
Neurocomputing | 1 |
| 2025 | LMT++: Adaptively Collaborating LLMs With Multi-Specialized Teachers for Continual VQA in Robotic Surgical VideosabstractVisual question answering (VQA) plays a vital role in advancing surgical education. However, due to the privacy concern of patient data, training VQA model with previously used data becomes restricted, making it necessary to use the exemplar-free continual learning (CL) approach. Previous CL studies in the surgical field neglected two critical issues: i) significant domain shifts caused by the wide range of surgical procedures collected from various sources, and ii) the data imbalance problem caused by the unequal occurrence of medical instruments or surgical procedures. This paper addresses these challenges with a multimodal large language model (LLM) and an adaptive weight assignment strategy. First, we developed a novel LLM-assisted multi-teacher CL framework (named LMT++), which could harness the strength of a multimodal LLM as a supplementary teacher. The LLM's strong generalization ability, as well as its good understanding of the surgical domain, help to address the knowledge gap arising from domain shifts and data imbalances. To incorporate the LLM in our CL framework, we further proposed an innovative approach to process the training data, which involves the conversion of complex LLM embeddings into logits value used within our CL training framework. Moreover, we design an adaptive weight assignment approach that balances the generalization ability of the LLM and the domain expertise of conventional VQA models obtained in previous model training processes within the CL framework. Finally, we created a new surgical VQA dataset for model evaluation. Comprehensive experimental findings on these datasets show that our approach surpasses state-of-the-art CL methods. Yuyang Du 0001, Kexin Chen 0003, Yue Zhan, Chang Han Low, Mobarakol Islam, Yueming Jin, Guangyong Chen, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 3 |
| 2024 | Hybrid Perturbation Strategy for Semi-Supervised Crowd CountingabstractA simple yet effective semi-supervised method is proposed in this paper based on consistency regularization for crowd counting, and a hybrid perturbation strategy is used to generate strong, diverse perturbations, and enhance unlabeled images information mining. The conventional CNN-based counting methods are sensitive to texture perturbation and imperceptible noises raised by adversarial attack, therefore, the hybrid strategy is proposed to combine a spatial texture transformation and an adversarial perturbation module to perturb the unlabeled data in the semantic and non-semantic spaces, respectively. Moreover, a cross-distribution normalization technique is introduced to address the model optimization failure caused by BN layer in the strong perturbation, and to stabilize the optimization of the learning model. Extensive experiments have been conducted on the datasets of ShanghaiTech, UCF-QNRF, NWPU-Crowd, and JHU-Crowd++. The results demonstrate that the proposed semi-supervised counting method performs better over the state-of-the-art methods, and it shows better robustness to various perturbations. Xin Wang 0135, Yue Zhan, Yang Zhao 0014, Tangwen Yang, Qiuqi Ruan |
IEEE Trans. Image Process. | 2 |
| 2024 | TG-Pose: Delving Into Topology and Geometry for Category-Level Object Pose EstimationabstractCategory-level 6D object pose estimation aims to estimate the pose and size of unseen objects with known categories. Existing methods mainly focus on capturing geometric features to handle shape variations, and are prone to failure in occlusion and noisy environments. In this paper, we propose TG-Pose, a unified pose estimation framework that delves into topology and geometry to deal with the above issues. To exploit topological properties, we first propose a topological feature predictor and a topological label generator to dig into the underlying structural details from encoded features using persistent homology. Then, the topological and geometric features are employed to facilitate the symmetry reconstruction of the original point cloud to obtain a reliable and coherent object shape, which, in turn, guides the pose estimation. For each object category, we construct geometric and topological templates by leveraging inherent intra-class similarities. These templates enhance the reliability of pose estimation and the completeness of object structure through geometric alignment and topological guidance, especially when handling incomplete objects. Moreover, a pose-aware enhancement strategy is designed to enhance the encoder in learning pose-sensitive features and robustness to noisy point clouds. Experimental results show that TG-Pose outperforms the state-of-the-art solutions on public benchmarks and achieves better generalization in real-world datasets. Project Page https://sites.google.com/view/tg-pose. Yue Zhan, Xin Wang 0135, Lang Nie, Yang Zhao 0014, Tangwen Yang, Qiuqi Ruan |
IEEE Trans. Multim. | 1 |
| 2023 | Semi-Supervised Crowd Counting With Spatial Temporal Consistency and Pseudo-Label FilterabstractSemi-supervised crowd counting (SSCC) aims to learn a crowd counting model with limited labeled images and a large number of unlabeled images. Previous works leverage unlabeled images by pseudo-labeling and spatial consistency regularization paradigms, which frequently adopt teacher-student frameworks. However, their performances are readily degraded due to the inconsistent and unreliable pseudo density map in complex crowd scenes. Here, we argue that the SSCC performance can be significantly improved by reducing the over-fitting of the incorrect pseudo labels, and a novel spatial-temporal consistency framework, named STC-Crowd, is proposed. Under different spatial perturbations, spatial consistency enables the counting model to output consistent predictions for the same crowd image. Temporal consistency generates similar feature embedding over adjacent training stages, alleviating the inconsistent issues of pseudo density maps generated for the same image over time. To store the temporal feature embedding of different density levels for temporal consistency, a dynamic temporal knowledge memory (DTKM) is deliberately designed, and considerably reduces the storage cost. Besides, a pseudo-label filter (PLF) mechanism is used to alleviate the negative impact of incorrect pseudo density maps, by reducing the supervision weights of unreliable pseudo labels with high uncertainty. Extensive experiments on four benchmark datasets show that our method obtains competent performance against leading SSCC methods, and especially works better on limited labeled images. Xin Wang 0135, Yue Zhan, Yang Zhao 0014, Tangwen Yang, Qiuqi Ruan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | Formal Validation for Natural Language Programming using Hierarchical Finite State Automata
Yue Zhan, Michael S. Hsiao |
ICAART (1) | 1 |
| 2020 | A Context Knowledge Map Guided Coarse-to-Fine Action RecognitionabstractHuman actions involve a wide variety and a large number of categories, which leads to a big challenge in action recognition. However, according to similarities on human body poses, scenes, interactive objects, human actions can be grouped into some semantic groups, i.e sports, cooking, etc. Therefore, in this paper, we propose a novel approach which recognizes human actions from coarse to fine. Taking full advantage of contributions from high-level semantic contexts, a context knowledge map guided recognition method is designed to realize the coarse-to-fine procedure. In the approach, we define semantic contexts with interactive objects, scenes and body motions in action videos, and build a context knowledge map to automatically define coarse-grained groups. Then fine-grained classifiers are proposed to realize accurate action recognition. The coarse-to-fine procedure narrows action categories in target classifiers, so it is beneficial to improving recognition performance. We evaluate the proposed approach on the CCV, the HMDB-51, and the UCF101 database. Experiments verify its significant effectiveness, on average, improving more than 5% of recognition precisions than current approaches. Compared with the state-of-the-art, it also obtains outstanding performance. The proposed approach achieves higher accuracies of 93.1%, 95.4% and 74.5% in the CCV, the UCF-101 and the HMDB51 database, respectively. Yanli Ji, Yue Zhan, Yang Yang 0002, Xing Xu 0001, Fumin Shen, Heng Tao Shen |
IEEE Trans. Image Process. | 2 |
| 2016 | Control and experimental validation of robot-assisted automatic measurement system for Multi-Stud Tensioning Machine (MSTM)abstractMulti-Stud Tensioning Machine (MSTM) is a specialized equipment used to open/seal the cover of the Reactor Pressure Vessel (RPV) during nuclear power plant maintenance. The tensioning residual values of the 58 studs are monitored for procedure evaluation. It is time-consuming for human operators to place the measurement meters into working positions. In order to reduce labor intensity and eliminate radiation exposure time, we develop a robot-assisted automatic measurement system to achieve meter placement and real-time data monitoring. The Field Programmable Gate Array (FPGA)-based distributed control scheme realizes high-speed data acquisition and coordinated control of the 58 node robots. The control software performs data analysis and sends emergency stop signals to the MSTM control PLC. The proposed system is validated in China Nuclear Power Technology Research Institute. Total operation time decreases from over 580 s to less than 120 s. Xingguang Duan, Tengfei Cui, Yue Zhan |
ICRA | 6 |