VLDB 2026 Research / reviewers in the wild / expert
Zhiyuan Zhao 0005
dblp:93/5901-5
· DBLP profile ↗
18ranked-venue papers
2as first author
17since 2021 · last 2026
0000-0001-7666-4020ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Exploring Efficient Open-Vocabulary Segmentation in the Remote SensingabstractOpen-Vocabulary Remote Sensing Image Segmentation (OVRSIS), an emerging task that adapts Open-Vocabulary Segmentation (OVS) to the remote sensing (RS) domain, remains underexplored due to the absence of a unified evaluation benchmark and the domain gap between natural and RS images. To bridge these gaps, we first establish a standardized OVRSIS benchmark (OVRSISBench) based on widely-used RS segmentation datasets, enabling consistent evaluation across methods. Using this benchmark, we comprehensively evaluate several representative OVS/OVRSIS models and reveal their limitations when directly applied to remote sensing scenarios. Building on these insights, we propose RSKT-Seg, a novel open-vocabulary segmentation framework tailored for remote sensing. RSKT-Seg integrates three key components: (1) a Multi-Directional Cost Map Aggregation (RS-CMA) module that captures rotation-invariant visual cues by computing vision-language cosine similarities across multiple directions; (2) an Efficient Cost Map Fusion (RS-Fusion) transformer, which jointly models spatial and semantic dependencies with a lightweight dimensionality reduction strategy; and (3) a Remote Sensing Knowledge Transfer (RS-Transfer) module that injects pre-trained knowledge and facilitates domain adaptation via enhanced upsampling. Extensive experiments on the benchmark show that RSKT-Seg consistently outperforms strong OVS baselines by +3.8 mIoU and +5.9 mACC, while achieving 2× faster inference through efficient aggregation. Bingyu Li 0002, Haocheng Dong, Da Zhang 0010, Zhiyuan Zhao 0005, Hao Sun 0038, Junyu Gao 0001 |
AAAI | 4 |
| 2026 | FusAD: Time-Frequency Fusion with Adaptive Denoising for General Time Series AnalysisabstractTime series analysis plays a vital role in fields such as finance, healthcare, industry, and meteorology, underpinning key tasks including classification, forecasting, and anomaly detection. Although deep learning models have achieved remarkable progress in these areas in recent years, constructing an efficient, multi-task compatible, and generalizable unified framework for time series analysis remains a significant challenge. Existing approaches are often tailored to single tasks or specific data types, making it difficult to simultaneously handle multi-task modeling and effectively integrate information across diverse time series types. Moreover, real-world data are often affected by noise, complex frequency components, and multi-scale dynamic patterns, which further complicate robust feature extraction and analysis. To ameliorate these challenges, we propose FusAD, a unified analysis framework designed for diverse time series tasks. FusAD features an adaptive time-frequency fusion mechanism, integrating both Fourier and Wavelet transforms to efficiently capture global-local and multi-scale dynamic features. With an adaptive denoising mechanism, FusAD automatically senses and filters various types of noise, highlighting crucial sequence variations and enabling robust feature extraction in complex environments. In addition, the framework integrates a general information fusion and decoding structure, combined with masked pre-training, to promote efficient learning and transfer of multi-granularity representations. Extensive experiments demonstrate that FusAD consistently outperforms state-of-the-art models on mainstream time series benchmarks for classification, forecasting, and anomaly detection tasks, while maintaining high efficiency and scalability. Code is available at https://github.com/zhangda1018/FusAD. Da Zhang 0010, Bingyu Li 0002, Zhiyuan Zhao 0005, Feiping Nie 0001, Junyu Gao 0001, Xuelong Li 0001 |
ICDE | 3 |
| 2026 | Cross-attention multi-scale state space model for remaining useful life prediction of aircraft engines
Da Zhang 0010, Bingyu Li 0002, Zhiyuan Zhao 0005, Junyu Gao 0001, Xuelong Li 0001 |
Adv. Eng. Informatics | 4 |
| 2026 | Quantum-inspired interpretable deep learning architecture for text sentiment analysis
Bingyu Li 0002, Da Zhang 0010, Zhiyuan Zhao 0005, Yuan Yuan 0001, Junyu Gao 0001, Xuelong Li 0001 |
Neural Networks | 3 |
| 2026 | Dynamic proxy domain generalizes the crowd localization by better binary segmentation
Junyu Gao 0001, Da Zhang 0010, Qiyu Wang, Zhiyuan Zhao 0005, Xuelong Li 0001 |
Pattern Recognit. | 4 |
| 2025 | LLMs Caught in the Crossfire: Malware Requests and Jailbreak ChallengesabstractThe widespread adoption of Large Language Models (LLMs) has heightened concerns about their security, particularly their vulnerability to jailbreak attacks that leverage crafted prompts to generate malicious outputs.While prior research has been conducted on general security capabilities of LLMs, their specific susceptibility to jailbreak attacks in code generation remains largely unexplored.To fill this gap, we propose MalwareBench, a benchmark dataset containing 3,520 jailbreaking prompts for malicious code-generation, designed to evaluate LLM robustness against such threats.Mal-wareBench is based on 320 manually crafted malicious code generation requirements, covering 11 jailbreak methods and 29 code functionality categories.Experiments show that mainstream LLMs exhibit limited ability to reject malicious code-generation requirements, and the combination of multiple jailbreak methods further reduces the model's security capabilities: specifically, the average rejection rate for malicious content is 60.93%, dropping to 39.92% when combined with jailbreak attack algorithms.Our work highlights that the code security capabilities of LLMs still pose significant challenges. Zhiyuan Zhao 0005, Junyu Gao 0001, Xuelong Li 0001 |
ACL (1) | 3 |
| 2025 | KAID: Knowledge-Aware Interactive Distillation for Vision-Language ModelsabstractVision-Language Models (VLMs) such as CLIP have demonstrated outstanding performance in cross-modal tasks, but the prohibitive computational cost hinders practical deployment. Although Knowledge Distillation (KD) provides a promising compression paradigm, most existing methods rely heavily on feature imitation and contrastive relations without explicit fine-grained alignment. Additionally, they do not fully leverage the multimodal interaction knowledge from the teacher model, restricting cross-modal semantic alignment. To address these challenges, we propose KAID, a Knowledge-Aware Interactive Distillation method for VLMs. Specifically, we first pretrain a large CLIP teacher model with domain few-shot labels and store text features as category vectors. Then, an Image Feature Matching (IFM) module is introduced to calculate the feature distribution of teacher-student models with improved cosine similarity, which achieves hierarchical knowledge transfer from global to local levels and enhances the fine-grained perception of student model through joint optimization. Moreover, a Pixel-Wise Alignment (PWA) module is constructed between the teacher's text features and the student's image features, employing a cross-modal attention mechanism to establish semantic associations, while a Text-guided Pixel alignment Loss function (TPloss) is concurrently designed to enhance the student's comprehension capabilities. Ultimately, the well-trained student model is used for inference. Extensive experiments on 11 datasets validate the effectiveness of our method. Specifically, our method achieves average improvements of 2.14% and 2.40% on the base and new classes across these datasets. Da Zhang 0010, Bingyu Li 0002, Zhiyuan Zhao 0005, Junyu Gao 0001, Xuelong Li 0001 |
ACM Multimedia | 4 |
| 2025 | Secure Tug-of-War (SecTOW): Iterative Defense-Attack Training with Reinforcement Learning for Multimodal Model SecurityabstractThe rapid advancement of multimodal large language models (MLLMs) has led to breakthroughs in various applications, yet their security remains a critical challenge. One pressing issue involves unsafe image-query pairs-jailbreak inputs specifically designed to bypass security constraints and elicit unintended responses from MLLMs. Compared to general multimodal data, such unsafe inputs are relatively sparse, which limits the diversity and richness of training samples available for developing robust defense models. Meanwhile, existing guardrail-type methods rely on external modules to enforce security constraints but fail to address intrinsic vulnerabilities within MLLMs. Traditional supervised fine-tuning (SFT), on the other hand, often over-refuses harmless inputs, compromising general performance. Given these challenges, we propose Secure Tug-of-War (SecTOW), an innovative iterative defense-attack training method to enhance the security of MLLMs. SecTOW consists of two modules: a defender and an auxiliary attacker, both trained iteratively using reinforcement learning (GRPO). During the iterative process, the attacker identifies security vulnerabilities in the defense model and expands jailbreak data. The expanded data are then used to train the defender, enabling it to address identified security vulnerabilities. We also design reward mechanisms used for GRPO to simplify the use of response labels, reducing dependence on complex generative labels and enabling the efficient use of synthetic data. Additionally, a quality monitoring mechanism is used to mitigate the defender's over-refusal of harmless inputs and ensure the diversity of the jailbreak data generated by the attacker. Experimental results on safety-specific and general benchmarks demonstrate that SecTOW significantly improves security while preserving general performance. Warning: This paper contains offensive and unsafe content. Muzhi Dai, Zhiyuan Zhao 0005, Junyu Gao 0001, Hao Sun 0038, Xuelong Li 0001 |
ACM Multimedia | 3 |
| 2025 | From Captions to Rewards (CaReVL): Leveraging Large Language Model Experts for Enhanced Reward Modeling in Large Vision-Language ModelsabstractAligning large vision-language models (LVLMs) with human preferences is challenging due to the scarcity of fine-grained, high-quality, and multimodal preference data without human annotations. Existing methods relying on direct distillation often struggle with low-confidence data, leading to suboptimal performance. To address this, we propose (CaReVL), a novel method for preference reward modeling by reliably using both high- and low-confidence data. First, a cluster of auxiliary expert models (textual reward models) innovatively leverages image captions as weak supervision signals to filter high-confidence data. The high-confidence data are then used to fine-tune the LVLM. Second, low-confidence data are used to generate diverse preference samples using the fine-tuned LVLM. These samples are then scored and selected to construct reliable chosen-rejected pairs for further training. (CaReVL) achieves performance improvements over traditional distillation-based methods on VL-RewardBench and MLLM-as-a-Judge benchmark, demonstrating its effectiveness. Muzhi Dai, Jiashuo Sun, Zhiyuan Zhao 0005, Junyu Gao 0001, Xuelong Li 0001 |
ACM Multimedia | 3 |
| 2025 | StitchFusion: Weaving Any Visual Modalities to Enhance Multimodal Semantic SegmentationabstractMultimodal semantic segmentation shows significant potential for enhancing segmentation accuracy in complex scenes. However, current methods often incorporate specialized feature fusion modules tailored to specific modalities, thereby restricting input flexibility and increasing the number of training parameters. To address these challenges, we propose StitchFusion, a straightforward yet effective modal fusion framework that integrates large-scale pre-trained models directly as encoders and feature fusers. This approach facilitates comprehensive multi-modal and multi-scale feature fusion, accommodating any visual modal inputs. Specifically, our framework achieves modal integration during encoding by sharing multi-modal visual information. To enhance information exchange across modalities, we introduce a multi-directional Modality Adapter module (MoA) to enable cross-modal information transfer during encoding. By leveraging MoA to propagate multi-scale information across pre-trained encoders during the encoding process, StitchFusion achieves multi-modal visual information integration during encoding. Extensive comparative experiments demonstrate that our model achieves state-of-the-art performance on four multi-modal segmentation datasets with minimal additional parameters. Furthermore, the experimental integration of MoA with existing Feature Fusion Modules (FFMs) highlights their complementary nature. Our anonymous code is https://anonymous.4open.science/r/StitchFusion_V2-E777 Bingyu Li 0002, Da Zhang 0010, Zhiyuan Zhao 0005, Junyu Gao 0001, Xuelong Li 0001 |
ACM Multimedia | 3 |
| 2025 | SVGen: Interpretable Vector Graphics Generation with Large Language ModelsabstractScalable Vector Graphics (SVG) has become an indispensable technology in front-end development and UI/UX design, due to its inherent advantages in scalability, editability, and rendering efficiency. In the creation of vector graphics, while expressing creative concepts is straightforward, translating them into precise digital artworks is often challenging and time-consuming. To overcome this technical bottleneck and achieve intelligent conversion from concept to final product, we have constructed SVG-1M, a large-scale dataset of high-quality SVG samples with paired textual descriptions. Through innovative data augmentation and annotation processes, we built precisely aligned ''Text instruction-SVG code'' training pairs, with a subset enhanced by Chain-of-Thought (CoT) annotations. This provides rich semantic supervision signals for model learning. Based on this dataset, we propose SVGen, an end-to-end generative model capable of directly converting natural language descriptions into SVG code. This design addresses the challenges of generating semantically accurate vector graphics while preserving complete structural information. We explored various training strategies and introduced a progressive curriculum learning approach, optimized with reinforcement learning algorithms. Notably, this study innovatively applies the CoT paradigm to vector graphics generation, effectively enhancing both the accuracy and interpretability of SVG synthesis. Experimental validation demonstrates that SVGen exhibits significant advantages over general large models in terms of SVG generation quality, while also surpassing optimization-based rendering methods in generation efficiency. The proposed method enables intelligent conversion between natural language and vector graphics, enabling novel workflows like real-time AI-assisted design iteration. Code, model, and data is released at: https://github.com/gitcat-404/SVGen Zhiyuan Zhao 0005, Yuandong Liu, Da Zhang 0010, Junyu Gao 0001, Hao Sun 0038, Xuelong Li 0001 |
ACM Multimedia | 2 |
| 2025 | PUO-Bench: A Panel Understanding and Operation Benchmark with A Privacy-Preserving FrameworkabstractRecent advancements in Vision-Language Models (VLMs) have enabled GUI agents to leverage visual features for interface understanding and operation in the digital world. However, limited research has addressed the interpretation and interaction with control panels in real-world settings. To bridge this gap, we propose the Panel Understanding and Operation (PUO) benchmark, comprising annotated panel images from appliances and associated vision-language instruction pairs. Experimental results on the benchmark demonstrate significant performance disparities between zero-shot and fine-tuned VLMs, revealing the lack of PUO-specific capabilities in existing language models. Furthermore, we introduce a Privacy-Preserving Framework (PPF) to address privacy concerns in cloud-based panel parsing and reasoning. PPF employs a dual-stage architecture, performing panel understanding on edge devices while delegating complex reasoning to cloud-based LLMs. Although this design introduces a performance trade-off due to edge model limitations, it eliminates the transmission of raw visual data, thereby mitigating privacy risks. Overall, this work provides foundational resources and methodologies for advancing interactive human-machine systems and robotic field in panel-centric applications. Wei Lin 0018, Yiwei Zhou, Zhiyuan Zhao 0005, Junyu Gao 0001, Antoni B. Chan, Xuelong Li 0001 |
NeurIPS | 5 |
| 2025 | U3M: Unbiased Multiscale Modal Fusion Model for Multimodal Semantic Segmentation
Bingyu Li 0002, Da Zhang 0010, Zhiyuan Zhao 0005, Junyu Gao 0001, Xuelong Li 0001 |
Pattern Recognit. | 3 |
| 2025 | Combining SAM With Limited Data for Change Detection in Remote SensingabstractChange detection is a critical task in the remote sensing image (RSI) analysis, widely used in fields such as land cover change and urban planning. With the introduction of foundational models like SAM in computer vision (CV) tasks, their advantages in zero-shot and interactive segmentation have enabled rapid application across diverse visual scenarios. Current research in change detection focuses on designing learnable plug-in modules and fine-tuning foundational models using large annotated data. However, constructing comprehensive datasets and designing effective additional modules pose significant challenges, leading to high costs. To address these issues, we propose a model named Meta-CD for remote sensing change detection (RSCD) with limited data. By introducing a simple fine-tuning module, this model is trained on limited datasets and quickly adapts to change detection tasks. Specifically, we integrate an additional CNN as an adapter with the foundational model FastSAM. Initially, we freeze the parameters of FastSAM and train only the parameters of the introduced adapter and decoder to generate change confidence maps. Subsequently, to enhance the quality of change detection, we introduce a novel pixel-level binarization module that learns the threshold for each pixel in the original image. This module combines the thresholds with the confidence maps to output binary change detection maps, filtering out invalid change pixels. Experimental results demonstrate that our method outperforms other competing approaches on limited datasets and has great zero-shot learning ability. Our code is available at Meta-CD. Junyu Gao 0001, Da Zhang 0010, Lichen Ning, Zhiyuan Zhao 0005, Xuelong Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Integrating SAM With Feature Interaction for Remote Sensing Change DetectionabstractVision foundation models (VFMs) have rapidly gained application across various visual scenarios due to their robust universality and generalization capabilities. However, when directly applied to remote sensing images (RSIs), their performance often falls short owing to the unique inherent imaging characteristics. Moreover, these models typically suffer from inadequate feature extraction capabilities and unclear boundary detection because of the lack of specialized knowledge in the remote sensing (RS) field. To ameliorate these issues, we propose SFCD-Net, a novel network integratingSAM withfeature interaction for RSchangedetection. To be specific, we first introduce a parameter-efficient fine-tuning (PEFT) method that allows the model to learn domain-specific knowledge, thereby enhancing its fine-grained feature extraction capability. Second, an innovative bitemporal feature interaction (BFI) module is designed to improve the model’s sensitivity to changes. Finally, we use the boundary loss function (BLF) to enhance the model’s ability to process boundary details, thereby improving its performance in recognizing boundaries and small targets. Through a series of ablation studies and comparative experiments, we demonstrate that the proposed SFCD-Net significantly improves model adaptability in RS tasks under limited computational resources, outperforming existing models. Da Zhang 0010, Lichen Ning, Zhiyuan Zhao 0005, Junyu Gao 0001, Xuelong Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Deformable Density Estimation via Adaptive RepresentationabstractCrowd counting is the basic task of crowd analysis and it is of great significance in the field of public safety. Therefore, it receives more and more attention recently. The common idea is to combine the crowd counting task with convolutional neural networks to predict the corresponding density map, which is generated by filtering the dot labels with specific Gaussian kernels. Although the counting performance is promoted by the newly proposed networks, they all suffer one conjunct problem, which is due to the perspective effect, there is significant scale contrast among targets in different positions within one scene, but the existing density maps can not represent this scale change well. To address the prediction difficulties caused by target scale variation, we propose a scale-sensitive crowd density map estimation framework, which focuses on dealing with target scale change from density map generation, network design, and model training stage. It consists of the Adaptive Density Map (ADM), Deformable Density Map Decoder (DDMD), and Auxiliary Branch. To be specific, the Gaussian kernel size variates adaptively based on target size to generate ADM that contains scale information for each specific target. DDMD introduces the deformable convolution to fit the Gaussian kernel variation and boosts the model's scale sensitivity. The Auxiliary Branch guides the learning of deformable convolution offsets during the training phase. Finally, we construct experiments on different large-scale datasets. The results show the effectiveness of the proposed ADM and DDMD. Furthermore, the visualization demonstrates that deformable convolution learns the target scale variation. Zhiyuan Zhao 0005, Xuelong Li 0001 |
IEEE Trans. Image Process. | 1 |
| 2022 | ABSSNet: Attention-Based Spatial Segmentation Network for Traffic Scene UnderstandingabstractThe location information of road and lane lines is the supremely important thing for the automatic drive and auxiliary drive. The detection accuracy of these two elements dramatically affects the reliability and practicality of the whole system. In real applications, the traffic scene can be very complicated, which makes it particularly challenging to obtain the precise location of road and lane lines. Commonly used deep learning-based object detection models perform pretty well on the lane line and road detection tasks, but they still encounter false detection and missing detection frequently. Besides, existing convolution neural network (CNN) structures only pay attention to the information flow between layers, while it cannot fully utilize the spatial information inside the layers. To address those problems, we propose an attention-based spatial segmentation network for traffic scene understanding. We use the convolutional attention module to improve the network's understanding capacity of spatial location distribution. Spatial CNN (SCNN) obtains through the information flow within one single convolutional layer and improves the spatial relationship modeling ability of the network. The experimental results demonstrate that this method effectively improves the neural network's application ability of the spatial information, thereby improving the effect of traffic scene understanding. Furthermore, a pixel-level road segmentation dataset called NWPU Road Dataset is built to help improve the process of traffic scene understanding. Xuelong Li 0001, Zhiyuan Zhao 0005, Qi Wang 0009 |
IEEE Trans. Cybern. | 2 |
| 2020 | Deep reinforcement learning based lane detection and localization
Zhiyuan Zhao 0005, Qi Wang 0009, Xuelong Li 0001 |
Neurocomputing | 1 |