VLDB 2026 Research / reviewers in the wild / expert
Xiaoyong Zhu
dblp:33/7023
· DBLP profile ↗
15ranked-venue papers
2as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Trustworthy machine learning · 53% Vision and language · 19% Language models and text generation · 10% | |
| Network and information security
2 papers |
Security and privacy of machine learning · 100% | |
| Databases, data mining, and information retrieval
1 paper |
Machine learning and data management · 50% Query processing and optimization · 50% |
Topics — the 14 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language › vision-language model
multimodal large language model |
2.0 | 3 | 2026 | Towards Visualization-of-Thought Jailbreak Attack against Large Visual Language Models · NeurIPS 2025 Inter: Mitigating Hallucination in Large Vision-Language Models by Interaction Guidance Sampling · ICCV 2025 USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models · ACL (1) 2026 |
Machine learning › Trustworthy machine learning
safety evaluation |
1.9 | 2 | 2026 | USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models · ACL (1) 2026 Chinese SafetyQA: A Safety Short-form Factuality Benchmark for Large Language Models · ACL (1) 2025 |
Security and privacy of machine learning
adversarial attack |
1.7 | 2 | 2025 | Towards Visualization-of-Thought Jailbreak Attack against Large Visual Language Models · NeurIPS 2025 HiddenDetect: Detecting Jailbreak Attacks against Multimodal Large Language Models via Monitoring Hidden States · ACL (1) 2025 |
Security and privacy of machine learning › adversarial attack
jailbreak attack |
1.7 | 2 | 2025 | Towards Visualization-of-Thought Jailbreak Attack against Large Visual Language Models · NeurIPS 2025 HiddenDetect: Detecting Jailbreak Attacks against Multimodal Large Language Models via Monitoring Hidden States · ACL (1) 2025 |
Machine learning › Trustworthy machine learning
robustness |
1.1 | 2 | 2025 | HiddenDetect: Detecting Jailbreak Attacks against Multimodal Large Language Models via Monitoring Hidden States · ACL (1) 2025 Towards Visualization-of-Thought Jailbreak Attack against Large Visual Language Models · NeurIPS 2025 |
Knowledge, reasoning and agents › Multi-agent systems
agent communication |
1.0 | 1 | 2026 | Enabling Agents to Communicate Entirely in Latent Space · ACL (1) 2026 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning › multi-agent communication
latent space communication |
1.0 | 1 | 2026 | Enabling Agents to Communicate Entirely in Latent Space · ACL (1) 2026 |
Machine learning › Trustworthy machine learning › generative model safety
multimodal large language model safety |
1.0 | 1 | 2026 | USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models · ACL (1) 2026 |
Natural language and speech › Language models and text generation
hallucination mitigation |
0.9 | 1 | 2025 | Inter: Mitigating Hallucination in Large Vision-Language Models by Interaction Guidance Sampling · ICCV 2025 |
Machine learning › Trustworthy machine learning › AI safety
jailbreak detection |
0.9 | 1 | 2025 | HiddenDetect: Detecting Jailbreak Attacks against Multimodal Large Language Models via Monitoring Hidden States · ACL (1) 2025 |
Machine learning › Trustworthy machine learning › AI safety
safety alignment |
0.9 | 1 | 2025 | Towards Visualization-of-Thought Jailbreak Attack against Large Visual Language Models · NeurIPS 2025 |
Machine learning and data management › data management for machine learning
feature store |
0.7 | 1 | 2023 | Optimizing Data Pipelines for Machine Learning in Feature Stores · Proc. VLDB Endow. 2023 |
Query processing and optimization › query optimization › join ordering
join optimization |
0.7 | 1 | 2023 | Optimizing Data Pipelines for Machine Learning in Feature Stores · Proc. VLDB Endow. 2023 |
Natural language and speech › Language models and text generation
large language model evaluation |
0.3 | 1 | 2025 | Chinese SafetyQA: A Safety Short-form Factuality Benchmark for Large Language Models · ACL (1) 2025 |
Methods — techniques the papers use, named apart from their topics
visualization-of-thought · 1.7hidden state monitoring · 1.7chain-of-thought · 1.7activation analysis · 1.7safety benchmarking · 1.0interaction guidance sampling · 0.9decoding strategies · 0.9benchmark construction · 0.9database-style optimization · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enabling Agents to Communicate Entirely in Latent SpaceabstractZhuoyun Du, Runze Wang, Huiyu Bai, Zouying Cao, Xiaoyong Zhu, Yu Cheng, Bo Zheng, Wei Chen, Haochao Ying. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zhuoyun Du, Huiyu Bai, Zouying Cao, Xiaoyong Zhu, Wei Chen 0001, Haochao Ying |
ACL (1) | 5 |
| 2026 | USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language ModelsabstractBaolin Zheng, Guanlin Chen, Qingyang Teng, Hongqiong Zhong, Yingshui Tan, Zhendong Liu, Weixun Wang, Jiaheng Liu, Jian Yang, Huiyun Jing, Jincheng Wei, Wenbo Su, Xiaoyong Zhu, Bo Zheng, Kaifu Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Baolin Zheng, Qingyang Teng, Hongqiong Zhong, Yingshui Tan, Weixun Wang, Jian Yang 0037, Huiyun Jing, Jincheng Wei, Wenbo Su, Xiaoyong Zhu, Bo Zheng 0007, Kaifu Zhang |
ACL (1) | 13 |
| 2026 | REVISION:Reflective Intent Mining and Online Reasoning Auxiliary for E-Commerce Visual Search System Optimization
Qiuyu Zhao, Zenghui Sun, Jinsong Lan, Xiaoyong Zhu, Bo Zheng 0007 |
ICDE | 5 |
| 2025 | HiddenDetect: Detecting Jailbreak Attacks against Multimodal Large Language Models via Monitoring Hidden StatesabstractThe integration of additional modalities increases the susceptibility of large vision-language models (LVLMs) to safety risks, such as jailbreak attacks, compared to their language-only counterparts. While existing research primarily focuses on post-hoc alignment techniques, the underlying safety mechanisms within LVLMs remain largely unexplored. In this work , we investigate whether LVLMs inherently encode safety-relevant signals within their internal activations during inference. Our findings reveal that LVLMs exhibit distinct activation patterns when processing unsafe prompts, which can be leveraged to detect and mitigate adversarial inputs without requiring extensive fine-tuning. Building on this insight, we introduce HiddenDetect, a novel tuning-free framework that harnesses internal model activations to enhance safety. Experimental results show that HiddenDetect surpasses state-of-the-art methods in detecting jailbreak attacks against LVLMs. By utilizing intrinsic safety-aware patterns, our method provides an efficient and scalable solution for strengthening LVLM robustness against multimodal threats. Our code and data will be released publicly. Yilei Jiang, Xinyan Gao, Tianshuo Peng, Yingshui Tan, Xiaoyong Zhu, Bo Zheng 0007, Xiangyu Yue 0001 |
ACL (1) | 5 |
| 2025 | Chinese SafetyQA: A Safety Short-form Factuality Benchmark for Large Language ModelsabstractYingshui Tan, Boren Zheng, Baihui Zheng, Kerui Cao, Huiyun Jing, Jincheng Wei, Jiaheng Liu, Yancheng He, Wenbo Su, Xiaoyong Zhu, Bo Zheng, Kaifu Zhang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Yingshui Tan, Boren Zheng, Baihui Zheng, Kerui Cao, Huiyun Jing, Jincheng Wei, Yancheng He, Wenbo Su, Xiaoyong Zhu, Bo Zheng 0007, Kaifu Zhang |
ACL (1) | 10 |
| 2025 | Inter: Mitigating Hallucination in Large Vision-Language Models by Interaction Guidance SamplingabstractHallucinations in large vision-language models (LVLMs) pose significant challenges for real-world applications, as LVLMs may generate responses that appear plausible yet remain inconsistent with the associated visual content. This issue rarely occurs in human cognition. We argue that this discrepancy arises from humans' ability to effectively leverage multimodal interaction information in data samples. Specifically, humans typically first gather multimodal information, analyze the interactions across modalities for understanding, and then express their understanding through language. Motivated by this observation, we conduct extensive experiments on popular LVLMs and obtained insights that surprisingly reveal human-like, though less pronounced, cognitive behavior of LVLMs on multimodal samples. Building on these findings, we further propose \textbf{INTER}: \textbf{Inter}action Guidance Sampling, a novel training-free algorithm that mitigate hallucinations without requiring additional data. Specifically, INTER explicitly guides LVLMs to effectively reapply their understanding of multimodal interaction information when generating responses, thereby reducing potential hallucinations. On six benchmarks including VQA and image captioning tasks, INTER achieves an average improvement of up to 3.4\% on five LVLMs compared to the state-of-the-art decoding strategy. The code will be released when the paper is accepted. Zenghui Sun, Lihua Jing, Jinsong Lan, Xiaoyong Zhu, Bo Zheng 0007 |
ICCV | 9 |
| 2025 | Towards Visualization-of-Thought Jailbreak Attack against Large Visual Language ModelsabstractAs Visual Language Models (VLMs) continue to evolve, they have demonstrated increasingly sophisticated logical reasoning capabilities and multimodal thought generation, opening doors to widespread applications. However, this advancement raises serious concerns about content security, particularly when these models process complex multimodal inputs requiring intricate reasoning. When faced with these safety challenges, the critical competition between logical reasoning and safety objectives of VLMs is often overlooked in previous works. In this paper, we introduce Visualization-of-Thought Attack (\textbf{VoTA}), a novel and automated attack framework that strategically constructs chains of images with risky visual thoughts to challenge victim models. Our attack provokes the inherent conflict between the model's logical processing and safety protocols, ultimately leading to the generation of unsafe content. Through comprehensive experiments, VoTA achieves remarkable effectiveness, improving the average attack success rate (ASR) by 26.71\% (from 63.70\% to 90.41\%) on 9 open-source and 6 commercial VLMs, compared to the state-of-the-art methods. These results expose a critical vulnerability: current VLMs struggle to maintain safety guarantees when processing insecure multimodal visualization-of-thought inputs, highlighting the urgency and necessity of enhancing safety alignment. Our code and dataset are available at
https://github.com/Hongqiong12/VoTA.
Content Warning: This paper contains harmful contents that may be offensive. Hongqiong Zhong, Qingyang Teng, Baolin Zheng, Yingshui Tan, Wenbo Su, Xiaoyong Zhu, Bo Zheng 0007, Kaifu Zhang |
NeurIPS | 9 |
| 2025 | Dual-Layer-Based Active Disturbance Rejection Decoupling Control for Lateral Stability of Power-Decoupled Agricultural Mobile Platforms Under Complex Operating Conditions
Xiaoyong Zhu, Yudan Cai, Lei Xu 0034, Lizhang Xu, Wen-Hua Chen 0001 |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2024 | Accuracy Assessment of GF-7 Geolocation Without GCPs Considering Atmospheric Refraction and Aberration of LightabstractConsidering that the atmospheric refraction and aberration of light (AOL) bend the path of light and thus affect the geolocation accuracy, a rigorous imaging geometric model with corrections of atmospheric refraction and AOL is proposed to improve the accuracy of geolocation for GaoFen-7 (GF-7) images without ground control points (GCPs). The correction of atmospheric refraction is calculated using a simplified two-layer atmospheric refraction model, while the correction of AOL is calculated using the satellite altitude and velocity. Then, the rigorous imaging geometric model is refined with both corrections. A total of 68 scenes of GF-7 satellite backward images in ten regions with different roll-off-nadir angles are selected to compare and detect the variation of accuracy with and without the corrections. The results show that the average geolocation accuracy is improved by 0.14 m in object space, and about 0.20 pixels in image space if the corrections of atmospheric refraction and AOL are taken into the geometric model. The average improvement of RMSEs is about 1.00 pixels and the ratio of improved scenes reaches 88.89% when the roll-off-nadir angle is larger than 15°, whereas there is no need to consider atmospheric refraction and AOL if the off-nadir angle is smaller than 5°. The methodology and results can provide support for the need for corrections for high-precision positioning and the promotion of GF-7 imagery. Xiaoyong Zhu, Xinming Tang, Wenmin Hu, Bin Liu 0049, Guo Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | AI meets UAVs: A survey on AI empowered UAV perception systems for precision agriculture
Jinya Su, Xiaoyong Zhu, Shihua Li 0001, Wen-Hua Chen 0001 |
Neurocomputing | 2 |
| 2023 | Optimizing Data Pipelines for Machine Learning in Feature StoresabstractData pipelines (i.e., converting raw data to features) are critical for machine learning (ML) models, yet their development and management is time-consuming. Feature stores have recently emerged as a new "DBMS-for-ML" with the premise of enabling data scientists and engineers to define and manage their data pipelines. While current feature stores fulfill their promise from a functionality perspective, they are resource-hungry---with ample opportunities for implementing database-style optimizations to enhance their performance. In this paper, we propose a novel set of optimizations specifically targeted for point-in-time join, which is a critical operation in data pipelines. We implement these optimizations on top of Feathr: a widely-used feature store, and evaluate them on use cases from both the TPCx-AI benchmark and real-world online retail scenarios. Our thorough experimental analysis shows that our optimizations can accelerate data pipelines by up to 3× over state-of-the-art baselines. Rui Liu 0002, Kwanghyun Park 0001, Fotis Psallidas, Xiaoyong Zhu, Jinghui Mo, Rathijit Sen, Matteo Interlandi, Konstantinos Karanasos, Yuanyuan Tian 0001, Jesús Camacho-Rodríguez |
Proc. VLDB Endow. | 4 |
| 2022 | Systematic Geolocation Errors of FengYun-3D MERSI-IIabstractGeolocation accuracy is a critical issue for remote sensing applications. To achieve subpixel accuracy, geolocation errors need to be systematically identified and corrected. In this study, we propose a geometric sensor model for FengYun-3D (FY-3D) MERSI-II, a second-generation visible (VIS)/infrared (IR) spectroradiometer, to generate the geolocation lookup table (GLT). The geometric sensor model retrieves the imaging rays from the focal plane to the K-mirrors, 45° scanning mirrors, the platform, and the earth’s surface. After refining the attitude errors with ground control points (GCPs), the rigorous sensor model can achieve subpixel geolocation accuracy. However, significant systematic geolocation errors were identified from the residuals, especially for the area with large view angles. To study the errors of MERSI-II, we proposed a homogenous coordinate in the focal plane. As proven by both theory and experiments, the attitudes were adjusted to a wrong value and introduced systematic errors when there were principal point errors. The pitch angle error of K-mirrors caused the oscillation in the flight direction. The principal distance error introduced line coordinate-related error in the flight direction. Meanwhile, the initial phase angle error between the K-mirror and 45° scanning mirrors caused the line coordinate-related errors in the scanning direction. After correcting all the above-mentioned errors, the systematic geolocation errors of MERSI-II were removed. With 23 independent datasets, the root mean square errors (RMSEs) of 250 m bands were approximately 0.4 pixels, 100 m at nadir. Zehua Cui, Xiuqing Hu, Xiaoyong Zhu |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2016 | A Penalized Spline-Based Attitude Model for High-Resolution Satellite ImageryabstractAttitude models play a prominent role in the geometric processing of high-resolution satellite imagery (HRSI). Because of the high accuracy of the matching algorithm, attitude oscillations can occur in HRSI. Various methods for correcting this attitude oscillation with parallax observations have been proposed. However, few researchers have attempted to model the oscillation from the attitude records or have taken noise into consideration. In this paper, a penalized spline-based attitude model is proposed, which can model the oscillation with piecewise and continuously differentiable polynomials and smooth out the attitude noise with a penalty function. The balance between the fitting accuracy and noise smoothing is controlled by a penalty parameter, which is estimated by generalized cross-validation. Given that the attitude error introduces distortions into sensor-corrected images, the band-to-band registration of multispectral images is used to validate the attitude model. Five multispectral data sets captured by ZiYuan-3 are used to demonstrate the effectiveness of the proposed method. Compared with third-degree polynomials and cubic spline interpolation, the penalized spline model delivers the best performance by limiting the misregistration caused by the attitude model to within 0.1 pixels. Zhengrong Zou, Guo Zhang 0001, Xiaoyong Zhu, Xinming Tang |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2014 | Geometric Accuracy Validation for ZY-3 Satellite ImageryabstractThe ZiYuan-3 surveying satellite (ZY-3) is a high-precision civilian satellite imaging sensor. Since its launch on January 9, 2012, it has been in operation for one and a half years. Although the initial postlaunch ZY-3 geometric accuracy was verified during an in-orbit operation period, on-orbit calibration was still necessary from time to time. This on-orbit calibration has vastly improved the location accuracy in planimetry for ZY-3 panchromatic images. This letter briefly describes the principle of on-orbit calibration and production processes of sensor-corrected products. Furthermore, block adjustment based on a rational function model test showed planimetric and vertical accuracy values of 10 m and 5 m, respectively, without ground control points (GCPs). The accuracy values improved to 3 m and 2 m, respectively, with a few GCPs. The statistics results are from ten different regions with independent checkpoints (ICPs). All accuracy values are the root-mean-square error of ICPs. Therefore, ZY-3 can be used for the generation of cartographic maps at the 1 : 50 000 scale and for revision and updates of 1 : 25 000 scale maps. Compared with other mainstream high-resolution satellite images of the same ground resolution, ZY-3's geometric accuracy is almost the same and sometimes even better. Taoyang Wang, Guo Zhang 0001, DeRen Li, Xinming Tang, Yonghua Jiang 0001, Xiaoyong Zhu |
IEEE Geosci. Remote. Sens. Lett. | 7 |
| 2012 | Electromagnetic performances analysis of a new magnetic-planetary-geared permanent magnet brushless machine for hybrid electric vehiclesabstractThis paper proposes a new magnetic-planetary-geared permanent magnet brushless (MPG-PMBL) machine for hybrid electric vehicles. The key is to integrate a magnetic planetary gear into permanent magnet brushless machine. Thus, the resulting machine can achieve both power split and mixing flexibly, and therefore, can obtain ultralow emission and high fuel efficiency at different operational modes. Moreover, by incorporating the concept of non-contact transmission mechanisms of magnetic gear, the proposed machine possesses the merits of small size, light weight, reliability, reduced maintenance and inherent overload protection. By using the finite element method, the electromagnetic performances are analyzed. Finally, multi-operational modes are analyzed to evaluate the proposed system. Xiaoyong Zhu, Lingting Kong, Wenxiang Zhao, Li Quan 0001, Ming Cheng 0001 |
IECON | 1 |