VLDB 2026 Research / reviewers in the wild / expert
Ruiyuan Zhang
dblp:206/0708
· DBLP profile ↗
21ranked-venue papers
4as first author
20since 2021 · last 2026
0000-0003-2022-7387ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 3 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 8 since 2021Databases, data management, data science and information retrieval · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Active Multi-source Domain Adaptation for Multimodal Fake News DetectionabstractMultimodal fake news detection plays a crucial role in combating online misinformation. The inherent domain diversity of news in the real world has driven the development of cross-domain detection methods. However, these detection methods either suffer from significant performance degradation due to semantic and deception pattern shifts between the training (source) and test (target) domains or heavily rely on annotated labels. To address the problems, we propose ADOSE, an active multi-source domain adaptation framework for multimodal fake news detection which actively annotates a small subset of target samples to improve detection performance. Specifically, for domain shifts, we design a multi-expert classifier network based on refined features to comprehensively capture and adapt to the semantic space and deception patterns of news across different domains. To maximize adaptation performance with limited annotation cost, we propose a least-disagree uncertainty selector equipped with a diversity calculator for selecting the most informative samples. The selector leverages the uncertainty of inconsistent predictions before and after perturbations by multiple classifiers as an indicator of unfamiliar samples. It further incorporates diversity scores derived from multi-view features to ensure the chosen samples achieve maximal coverage of target domain features. The extensive experiments on multiple datasets show that ADOSE outperforms existing domain adaptation methods by 2.45% ~ 9.1%, indicating the superiority of our model. Mengze Li 0001, Yue Cui 0001, Ruiyuan Zhang, Hanghui Guo, Shimin Di, Ziyi Liu 0005, Jia Zhu 0003, Jiajie Xu 0001 |
AAAI | 6 |
| 2026 | Faithful in Steps: Improving Generalization and Citation in RAG via Query DecompositionabstractRetrieval-augment generation is a prevalent strategy to mitigate hallucinations of LLMs. The attributable RAG (RAGQ) generates quotes for its answers. The quotes indicate which input contexts support the RAG to derive the answers, enhancing the answer's verifiability and trustworthiness. However, existing RAGQs exhibit significant degradation when dealing with questions that require multi-hop reasoning and multi-modal understanding, suffering from over-citation, implicit entity identification failure, and poor generalization. In this paper, we propose a novel RAGQ framework, namely QDRAG. QDRAG breaks down the input question into atomic subquestions to identify the implicit entities. Then, the reranker prunes context distractors to eliminate the downstream over-citation. To facilitate query decomposition, we propose two zero-shot approaches: QD-C and QD-R, which guide the QD MLLM to decompose the question based on context knowledge and retrieval rewards, respectively. One interesting finding is that finetuning on the QD task shows better generalizability compared to directly finetuning on the downstream RAGQ task. Experiments on four multi-modal QA benchmarks demonstrate QDRAG's efficacy in grounding answers and generating faithful citations. The framework significantly outperforms all the baselines on both in-domain and out-of-domain tests, even surpassing Gemini-Pro. Zhongying Ru, Shimin Di, Ruiyuan Zhang, Xiaofang Zhou 0001 |
AAAI | 5 |
| 2026 | PFAvatar: Pose-Fusion 3D Personalized Avatar Reconstruction from Real-World Outfit-of-the-Day PhotosabstractWe propose PFAvatar (Pose-Fusion Avatar), a new method that reconstructs high-quality 3D avatars from Outfit of the Day (OOTD) photos, which exhibit diverse poses, occlusions, and complex backgrounds. Our method consists of two stages: (1) fine-tuning a pose-aware diffusion model from few-shot OOTD examples and (2) distilling a 3D avatar represented by a neural radiance field (NeRF). In the first stage, unlike previous methods that segment images into assets (e.g. garments, accessories) for 3D assembly, which is prone to inconsistency, we avoid decomposition and directly model the full-body appearance. By integrating a pre-trained ControlNet for pose estimation and a novel Condition Prior Preservation Loss (CPPL), our method enables end-to-end learning of fine details while mitigating language drift in few-shot training. Our method completes personalization in just 5 minutes, achieving a 48x speed-up compared to previous approaches. In the second stage, we introduce a NeRF-based avatar representation optimized by canonical SMPL-X space sampling and Multi-Resolution 3D-SDS. Compared to mesh-based representations that suffer from resolution-dependent discretization and erroneous occluded geometry, our continuous radiance field can preserve high-frequency textures (e.g., hair) and handle occlusions correctly through transmittance. Experiments demonstrate that PFAvatar outperforms state-of-the-art methods in terms of reconstruction fidelity, detail preservation, and robustness to occlusions/truncations, advancing practical 3D avatar generation from real-world OOTD albums. In addition, the reconstructed 3D avatars support downstream applications such as virtual try-on, animation, and human video reenactment, further demonstrating the versatility and practical value of our approach. Dianbing Xi, Guoyuan An, Jingsen Zhu, Ruiyuan Zhang, Jiayuan Lu, Yuchi Huo, Rui Wang 0004 |
AAAI | 6 |
| 2025 | KPL: Training-Free Medical Knowledge Mining of Vision-Language ModelsabstractVisual Language Models such as CLIP excel in image recognition due to extensive image-text pre-training. However, applying the CLIP inference in zero-shot classification, particularly for medical image diagnosis, faces challenges due to: 1) the inadequacy of representing image classes solely with single category names; 2) the modal gap between the visual and text spaces generated by CLIP encoders. Despite attempts to enrich disease descriptions with large language models, the lack of class-specific knowledge often leads to poor performance. In addition, empirical evidence suggests that existing proxy learning methods for zero-shot image classification on natural image datasets exhibit instability when applied to medical datasets. To tackle these challenges, we introduce the Knowledge Proxy Learning (KPL) to mine knowledge from CLIP. KPL is designed to leverage CLIP's multimodal understandings for medical image classification through Text Proxy Optimization and Multimodal Proxy Learning. Specifically, KPL retrieves image-relevant knowledge descriptions from the constructed knowledge-enhanced base to enrich semantic text proxies. It then harnesses input images and these descriptions, encoded via CLIP, to stably generate multimodal proxies that boost the zero-shot classification performance. Extensive experiments conducted on both medical and natural image datasets demonstrate that KPL enables effective zero-shot image classification, outperforming all baselines. These findings highlight the great potential in this paradigm of mining knowledge from CLIP for medical image classification and broader areas. Tianxiang Hu, Jiawei Du 0002, Ruiyuan Zhang, Joey Tianyi Zhou, Zuozhu Liu |
AAAI | 4 |
| 2025 | LegalReasoner: Step-wised Verification-Correction for Legal Judgment ReasoningabstractWeijie Shi, Han Zhu, Jiaming Ji, Mengze Li, Jipeng Zhang, Ruiyuan Zhang, Jia Zhu, Jiajie Xu, Sirui Han, Yike Guo. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Jiaming Ji, Mengze Li 0001, Ruiyuan Zhang, Jia Zhu 0003, Jiajie Xu 0001, Sirui Han, Yike Guo |
ACL (1) | 6 |
| 2025 | Efficient Time-Dependent Shortest Path Finding on Cargo NetworkabstractSurging e-commerce and global trade necessitate highly efficient cargo terminal operations. Modern automated terminals, crucial for supply chains, employ complex networks of static and movable equipment. This integration introduces a core challenge: movable equipment creates dynamic connectivity and state-dependent travel times, rendering classic shortest path algorithms based on static edge weights ineffective. Unlike typical time-dependent problems driven by external factors such as traffic congestion or fixed schedules, our dynamics stem from internal equipment state, presenting a unique optimization challenge. We address the problem of finding optimal cargo routes within these dynamic environments. We propose a novel approach by modeling the terminal as a cargo network, where virtual edges induced by movable equipment are explicitly materialised and edge costs reflect the status of the real-time equipment. We propose an efficient Dijkstra's-based algorithm to solve the cargo routing problem within this framework considering the system dynamics. The primary contributions of this paper are this novel modeling technique for dynamic terminals and the adapted algorithm for optimal routing, offering significant benefits for logistics optimization and automated warehouse design. Experimental results demonstrate that our approach significantly reduces cargo travel times compared to baseline methods, offering substantial improvements for logistics efficiency in automated terminals. Elton Chun-Chai Li, Ziyi Liu 0005, Ruiyuan Zhang, Sean Shing Fung Lau, Yehong Xu, Xiaofang Zhou 0001 |
IEEE Big Data | 3 |
| 2025 | DIDS: Domain Impact-aware Data Sampling for Large Language Model TrainingabstractWeijie Shi, Jipeng Zhang, Yaguang Wu, Jingzhi Fang, Shibo Zhang, Yao Zhao, Hao Chen, Ruiyuan Zhang, Yue Cui, Jia Zhu, Sirui Han, Jiajie Xu, Xiaofang Zhou. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Yaguang Wu, Jingzhi Fang, Ruiyuan Zhang, Yue Cui 0001, Jia Zhu 0003, Sirui Han, Jiajie Xu 0001, Xiaofang Zhou 0001 |
EMNLP | 8 |
| 2025 | Leveraging Pretrained Diffusion Models for Zero-Shot Part Assemblyabstract3D part assembly aims to understand part relationships and predict their 6-DoF poses to construct realistic 3D shapes, addressing the growing demand for autonomous assembly, which is crucial for robots. Existing methods mainly estimate the transformation of each part by training neural networks under supervision, which requires a substantial quantity of manually labeled data. However, the high cost of data collection and the immense variability of real-world shapes and parts make traditional methods impractical for large-scale applications. In this paper, we propose first a zero-shot part assembly method that utilizes pre-trained point cloud diffusion models as discriminators in the assembly process, guiding the manipulation of parts to form realistic shapes. Specifically, we theoretically demonstrate that utilizing a diffusion model for zero-shot part assembly can be transformed into an Iterative Closest Point (ICP) process. Then, we propose a novel pushing-away strategy to address the overlap parts, thereby further enhancing the robustness of the method. To verify our work, we conduct extensive experiments and quantitative comparisons to several strong baseline methods, demonstrating the effectiveness of the proposed approach, which even surpasses the supervised learning method. The code has been released on https://github.com/Ruiyuan-Zhang/Zero-Shot-Assembly. Ruiyuan Zhang, Qi Wang 0111, Yuchi Huo, Chao Wu 0001 |
IJCAI | 1 |
| 2025 | CargoFlow: A Comprehensive System for Congestion Detection and Root Cause Analysis in Cargo HandlingabstractCongestion remains one of the most critical yet inadequately addressed challenges in air cargo material handling systems. The industry currently faces two key limitations: (1) traditional definitions of congestion fail to meet analytical requirements, and (2) root cause analysis relies heavily on manual expert judgment. To address these issues, this paper introduces CargoFlow, a system co-developed with industry partners that integrates advanced search algorithms, large language models (LLMs), and a 3D interactive data visualization interface. The system enables efficient and accurate detection of complex congestion patterns, including those imperceptible to human operators. CargoFlow employs optimized search algorithms to detect intricate congestion patterns efficiently, overcoming the limitations of human-driven inspection. By leveraging advanced LLMs with sophisticated prompt engineering, the system min-imizes human effort in diagnosing root causes. The proposed solution has been deployed and rigorously tested in real-world industrial environments, demonstrating its practical effectiveness. Elton Chun-Chai Li, Ruiyuan Zhang, Yichen Ren, Sean Shing Fung Lau, Morgan Xian Biao Hiew, Yan Nei Law |
INDIN | 2 |
| 2024 | Scalable Geometric Fracture Assembly via Co-creation Space among AssemblersabstractGeometric fracture assembly presents a challenging practical task in archaeology and 3D computer vision. Previous methods have focused solely on assembling fragments based on semantic information, which has limited the quantity of objects that can be effectively assembled. Therefore, there is a need to develop a scalable framework for geometric fracture assembly without relying on semantic information. To improve the effectiveness of assembling geometric fractures without semantic information, we propose a co-creation space comprising several assemblers capable of gradually and unambiguously assembling fractures. Additionally, we introduce a novel loss function, i.e., the geometric-based collision loss, to address collision issues during the fracture assembly process and enhance the results. Our framework exhibits better performance on both PartNet and Breaking Bad datasets compared to existing state-of-the-art frameworks. Extensive experiments and quantitative comparisons demonstrate the effectiveness of our proposed framework, which features linear computational complexity, enhanced abstraction, and improved generalization. Our code is publicly available at https://github.com/Ruiyuan-Zhang/CCS. Ruiyuan Zhang, Zexi Li 0001, Hao Dong 0003, Jie Fu 0001, Chao Wu 0001 |
AAAI | 1 |
| 2024 | DeepTreeSketch: Neural Graph Prediction for Faithful 3D Tree Modeling from SketchesabstractWe present DeepTreeSketch, a novel AI-assisted sketching system that enables users to create realistic 3D tree models from 2D freehand sketches. Our system leverages a tree graph prediction network, TGP-Net, to learn the underlying structural patterns of trees from a large collection of 3D tree models. The TGP-Net simulates the iterative growth of botanical trees and progressively constructs the 3D tree structures in a bottom-up manner. Furthermore, our system supports a flexible sketching mode for both precise and coarse control of the tree shapes by drawing branch strokes and foliage strokes, respectively. Combined with a procedural generation strategy, users can freely control the foliage propagation with diverse and fine details. We demonstrate the expressiveness, efficiency, and usability of our system through various experiments and user studies. Our system offers a practical tool for 3D tree creation, especially for natural scenes in games, movies, and landscape applications. Fangyuan Tu, Ruiyuan Zhang, Zhanglin Cheng, Naoto Yokoya |
CHI | 4 |
| 2024 | Efficient Approximate Maximum Inner Product Search Over Sparse VectorsabstractThe maximum inner product search (MIPS) problem in high-dimensional vector spaces has various applications, primarily driven by the success of deep neural network-based embedding models. Existing MIPS methods designed for dense vectors using approximate techniques like locality-sensitive hashing (LSH) have been well studied, but they are not efficient and effective for searching sparse vectors due to the near-orthogonality among the sparse vectors. The solutions to MIPS over sparse vectors rely heavily on inverted lists, resulting in poor query efficiency, particularly when dealing with large-scale sparse datasets. In this paper, we introduce SOSIA, a novel framework specifically tailored to address these limitations. To handle sparsity, we propose the SOS transformation, which converts sparse vectors into a binary space while providing an unbiased estimator of the inner product between any two vectors. Additionally, we develop a minHash-based index to enhance query efficiency. We provide a theoretical analysis on the query quality of SOSIA and present extensive experiments on real-world sparse datasets to validate its effectiveness. The experimental results demonstrate its superior performance in terms of query efficiency and accuracy compared to existing methods. Xi Zhao 0006, Zhonghan Chen, Kai Huang 0011, Ruiyuan Zhang, Bolong Zheng, Xiaofang Zhou 0002 |
ICDE | 4 |
| 2024 | LDPGuard: Defenses Against Data Poisoning Attacks to Local Differential Privacy ProtocolsabstractThe protocols that satisfy Local Differential Privacy (LDP) enable untrusted third parties to collect aggregate information about a population without disclosing each user's privacy. In particular, each user locally encodes and perturbs his private data before sending it to the data collector, who aggregates and estimates the statistics about the population based on the collected perturbed values from individuals. Owing to their growing importance, LDP protocols have been widely studied and deployed in real-world scenarios (eg Chrome and Windows). However, as data poisoning attacks may be injected by attackers who introduce many fake users, the utility of the statistics is heavily poisoned. In this paper, we present a generic and extensible framework called LDPGuard to address the problem. LDPGuard provides effective defenses against data poisoning attacks to LDP protocols for frequency estimation, a basic query of most data analytics tasks. In particular, it first precisely estimates the percentage of fake users and then provides adversarial schemes to defend against particular data poisoning attacks. Experimental study on real-world and synthetic datasets demonstrates the superiority of LDPGuard compared to existing techniques. Kai Huang 0011, Gaoya Ouyang, Qingqing Ye 0001, Haibo Hu 0001, Bolong Zheng, Xi Zhao 0006, Ruiyuan Zhang, Xiaofang Zhou 0001 |
IEEE Trans. Knowl. Data Eng. | 7 |
| 2023 | Score-PA: Score-based 3D Part Assembly
Junfeng Cheng, Mingdong Wu, Ruiyuan Zhang, Guanqi Zhan, Chao Wu 0001, Hao Dong 0003 |
BMVC | 3 |
| 2023 | ST-MoE: Spatio-Temporal Mixture-of-Experts for Debiasing in Traffic PredictionabstractThe pervasiveness of GPS-enabled devices and wireless communication technologies results in a proliferation of traffic data in intelligent transportation systems, where traffic prediction is often essential to enable reliability and safety. Many recent studies target traffic prediction using deep learning techniques. They model spatio-temporal dependencies among traffic states by deep learning and achieve good overall performance. However, existing studies ignore the bias on traffic prediction models, which refers to non-uniformed performance distribution across road segments, especially the significantly poor prediction results on certain road segments. To solve this issue, we propose a framework named spatio-temporal mixture-of-experts (ST-MoE) that aims to eliminate the bias on traffic prediction. In general, we refer to any traffic prediction model as the based model, and adopt the proposed ST-MoE framework as a plug-in to debias. ST-MoE uses stacked convolution-based networks to learn spatio-temporal representations of individual patterns of road segments and then adaptively assigns appropriate expert layers (sub-networks) to different patterns through a spatio-temporal gating network. To this end, the patterns can be distinguished, and biased performance among road segments can be eliminated by experts tailored for specific patterns, which also further improves the overall prediction accuracy of the base model. Extensive experimental results on various base models and real-world datasets prove the effectiveness of ST-MoE. Shuhao Li 0001, Yue Cui 0001, Yan Zhao 0008, Weidong Yang 0001, Ruiyuan Zhang, Xiaofang Zhou 0001 |
CIKM | 5 |
| 2023 | A Learned Cuckoo Filter for Approximate Membership Queries over Variable-sized Sliding Windows on Data StreamsabstractDesigning a space-efficient data structure to answer membership queries while ensuring high accuracy and real-time response is a challenging task in the field of stream processing. Many techniques have been developed to answer these queries in a sliding windows manner. However, assuming the user will conduct the query with the presupposed window size is not always practical. In this paper, we introduce a novel data structure called Learned Cuckoo Filter (LCF). It can provide satisfactory results for the approximate membership query on data streams, regardless of the user-defined query windows. LCF operates by adaptively maintaining cuckoo filters with the assistance of a well-trained oracle that learned the frequency feature of the data within the stream. To further enhance memory utilization, we develop a compact version of LCF (denoted by LCF_C), which selectively removes redundant information to reduce space consumption without compromising query accuracy. Furthermore, we conduct a thorough theoretical analysis of query accuracy and provide detailed guidelines for optimal parameter selection (denoted by LCF_O). Extensive experimental studies on synthetic and real-world datasets demonstrate the superiority of the proposed methods in terms of both space consumption and accuracy. Compared to the state-of-the-art algorithms, LCF_O can reduce up to 61% of space cost at the same error level, and achieve up to 12× improved accuracy with the same space cost. Tingyun Yan, Ruiyuan Zhang, Kai Huang 0011, Bolong Zheng, Xiaofang Zhou 0001 |
Proc. ACM Manag. Data | 3 |
| 2023 | VisualNeo: Bridging the Gap between Visual Query Interfaces and Graph Query EnginesabstractVisual Graph Query Interfaces (VQIs) empower non-programmers to query graph data by constructing visual queries intuitively. Devising efficient technologies in Graph Query Engines (GQEs) for interactive search and exploration has also been studied for years. However, these two vibrant scientific fields are traditionally independent of each other, causing a vast barrier for users who wish to explore the full-stack operations of graph querying. In this demonstration, we propose a novel VQI system built upon Neo4j called VisualNeo that facilities an efficient subgraph query in large graph databases. VisualNeo inherits several advanced features from recent advanced VQIs, which include the data-driven gui design and canned pattern generation. Additionally, it embodies a database manager module in order that users can connect to generic Neo4j databases. It performs query processing through the Neo4j driver and provides an aesthetic query result exploration. Kai Huang 0011, Houdong Liang, Chongchong Yao, Xi Zhao 0006, Yue Cui 0001, Ruiyuan Zhang, Xiaofang Zhou 0001 |
Proc. VLDB Endow. | 7 |
| 2022 | SGW-Based Multi-task Learning in Vision Tasks
Ruiyuan Zhang, Yuyao Chen, Dianbing Xi, Yuchi Huo, Chao Wu 0001 |
ACCV (4) | 1 |
| 2022 | WDIBS: Wasserstein deterministic information bottleneck for state abstraction to balance state-compression and performance
Xianchao Zhu, Tianyi Huang, Ruiyuan Zhang, William Zhu 0001 |
Appl. Intell. | 3 |
| 2022 | MDMD options discovery for accelerating exploration in sparse-reward domains
Xianchao Zhu, Ruiyuan Zhang, William Zhu 0001 |
Knowl. Based Syst. | 2 |
| 2020 | Data-Driven Photovoltaic Generation Forecasting Based on a Bayesian Network With Spatial-Temporal Correlation AnalysisabstractSpatiotemporal analysis has been recognized as one of the most promising techniques to improve the accuracy of photovoltaic (PV) generation forecasts. In recent years, PV generation data of a number of PV systems distributed in a geographical locale have become increasingly available. This paper conducts a thorough investigation of the spatial-temporal correlation amongst PV generation data of distributed PV systems. PV generation data of different PV systems located at different sites may exhibit similar time varying patterns. To quantify such spatial correlation, a suitable spatial similarity metric is chosen and its applicability is examined. To evaluate the temporal correlations amongst the PV generation data collected from distributed PV systems, a shape-based distance metric is proposed. A data-driven inference model, built on a Bayesian network, is developed for a very short-term PV generation forecast (less than 30 min). The model utilizes historic PV generation and weather data, and incorporates the abovementioned spatial similarity and temporal correlation to support the PV output forecast. The experiment results show that the proposed method achieves a promising performance compared to a number of baseline methods. Ruiyuan Zhang, Hui Ma 0007, Wen Hua, Tapan Kumar Saha, Xiaofang Zhou 0001 |
IEEE Trans. Ind. Informatics | 1 |