Yuxiang Wang 0014

dblp:62/1637-14 · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
7since 2021 · last 2026
0009-0006-2943-500XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Federated Retrieval Over Embedding-Heterogeneous Vector Databases
Yuxiang Wang 0014, Yongxin Tong, Zimu Zhou, Ziyuan He, Ruixi Hu
ICDE1
2026 FedMosaic: Federated Retrieval-Augmented Generation via Parametric Adapters
abstract
Retrieval-Augmented Generation (RAG) enhances Large Language Models (LLMs) by grounding generation in external knowledge to improve factuality and reduce hallucinations. Yet most deployments assume a centralized corpus, which is infeasible in privacy-aware domains where knowledge remains siloed. This motivates federated RAG (FedRAG), where a central LLM server collaborates with distributed silos without sharing raw documents. In-context RAG violates this requirement by transmitting verbatim documents, whereas parametric RAG encodes documents into light weight adapters that merge with a frozen LLM at inference, avoiding raw-text exchange. We adopt the parametric approach but face two unique challenges induced by FedRAG: high storage and communication from per-document adapters, and destructive aggregation caused by indiscriminately merging multiple adapters. We present FedMosaic, the first federated RAG framework built on parametric adapters. FedMosaic clusters semantically related documents into multi-document adapters with document-specific masks to reduce overhead while preserving specificity, and performs selective adapter aggregation to combine only relevance-aligned, non-conflicting adapters. Experiments show that FedMosaic achieves an average 10.9% higher accuracy than state-of-the-art methods in four categories, while lowering storage costs by 78.8% to 86.3% and communication costs by 91.4%, and never sharing raw documents.
Zhilin Liang, Yuxiang Wang 0014, Zimu Zhou, Hainan Zhang 0001, Boyi Liu 0002, Yongxin Tong
SIGIR2
2026 Efficient shapley-based data valuation for federated trajectories
Yuxiang Wang 0014, Shuyuan Li, Yongxin Tong, Shuyue Wei 0001, Zimu Zhou
Frontiers Comput. Sci.1
2025 Timestamp Approximate Nearest Neighbor Search Over High-Dimensional Vector Data
abstract
Unstructured data, such as images and texts, are increasingly represented as high-dimensional vectors for emerging AI applications like retrieval-augmented generation. A key operation in these applications is querying for vectors that are both semantically similar and temporally relevant. This operation can be formulated as Timestamp Approximate Nearest Neighbor Search (TANNS), where both the vectors and the query incorporate temporal attributes, aiming to retrieve the approximate nearest neighbors valid at the given timestamp. A naive solution is to create separate indexes for each timestamp, which enables accurate and fast searches but incurs high update latency and excessive storage demands. In this paper, we introduce the timestamp graph, a novel structure that supports rapid index updates while minimizing storage costs. Exploiting the temporal locality of changes in valid vectors, our timestamp graph effectively manages a unified index across all historical timestamps, thereby substantially reducing storage overhead. Moreover, we design the historic neighbor tree, which further compresses the space complexity to that of a single-timestamp index. Extensive evaluations on four standard datasets show that our method achieves over 99% accuracy while improving the query efficiency by 4.4× to 138.1× than existing solutions.
Yuxiang Wang 0014, Ziyuan He, Yongxin Tong, Zimu Zhou, Yiman Zhong
ICDE1
2025 FedMetro: Efficient Metro Passenger Flow Prediction via Federated Graph Learning
abstract
Metro passenger flow prediction is crucial for effective urban transportation management. However, its practical adoption is hindered by data silos from distributed automatic fare collection (AFC) systems, compromising prediction accuracy. While federated graph learning facilitates privacy-preserving collaboration, existing methods struggle with the unique challenges of cross-line metro passenger flow prediction, particularly in handling time-evolving spatial correlations and heterogeneous temporal correlations. To address these challenges, we present FedMetro, a novel metro passenger flow prediction system based on federated graph learning. We introduce a federated dynamic graph learning approach with cross-attention mechanisms to capture spatial-temporal correlations in passenger flow. Additionally, we propose a dynamic mask-based communication compression method to mitigate communication bottlenecks in federated inference. Extensive evaluations on three real-world metro AFC datasets demonstrate that FedMetro significantly outperforms baseline methods, achieving up to 17.08% higher accuracy while reducing federated inference communication overhead by 77.99%. Practical deployments further confirm its effectiveness in delivering accurate station-level predictions across metro lines. Our code is available at https://github.com/AlexMufeng/FedMetro.
Tianlong Zhang, Xiaoxi He, Yuxiang Wang 0014, Yi Xu 0013, Rendi Wu, Yongxin Tong
KDD (2)3
2024 An Experimental Study on Federated Equi-Joins
abstract
Data federation has emerged as a novel database system enabling collaborative queries across mutually distrusted data owners. Federated equi-join, a commonly used operation in data federation, combines relations from distinct data owners while preserving their data privacy. Due to the wide applications of this query, many solutions to federated equi-joins have been proposed. However, it is still challenging for practitioners to choose the most appropriate algorithm due to various reasons, including incomplete evaluation protocols (e.g., lack of evaluating multi-way equi-joins), under-explored performance metric (main memory usage), and absence of a standardized comparison. Motivated by this reason, this paper conducts a comprehensive experimental study and builds a new benchmark, called${\sf FEJ-Bench}$, for federated equi-joins. The experimental study and the benchmark consist of eight state-of-the-art algorithms and five datasets. Our evaluation reveals the query efficiency ranking, its impact factors, and potential research opportunities. Finally, we open-source${\sf FEJ-Bench}$on GitHub, which is the first benchmark for federated equi-joins. Our findings aim to guide researchers and practitioners in deploying federated equi-joins in practice.
Shuyuan Li, Yuxiang Zeng, Yuxiang Wang 0014, Yiman Zhong, Zimu Zhou, Yongxin Tong
IEEE Trans. Knowl. Data Eng.3
2024 Efficient and Private Federated Trajectory Matching
abstract
Federated Trajectory Matching (FTM) is gaining increasing importance in big trajectory data analytics, supporting diverse applications such as public health, law enforcement, and emergency response. FTM retrieves trajectories that match with a query trajectory from a large-scale trajectory database, while safeguarding the privacy of trajectories in both the query and the database. A naive solution to FTM is to process the query through Secure Multi-party Computation (SMC) across the entire database, which is inherently secure yet inevitably slow due to the massive secure operations. A promising acceleration strategy is to filter irrelevant trajectories from the database based on the query, thus reducing the SMC operations. However, a key challenge is how to publish the query in a way that both preserves privacy and enables efficient trajectory filtering. In this paper, we design${\sf GIST}$, a novel framework for efficient Federated Trajectory Matching.${\sf GIST}$is grounded in Geo-Indistinguishability, a privacy criterion dedicated to locations. It employs a new privacy mechanism for the query that facilitates efficient trajectory filtering. We theoretically prove the privacy guarantee of the mechanism and the accuracy of the filtering strategy of${\sf GIST}$. Extensive evaluations on five real datasets show that${\sf GIST}$is significantly faster and incurs up to 2 orders of magnitude lower communication cost than the state-of-the-arts.
Yuxiang Wang 0014, Yuxiang Zeng, Shuyuan Li, Yuanyuan Zhang 0013, Zimu Zhou, Yongxin Tong
IEEE Trans. Knowl. Data Eng.1