Xizhe Yin

dblp:167/5787 · DBLP profile ↗
← Back
14ranked-venue papers
7as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 3 first-author · 7 since 2021Software engineering, systems software and programming languages · 4 · 2 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 4 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 GAAF: Fast and Scalable Graph-based Vector Similarity Search with Any-Match Label Filtering
abstract
In many practical scenarios, vector retrieval is frequently coupled with keyword constraints, particularly under Any-Match semantics. Filtered Approximate Nearest Neighbor Search (Filtered ANNS) has emerged as a widely adopted solution. Within this domain, state-of-the-art methods often utilize graph-based indices that enforce constraints via runtime filtering on a monolithic graph. However, real-world label skew degrades this monolithic design: frequent labels waste computation on largely valid neighborhoods, while rare labels suffer from graph sparsity in locating limited candidates. To address this, we propose GAAF, a frequency-aware Graph Ensemble framework that decouples the handling of high- and low-frequency labels. GAAF partitions the dataset into specialized graphs: utilizing dedicated indexes for high-frequency labels to eliminate redundant comparisons, while consolidating the rest of the labels into shared graphs to restore connectivity. Leveraging the fine-grained control afforded by this ensemble, we introduce NUMA-aware data placement to minimize remote access, and Adaptive Inter-graph Pruning to bypass redundant traversals. Experiments on diverse datasets demonstrate that GAAF significantly outperforms state-of-the-art baselines. © 2026 Copyright held by the owner/author(s).
Mengyang Ma, Xizhe Yin, Junqiao Qiu
ICS2
2026 UVVs: Identifying Unchanged Vertex Values in Evolving Graphs via Intersection-Union Analysis
Mahbod Afarin, Xizhe Yin, Zhijia Zhao 0001, Nael B. Abu-Ghazaleh, Rajiv Gupta 0001
IPDPS3
2025 PIE: Enabling Fast and Scalable Incremental Evolving Graph Analytics on Persistent Memory
abstract
Graph processing is crucial for unstructured-data-driven applications in various domains.In recent years, there has been a growing need to perform real-time analytics on largescale evolving graphs, which involves evaluating a graph query on a sequence of snapshots within a given time window.Some prior studies have explored utilizing persistent memory (PM) technologies, such as non-volatile memory, for efficient evolving graph analytics.However, the latest incremental processing designs fail to fully exploit the PM potential, suffering from severe read and write amplification during update ingestion and query evaluation.In this paper, we develop PIE, a PM-based incremental processing framework for fast and scalable evolving graph analytics.We first observe that leveraging CommonGraph, a recently proposed DRAM-based incremental approach that transforms costly deletions into additions, can significantly improve efficiency for evolving graph analytics in PM, although the direct adaptation introduces significant PM access inefficiencies.To enable PM-friendly incremental processing, PIE introduces a logical graph view abstraction that is detached from the physical storage to avoid extra PM writes, and a
Yunmo Zhang, Jiacheng Huang 0002, Xizhe Yin, Junqiao Qiu, Hong Xu 0001, Chun Jason Xue
ICS3
2025 PANNS: Enhancing Graph-based Approximate Nearest Neighbor Search through Recency-aware Construction and Parameterized Search
abstract
Approximate Nearest-Neighbor Search (ANNS) has become the standard querying method in vector databases, especially with the recent surge in large-scale, high-dimensional data driven by LLM-based applications. Recently, graph-based ANNS has shown improved throughput by constructing a graph from the dataset, with edges representing the distances between data points, and using best-first or beam search algorithms for query evaluation.
Xizhe Yin, Zhijia Zhao 0001, Rajiv Gupta 0001
PPoPP1
2025 An Empirical Study of Bugs in the rustc Compiler
abstract
Rust is gaining popularity for its well-known memory safety guarantees and high performance, distinguishing it from C/C++ and JVM-based languages. Its compiler, rustc , enforces these guarantees through specialized mechanisms such as trait solving, borrow checking, and specific optimizations. However, Rust’s unique language mechanisms introduce complexity to its compiler, resulting in bugs that are uncommon in traditional compilers. With Rust’s increasing adoption in safety-critical domains, understanding these language mechanisms and their impact on compiler bugs is essential for improving the reliability of both rustc and Rust programs. Such understanding could provide the foundation for developing more effective testing strategies tailored to rustc . Improving the quality of rustc testing is essential for enhancing compiler reliability, which in turn strengthens the safety and correctness of all Rust programs, as compiler bugs can silently propagate into every compiled program. Yet, we still lack a large-scale, detailed, and in-depth study of rustc bugs. To bridge this gap, this work presents a comprehensive and systematic study of rustc bugs, specifically those originating in semantic analysis and intermediate representation (IR) processing, which are stages that implement essential Rust language features such as ownership and lifetimes. Our analysis examines issues and fixes reported between 2022 and 2024, with a manual review of 301 valid issues. We categorize these bugs based on their causes, symptoms, affected compilation stages, and test case characteristics. Additionally, we evaluate existing rustc testing tools to assess their effectiveness and limitations. Our key findings include: (1) rustc bugs primarily arise from Rust’s type system and lifetime model, with frequent errors in the High-Level Intermediate Representation (HIR) and Mid-Level Intermediate Representation (MIR) modules due to complex checkers and optimizations; (2) bug-revealing test cases often involve unstable features, advanced trait usages, lifetime annotations, standard APIs, and specific optimization levels; (3) while both valid and invalid programs can trigger bugs, existing testing tools struggle to detect non-crash errors, underscoring the need for further advancements in rustc testing.
Yang Feng 0003, Yunbo Ni, Shaohua Li 0002, Xizhe Yin, Qingkai Shi, Baowen Xu, Zhendong Su 0001
Proc. ACM Program. Lang.5
2024 IncBoost: Scaling Incremental Graph Processing for Edge Deletions and Weight Updates
abstract
Incremental query evaluation is key to efficiently processing rapidly changing graph data. By focusing on the parts of the query results affected by updates, it avoids unnecessary computations, allowing for faster query evaluation. While this technique works well in the cases of edge insertions, its benefit quickly diminishes when the volumes of edge deletions and edge weight updates increases.
Xizhe Yin, Zhijia Zhao 0001, Rajiv Gupta 0001
SoCC1
2024 FRIES: Fuzzing Rust Library Interactions via Efficient Ecosystem-Guided Target Generation
abstract
Rust has been extensively used in software development in the past decades due to its memory safety mechanisms and gradually matured ecosystems. Enhancing the quality of Rust libraries is critical to Rust ecosystems as the libraries are often the core component of software systems. Nevertheless, we observe that existing approaches fall short in testing Rust API interactions - they either lack a Rust ownership-compliant API testing method, fail to handle the large search space of function dependencies, or are limited by pre-selected codebases, resulting in inefficiencies in finding errors. To address these issues, we propose a fuzzing technique, namely FRIES, that efficiently synthesizes and tests complex API interactions to identify defects in Rust libraries, and therefore promises to significantly improve the quality of Rust libraries. Behind our approach, a key technique is to traverse a weighted API dependency graph, which encodes not only syntactic dependency between functions but also the common usage patterns mined from the Rust ecosystem that reflect the programmer’s thinking. Combined with our efficient generation algorithm, such a graph structure significantly reduces the search space and lets us focus on finding hidden bugs in common application scenarios. Meanwhile, an ownership assurance algorithm is specially designed to ensure the validity of the generated Rust programs, notably improving the success rate of compiling fuzz targets. Experimental results demonstrate that this technique can indeed generate high-quality fuzz targets with minimal computational resources, while more efficiently discovering errors that have a greater impact on actual development, thereby mitigating the impact on the robustness of programs in the Rust ecosystem. So far, FRIES has identified 130 bugs, including 84 previously unknown bugs, in 20 well-known latest versions of Rust libraries, of which 54 have been confirmed.
Xizhe Yin, Yang Feng 0003, Qingkai Shi, Hongwang Liu, Baowen Xu
ISSTA1
2023 Glign: Taming Misaligned Graph Traversals in Concurrent Graph Processing
abstract
In concurrent graph processing, different queries are evaluated on the same graph simultaneously, sharing the graph accesses via the memory hierarchy. However, different queries may traverse the graph differently, especially for those starting from different source vertices. When these graph traversals are ”misaligned”, the benefits of graph access sharing can be seriously compromised. As more concurrent queries are added to the evaluation batch, the issue tends to become even worse.
Xizhe Yin, Zhijia Zhao 0001, Rajiv Gupta 0001
ASPLOS (1)1
2021 Tripoline: generalized incremental graph processing via graph triangle inequality
abstract
For compute-intensive iterative queries over a streaming graph, it is critical to evaluate the queries continuously and incrementally for best efficiency. However, the existing incremental graph processing requires a priori knowledge of the query (e.g., the source vertex of a vertex-specific query); otherwise, it has to fall back to the expensive full evaluation that starts from scratch.
Xiaolin Jiang 0002, Chengshuo Xu, Xizhe Yin, Zhijia Zhao 0001, Rajiv Gupta 0001
EuroSys3
2020 TLC: temporal logic of distributed components
abstract
Distributed systems are critical to reliable and scalable computing; however, they are complicated in nature and prone to bugs. To manage this complexity, network middleware has been traditionally built in layered stacks of components.We present a novel approach to compositional verification of distributed stacks to verify each component based on only the specification of lower components. We present TLC (Temporal Logic of Components), a novel temporal program logic that offers intuitive inference rules for verification of both safety and liveness properties of functional implementations of distributed components. To support compositional reasoning, we define a novel transformation on the assertion language that lowers the specification of a component to be used as a subcomponent. We prove the soundness of TLC and the lowering transformation with respect to a novel operational semantics for stacks of composed components in partially synchronous networks. We successfully apply TLC to compose and verify a stack of fundamental distributed components.
Jeremiah Griffin, Mohsen Lesani, Narges Shadab, Xizhe Yin
Proc. ACM Program. Lang.4
2017 A novel WiFi-based indoor localization system
abstract
This paper proposes a novel Wi-Fi based indoor localization system. Specially designed Wi-Fi beacons are set up to detect the real time signal strengths of Wi-Fi access points and send this data to a server. By establishing the distances along flat planes between beacons and a mobile tag, the location of the mobile tag is estimated by finding the most likely intersection between the planes corresponding to the tag. This proposed approach eliminates several bottlenecks that affect time and cost efficiency, as the localization technique developed in this project uses the relative signal strengths of the smartphone compared to signal strengths recorded at beacon devices. This eliminates the requirements of having control of and location knowledge of Wi-Fi access points. Preliminary experimentation results show that the proposed approach can achieve a localization accuracy comparable to a fingerprint approach.
Gary Shen, Xizhe Yin, Xianbin Wang 0001, Carl Shen
CSCWD2
2016 Incremental clustering for human activity detection based on phone sensor data
abstract
This paper presents our recent work on human activity detection based on smart phone sensors and incremental clustering algorithms. The proposed unsupervised (clustering) activity detection scheme works in an incremental manner, which contains two stages. In the first stage, streamed sensor data will be processed. A single-pass clustering algorithm is used in order to generate pre-clustered results for the next stage. In the second stage, pre-clustered results will be refined to form the final clusters, which means the clusters are built incrementally adding one cluster at a time. Experiments on phone sensors data of five basic human activities show that the proposed scheme could get comparable results with traditional clustering algorithms but working in a streaming and incremental manner, which is promising for automatic annotated data collection.
Xizhe Yin, Weiming Shen 0001, Xianbin Wang 0001
CSCWD1
2016 Mitigating sensor differences for phone-based human activity recognition
abstract
This paper presents our recent work on the analyses of smart phone sensor data collected for the human activity recognition (HAR), with the objective to develop more accurate activity recognition systems independent of smart phone models. We identify the multi-device scenario and present the impairments of different smartphone embedded sensor models on HAR applications. Outlier removal, interpolation, and filters in the preprocessing stage are proposed as mitigating techniques. Based on datasets collected from four distinct smartphones, the proposed mitigating methods show positive effects on 10-fold cross validation, device-to-device validation, and leave-one-out validation. Improved performance for smartphone based human activity recognition is observed.
Xizhe Yin, Gary Shen, Xianbin Wang 0001, Weiming Shen 0001
SMC1
2015 Human activity detection based on multiple smart phone sensors and machine learning algorithms
abstract
This paper presents our recent work on human activity detection based on smart phone embedded sensors and learning algorithms. The proposed human activity detection system recognizes human activities including walking, running, and sitting. While walking and running can be recorded as daily fitness activities, falling will also be detected as anomalous situations and alerting messages can be sent as needed. Embedded sensors including a tri-axial accelerometer, tri-axial linear accelerometer, gyroscope sensor, and orientation sensors are used for motion data collection. A two-stage data analysis approach is used for prediction model generation: short period statistical analysis (max, min, mean, and standard deviation) and long period data analysis using machine learning. The system is implemented in an Android smart phone platform.
Xizhe Yin, Weiming Shen 0001, Jagath Samarabandu, Xianbin Wang 0001
CSCWD1