VLDB 2026 Research / reviewers in the wild / expert
Zimu Yuan
dblp:91/10236
· DBLP profile ↗
16ranked-venue papers
6as first author
3since 2021 · last 2022
0000-0002-9494-7478ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 8 · 6 first-authorSoftware engineering, systems software and programming languages · 4 · 2 since 2021Security and privacy · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | A Sanitizer-centric Analysis to Detect Cross-Site Scripting in PHP ProgramsabstractA large number of PHP applications suffer from Cross-Site Scripting (XSS) attacks every year. Static taint analysis is a prevalent way to detect taint-style vulnerabilities like XSS. However, the precision of current tools suffers severely due to dynamic features of PHP programs and the incomplete recognition of user-defined sanitizers, which lead to false negatives and a large number of false positives. In this paper, we present PAT, a PHP static Analysis Tool for effective XSS vulnerability detection. A new concept of “inner” source and sink is introduced for the first time to shorten the taint paths needed to be traced statically, which therefore mitigates the broken path problem induced by dynamic language features to a certain extent. A sanitizer-centric approach is proposed to automatically identify them. Moreover, PAT leverages both data flow analysis and NLP technique to accurately identify user-defined sanitizers with a precision 82.8%. Lastly, PAT performs a classical taint analysis with the enhanced taint specifications (i.e., sources, sinks and sanitizers). Evaluations on 5 large, real-world PHP web applications and 5 popular WordPress plugins show that PAT performs better in XSS detection compared with 3 existing tools. Besides, 8 zero-day bugs are detected and confirmed by the developers. He Su, Huina Chao, Feng Li 0045, Zimu Yuan, Wei Huo 0005 |
ISSRE | 5 |
| 2021 | VIVA: Binary Level Vulnerability Identification via Partial SignatureabstractBinary level code clone detection techniques have been used to identify 1-day vulnerabilities in software. It collects functions with known vulnerabilities and searches for similar functions in the target system. However, existing approaches are limited to detect the same vulnerabilities in different binaries. They can hardly find new recurring vulnerabilities, which share similar logic. Moreover, they only focus on improving the accuracy of binary function matching algorithms while overlooking the presence of security patches, which results in high false-positive rates and requires significant effort to verify the results.To this end, we propose VIVA, a binary level vulnerability and patch semantic summarization and matching tool for accurate recurring vulnerability detection. It uses novel binary program slicing techniques with the aid of pseudo-code trace refinement to generate partial vulnerability and patch signatures, which capture the semantics. It matches the signatures with pre-filtering to efficiently detect 1-day and recurring vulnerabilities. The experimental results show that VIVA outperforms other source code and binary matching tools with a precision of 100% for 1-day vulnerabilities and 87.6% for recurring vulnerabilities and good performance (28.58s per signature search in 4M functions). It detects 92 new vulnerabilities in different series and different versions of real-world projects, with 11 exist without fixing in the latest version. Yang Xiao 0011, Zhengzi Xu, Chendong Yu, Longquan Liu, Zimu Yuan, Yang Liu 0003, Aihua Piao, Wei Huo 0005 |
SANER | 7 |
| 2021 | B2SMatcher: fine-Grained version identification of open-Source software in binary filesabstractAbstract Codes of Open Source Software (OSS) are widely reused during software development nowadays. However, reusing some specific versions of OSS introduces 1-day vulnerabilities of which details are publicly available, which may be exploited and lead to serious security issues. Existing state-of-the-art OSS reuse detection work can not identify the specific versions of reused OSS well. The features they selected are not distinguishable enough for version detection and the matching scores are only based on similarity.This paper presents B2SMatcher, a fine-grained version identification tool for OSS in commercial off-the-shelf (COTS) software. We first discuss five kinds of version-sensitive code features that are trackable in both binary and source code. We categorize these features into program-level features and function-level features and propose a two-stage version identification approach based on the two levels of code features. B2SMatcher also identifies different types of OSS version reuse based on matching scores and matched feature instances. In order to extract source code features as accurately as possible, B2SMatcher innovatively uses machine learning methods to obtain the source files involved in the compilation and uses function abstraction and normalization methods to eliminate the comparison costs on redundant functions across versions. We have evaluated B2SMatcher using 6351 candidate OSS versions and 585 binaries. The result shows that B2SMatcher achieves a high precision up to 89.2% and outperforms state-of-the-art tools. Finally, we show how B2SMatcher can be used to evaluate real-world software and find some security risks in practice. Gu Ban, Yang Xiao 0011, Xinhua Li, Zimu Yuan, Wei Huo 0005 |
Cybersecur. | 5 |
| 2020 | MVP: Detecting Vulnerabilities using Patch-Enhanced Vulnerability Signatures
Yang Xiao 0011, Bihuan Chen 0001, Chendong Yu, Zhengzi Xu, Zimu Yuan, Feng Li 0045, Binghong Liu, Yang Liu 0003, Wei Huo 0005, Wenchang Shi |
USENIX Security Symposium | 5 |
| 2019 | Event Detection on Unreliable Distributed Storage NetworkabstractEvent uncertainty problem happens when a distributed storage system is deployed under unreliable network environment. To address the uncertainty problem, this paper proposes an event detection scheme based on indeterminate subevent-to-event belonging relationship. For a target event, the scheme decomposes its subevent correlation into a group of total order relations, and matches them with the subevent stream to further determine its possible occurrence time range. With the occurrence range, this paper calculates event occurrence confidence based on transition probability in match components. Zimu Yuan, Wei Li 0008 |
GLOBECOM | 1 |
| 2019 | B2SFinder: Detecting Open-Source Software Reuse in COTS SoftwareabstractCOTS software products are developed extensively on top of OSS projects, resulting in OSS reuse vulnerabilities. To detect such vulnerabilities, finding OSS reuses in COTS software has become imperative. While scalable to tens of thousands of OSS projects, existing binary-to-source matching approaches are severely imprecise in analyzing COTS software products, since they support only a limited number of code features, compute matching scores only approximately in measuring OSS reuses, and neglect the code structures in OSS projects. We introduce a novel binary-to-source matching approach, called B2SFINDER1, to address these limitations. First of all, B2SFINDER can reason about seven kinds of code features that are traceable in both binary and source code. In order to compute matching scores precisely, B2SFINDER employs a weighted feature matching algorithm that combines three matching methods (for dealing with different code features) with two importance-weighting methods (for computing the weight of an instance of a code feature in a given COTS software application based on its specificity and occurrence frequency). Finally, B2SFINDER identifies different types of code reuses based on matching scores and code structures of OSS projects. We have implemented B2SFINDER using an optimized data structure. We have evaluated B2SFINDER using 21991 binaries from 1000 popular COTS software products and 2189 candidate OSS projects. Our experimental results show that B2SFINDER is not only precise but also scalable. Compared with the state ofthe art, B2SFINDER has successfully found up to 2.15× as many reuse cases in 53.85 seconds per binary file on average. We also discuss how B2SFINDER can be leveraged in detecting OSS reuse vulnerabilities in practice. Muyue Feng, Zimu Yuan, Feng Li 0045, Gu Ban, He Su, Chendong Yu, Jiahuan Xu, Aihua Piao, Jingling Xue, Wei Huo 0005 |
ASE | 2 |
| 2019 | Open-Source License Violations of Binary Software at Large ScaleabstractOpen-source licenses are widely used in open-source projects. However, developers using or modifying the source code of open-source projects do not always strictly follow the licenses. GPL and AGPL, two of the most popular copyleft licenses, are most likely to be violated, because they require developers to open-source the entire project if any code under GPL/AGPL protection is included whether modified or not. There are few license violation detectors focusing on binary software, owning to the challenge of mapping binary code to source code efficiently and accurately at large scale. In this paper, we propose a scalable and fully-automated system to check open-source license violation of binary software at large scale. We match source code to binary code by analyzing file attributes of executable files and code features that are not affected by compilation and could vary between projects. Moreover, to break the barrier of large-scale analysis, we introduce an automatic extractor to parse executable files from installation packages that are broadly available in software download sites. In empirical experiments of binary-to-source mapping, we have got a remarkable high accuracy of 99.5% and recall of 95.6% without significant loss of precision. Besides, 2270 pairs of binary-to-source mapping relationships are discovered, with 110 license violations of GPL and AGPL licenses related to 7.2% of the 1000 real-world binary software projects. Muyue Feng, Weixuan Mao, Zimu Yuan, Yang Xiao 0011, Gu Ban, Jiahuan Xu, He Su, Binghong Liu, Wei Huo 0005 |
SANER | 3 |
| 2018 | Error Analysis on RSS Range-Based Localization Based on General Log-Distance Path Loss ModelabstractReceived Signal Strength (RSS) is considered to be a promising measurement for indoor positioning. Many RSS range-based localization methods have been proposed due to the convenience and low cost of RSS measurements. However, a fundamental problem has not been answered, that is, how accurate are these methods? We think a key reason leading to this situation is the inappropriate assumption on RSS range models and measurement errors, which results in oversimplified analysis on those methods. In this paper, we use a more general range model and recognize the Generalized Least Square (GLS) method as an optimal estimator whose estimation error equals to the Cramer-Rao lower bound (CRLB). Through mathematical, techniques, we derive the analytic expression of the localization error for the GLS method, which reveals the key factors that affect the localization accuracy. Further studies on the minimal localization error disclose the proportional relationship between the localization accuracy and the above key factors. Wei Li 0008, Zimu Yuan, Wei Zhao 0001 |
MASS | 2 |
| 2015 | An Approach for Mitigating Potential Threats in Practical SSO Systems
Liang Yang 0002, Zimu Yuan, Rui Zhang 0016, Rui Xue 0001 |
Inscrypt | 3 |
| 2015 | CIUV: Collaborating information against unreliable viewsabstractIn many real world applications, the information of an object can be obtained from multiple sources. The sources may provide different point of views based on their own origin. As a consequence, conflicting pieces of information are inevitable, which gives rise to a crucial problem: how to find the truth from these conflicts. Many truth-finding methods have been proposed to resolve conflicts based on information trustworthy (i.e. more appearance means more trustworthy) as well as source reliability. However, the factor of men's involvement, i.e., information may be falsified by men with malicious intension, is more or less ignored in existing methods. Collaborating the possible relationship between information's origins and men's participation are still not studied in research. To deal with this challenge, we propose a method - Collaborating Information against Unreliable Views (CIUV) - in dealing with men's involvement for finding the truth. CIUV contains 3 stages for interactively mitigating the impact of unreliable views, and calculate the truth by weighting possible biases between sources. We theoretically analyze the error bound of CIUV, and conduct intensive experiments on real dataset for evaluation. The experimental results show that CIUV is feasible and has the smallest error compared with other methods. Zimu Yuan, Zhiwei Xu 0002, Guojie Li |
ISCC | 1 |
| 2015 | A New ETL Approach Based on Data Virtualization
Shusheng Guo, Zimu Yuan, Aobing Sun, Qiang Yue 0001 |
J. Comput. Sci. Technol. | 2 |
| 2013 | A two-tier positioning algorithm for wireless networks with diverse measurement typesabstractA variety of measurement methods have been developed to obtain positions of nodes in wireless networks. However, most of existing positioning methods can only use limited types of measurement. These algorithms fail to exploit arbitrary measurement methods to improve the accuracy of positioning. In this paper, we propose a high-accuracy positioning algorithm which can utilize unlimited types of measurement. Our algorithm first selects a set of nodes equipped with more accurate measurement mechanisms to locate nodes whose positions are unknown. By this selection, a system of heterogenous equations is built. Then, our algorithm uses the classical genetic method to solve equations which commonly are hard to solve. The comprehensive simulation shows that our algorithm significantly improves the accuracy of positioning compared with existing methods. Zimu Yuan, Wei Li 0008 |
GLOBECOM | 1 |
| 2012 | An efficient hybrid localization scheme for Heterogeneous Wireless NetworksabstractThe ability to track and locate physical entities is a fundamental requirement for Cyber-Physical Systems (CPSs), especially in an ad-hoc wireless environment. In Heterogeneous Wireless Networks (HWNs), hybrid localization schemes are needed due to the coexistence of both accurate and coarse measurement mechanisms. However, current localization schemes cannot fully satisfy HWNs' accuracy requirements. Therefore, we propose a universal measurement metric called Direct Proportion Distance (DPD) that can leverage most existing measurement mechanisms such as TOA/TDOA, RSS, AOA, Link Diagnosis (LD) and Signal Coverage Detection (SCD). We also prove that DPD is directly proportional to the physical distance between two wireless nodes. Based on this metric, we present three new localization algorithms and compare them with classical methods. The experiments verify that our method performs better than previous localization algorithms when both accurate and coarse measurements are fully utilized. Zimu Yuan, Wei Li 0008, Adam C. Champion, Wei Zhao 0001 |
GLOBECOM | 1 |
| 2012 | A new routing scheme based on adaptive selection of geographic directionsabstractGeographic routing is recognized as an appealing approach to achieve efficient communications with low computational complexity and space cost. In order to apply this technology in Cyber-Physical Systems (CPSs), a comprehensive consideration must be given to performance issues such as throughput, delay, and load balance. In this paper, we provide a new routing scheme based on forwarding packets to multiple geographic directions. The proposed routing protocols are studied and analyzed theoretically. Theoretical bounds of throughput, delays and space cost are presented. Simulations show that our method performs more efficiently than traditional geographic routing schemes in terms of throughput, delay, and load balance with acceptable space cost. Our experiments also verify the tradeoff between performance metrics. Zimu Yuan, Wei Li 0008 |
GLOBECOM | 1 |
| 2011 | History-Aware Adaptive Backoff for Neighbor Discovery in Wireless NetworksabstractThe ability of discovering neighboring nodes, namely neighbor discovery, is essential for the self-organization of wireless ad hoc networks. In this paper, we propose a history-aware adaptive back off algorithm for neighbor discovery assuming collision detection and feedback mechanisms. Given successful discovery feedback, undiscovered nodes can adjust their contention window. With collision feedback and historical information, only transmission nodes enter the re-contention process, and decrease their contention window to accelerate neighbor discovery process after collision. Then, we give theoretical analysis of our algorithm on the discovery time and energy consumption, and derive the optimal size of contention windows by two rounds of optimization. Finally, we validate our theoretical analysis by simulations, and show the performance improvement over existing algorithms. Zimu Yuan, Lizhao You, Wei Li 0008, Biao Chen 0002, Zhiwei Xu 0002 |
MSN | 1 |
| 2011 | ALOHA-like neighbor discovery in low-duty-cycle wireless sensor networksabstractNeighbor discovery is an essential step for the self-organization of wireless sensor networks. Many algorithms have been proposed for efficient neighbor discovery. However, most of those algorithms need nodes to keep active during the process of neighbor discovery, which might be difficult for low-duty-cycle wireless sensor networks in many real deployments. In this paper, we investigate the problem of neighbor discovery in low-duty-cycle wireless sensor networks. We give an ALOHA-like algorithm and analyze the expected time to discover all n - 1 neighbors for each node. By reducing the analysis to the classical K Coupon Collector's Problem, we show that the upper bound is ne(log2n + (3 log2n - 1) log2log2n + c) with high probability, for some constant c, where e is the base of natural logarithm. Furthermore, not knowing number of neighbors leads to no more than a factor of two slowdown in the algorithm performance. Then, we validate our theoretical results by extensive simulations, and explore the performance of different algorithms in duty-cycle and non-duty-cycle networks. Finally, we apply our approach to analyze the scenario of unreliable links in low-duty-cycle wireless sensor networks. Lizhao You, Zimu Yuan, Panlong Yang, Guihai Chen |
WCNC | 2 |