Yifan Tian

dblp:192/9839 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
5since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 3 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorSystems, architecture and hardware · 1Computer networks · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 CShard: Blockchain Sharding via Repairable Fountain Codes and the Paradigm for Sharding Design
abstract
Sharding is an important solution to improve the scalability of blockchain. The basic idea of blockchain sharding is to separate transactions among multiple disjoint shards processing in parallel to maximize system performance. The current sharding protocols mainly rely on node rotation (namely, node allocation and migration) randomly among shards periodically to ensure security, which is often considered the most challenging when developing a sharding system. However, (1) if a shard or multiple shards are corrupted, the sharding system will not be available anymore; (2) to avoid shard corruption, the demand of large size of each shard limits the throughput performance of sharding protocols. To solve (1), we introduce a blockchain sharding protocol with scale-out transaction processing capacity called CShard. The main idea of CShard is to use repairable fountain codes (RFCs), an information coding method with the locality feature, to innovate sharding design. By adjusting encoding parameters, topological associations among shards are constructed, which are then utilized to define the verification logic for transactions. The blocks of corrupted shard(s) can be recovered through decoding by its corresponding shard group(s), and the sharding system is still secure and available. Our approach utilizes encoding techniques to build a general architecture of a sharding system, establishing a paradigm of “encoding as verification” and showcases new horizons in the field of blockchain sharding. To solve (2), we propose the ghost reporter mechanism that gives all nodes chances to verify a transaction by submitting reports in the sharding network. The mechanism brings two direct benefits. Firstly, it provides the way to detect corrupted shards and recover the blocks by RFCs; the second is to make the number of nodes in a single shard smaller, which solves the limitation of existing sharding schemes that usually require a larger number of nodes in a single shard to ensure security. In principle, this mechanism can also be applicable to the known and even unknown sharding protocols for its generality.
Yifan Tian, Butian Huang, Xiaosong Zhang 0001
IEEE Trans. Inf. Forensics Secur.2
2025 Argus: Resilience-Oriented Safety Assurance Framework for End-to-End ADSs
abstract
End-to-end autonomous driving systems (ADSs), with their strong capabilities in environmental perception and generalizable driving decisions, are attracting growing attention from both academia and industry. However, once deployed on public roads, ADSs are inevitably exposed to diverse driving hazards that may compromise safety and degrade system performance. This raises a strong demand for resilience of ADSs, particularly the capability to continuously monitor driving hazards and adaptively respond to potential safety violations, which is crucial for maintaining robust driving behaviors in complex driving scenarios.To bridge this gap, we propose a resilience-oriented runtime framework, named Argus, to mitigate the driving hazards, thus preventing potential safety violations and improving the driving performance of an ADS. Argus continuously monitors the trajectories generated by the ADS for potential hazards and, whenever the EGO vehicle is deemed unsafe, seamlessly takes control via a hazard mitigator. We integrate Argus with three state-of-the-art end-to-end ADSs, i.e., TCP, UniAD and VAD. Our evaluation has demonstrated that Argus effectively and efficiently enhances the resilience of ADSs, improving the driving score of ADSs by 150.30% on average, and preventing 64.38% of the violations, with little additional time overhead.
Dingji Wang, You Lu 0005, Bihuan Chen 0001, Shuo Hao, Haowen Jiang, Yifan Tian, Xin Peng 0001
ASE6
2024 DiaVio: LLM-Empowered Diagnosis of Safety Violations in ADS Simulation Testing
abstract
Simulation testing has been widely adopted by leading companies to ensure the safety of autonomous driving systems (ADSs). Anumber of scenario-based testing approaches have been developed to generate diverse driving scenarios for simulation testing, and demonstrated to be capable of finding safety violations. However, there is no automated way to diagnose whether these violations are caused by the ADS under test and which category these violations belong to. As a result, great effort is required to manually diagnose violations. To bridge this gap, we propose DiaVio to automatically diagnose safety violations in simulation testing by leveraging large language models (LLMs). It is built on top of a new domain specific language (DSL) of crash to align real-world accident reports described in natural language and violation scenarios in simulation testing. DiaVio fine-tunes a base LLM with real-world accident reports to learn diagnosis capability, and uses the fine-tuned LLM to diagnose violation scenarios in simulation testing. Our evaluation has demonstrated the effectiveness and efficiency of DiaVio in violation diagnosis.
You Lu 0005, Yifan Tian, Yuyang Bi, Bihuan Chen 0001, Xin Peng 0001
ISSTA2
2024 CRSP: Emulating Human Cooperative Reasoning for Intelligible Story Point Estimation
abstract
Software effort estimation plays a critical role in software project development. Inaccurate cost estimation can impact progress and result in budget overruns. The story point estimation technique is a commonly used practice in agile software development for the estimation of software development effort. It allows for the evaluation of relative task workloads by analyzing task titles and descriptions. In previous studies, researchers have mainly focused on providing story point estimation results by task titles. However, in practical scenarios, users are often unable to provide task titles as precise as those found in the training dataset, leading to inaccurate estimation results. To address this problem, we propose a Cooperative Reasoning Story Point estimation method (CRSP). We approach the estimation problem as a question-and-answer challenge, addressing it through a framework of model construction, Monte Carlo Tree search, and model inference. In the model construction phase, we train a generator responsible for generating problem-solving reasoning paths and employ verifier to score the quality of these reasoning paths. During the Monte Carlo Tree search stage, we execute MCTS using generator and verifier to generate candidate solutions. In the final model inference phase, we employ a solver to derive the ultimate answer. To evaluate the effectiveness of CRSP, we modify and adapt the well-known JIRA dataset to make it more compatible with the input format of the question-answering model. The new JIRA dataset contains 21,082 issues from 16 open-source software projects. Across 16 open-source projects, the mean absolute error of CRSP is lower than other baseline methods. In contrast to the traditional regression and classification methods, we pioneer the use of question-and-answer method to address the issue, opening up new directions for future research.
Wanjiang Han, Zhuoyan Han, Yifan Tian, Longzheng Chen, Ren Han
ICPC4
2021 Low-Latency Privacy-Preserving Outsourcing of Deep Neural Network Inference
abstract
Efficiently supporting inference tasks of deep neural network (DNN) on the resource-constrained Internet-of-Things (IoT) devices has been an outstanding challenge for emerging smart systems. To mitigate the burden on IoT devices, one prevalent solution is to outsource DNN inference tasks to the public cloud. However, this type of “cloud-backed” solutions can cause privacy breach since the outsourced data may contain sensitive information. For privacy protection, the research community has resorted to advanced cryptographic primitives to support DNN inference over encrypted data. Nevertheless, these attempts are limited by the real-time performance due to the heavy IoT computational overhead brought by cryptographic primitives. In this article, we proposed an edge computing-assisted framework to boost the efficiency of DNN inference tasks on IoT devices, which also protects the privacy of IoT data to be outsourced. In our framework, the most time-consuming DNN layers are outsourced to edge computing devices. The IoT device only processes compute-efficient layers and fast encryption/decryption. Thorough security analysis and numerical analysis are carried out to show the security and efficiency of the proposed framework. Our analysis results indicate a 99%+ outsourcing rate of DNN operations for IoT devices. Experiments on AlexNet show that our scheme can speed up DNN inference for 40.6× with a 96.2% energy saving for IoT devices.
Yifan Tian, Laurent Njilla, Shucheng Yu
IEEE Internet Things J.1
2020 Benchmarking HOAP for Scalable Document Data Management: A First Step
abstract
Enterprises today are becoming ever more reliant on real-time information and analytics for steering and optimizing their businesses. As a result, database system architectures with hybrid data management support - known as HTAP (Hybrid Transactional/ Analytical Processing) or HOAP (Hybrid Operational/Analytical Processing) support - are appearing and increasingly gaining traction in both the commercial and research sectors. Hybrid platforms first appeared in the relational world, and they are often linked in that world to concurrent high-end server technology trends such as columnar storage and mainmemory data management.This paper focuses on hybrid platforms, but in a very different world - in the document data management, or NoSQL, world. In this work, we report on a first effort to characterize the hybrid performance of a scalable document database system that purports to provide what one might call "HOAP for JSON". We have borrowed from and extended the TPC-C benchmark to study the performance of Couchbase Server, a horizontally scalable NoSQL platform that offers HOAP via the combination of its Data/Index/Query and Analytics Services. Our results attest to the importance of architecting a NoSQL platform for HOAP, both in terms of its approach(es) to query processing and its provision of performance isolation for the operational and analytical components of a mixed workload. We share our initial results, the insights that we have gained thus far, and our thoughts on future work related to benchmarking such systems.
Yifan Tian, Michael J. Carey 0001, Ian Maxon
IEEE BigData1
2019 Edge-Assisted CNN Inference over Encrypted Data for Internet of Things
Yifan Tian, Shucheng Yu, Yantian Hou, Houbing Song
SecureComm (1)1
2019 Efficient privacy-preserving authentication framework for edge-assisted Internet of Drones
Yifan Tian, Houbing Song
J. Inf. Secur. Appl.1
2019 Practical Privacy-Preserving MapReduce Based K-Means Clustering Over Large-Scale Dataset
abstract
Clustering techniques have been widely adopted in many real world data analysis applications, such as customer behavior analysis, targeted marketing, digital forensics, etc. With the explosion of data in today's big data era, a major trend to handle a clustering over large-scale datasets is outsourcing it to public cloud platforms. This is because cloud computing offers not only reliable services with performance guarantees, but also savings on in-house IT infrastructures. However, as datasets used for clustering may contain sensitive information, e.g., patient health information, commercial data, and behavioral data, etc, directly outsourcing them to public cloud servers inevitably raise privacy concerns. In this paper, we propose a practical privacy-preserving K-means clustering scheme that can be efficiently outsourced to cloud servers. Our scheme allows cloud servers to perform clustering directly over encrypted datasets, while achieving comparable computational complexity and accuracy compared with clusterings over unencrypted ones. We also investigate secure integration of MapReduce into our scheme, which makes our scheme extremely suitable for cloud computing environment. Thorough security analysis and numerical analysis carry out the performance of our scheme in terms of security and efficiency. Experimental evaluation over a 5 million objects dataset further validates the practical performance of our scheme.
Yifan Tian
IEEE Trans. Cloud Comput.2