Zixia Liu

dblp:167/2203 · DBLP profile ↗
← Back
14ranked-venue papers
3as first author
7since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 1 first-authorComputer networks · 2 · 2 since 2021Security and privacy · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Your Outer Appearance Mirrors Your Inner Self: Exploiting Unobservable Node Internals to Deanonymize Uploaders in Freenet
Yonghuan Xu, Ming Yang 0001, Shan Wang 0008, Xiaodan Gu, Zixia Liu, Zhen Ling 0001
INFOCOM5
2026 A Tor-Based Anonymous Network Covert Channel
abstract
Network Covert Channels (NCC) enhance covertness by concealing the existence of information transmission. However, traditional NCCs remain vulnerable to traffic analysis. Once NCC is detected, adversaries can breach anonymity by uncovering users' network identities and even communication relationships. While certain indirect NCCs offer limited anonymity to protect the identity of at most one party and the relationship, this level proves insufficient. This paper proposes ANCC, an innovative Anonymous Network Covert Channel that is the first to achieve comprehensive anonymity for the sender, the receiver and the communication relationship. By leveraging the Tor network's Hidden Service Directories (HSDirs) as intermediate nodes, Tor-based ANCC modulates covert information through the publication and retrieval statuses of hidden services distributed on multiple HSDirs. This mechanism allows ANCC traffic to blend seamlessly into legitimate Tor traffic, ensuring both robust covertness and high-level anonymity. Theoretical analysis demonstrates that even against a powerful adversary compromising fifty intermediate nodes, the detection probability remains below 0.25%, with the risk of identity or relationship exposure staying negligible (under 0.0021% and 0.00002% respectively). Additionally, the multiple HSDirs supporting parallel transmission enhance the channel capacity and error correction encoding strengthens the robustness. Extensive evaluation within the real-world Tor network demonstrates a transmission accuracy exceeding 99.6% and a channel capacity of around 3 Kbps, proving its effectiveness for practical applications.
Ming Yang 0001, Zhen Ling 0001, Zixia Liu, Changwei Cao, Shan Wang 0008, Xinwen Fu
IEEE Trans. Dependable Secur. Comput.4
2025 The Cost of Performance: Breaking ThreadX with Kernel Object Masquerading Attacks
Xinhui Shao, Zhen Ling 0001, Yue Zhang 0025, Huaiyu Yan, Yumeng Wei, Zixia Liu, Junzhou Luo, Xinwen Fu
USENIX Security Symposium7
2025 Do Not Trust What They Tell: Exposing Malicious Accomplices in Tor via Anomalous Circuit Detection
abstract
The Tor network, while offering anonymity through traffic routing across volunteer-operated nodes, remains vulnerable to attacks that aim to deanonymize users by correlating traffic patterns between colluded entry and exit nodes in circuits. This paper presents a novel approach for detecting anomalous circuits in the Tor network, and for the first time provides a more comprehensive identification of potential malicious accomplice nodes in Tor by taking roles of nodes in anomalous circuits into consideration. Our method strategically utilizes modified middle nodes to capture traffic data, followed by a novel circuit classification based on traffic patterns to pinpoint concerned circuits. Two kinds of anomalies are identified: routing anomalies and usage anomalies, that respectively represent the anomalies with explicit or implicit violation of Tor's circuit construction guidelines. This leads to a successful revealing of totally 1,960 anomalous nodes in Tor. Furthermore, we apply clustering analysis with considering corresponding anomalous circuits and other key characteristics to the detected anomalous nodes, revealing potential hidden organizations behind these nodes that can threaten the network's security. Our findings highlight the necessity for the Tor project to adopt targeted mitigation strategies to enhance overall network security and privacy.
Yixuan Yao, Ming Yang 0001, Zixia Liu, Kai Dong 0001, Xiaodan Gu, Chunmian Wang
WWW3
2024 A De-anonymization Attack against Downloaders in Freenet
abstract
Freenet is a well-known anonymous communication system that enables file sharing among users. It employs a probabilistic hops-to-live (HTL) decrement approach to hide the originator among nodes in a multi-hop path. Therefore, all nodes shall exhibit identical behaviors to preserve anonymity. However, we discover that the path folding mechanism in Freenet violates this principle due to behavior discrepancy between downloaders and intermediate nodes. The path folding mechanism is designed to optimize the network topology of Freenet. A delayed path folding message by a successor node may incur a timeout event at its predecessor, and an intermediate node reacts differently to such timeout with a downloader. Therefore, malicious nodes can deliberately trigger the timeout event to identify downloaders. The complex implementation of the path folding timeout detection mechanism in Freenet complicates our de-anonymization attack. We thoroughly analyze the underlying cause and develop three strategies to manipulate three types of messages respectively at the malicious node, minimizing the false positive rate. We conduct extensive real-world experiments to verify the feasibility and effectiveness of our attack. They show that our attack achieves a true positive rate of 100% and false positive rate of near 0% under two different Freenet download modes.
Yonghuan Xu, Ming Yang 0001, Zhen Ling 0001, Zixia Liu, Xiaodan Gu
INFOCOM4
2023 Contrastive JS: A Novel Scheme for Enhancing the Accuracy and Robustness of Deep Models
abstract
Deep learning technologies have been applied in various computer vision tasks in recent years. However, deep models suffer performance decay when some unforeseen data are contained in the testing dataset. Although data enhancement techniques can alleviate this dilemma, the diversity of real data is too tremendous to simulate. To tackle this challenge, we study a scheme for improving the robustness and efficiency of the deep network training process in visual tasks. Specifically, first, we build positive and negative sample pairs based on a class-sensitive strategy. Then, we construct a feature-consistent learning strategy based on contrastive learning to constrain the representations of interclass features while paying attention to the intraclass features. To extend the effect of the consistent strategy, we propose a novel contrastive Jensen-Shannon divergence consistency loss (JS loss) to restrict the probability distributions of different sample pairs. The proposed scheme successfully enhances the robustness and accuracy of the utilized model. We validated our approach by conducting extensive experiments in the domains of model robustness and few-shot object detection (FSOD). The results showed that the proposed method achieved remarkable gains over state-of-the-art (SOTA) methods. We obtained a 3.2% average improvement over the best-performing FSOD method.
Weiwei Xing, Zixia Liu, Weibin Liu, Shunli Zhang 0005, Liqiang Wang 0001
IEEE Trans. Multim.3
2021 SODA: A Semantics-Aware Optimization Framework for Data-Intensive Applications Using Hybrid Program Analysis
abstract
In the era of data explosion, a growing number of data-intensive computing frameworks, such as Apache Hadoop and Spark, have been proposed to handle the massive volume of unstructured data in parallel. Since programming models provided by these frameworks allow users to specify complex and diversified user-defined functions (UDFs) with predefined operations, the grand challenge of tuning up entire system performance arises if programmers do not fully understand the semantics of code, data, and runtime systems. In this paper, we design a holistic semantics-aware (optimization for data-intensive applications using hybrid program analysis (SODA) to assist programmers to tune performance issues. SODA is a two-phase framework: the offline phase is a static analysis that analyzes code and performance profiling data from the online phase of prior executions to generate a parameterized and instrumented application; the online phase is a dynamic analysis that keeps track of the application's execution and collects runtime information of data and system. Extensive experimental results on four real-world Spark applications show that SODA can gain up to 60%, 10%, 8%, faster than its original implementation, with the three proposed optimization strategies, i.e., cache management, operation reordering, and element pruning, respectively.
BingBing Rao, Zixia Liu, Hong Zhang 0047, Siyang Lu, Liqiang Wang 0001
CLOUD2
2020 Deep Reinforcement Learning based Elasticity-compatible Heterogeneous Resource Management for Time-critical Computing
abstract
Rapidly generated data and the amount magnitude of data analytical jobs pose great pressure to the underlying computing facilities. A distributed multi-cluster computing environment such as a hybrid cloud consequently raises its necessity due to its advantages in adapting geographically distributed and potentially cloud-based computing resources. Different clusters forming such an environment could be heterogeneous and may be resource-elastic as well. From analytical perspective, in accordance with increasing needs on streaming applications and timely analytical demands, many data analytical jobs nowadays are time-critical in terms of their temporal urgency. And the overall workload of the computing environment can be hybrid to contain both time-critical and general applications. These all call for an efficient resource management approach capable to apprehend both computing environment and application features.
Zixia Liu, Liqiang Wang 0001, Gang Quan
ICPP1
2018 A Reinforcement Learning Based Resource Management Approach for Time-critical Workloads in Distributed Computing Environment
abstract
Many data analyzing applications highly rely on timely response from execution, and are referred as time-critical data analyzing applications. Due to frequent appearing of gigantic amount of data and analytical computations, running them on large scale distributed computing environments is often advantageous. The workload of big data applications is often hybrid, i.e., contains a combination of time-critical and regular non-time-critical applications. Resource management for hybrid workloads in complex distributed computing environment is becoming more critical and needs more studies. However, it is difficult to design rule-based approaches best suited for such complex scenarios because many complicated characteristics need to be taken into account.Therefore, we present an innovative reinforcement learning (RL) based resource management approach for hybrid workloads in distributed computing environment. We utilize neural networks to capture desired resource management model, use reinforcement learning with designed value definition to gradually improve the model and use ε-greedy methodology to extend exploration along the reinforcement process. The extensive experiments show that our obtained resource management solution through reinforcement learning is able to greatly surpass the baseline rule-based models. Specifically, the model is good at reducing both the missing deadline occurrences for time-critical applications and lowering average job delay for all jobs in the hybrid workloads. Our reinforcement learning based approach has been demonstrated to be able to provide an efficient resource manager for desired scenarios.
Zixia Liu, Hong Zhang 0047, BingBing Rao, Liqiang Wang 0001
IEEE BigData1
2018 Tuning Performance of Spark Programs
abstract
Along with the explosive growth of data, there is a great demand to speedup the ability to process them. Although there are several platforms such as Spark that have made analysis easier to developers, the performance tuning for such platforms meanwhile becomes complex. In this paper, we propose an efficient performance optimization engine called Hedgehog to evaluate the performance based on "Law of Diminishing Marginal Utility" and give an optimal configuration setting. The initial experiments show that our optimization can gain 19.6% performance improvement compared to the naive configuration by tuning only 3 parameters.
Hong Zhang 0047, Zixia Liu, Liqiang Wang 0001
IC2E2
2018 Improving the Improved Training of Wasserstein GANs: A Consistency Term and Its Dual Effect
Wei Xiang 0007, Boqing Gong, Zixia Liu, Wei Lu 0010, Liqiang Wang 0001
ICLR (Poster)3
2017 Hierarchical Spark: A Multi-Cluster Big Data Computing Framework
abstract
Nowadays, with the increasing burst of newly generated data everyday, as well as the vast expanding needs for corresponding data analyses, grand challenges have been brought to big data computing platforms. Computing resources in a single cluster are often not able to fulfill the computing capability needs. The requests of distributed computing resources are dramatically arising. In addition, with increasing popularity of cloud computing platforms, many organizations with data security concerns are more favor to hybrid cloud, a multi-cluster environment composed by both public cloud and private cloud in purpose of keeping sensitive data local. All these scenarios show great necessity of migrating big data computing to multi-cluster environment. In this paper, we present a hierarchical multi-cluster big data computing framework built upon Apache Spark. Our framework supports combination of heterogeneous Spark computing clusters. With an integrated controller within the framework, it also facilitates ability for submitting, monitoring, executing of Spark workflow. Our experimental results show that the proposed framework not only enables possibility of distributing Spark workflow throughout multiple clusters, but also provides significant performance improvement compared to single cluster environment by optimizing utilization of multi-cluster computing resources.
Zixia Liu, Hong Zhang 0047, Liqiang Wang 0001
CLOUD1
2016 Migrating GIS Big Data Computing from Hadoop to Spark: An Exemplary Study Using Twitter
abstract
Recent research has demonstrated that social media could provide valuable spatio-temporal data about users activities. However, information extraction and computation from big amount of data pose various challenges. To effectively process massive datasets, several platforms have been developed. Our previous study [20] explored Hadoop-based cloud computing for processing big amount of social media data [9] to study geographic distributions of social media users. In this paper, we investigate an emerging system named Spark and present a timely pilot experience on geospatial big data research. In our study, Spark has been utilized to perform some classic geospatial analyses like K-Nearest Neighbors (KNN), geographic mean and median points, and the distribution of the median points. Our design is tested on an Amazon EC2 cluster. An exemplary study using 60GB, 120GB and 180GB Twitter data has demonstrated the performance achievements by migrating computing tasks from Hadoop to Spark. In our experiments, the Spark-based solution can be up to 2.3x faster than the Hadoop-based solution due to its in-memory processing and coarse-grained resource allocation strategy. In the paper, we also discuss optimization strategies on using Spark for different geospatial computing tasks.
Hong Zhang 0047, Zixia Liu, Liqiang Wang 0001
CLOUD3
2015 Dart: A Geographic Information System on Hadoop
abstract
In the field of big data research, analytics on spatio-temporal data from social media is one of the fastest growing areas and poses a major challenge on research and application. An efficient and flexible computing and storage platform is needed for users to analyze spatio-temporal patterns in huge amount of social media data. This paper introduces a scalable and distributed geographic information system, called Dart, based on Hadoop and HBase. Dart provides a hybrid table schema to store spatial data in HBase so that the Reduce process can be omitted for operations like calculating the mean center and the median center. It employs reasonable pre-splitting and hash techniques to avoid data imbalance and hot region problems. It also supports massive spatial data analysis like K-Nearest Neighbors (KNN) and Geometric Median Distribution. In our experiments, we evaluate the performance of Dart by processing 160 GB Twitter data on an Amazon EC2 cluster. The experimental results show that Dart is very scalable and efficient.
Hong Zhang 0047, Zixia Liu, Liqiang Wang 0001
CLOUD3