Wenzhi Chen

dblp:70/2079 · DBLP profile ↗
← Back
96ranked-venue papers
3as first author
56since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 36 · 2 first-author · 16 since 2021Security and privacy · 18 · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 13 · 10 since 2021Computer networks · 8 · 8 since 2021Software engineering, systems software and programming languages · 6 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-authorTheory of computation · 1
YearPublicationVenuePosition
2026 Enhancing Meme Emotion Understanding with Multi-Level Modality Enhancement and Dual-Stage Modal Fusion
abstract
With the rapid rise of social media and Internet culture, memes have become a popular medium for expressing emotional tendencies. This has sparked growing interest in Meme Emotion Understanding (MEU), which aims to classify the emotional intent behind memes by leveraging their multimodal contents. While existing efforts have achieved promising results, two major challenges remain: (1) a lack of fine-grained multimodal fusion strategies, and (2) insufficient mining of memes' implicit meanings and background knowledge. To address these challenges, we propose MemoDetector, a novel framework for advancing MEU. First, we introduce a four-step textual enhancement module that utilizes the rich knowledge and reasoning capabilities of Multimodal Large Language Models (MLLMs) to progressively infer and extract implicit and contextual insights from memes. These enhanced texts significantly enrich the original meme contents and provide valuable guidance for downstream classification. Next, we design a dual-stage modal fusion strategy: the first stage performs shallow fusion on raw meme image and text, while the second stage deeply integrates the enhanced visual and textual features. This hierarchical fusion enables the model to better capture nuanced cross-modal emotional cues. Experiments on two datasets, MET-MEME and MOOD, demonstrate that our method consistently outperforms state-of-the-art baselines. Specifically, MemoDetector improves F1 scores by 4.3% on MET-MEME and 3.4% on MOOD. Further ablation studies and in-depth analyses validate the effectiveness and robustness of our approach, highlighting its strong potential for advancing MEU.
Wenlong Meng, Zhenyuan Guo, Chengkun Wei, Wenzhi Chen
AAAI5
2026 Zephyr: A Zero-loss and Tranparent TLS Connection Migration Framework
abstract
While essential for stateful modern workloads like Large Language Model agents and IoT services, long-lived connections impede cloud infrastructure agility by complicating maintenance and load balancing. Existing connection migration solutions either lack support for industrial-grade encrypted traffic or fail to prevent packet loss during handover in active production environments. To address this gap, we propose Zephyr, a zero-loss and transparent TLS connection migration framework for cross-node migration between servers with different addresses. Zephyr ensures transport-layer consistency by orchestrating an eBPF-based packet buffering mechanism to safely intercept in-flight data. At the application layer, rather than deeply modifying standard TLS libraries, Zephyr creatively reuses the native session resumption mechanism via a “fake client” strategy to reconstruct complex cryptographic states without client involvement. Implemented in widely-used industrial stacks (Nginx and OpenSSL), Zephyr achieves connection migration with approximately 4.1 ms downtime and strict zero packet loss. This approach enables seamless infrastructure optimization without disrupting cloud services.
Chengcheng Yu, Yueshang Zuo, Enge Song, Shaokai Zhang, Jiangu Zhao, Tian Pan 0001, Yang Song 0031, Xing Li 0007, Rong Wen, Chengkun Wei, Shunmin Zhu, Wenzhi Chen
APNet13
2026 Ada-Store: An Adaptive Load-aware Hybrid Storage Architecture for Bursty Workloads
abstract
The adoption of quad-level cell (QLC) flash memory in NVMe solid-state drives (SSDs) offers cloud vendors high capacity and cost efficiency, but its limited performance and endurance remain critical challenges. Hybrid storage architectures (e.g., Cloud Storage Acceleration Layer (CSAL)) combine high-performance and high-density SSDs to balance performance and cost. They typically take the former as the write cache. Meanwhile, they proactively migrate data from high-performance SSDs to high-density SSDs, a process also known as compaction. However, they still face two key issues in production environments. First, they suffer from severe performance degradation under burst I/O traffic due to continuous garbage collection (GC) for high-density SSDs. Second, they adopt a fixed-size simple moving average on the compaction space reclamation rate as the upper limit of user I/O bandwidth, which can lead to user I/O performance fluctuation or poor responsiveness to rapid changes in compaction throughput.
Luyang Ni, Jiexiong Xu, Yiquan Chen, Wenzhi Chen
CF5
2026 I-POP: Ignite Positive Prefetchers
abstract
Hardware prefetching is a well-established technique for bridging the processor-memory performance gap. To improve cache miss coverage, modern processors often integrate multiple prefetchers. However, multi-prefetcher systems without proper management often suffer from suboptimal performance due to a surge of useless prefetches. Several techniques have been proposed to select appropriate prefetchers for issuing requests, but they all face limitations. Specifically, existing (1) static schemes lack feedback regulation mechanisms and suffer from inflexible prefetcher selections; (2) reinforcement learning (RL)based schemes incur high overhead and suffer from adjustment lag; and (3) performance-counter-based schemes rely on inefficient runtime metrics that fail to accurately and clearly reflect a prefetcher's true impact on performance. In this paper, we propose I-POP, a high-performance and lowoverhead prefetcher management scheme for multi-prefetcher systems. I-POP introduces a novel runtime metric, Prefetch Effectiveness (PE), which aggregates each prefetch request's beneficial and harmful effects to precisely quantify the impact of a prefetcher on performance, effectively overcoming the limitations of prior metrics. To compute and leverage this metric, I-POP incorporates two key components: the Metric Collector, which periodically calculates each prefetcher's PE, and the Control Engine, which dynamically manages all prefetchers based on their PE values. Specifically, I-POP ignites (enables) prefetchers with positive PE values, adaptively tuning their aggressiveness, and disables those with non-positive PE. We evaluated I-POP on numerous workloads, and the results show I-POP outperforms two state-of-the-art approaches, Bandit and Alecto, by$\mathbf{4. 2 \%}$and 3.5 % across three benchmark suites in a single-core system, and 6.6 % and 8.6 % in a 16 -core system, while incurring only 1.46 KB of storage overhead.
Yiquan Lin, Wenhai Lin, Yiquan Chen, Jiexiong Xu, Shishun Cai, Jiarong Ye, Zonghui Wang, Wenzhi Chen
HPCA8
2026 Lmte: Putting the "Reasoning" into WAN Traffic Engineering with Language Models
Xinyu Yuan, Yan Qiao 0001, Zonghui Wang, Meng Li 0006, Wenzhi Chen
INFOCOM5
2026 Spillway: Orchestrating DPU and Host into a Unified vSwitching Fabric
abstract
The transition to Data Processing Unit (DPU)-centric architectures has become the de-facto standard in modern cloud networks, enabling infrastructure offload and improved host resource utilization. However, the fixed hardware limits of DPUs increasingly fail to keep pace with the rapid growth of host compute density and network-intensive workloads. As a result, when DPU resources are saturated, host compute capacity often remains underutilized due to insufficient network provisioning.
Xiaochong Jiang, Yilong Lv, Naixuan Guan, Qiming Zhao, Sihan Fu, Xuyang Ge, Denghui Wu, Yibin Shen, Guochun Hong, Yijian Dong, Yiquan Chen, Shaoliang An, Zhixiong Guo, Yisong Qiao, Hongwei Ding 0004, Shize Zhang, Rong Wen, Yang Song 0031, Zhigang Zong, Xing Li 0007, Chengkun Wei, Shunmin Zhu, Wenzhi Chen
SIGCOMM36
2026 Distribution-Aligned Synthetic Text Generation via Tail-Aware Enhancement
abstract
Recent advances in generative AI have popularized synthetic content for training, offering a practical alternative to costly data curation while addressing privacy concerns. However, accumulating evidence shows that the indiscriminate reuse of synthetic data can induce model collapse—a degenerative process that contracts the learned distribution and erodes rare features. For instance, when models are iteratively trained on their own synthetic outputs, the upper tail of the perplexity distribution substantially compresses, with high-percentile values dropping by nearly half—a clear indicator of severe diversity loss.
Xiaoyuan Liu 0002, Wubing Wang, Wenzhi Chen, Huaikang Fang, Lifeng Tao
WWW6
2026 Boosting Large Language Models for Mental Manipulation Detection via Data Augmentation and Distillation
abstract
Mental manipulation on social media poses a covert yet serious threat to individuals' psychological well-being and the integrity of online interactions. Detecting such behavior is challenging due to the difficult-to-annotate training data, its highly covert and multi-turn nature, and the lack of real-world datasets. To address these challenges, we propose MentalMAD, a framework that enhances large language models for mental manipulation detection. Our approach consists of three key components: EvoSA, an annotation-free data augmentation method that combines evolutionary operations with speech-act-aware prompting; teacher-model-generated complementary-task supervision; and Complementary-Convergent Distillation, a phase-wise strategy for transferring manipulation-specific knowledge to student models. We then constructed the ReaMent dataset, comprising 5,000 real-world-sourced dialogues. Extensive experiments show that MentalMAD improves accuracy by 14.0%, macro-F1 by 27.3%, and weighted F1 by 15.1% over the strongest baseline. The code and the dataset are publicly available at https://github.com/Yuansheng-Gao/MentalMAD.
Yuansheng Gao, Bin Li 0083, Jixiang Luo, Zonghui Wang, Wenzhi Chen
WWW7
2026 EditCoT: A Stepwise Chain-of-Thought Reasoning Framework for Multi-Intent Text Revision
abstract
Text revision is necessary to harness the written-text following human-acceptable requirements. Multi-intent text revision, however, requires all potential textual defects to be addressed in the same computational model, which poses a new challenge to the traditional single-intent-based text revision modeling approach. Conventional approaches often rely on models tailored to specific edit intents, limiting their ability to address diverse or unseen edit intents. Inspired by the reasoning strengths of Large Language Models (LLMs), we introduce EditCoT, a novel framework for multi-intent text revision. EditCoT breaks down the revision process into sequential reasoning steps, each targeting a specific text defect. The structured approach can enhance LLMs’ editing capabilities by enabling precise, intent-specific revisions within a unified model. We evaluate the effect of EditCoT on multi-/single-intent text revision tasks. For multi-intent tasks, EditCoT achieves state-of-the-art results, with a SARI score of 65.80 and a BERTScore of 88.27. For single-intent tasks, EditCoT, paired with GPT-o1, presents a competitive performance compared with specifically fine-tuned models. Furthermore, when combined with GPT-o1 or DeepSeek, EditCoT demonstrates impressive transferability to new edit intents via custom edit-chains. Overall, this study offers an effective framework for modeling and resolving text editing tasks, contributing a multi-intent dataset and an augmented single-intent dataset to support the community in advancing text revision research.
Xu Li 0032, Chengkun Wei, Wenzhi Chen
ACM Trans. Asian Low Resour. Lang. Inf. Process.3
2026 Dialogue Injection Attack: Jailbreaking LLMs Through Context Manipulation
abstract
Large language models (LLMs) have demonstrated significant utility in a wide range of applications; however, their deployment is plagued by security vulnerabilities, notably jailbreak attacks. These attacks manipulate LLMs to generate harmful or unethical content by crafting adversarial prompts. While much of the current research on jailbreak attacks has focused on single-turn interactions, it has largely overlooked the impact of historical dialogues on model behavior. Although recent studies have explored multi-turn jailbreak attacks, they generally assume that the attacker can only manipulate the user prompt. In contrast, we highlight that an attacker can also control the model’s previous outputs. To this end, we introduce DIA, a new paradigm that leverages fabricated dialogue history to enhance jailbreak effectiveness. DIA operates in a black-box setting, requiring only access to the chat API or knowledge of the LLM’s chat template. We propose two methods for constructing adversarial historical dialogues: one adapts gray-box prefilling attacks, and the other exploits deferred responses. Our experiments demonstrate that DIA achieves state-of-the-art attack success rates on recent LLMs, including Llama-3.1 and GPT-4o. Additionally, we show that DIA can bypass 6 different defense mechanisms, highlighting its robustness.
Wenlong Meng, Wendao Yao, Zhenyuan Guo, Yuwei Li 0002, Chengkun Wei, Wenzhi Chen
IEEE Trans. Inf. Forensics Secur.7
2026 Learning-Based Sketches for Frequency Estimation in Data Streams Without Ground Truth
abstract
Estimating the frequency of items on the high-volume, fast data stream has been extensively studied in many areas, such as database and network measurement. Traditional sketches provide only coarse estimates under strict memory constraints. Although some learning-augmented methods have emerged recently, they typically rely on offline training with real frequencies or/and labels, which are often unavailable. Moreover, these methods suffer from slow update speeds, limiting their suitability for real-time processing despite offering only marginal accuracy improvements. To overcome these challenges, we propose UCL-sketch, a practical learning-based paradigm for per-key frequency estimation. Our design introduces two key innovations: (i) an online training mechanism based on equivalent learning that requires no ground truth (GT), and (ii) a highly scalable architecture leveraging logically structured estimation buckets to scale to real-world data stream. The UCL-sketch, which utilizes compressive sensing (CS), converges to an estimator that provably yields an error bound far lower than that of prior works, without sacrificing the speed of processing. Extensive experiments on both real-world and synthetic datasets demonstrate that our approach outperforms previously proposed approaches regarding per-key accuracy and distribution. Notably, under extremely tight memory budgets, its quality almost matches that of an (infeasible) omniscient oracle. Moreover, compared to the existing equation-based sketch, UCL-sketch achieves an average decoding speedup of nearly 500 times.
Xinyu Yuan, Yan Qiao 0001, Meng Li 0006, Zhenchun Wei, Cuiying Feng, Zonghui Wang, Wenzhi Chen
IEEE Trans. Knowl. Data Eng.7
2026 Distributed Rate Limiting Under Decentralized Cloud Networks
abstract
The rapid expansion of cloud applications has led to unprecedented increases in network traffic volume, diversity, and complexity. As Cloud Service Providers (CSPs) adopt decentralized, geographically distributed data centers, effective traffic management across these environments has become critical. Distributed Rate Limiting (DRL) has emerged as an essential tool to manage the complex traffic dynamics of decentralized networks, yet traditional centralized rate limiting methods fall short, facing limitations in scalability, adaptability to bursty traffic, and efficiency. This paper presents C3PDAR (Cloud Control with Constant Probabilities and Dynamic Adjustment Range), a novel DRL algorithm tailored for decentralized cloud infrastructures. C3PDAR introduces three key innovations: (1) CPS-BPS DualPoint Rate Limiting and Parent-Child Token Bucket mechanisms, which effectively mitigate burst traffic and short-lived connections while improving bandwidth fairness and inter-tenant isolation; (2) A vSwitch-CGW Cascade Rate Limiting architecture, which reduces CPU overhead in CGW clusters and accelerates convergence by 42%–78%; (3) Virtual Extensible Local Area Network (VXLAN) Padding scheme, which embeds rate-limiting information in existing traffic instead of transmitting new data packets, reducing the communication overhead of the C3PDAR algorithm by over 40%. By integrating these advancements, C3PDAR delivers a scalable, robust solution that outperforms traditional DRL approaches in performance, fault tolerance, and resource efficiency. C3PDAR uniquely empowers CSPs to manage complex, high-volume traffic dynamics in decentralized cloud environments, offering both theoretical insights and practical optimizations for next-generation network control.
Tianyu Xu 0007, Lilong Chen, Xiaochong Jiang, Liming Ye, Yilong Lv, Chenhao Jia, Yongwang Wu, Zhigang Zong, Xing Li 0007, Bingqian Lu, Shunmin Zhu, Chengkun Wei, Wenzhi Chen
IEEE Trans. Mob. Comput.16
2025 Facial Authentication Security Evaluation Against Deepfake Attacks in Mobile Apps
Chuer Yu, Siyi Xia, Zonghui Wang, Lirong Fu, Zhiyuan Wan, Yandong Gao, Wenzhi Chen
ACISP (3)10
2025 OS2G: A High-Performance DPU Offloading Architecture for GPU-based Deep Learning with Object Storage
abstract
Object storage is increasingly attractive for deep learning (DL) applications due to its cost-effectiveness and high scalability. However, it exacerbates CPU burdens in DL clusters due to intensive object storage processing and multiple data movements. Data processing unit (DPU) offloading is a promising solution, but naively offloading the existing object storage client leads to severe performance degradation. Besides, only offloading the object storage client still involves redundant data movements, as data must first transfer from the DPU to the host and then from the host to the GPU, which continues to consume valuable host resources.
Zhen Jin 0008, Yiquan Chen, Mingxu Liang, Guoju Fang, Keyao Zhang, Jiexiong Xu, Wenhai Lin, Yiquan Lin, Shushu Zhao, Wenkai Shi, Zhenhua He, Shishun Cai, Wenzhi Chen
ASPLOS (2)15
2025 BatchZK: A Fully Pipelined GPU-Accelerated System for Batch Generation of Zero-Knowledge Proofs
abstract
Zero-knowledge proof (ZKP) is a cryptographic primitive that enables one party to prove the validity of a statement to other parties without disclosing any secret information. With its widespread adoption in applications such as blockchain and verifiable machine learning, the demand for generating zero-knowledge proofs has increased dramatically. In recent years, considerable efforts have been directed toward developing GPU-accelerated systems for proof generation. However, these previous systems only explored efficiently generating a single proof by reducing latency rather than batch generation to provide high throughput.
Tao Lu 0015, Yuxun Chen, Zonghui Wang, Xiaohang Wang 0001, Wenzhi Chen, Jiaheng Zhang
ASPLOS (1)5
2025 rInfer: A Generic and High-Performance Framework for Remote Inference with Heterogeneous Accelerators
abstract
Inference applications leverage various heterogeneous accelerators, including GPUs, TPUs, and FPGAs, to achieve remarkable performance. Remote inference, which enables applications and accelerators to reside on different nodes, can effectively address low hardware resource utilization and enhance flexibility in resource allocation. However, existing APIremoting solutions face compatibility issues, as they demand extensive development efforts to support diverse heterogeneous accelerators and adapt to rapidly evolving device runtimes. On the other hand, current inference service systems demand cumbersome cross-server configuration and suffer from suboptimal data transmission performance, overlooking data copy overhead in the network stack and not leveraging high-performance RDMA. In this paper, we present rInfer, a generic and highperformance framework for remote inference with heterogeneous accelerators. The key idea of rInfer is to abstract the entire remote inference process and optimize network transmission. On the client node, rInfer offers users an inference-oriented rInfer Device and an ease-of-use rInfer programming model, shielding users from the complexities of server configuration and network communication. On the server node, rInfer integrates with diverse inference frameworks, such as TensorRT and Torch, ensuring high compatibility with various accelerators. Furthermore, we optimize network data transmission by eliminating the need for serialization and deserialization in Protobuf to minimize data copying and adopt RDMA to enhance data transfer efficiency. Experimental results demonstrate that rInfer achieves performance nearly equivalent to local execution for the BERT model, with only a minor difference of 1.5 %. Compared to TorchServe, rInfer-TCP achieves a 51.6 % reduction in execution time for the ResNet models and consumes 37.3% less CPU and 61.2% less memory bandwidth on the server node.
Zhen Jin 0008, Yiquan Chen, Yin Du, Keyao Zhang, Jiexiong Xu, Wenhai Lin, Jingchang Qin, Kanghua Fang, Wenzhi Chen
CCGrid10
2025 Pre-training CLIP against Data Poisoning with Optimal Transport-based Matching and Alignment
abstract
Recent studies have shown that Contrastive Language-Image Pre-training (CLIP) models are threatened by targeted data poisoning and backdoor attacks due to massive training imagecaption pairs crawled from the Internet.Previous defense methods correct poisoned imagecaption pairs by matching a new caption for each image.However, the matching process relies solely on the global representations of images and captions, overlooking fine-grained features of visual and textual features.It may introduce incorrect image-caption pairs and harm the CLIP pre-training.To address their limitations, we propose an Optimal Transportbased framework to reconstruct image-caption pairs, named OTCCLIP.We propose a new optimal transport-based distance measure between fine-grained visual and textual feature sets and re-assign new captions based on the proposed optimal transport distance.Additionally, to further reduce the negative impact of mismatched pairs, we encourage the inter-and intra-modality fine-grained alignment by employing optimal transport-based objective functions.Our experiments demonstrate that OTC-CLIP can successfully decrease the attack success rates of poisoning attacks.Also, compared to previous methods, OTCCLIP significantly improves CLIP's zero-shot and linear probing performance trained on poisoned datasets.
Kuofeng Gao, Jiawang Bai, Leo Yu Zhang, Zonghui Wang, Shouling Ji, Wenzhi Chen
EMNLP8
2025 NVMePass: A Lightweight, High-performance and Scalable NVMe Virtualization Architecture with I/O Queues Passthrough
abstract
Most data-intensive applications currently run on NVMe storage, and virtualization is essential in cloud computing. Existing NVMe virtualization technologies include software-based and hardware-assisted. Virtio suffers from severe performance degradation, and polling-based solutions consume too many valuable CPU resources. Hardware-assisted solutions provide high performance and no CPU usage but have the challenges of developing dedicated hardware.In this paper, we propose NVMePass, a novel software-hardware co-design NVMe passthrough virtualization architecture designed to achieve high performance and no CPU overhead while maintaining high scalability. The key ideas of NVMePass are NVMe I/O queues passthrough for VMs and a mechanism to ensure security. The NVMePass supports DMA and interrupts remapping for VMs without hypervisor involvement, eliminating virtualization overhead and providing near-native performance. Isolation is achieved by I/O queues and logical block address resources exclusively allocated to VMs. We propose NVMe Resource Domain (NRD) and implement it in the NVMe controller to intercept illegal I/O requests. Thus, isolation and security are fully achieved. Results from our experiments show that NVMePass can provide comparable performance to VFIO, with an IOPS of $\mathbf{1 0 0. 1 \% - 1 0 0. 5 \%}$ of VFIO. Furthermore, compared to SPDK-Vhost, NVMePass achieves $\mathbf{4 0. 0 \%}$ lower latency when running 150 VMs, and NVMePass has an improvement of $\mathbf{6 8. 0 \%}$ OPS performance in a real-world application when running 100 VMs.
Yiquan Chen, Zhen Jin 0008, Jiexiong Xu, Hao Yu 0016, Wenhai Lin, Kanghua Fang, Keyao Zhang, Chengkun Wei, Yuan Xie 0001, Wenzhi Chen
HPCA14
2025 Bidirectional Reference Image Quality Assessment via Content-Quality Correlation Modeling
abstract
The emphasis on no-reference image quality assessment has often overshadowed the significance of Full-Reference Image Quality Assessment (FR-IQA), which generally better reflects human contrastive perception mechanism. However, FRIQA presents challenges in obtaining content-aligned reference images. To tackle these issues, a novel Bidirectional Reference Image Quality Assessment (BRIQA) method is proposed, centering on leveraging bidirectional reference images and content-quality correlation modeling. First, triplets of content-aligned low-quality and content-non-aligned high-quality reference images are generated using two easily accessible approaches. To prevent the extraction of redundant information, two feature extractors pretrained through unsupervised contrastive learning are utilized to independently extract content and quality features for the triplet images. Then, an attention-mixer is introduced to further mine quality difference information and enhance content feature. Finally, a content-quality correlation modeler is proposed to model the relationship between quality differences and visual contents. Experimental results on benchmark datasets demonstrate that the BRIQA outperforms existing state-of-the-art methods.
Bo Hu 0008, Wenzhi Chen, Chunyi Li 0001, Jiaxu Leng, Weisheng Li 0001, Xinbo Gao 0001
ICASSP2
2025 An Inversion-Based Measure of Memorization for Diffusion Models
Zhe Ma 0002, Qingming Li, Xuhong Zhang 0002, Tianyu Du, Ruixiao Lin, Zonghui Wang, Shouling Ji, Wenzhi Chen
ICCV8
2025 MineShark: Cryptomining Traffic Detection at Scale
Shaoke Xi, Tianyi Fu, Kai Bu, Zhihua Chang, Wenzhi Chen, Zhou Ma, Chongjie Chen, Yongsheng Shen, Kui Ren 0001
NDSS6
2025 GradEscape: A Gradient-Based Evader Against AI-Generated Text Detectors
Wenlong Meng, Shuguo Fan, Chengkun Wei, Min Chen 0032, Yuwei Li 0002, Zhikun Zhang 0001, Wenzhi Chen
USENIX Security Symposium8
2025 PRSA: Prompt Stealing Attacks against Real-World Prompt Services
Yong Yang 0017, Changjiang Li, Qingming Li, Oubo Ma, Zonghui Wang, Yandong Gao, Wenzhi Chen, Shouling Ji
USENIX Security Symposium8
2025 DC-SGD: Differentially Private SGD With Dynamic Clipping Through Gradient Norm Distribution Estimation
abstract
Differentially Private Stochastic Gradient Descent (DP-SGD) is a widely adopted technique for privacy-preserving deep learning. A critical challenge in DP-SGD is selecting the optimal clipping threshold C, which involves balancing the trade-off between clipping bias and noise magnitude, incurring substantial privacy and computing overhead during hyperparameter tuning. In this paper, we propose Dynamic Clipping DP-SGD (DC-SGD), a framework that leverages differentially private histograms to estimate gradient norm distributions and dynamically adjust the clipping thresholdC. Our framework includes two novel mechanisms: DC-SGD-P and DC-SGD-E. DC-SGD-P adjusts the clipping threshold based on a percentile of gradient norms, while DC-SGD-E minimizes the expected squared error of gradients to optimizeC. These dynamic adjustments significantly reduce the burden of hyperparameter tuningC. The extensive experiments on various deep learning tasks, including image classification and natural language processing, show that our proposed dynamic algorithms achieve up to 9 times acceleration on hyperparameter tuning than DP-SGD. And DC-SGD-E can achieve an accuracy improvement of 10.62% on CIFAR10 than DP-SGD under the same privacy budget of hyperparameter tuning. We conduct rigorous theoretical privacy and convergence analyses, showing that our methods seamlessly integrate with the Adam optimizer. Our results highlight the robust performance and efficiency of DC-SGD, offering a practical solution for differentially private deep learning with reduced computational overhead and enhanced privacy guarantees.
Chengkun Wei, Weixian Li, Chen Gong 0005, Wenzhi Chen
IEEE Trans. Inf. Forensics Secur.4
2025 Invisible-Face: Rethinking Facial Attribute Privacy in Social Media Photo Sharing
abstract
As social media gains popularity, users frequently share personal photos without recognizing the risks of exposing their faces to advanced facial attribute detection technologies. These technologies can extract sensitive attributes such as age, race, sexual orientation, and potential health information from facial images, raising significant privacy concerns. Despite the availability of various anonymization techniques, our research reveals that current methods inadequately protect facial attribute privacy. They often fail to balance effectiveness and utility, underscoring the pressing need for more robust solutions in today’s pervasive photo-sharing culture. To remedy this gap, we introduce Invisible-Face, a tool designed to safeguard users’ facial attribute privacy using advanced adversarial perturbation techniques. Invisible-Face uses local, directional, and resilient perturbation generative strategies to obfuscate multiple facial attributes effectively, thus ensuring privacy while retaining the utility of the facial images. Our comprehensive evaluation across various datasets and model architectures shows that Invisible-Face significantly outperforms existing privacy-preserving methods in terms of effectiveness while maintaining high image naturalness. Furthermore, our extensive real-world evaluations on four popular MLaaS platforms—Baidu Brain, Tencent Cloud, Aliyun, and Face++—reveal that Invisible-Face achieves comparable privacy protection results while preserving the visual naturalness of images, outperforming existing methods. These findings boost public awareness about the importance of facial attribute privacy and urge online social platforms to improve their protection measures.
Yong Yang 0017, Changjiang Li, Xuhong Zhang 0002, Zonghui Wang, Shouling Ji, Wenzhi Chen
IEEE Trans. Inf. Forensics Secur.8
2025 Beehive: Decentralised High-Frequency Small Tasks Scheduling in Large Clusters
abstract
Data centers struggle with growing cluster sizes and rising submissions of short-lived, high-frequency tasks that cause performance bottlenecks in task scheduling. Existing centralized and distributed scheduling systems fall short in meeting performance requirements due to computational overload on the scheduler, cluster state management overhead, and scheduling conflicts. To address these challenges, this paper introduces Beehive, a novel lightweight decentralized scheduling framework. In Beehive, each cluster node can schedule tasks within its local neighborhood, effectively reducing resource management overhead and scheduling conflicts. Moreover, all nodes are interconnected in a small-world network, an efficient structure that allows tasks to access resources across the entire cluster through global routing. This lightweight design enables Beehive to scale efficiently, supporting over 10,000 nodes and up to 80,000 task submissions per second without causing single-node scheduling bottlenecks. Experimental results demonstrate that Beehive significantly reduces scheduling latency. Specifically, 99% of tasks are scheduled within 100 milliseconds, and scheduling throughput can increase linearly with the number of nodes. Compared to existing centralized and distributed scheduling frameworks, Beehive substantially alleviates scheduling bottlenecks, particularly for high-frequency, short-lived tasks
Yuxia Cheng, Tongkai Yang, Antong Yu, Wenzhi Chen
IEEE Trans. Parallel Distributed Syst.7
2024 CMDRL: A Markovian Distributed Rate Limiting Algorithm in Cloud Networks
abstract
As cloud networks continue to evolve, network traffic has experienced an exponential increase. The network architecture is progressively adopting a distributed structure to address this challenge. This architecture extensively utilizes technologies like gateway clusters and Equal-Cost Multi-Path (ECMP) routing, enabling traffic from individual tenants to be routed through multiple pathways. As a result, distributed rate limiting (DRL) has emerged as an essential aspect. Nonetheless, the shift from centralized to DRL has encountered obstacles, with the associated algorithms grappling with simplicity, precision, and applicability issues. Consequently, our research seeks to reconceptualize the issue of DRL from a theoretical standpoint to discover a more holistic and efficacious solution.
Lilong Chen, Xiaochong Jiang, Tianyu Xu 0007, Xing Li 0007, Bingqian Lu, Chengkun Wei, Wenzhi Chen
APNet9
2024 CINDA: Don't Ignore Instructions When Cloning Memory Access Behavior
abstract
Existing workload cloning methods suffer from low accuracy as they primarily focus on data access patterns and ignore instruction access. This limitation reduces the accuracy of shared L2 cache design exploration and impedes processor designers from optimizing Icache and ITLB designs. In this paper, we propose CINDA, a novel workload cloning technique that can Capture both INstruction and DAta access patterns of applications. In particular, CINDA separates the instruction and data traces of applications to generate proxy instruction and proxy data traces, subsequently merging them. The results show that CINDA can accurately replicate memory access behavior with 99.1%, 99.9%, and 96.2% accuracy in replicating L1 Icache, ITLB and L2 cache performance, respectively. Furthermore, CINDA outperforms the state-of-the-art methods by reducing 7.7% L2 cache miss error.
Wenhai Lin, Yiquan Chen, Jiexiong Xu, Zhen Jin 0008, Peiyu Liu 0003, Shishun Cai, Yuzhong Zhang, Jingchang Qin, Yiquan Lin, Wenzhi Chen
CCGrid10
2024 BlueJay: A Platform to Quantifying the Impact of Memory Latency on Datacenter Application Performance
abstract
Understanding the impact of memory latency on datacenter application performance can provide decision support to memory subsystem designers. Currently, various methods are available to quantify this impact, including cycle-accurate simulators, memory-level parallelism models, and software delay injection techniques. However, these methods suffer from several limitations, such as slow simulation speed, inaccuracy, and insufficient compatibility that requires application modification.This paper proposes BlueJay, a novel platform to quantify the impact of memory latency on the end-to-end performance of datacenter applications, avoiding slow simulation and providing high accuracy and compatibility. The key idea of BlueJay is to control the consumed memory bandwidth and read/write ratio, thereby manipulating memory latency to achieve quantification. Experiment shows that BlueJay provides accurate quantification with an average error of 3.04%. In addition, we built regression models for five applications deployed at scale in Alibaba data centers. The results reveal that a 10 ns increase in memory latency results in a performance decrease of 2.61%-3.31% for enterprise Java applications and databases, while the elastic block storage service experiences a more modest performance decrease of 0.73%-0.92%.
Jingchang Qin, Yiquan Chen, Shishun Cai, Wenhai Lin, Jiexiong Xu, Zhen Jin 0008, Lifa Cao, Yuzhong Zhang, Wenzhi Chen
CCGrid11
2024 LightPool: A NVMe-oF-based High-performance and Lightweight Storage Pool Architecture for Cloud-Native Distributed Database
abstract
Emerging cloud-native distributed databases rely on local NVMe SSDs to provide high-performance and highavailable data services to many cloud applications. However, the database clusters suffer from low utilization of local storage because of the imbalance between CPU and storage capacities within each node. For instance, the OceanBase distributed database cluster, with hundreds of PB local storage capacity, only utilizes around 40% of its local storage. Although disaggregated storage (EBS) can enhance storage utilization by provisioning the CPU and storage independently on demand, they suffer from performance bottlenecks and high costs. In this paper, we propose LightPool, a high-performance and lightweight storage pool architecture large-scale deployed in the OceanBase clusters, enhancing storage resource utilization. The key idea of LightPool is aggregating cluster storage into a storage pool and enabling unified management. In particular, LightPool adopts NVMe-oF to enable high-performance storage resource sharing among cluster nodes and integrate the storage pool with Kubernetes to achieve flexible management and allocation of storage resources. Furthermore, we design the hot-upgrade and hot-migration mechanisms to enhance the availability of LightPool. We have deployed LightPool on over 8500 nodes in production clusters. Statistics show that LightPool can improve storage resource utilization from about 40% to 65%. Experimental results show that the extra latency from LightPool is only about 2.1 μs compared to local storage. Compared to OpenEBS, LightPool enhances bandwidth up to 190.9% in microbenchmarks and throughput up to 6.9% in real-world applications. LightPool is the best practice to deploy NVMe-oF (NVMe/TCP) in the production environment. We also discuss important lessons and experiences learned from the development of LightPool.
Jiexiong Xu, Yiquan Chen, Wenhui Shi, Guoju Fang, Huasheng Liao, Zhen Jin 0008, Wenzhi Chen
HPCA12
2024 ShardingSim: A Modular Committee-Based Sharding Blockchain Simulator
abstract
Blockchain performance is crucial in research, with sharding emerging as an effective solution for scalability. By dividing the network into smaller shards, sharding facilitates faster transaction processing. However, there is currently a lack of effective simulators for modeling sharding blockchain performance and assessing shard load balancing. In this paper, we introduce ShardingSim, a modular, committee-based sharding blockchain simulator. ShardingSim simulates various sharding configurations and evaluates their performance across diverse network conditions and transaction datasets. We present a use case by modeling and simulating RapidChain with ShardingSim: simulation results show that RapidChain’s performance improves with more shards under historical Bitcoin transaction datasets; however, such enhancement is absent in scenarios with uneven transaction distributions, highlighting the need for additional load balancing methods. Through ShardingSim, we can simulate committee-based sharding blockchain, evaluateits performance, and identify potential load-balancing issues between shards.
Yuehua Wu, Feihu Yan, Wenzhi Chen
ICBC4
2024 Pluggable Watermarking of Deepfake Models for Deepfake Detection
Xuhong Zhang 0002, Qinying Wang, Kangming Liang, Zonghui Wang, Shouling Ji, Wenzhi Chen
IJCAI7
2024 LMSanitator: Defending Prompt-Tuning Against Task-Agnostic Backdoors
Chengkun Wei, Wenlong Meng, Zhikun Zhang 0001, Min Chen 0032, Minghu Zhao, Wenjing Fang, Lei Wang 0152, Wenzhi Chen
NDSS9
2024 Triton: A Flexible Hardware Offloading Architecture for Accelerating Apsara vSwitch in Alibaba Cloud
abstract
Apsara vSwitch (AVS) is a per-host deployed forwarding component for instance network connectivity in the Alibaba Cloud. To meet the growing performance demands, we accelerated AVS by adopting the most widely used "Sep-path" offloading architecture, which introduces a separate hardware data path to speed up popular traffic. However, the deployment results prove that it is difficult to bridge the gap in performance and programming flexibility of the software and hardware data paths, resulting in unpredictable performance and low iteration velocity.
Xing Li 0007, Xiaochong Jiang, Lilong Chen, Yi Wang 0004, Chao Wang 0128, Chao Xu 0017, Yilong Lv, Taotao Wu, Haifeng Gao, Yisong Qiao, Hongwei Ding 0004, Yijian Dong, Jianming Song, Jianyuan Lu, Chengkun Wei, Wenzhi Chen, Qinming He, Shunmin Zhu
SIGCOMM22
2024 Performance Characterization of SmartNIC NVMe-over-Fabrics Target Offloading
abstract
The NVMe-over-Fabrics (NVMe-oF) is gaining popularity in cloud data centers as a remote storage protocol for accessing NVMe storage devices across servers. With the rapid increase in throughput of the NVMe storage devices, the NVMe-oF stack consumes a significant amount of valuable CPU resources. To release these valuable computing power to other tasks, many smartNICs now support NVMe-oF Target offloading. However, this emerging NVMe-oF offloading scheme's performance has not been fully investigated.
Jiexiong Xu, Yiquan Chen, Wenhai Lin, Yiquan Lin, Shushu Zhao, Wenzhi Chen
SYSTOR10
2024 Exploring ChatGPT's Capabilities on Vulnerability Management
Peiyu Liu 0003, Lirong Fu, Kangjie Lu, Xuhong Zhang 0002, Wenzhi Chen, Haiqin Weng, Shouling Ji, Wenhai Wang
USENIX Security Symposium7
2024 Spectral clustering with linear embedding: A discrete clustering method for large-scale data
Chenhui Gao, Wenzhi Chen, Feiping Nie 0001, Weizhong Yu, Zonghui Wang
Pattern Recognit.2
2024 PARS: A Pattern-Aware Spatial Data Prefetcher Supporting Multiple Region Sizes
abstract
Hardware data prefetching is a well-studied technique to bridge the processor-memory performance gap. Bit-pattern-based prefetchers are one of the most promising spatial data prefetchers that achieve substantial performance gains. In bit-pattern-based prefetchers, the region size is a crucial parameter, which denotes the memory size that can be recorded by a pattern or prefetched by a prediction. However, existing bit-pattern-based prefetchers only support one fixed region size. Our experiment shows that the fixed region size cannot meet the requirements for numerous applications and leads to suboptimal performance and high hardware overhead. In this article, we propose PARS, a pattern-aware spatial data prefetcher supporting multiple region sizes. The key idea of PARS is that it supports multiple region sizes, enabling it to simultaneously enhance application performance while reducing the hardware overhead. Moreover, PARS supports dynamically switching appropriate region sizes for different patterns through an adaptive RS-switching mechanism. We evaluated PARS on numerous workloads and results show that PARS provides an average performance improvement of 40.6% over a baseline with no data prefetchers and outperforms the two state-of-the-art prefetchers Bingo by 2.1% (up to 24.4%) and Pythia by 3.9% (up to 111.2%) in the single-core system. In the four-core system, PARS outperforms Bingo by 5.0% (up to 66.0%) and Pythia by 5.4% (up to 177.9%).
Yiquan Lin, Wenhai Lin, Jiexiong Xu, Yiquan Chen, Zhen Jin 0008, Jingchang Qin, Shishun Cai, Yuzhong Zhang, Zonghui Wang, Wenzhi Chen
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.11
2024 Diff-ID: An Explainable Identity Difference Quantification Framework for DeepFake Detection
abstract
In recent years, DeepFake technologies have seen widespread adoption in various domains, including entertainment and film production. However, they have also been maliciously employed for disseminating false information and engaging in video fraud. Existing detection methods often experience significant performance degradation when confronted with unknown forgeries or exhibit limitations when dealing with low-quality images. To address this challenge, we introduceDiff-ID, a novel approach designed to elucidate and quantify the identity loss induced by facial manipulations. When assessing the authenticity of an image,Diff-IDleverages a genuine image of the same individual as a reference and processes two images jointly. It aligns the reference image and the test image into the same identity-insensitive attribute feature space using a face-swapping generator. This alignment allows us to observe the identity disparities between the two images through the differences in the aligned generation pairs. Subsequently, we have developed a custom metric designed to quantify the identity loss relative to the reference image in the test image. This metric effectively distinguishes forgery images from the real ones. Extensive experiments have demonstrated the exceptional performance of our approach. It achieves a high level of detection accuracy on DeepFake images and showcases state-of-the-art generalization capabilities when confronted with previously unknown forgery methods. Moreover, it exhibits robustness even in the presence of image distortions.
Chuer Yu, Xuhong Zhang 0002, Yuxuan Duan, Senbo Yan, Zonghui Wang, Yang Xiang 0001, Shouling Ji, Wenzhi Chen
IEEE Trans. Dependable Secur. Comput.8
2024 MILG: Realistic lip-sync video generation with audio-modulated image inpainting
abstract
Existing lip synchronization (lip-sync) methods generate accurately synchronized mouths and faces in a generated video. However, they still confront the problem of artifacts in regions of non-interest (RONI), e.g. , background and other parts of a face, which decreases the overall visual quality. To solve these problems, we innovatively introduce diverse image inpainting to lip-sync generation. We propose Modulated Inpainting Lip-sync GAN (MILG), an audio-constraint inpainting network to predict synchronous mouths. MILG utilizes prior knowledge of RONI and audio sequences to predict lip shape instead of image generation , which can keep the RONI consistent. Specifically, we integrate modulated spatially probabilistic diversity normalization (MSPD Norm) in our inpainting network, which helps the network generate fine-grained diverse mouth movements guided by the continuous audio features. Furthermore, to lower the training overhead, we modify the contrastive loss in lip-sync to support small-batch-size and few-sample training. Extensive experiments demonstrate that our approach outperforms the existing state-of-the-art of image quality and authenticity while keeping lip-sync.
Xuhong Zhang 0002, Qinying Wang, Kangming Liang, Zonghui Wang, Shouling Ji, Wenzhi Chen
Vis. Informatics7
2023 HyQ: Hybrid I/O Queue Architecture for NVMe over Fabrics to Enable High- Performance Hardware Offloading
abstract
NVMe over Fabrics (NVMe-oF) has been widely applied as a remote storage protocol in cloud computing. The existing NVMe-oF software stack consumes a large number of CPU resources. Emerging devices, such as Smart NICs and DPUs, have supported hardware offloading of NVMe-oF to free these valuable CPU cores. However, NVMe-oF offloading capacity is always compromised because of limited hardware resources on design. Additionally, from thorough evaluations, we found that NVMe-oF inevitably suffers from severe performance degradation on complex application I/O patterns when using hardware offloading. It is challenging to achieve high performance and fully utilize NVMe-oF offloading simultaneously. In this paper, we propose HyQ, a novel hybrid I/O queue architecture for NVMe-oF, to achieve high performance while gaining the advantages of hardware offloading. HyQ realizes the coexistence of hardware offloading and software non-offloading queues, thus enabling the dynamic dispatching of I/O requests to appropriate processing queues according to user-defined I/O scheduling policies. Additionally, HyQ provides a request scheduling framework to support customized schedulers that select appropriate queues for I/O requests. In our evaluation, HyQ achieves up to 1.91x IOPS and 8.36x bandwidth performance improvement over the original hardware offloading scheme.
Yiquan Chen, Zhen Jin 0008, Jiexiong Xu, Guoju Fang, Wenhai Lin, Chengkun Wei, Wenzhi Chen
CCGrid10
2023 Securely Sampling Discrete Gaussian Noise for Multi-Party Differential Privacy
abstract
Differential Privacy (DP) is a widely used technique for protecting individuals' privacy by limiting what can be inferred about them from aggregate data. Recently, there have been efforts to implement DP using Secure Multi-Party Computation (MPC) to achieve high utility without the need for a trusted third party. One of the key components of implementing DP in MPC is noise sampling. Our work presents the first MPC solution for sampling discrete Gaussian, a common type of noise used for constructing DP mechanisms, which plays nicely with malicious secure MPC protocols.
Chengkun Wei, Ruijing Yu, Wenzhi Chen, Tianhao Wang 0001
CCS4
2023 DPMLBench: Holistic Evaluation of Differentially Private Machine Learning
abstract
Differential privacy (DP), as a rigorous mathematical definition quantifying privacy leakage, has become a well-accepted standard for privacy protection. Combined with powerful machine learning (ML) techniques, differentially private machine learning (DPML) is increasingly important. As the most classic DPML algorithm, DP-SGD incurs a significant loss of utility, which hinders DPML's deployment in practice. Many studies have recently proposed improved algorithms based on DP-SGD to mitigate utility loss. However, these studies are isolated and cannot comprehensively measure the performance of improvements proposed in algorithms. More importantly, there is a lack of comprehensive research to compare improvements in these DPML algorithms across utility, defensive capabilities, and generalizability.
Chengkun Wei, Minghu Zhao, Zhikun Zhang 0001, Min Chen 0032, Wenlong Meng, Wenzhi Chen
CCS8
2023 JACO: JAva Code Layout Optimizer Enabling Continuous Optimization without Pausing Application Services
abstract
Many Java applications in data centers suffer from severe processor pipeline frontend bottlenecks, which can be mitigated by profile-guided code layout optimizations (PGCLO). To maximize optimization opportunities, state-of-the-art PGCLO solutions adopt continuous optimization to ensure that the code layout consistently matches ever-changing application control flow characteristics. However, existing continuous optimizations inevitably pause the application to execute the new code completely, which leads to high response latency and significantly deteriorates user experience.In this paper, we propose JACO, a novel profile-guided Java code layout optimizer, enabling continuous optimization without pausing application services. The key idea of JACO is to enable the execution of both the old and new code simultaneously rather than completely switching to the new code. In particular, JACO is composed of three components: (1) A lightweight profiler captures the control flow information of the application and then generates an optimized function order. (2) A control flow switcher generates new code based on optimized function order and switches the application to execute the new code without pausing the application services. (3) A selective code reclaimer only frees the memory occupied by the inactive old code. We evaluated JACO on both open-source applications and real-world applications from a world-leading company. JACO achieved up to a 16.36% performance improvement for real-world applications. The state-of-the-art approach introduces up to 37.93x latency overhead that will interrupt application services, while JACO only introduces a negligible 7% latency overhead.
Wenhai Lin, Jingchang Qin, Yiquan Chen, Zhen Jin 0008, Jiexiong Xu, Yuzhong Zhang, Shishun Cai, Lirong Fu, Wenzhi Chen
CLUSTER10
2023 BM-Store: A Transparent and High-performance Local Storage Architecture for Bare-metal Clouds Enabling Large-scale Deployment
abstract
Bare-metal instances are crucial for high-value, mission-critical applications on the cloud. Tenants exclusively use these dedicated hardware resources. Local virtualized disks are essential for bare-metal instances to provide flexible and high-performance storage resources. Traditionally tenants can choose polling-based software virtualization techniques, but they consume too many valuable host CPU cores and suffer from performance degradation. Cloud vendors are hard to deploy existing hardware-assisted local storage solutions in bare-metal instances due to no access to the host OS to install customized drivers. Moreover, cloud vendors have difficulties managing and maintaining the local storage devices in bare-metal instances because hardware resources and host operating systems are completely utilized by tenants, then it will impact the availability of storage devices.This paper presents our design and experience with BM-Store, a novel high-performance hardware-assisted virtual local storage architecture for bare-metal clouds. BM-Store is transparent to the host that tenants are unaware of the underlying hardware architecture. Therefore, it can be deployed on a large scale in cloud vendors. BM-Store consists of two components: an FPGA-based BMS-Engine and an ARM-based BMS-Controller. The BMS-Engine accelerates the I/O path to enable high-performance virtual storage independent of disk devices without consuming any CPU resource on the host. The BMS-Controller is responsible for resource management and maintenance to achieve flexible and high available local storage. The results of the extensive experiments show that BM-Store can achieve near-native performance, which only introduces about 3 µs extra latency and average 4.0% throughput overhead to native disks. Compared to SPDK vhost, BM-Store achieves an average bandwidth improvement of 15.7% in microbenchmark and a maximum throughput enhancement of 13.4% in real-world applications.
Yiquan Chen, Jiexiong Xu, Chengkun Wei, Xulin Yu, Zeke Wang, Shuibing He, Wenzhi Chen
HPCA11
2023 Poster: Triton: Accelerating vSwitch with Flexibility through Hardware Assisting not Bypassing Software
abstract
The vSwitch, as a critical component for Virtual Machine (VM) network connectivity in cloud environments, has prompted increasing attention towards its forwarding performance. While software optimization schemes have limitations in meeting the expanding network capacity demands [11, 12, 15, 17, 18], hardware offloading architectures leveraging SoC, FPGA, and ASIC have been proposed to transfer the match-action workload [1, 3, 6, 7, 13, 16], addressing the growing need for network capacity.
Xing Li 0007, Xiaochong Jiang, Lilong Chen, Tianyu Xu 0007, Chao Xu 0017, Longbiao Xiao, Fengmin Shi, Yi Wang 0004, Taotao Wu, Yilong Lv, Hangfeng Gao, Yisong Qiao, Hongwei Ding 0004, Yijian Dong, Chengkun Wei, Shunmin Zhu, Wenzhi Chen
SIGCOMM20
2023 Achelous: Enabling Programmability, Elasticity, and Reliability in Hyperscale Cloud Networks
abstract
Cloud computing has witnessed tremendous growth, prompting enterprises to migrate to the cloud for reliable and on-demand computing. Within a single Virtual Private Cloud (VPC), the number of instances (such as VMs, bare metals, and containers) has reached millions, posing challenges related to supporting millions of instances with network location decoupling from the underlying hardware, high elastic performance, and high reliability. However, academic studies have primarily focused on specific issues like high-speed data plane and virtualized routing infrastructure, while existing industrial network technologies fail to adequately address these challenges.
Chengkun Wei, Xing Li 0007, Xiaochong Jiang, Tianyu Xu 0007, Taotao Wu, Chao Xu 0017, Yilong Lv, Haifeng Gao, Zeke Wang, Shunmin Zhu, Wenzhi Chen
SIGCOMM16
2023 How IoT Re-using Threatens Your Sensitive Data: Exploring the User-Data Disposal in Used IoT Devices
abstract
With the rapid technology evolution of the Internet of Things (IoT) and increasing user needs, IoT device re-using becomes more and more common nowadays. For instance, more than 300,000 used IoT devices are selling on Craigslist. During IoT re-using, sensitive data such as credentials and biometrics residing in these devices may face the risk of leakage if a user fails properly dispose of the data. Thus, a critical security concern is raised: do (or can) users properly dispose of the sensitive data in used IoT? To the best of our knowledge, it is still an unexplored problem that desires a systematic study.In this paper, we perform the first in-depth investigation on the user-data disposal of used IoT devices. Our investigation integrates multiple research methods to explore the status quo and the root causes of the user-data leakages with used IoT devices. First, we conduct a user study to investigate the user awareness and understanding of data disposal. Then, we conduct a large-scale analysis on 4,749 IoT firmware images to investigate user-data collection. Finally, we conduct a comprehensive empirical evaluation on 33 IoT devices to investigate the effectiveness of existing data disposal methods.Through the systematical investigation, we discover that IoT devices collect more sensitive data than users expect. Specifically, we detect 121,984 sensitive data collections in the tested firmware. Moreover, users usually do not or even cannot properly dispose of the sensitive data. Worse, due to the inherent characteristics of storage chips, 13.2% of the investigated firmware perform "shallow" deletion, which may allow adversaries to obtain sensitive data after data disposal. Given the large-scale IoT re-using, such leakage would cause a broad impact. We have reported our findings to world-leading companies. We hope our findings raise awareness of the failures of user-data disposal with IoT devices and promote the protection of users’ sensitive data in IoT devices.
Peiyu Liu 0003, Shouling Ji, Lirong Fu, Kangjie Lu, Xuhong Zhang 0002, Jingchang Qin, Wenhai Wang, Wenzhi Chen
SP8
2023 EduNER: a Chinese named entity recognition dataset for education research
Xu Li 0032, Chengkun Wei, Zhuoren Jiang, Wenlong Meng, Fan Ouyang, Wenzhi Chen
Neural Comput. Appl.7
2022 A robust scheme for copy detection of 3D object point clouds
Xuequan Lu, Wenzhi Chen
Neurocomputing3
2022 Subspace clustering by directly solving Discriminative K-means
Chenhui Gao, Wenzhi Chen, Feiping Nie 0001, Weizhong Yu, Feihu Yan
Knowl. Based Syst.2
2021 CPscan: Detecting Bugs Caused by Code Pruning in IoT Kernels
abstract
To reduce the development costs, IoT vendors tend to construct IoT kernels by customizing the Linux kernel. Code pruning is common in this customization process. However, due to the intrinsic complexity of the Linux kernel and the lack of long-term effective maintenance, IoT vendors may mistakenly delete necessary security operations in the pruning process, which leads to various bugs such as memory leakage and NULL pointer dereference. Yet detecting bugs caused by code pruning in IoT kernels is difficult. Specifically, (1) a significant structural change makes precisely locating the deleted security operations (DSO ) difficult, and (2) inferring the security impact of a DSO is not trivial since it requires complex semantic understanding, including the developing logic and the context of the corresponding IoT kernel.
Lirong Fu, Shouling Ji, Kangjie Lu, Peiyu Liu 0003, Xuhong Zhang 0002, Yuxuan Duan, Wenzhi Chen
CCS8
2021 IFIZZ: Deep-State and Efficient Fault-Scenario Generation to Test IoT Firmware
abstract
IoT devices are abnormally prone to diverse errors due to harsh environments and limited computational capabilities. As a result, correct error handling is critical in IoT. Implementing correct error handling is non-trivial, thus requiring extensive testing such as fuzzing. However, existing fuzzing cannot effectively test IoT error-handling code. First, errors typically represent corner cases, thus are hard to trigger. Second, testing error-handling code would frequently crash the execution, which prevents fuzzing from testing following deep error paths.In this paper, we propose IFIZZ, a new bug detection system specifically designed for testing error-handling code in Linux-based IoT firmware. IFIZZ first employs an automated binary-based approach to identify realistic runtime errors by analyzing errors and error conditions in closed-source IoT firmware. Then, IFIZZ employs state-aware and bounded error generation to reach deep error paths effectively. We implement and evaluate IFIZZ on 10 popular IoT firmware. The results show that IFIZZ can find many bugs hidden in deep error paths. Specifically, IFIZZ finds 109 critical bugs, 63 of which are even in widely used IoT libraries. IFIZZ also features high code coverage and efficiency, and covers 67.3% more error paths than normal execution. Meanwhile, the depth of error handling covered by IFIZZ is 7.3 times deeper than that covered by the state-of-the-art method. Furthermore, IFIZZ has been practically adopted and deployed in a worldwide leading IoT company. We will open-source IFIZZ to facilitate further research in this area.
Peiyu Liu 0003, Shouling Ji, Xuhong Zhang 0002, Qinming Dai, Kangjie Lu, Lirong Fu, Wenzhi Chen, Peng Cheng 0001, Wenhai Wang, Raheem A. Beyah
ASE7
2021 Deep Neural Network Based Noised Asian Speech Enhancement and Its Implementation on a Hearing Aid App
abstract
This article studies noised Asian speech enhancement based on the deep neural network (DNN) and its implementation on an app. We use the THCHS-30 speech dataset and the common noise dataset in daily life as training and testing data of the DNN. To stack the frequency data of multiple audio frames to improve the effect of speech enhancement, the system compares the best number of stacked frames during training and testing. At the same time, the influence of training rounds on the PESQ is compared, and the best number of rounds is obtained. On this basis, the best model is implemented on the hearing aid app, and the real-time performance of the device is tested. The experiment shows that based on the DNN, using an appropriate number of rounds for training and using an appropriate number of audio frames stacking to improve the speech enhancement effect, and transplanting this speech enhancement model to the hearing aid app, can effectively improve speech clarity and intelligibility within a reasonable time delay range.
Xiaoqian Fan, Wenzhi Chen, Quanfang Fan
ACM Trans. Asian Low Resour. Lang. Inf. Process.3
2021 Zweilous: A Decoupled and Flexible Memory Management Framework
abstract
Currently, with the booming growth of cloud computing, workloads from broad ranges of functions and demands are crammed into a single physical machine. They lay considerable stress on the need of evolution of the operating system underneath, especially the memory subsystem. Even enhancing large pages with main memory compression is not intuitively straightforward due to rigid rules imposed by the state-of-the-art manager Buddy System from the beginning of the design. To relieve the aforementioned problems and provide broader design space for system designers, we propose Zweilous, a clean slate physical memory management framework. It is self-contained, highly decoupled, and thus can co-exist with the vanilla memory manager. Separate self-contained metadata/functions guarantee a flexible extension with little modification to current frameworks. To show it is easy to add enhanced functions that accelerate the evolution of the memory management subsystem, we implement Hzmem, a new large page memory manager redesign enhanced with the function of main memory compression. Our method achieves competitive performance compared with native and virtualized large page support, effective memory size increased and fewer impacts on other parts of the operating system.
Guoxi Li, Wenzhi Chen, Yang Xiang 0001
IEEE Trans. Computers2
2021 OB-WSPES: A Uniform Evaluation System for Obfuscation-Based Web Search Privacy
abstract
Web search queries reveal extensive sensitive information about users’ interests and preferences to the search engines and eavesdroppers. Obfuscation-based private web search solutions automatically generate dummy queries and send the obfuscated queries to the search engine to hide users’ search intentions. Despite many obfuscation methods and tools have been developed, there is no practical system for evaluating their utility performance and the vulnerability against modern privacy attacks. In this article, we propose and develop OB-WSPES, a uniform evaluation system for obfuscation-based web search privacy, which allows researchers to conduct fair analysis and evaluation of existing or newly developed web search privacy protection/attack techniques. Leveraging OB-WSPES, we model the obfuscation activities and systematically implement and evaluate five obfuscation schemes and 10 modern web search attacks on the public AOL dataset. Our results demonstrate that, counter-intuitively, adding more fake queries to a user’s real data does not necessarily yield better privacy. The query utility of obfuscated queries declines with the increasing amount of dummy queries, while the application utility does not. We discuss the experimental results and point out the four important factors that affect the web search privacy and utility. Further, we propose possible directions for future research.
Chengkun Wei, Qinchen Gu, Shouling Ji, Wenzhi Chen, Zonghui Wang, Raheem A. Beyah
IEEE Trans. Dependable Secur. Comput.4
2020 Understanding the Security Risks of Docker Hub
Peiyu Liu 0003, Shouling Ji, Lirong Fu, Kangjie Lu, Xuhong Zhang 0002, Wei-Han Lee, Wenzhi Chen, Raheem A. Beyah
ESORICS (1)8
2020 Smart VM co-scheduling with the precise prediction of performance characteristics
Yuxia Cheng, Wenzhi Chen, Zonghui Wang, Zhongxian Tang, Yang Xiang 0001
Future Gener. Comput. Syst.2
2020 AsgLDP: Collecting and Generating Decentralized Attributed Graphs With Local Differential Privacy
abstract
A large amount of valuable information resides in a decentralized attributed social graph, where each user locally maintains a limited view of the graph. However, there exists a conflicting requirement between publishing an attributed social graph and protecting the privacy of sensitive information contained in each user's local data. In this paper, we aim to collect and generate attributed social graphs in a decentralized manner while providing local differential privacy (LDP) for the collected data. Existing LDP-based synthetic graph generation methods either fail to preserve important graph properties (such as modularity and clustering coefficient) due to excessive noise injection or are unable to process attribute data, thus limiting their adoption and applicability. To overcome these weaknesses, we propose AsgLDP, a novel technique to generate privacy-preserving attributed graph data while satisfying LDP. AsgLDP preserves various graph properties through carefully designing the injected noise and estimating the joint distribution of attribute data. There are two key steps in AsgLDP: 1) collecting and generating graph data while satisfying LDP, and 2) optimizing the privacy-utility tradeoff of the generated data while preserving general graph properties such as the degree distribution, community structure and attribute distribution. Through theoretical analysis as well as experiments over 6 real-world datasets, we demonstrate the effectiveness of AsgLDP in preserving general graph properties such as degree distribution, community structure and attributed community search, while rigorously satisfying LDP. We also show that AsgLDP achieves a superior balance between utility and privacy as compared to the state-of-the-art approaches.
Chengkun Wei, Shouling Ji, Changchang Liu, Wenzhi Chen, Ting Wang 0006
IEEE Trans. Inf. Forensics Secur.4
2019 Exit-Less Hypercall: Asynchronous System Calls in Virtualized Processes
Guoxi Li, Wenhai Lin, Wenzhi Chen
ICA3PP (1)3
2019 Prudent Practices for Designing Virtual Desktop Experiments
abstract
Virtual desktop technology aims at accessing a remote desktop by endpoint hardware.Great attention has been increasingly paid to virtual desktop since it can increase the utilization of computing resources and provide more flexible accesses.However, researchers have not yet come up with a comprehensive set of rigorous standards of experimental design and implementation in this field.Therefore, it is difficult to conduct prudent experiments, which is correct, real, and transparent.In this paper, we assess the experimental evaluations of recently published papers on desktop virtualization.We observe that most works can be further improved, due to the unsuitable experimental environment and the lack of descriptions of experimental settings.In this paper, in order to help researchers, reviewers, and readers, we propose several guidelines for designing correct, real, and transparent desktop virtualization experiment.
Peiyu Liu 0003, Wenzhi Chen, Zonghui Wang, Lirong Fu
SEKE2
2019 3D articulated skeleton extraction using a single consumer-grade depth camera
Xuequan Lu, Zhigang Deng 0001, Jun Luo 0001, Wenzhi Chen, Sai-Kit Yeung, Ying He 0001
Comput. Vis. Image Underst.4
2019 An online learned hough forest model based on improved multi-feature fusion matching for multi-object tracking
Wenzhi Chen
Multim. Tools Appl.2
2018 Unsupervised Articulated Skeleton Extraction From Point Set Sequences Captured by a Single Depth Camera
abstract
How to robustly and accurately extract articulated skeletons from point set sequences captured by a single consumer-grade depth camera still remains to be an unresolved challenge to date. To address this issue, we propose a novel, unsupervised approach consisting of three contributions (steps): (i) a non-rigid point set registration algorithm to first build one-to-one point correspondences among the frames of a sequence; (ii) a skeletal structure extraction algorithm to generate a skeleton with reasonable numbers of joints and bones; (iii) a skeleton joints estimation algorithm to achieve accurate joints. At the end, our method can produce a quality articulated skeleton from a single 3D point sequence corrupted with noise and outliers. The experimental results show that our approach soundly outperforms state of the art techniques, in terms of both visual quality and accuracy.
Xuequan Lu, Honghua Chen, Sai-Kit Yeung, Zhigang Deng 0001, Wenzhi Chen
AAAI5
2018 Towards a multilayered permission-based access control for extending Android security
abstract
Summary This paper discusses security issues on the user equipment, which is the “last mile” of social networks. One of the main Achilles' heel of social networks is not the organization of networks themselves, but the user devices, typically Android ones. The existing system of privileges makes it easy to infiltrate the network via applications installed on users' devices. Conventional signature‐based and static analysis methods are vulnerable. Access to privacy‐ and security‐relevant parts of the application programming interface is controlled by the corresponding permission in a manifest file. While requesting access to permissions, it may offer opportunities to malicious codes, which will cause security issues. Few works among permission analysis, however, pay attention to the prevention of permission leakage on both hardware and software frameworks. In this paper we tackle the challenge of providing our multilayered permission‐based security extension scheme on Android platforms. We propose a usage and access control model and an effective method of preventing permission leakage based on ARM TrustZone security extension mechanism. In contrast to previous work, the proposed security architecture provides a permission‐based mandatory access control on Android middleware, Linux kernel, and hardware layers. The evaluation results demonstrate the effectiveness of the proposed scheme in mitigating permission leakage vulnerabilities.
Liehui Jiang, Wenzhi Chen, Hongqi He, Shuiqiao Yang, Wei Liu 0006
Concurr. Comput. Pract. Exp.3
2018 Efficient cache resource aggregation using adaptive multi-level exclusive caching policies
Yuxia Cheng, Yang Xiang 0001, Wenzhi Chen, Houcine Hassan, Abdulhameed Alelaiwi
Future Gener. Comput. Syst.3
2018 Distributed shielded execution for transmissible cyber threats analysis
Yuxia Cheng, Qing Wu 0008, Wenzhi Chen
J. Parallel Distributed Comput.3
2018 GPF: GMM-Inspired Feature-Preserving Point Set Filtering
abstract
Point set filtering, which aims at reconstructing noise-free point sets from their corresponding noisy inputs, is a fundamental problem in 3D geometry processing. The main challenge of point set filtering is to preserve geometric features of the underlying geometry while at the same time removing the noise. State-of-the-art point set filtering methods still struggle with this issue: some are not designed to recover sharp features, and others cannot well preserve geometric features, especially fine-scale features. In this paper, we propose a novel approach for robust feature-preserving point set filtering, inspired by the Gaussian Mixture Model (GMM). Taking a noisy point set and its filtered normals as input, our method can robustly reconstruct a high-quality point set which is both noise-free and feature-preserving. Various experiments show that our approach can soundly outperform the selected state-of-the-art methods, in terms of both filtering quality and reconstruction accuracy.
Xuequan Lu, Honghua Chen, Sai-Kit Yeung, Wenzhi Chen, Matthias Zwicker
IEEE Trans. Vis. Comput. Graph.5
2017 Hzmem: New Huge Page Allocator with Main Memory Compression
Guoxi Li, Wenzhi Chen, Kui Su, Zhongyong Lu, Zonghui Wang
ICA3PP2
2017 Robust mesh denoising via vertex pre-filtering and L1-median normal filtering
Xuequan Lu, Wenzhi Chen, Scott Schaefer
Comput. Aided Geom. Des.2
2017 MBSA: a lightweight and flexible storage architecture for virtual machines
abstract
Summary With the advantages of extremely high access speed, low energy consumption, nonvolatility, and byte addressability, nonvolatile memory (NVM) device has already been setting off a revolution in storage field. Conventional storage architecture needs to be optimized or even redesigned from scratch to fully explore the performance potential of NVM device. However, most previous NVM‐related works only explore its low access latency and low energy consumption. Few works have been done to explore the appropriate way to use NVM device for improving virtual machine's storage performance. In this paper, we comprehensively evaluate and analyze conventional virtual machine's storage architecture. We find that, even with cutting‐edge optimization technologies, virtual machine can only achieve 30% of NVM device's original performance. Based on this observation, we propose a memory bus–based storage architecture, which we named MBSA. Memory bus–based storage architecture can greatly shorten the length of virtual machine's storage input/output stack and improve NVM device's use flexibility. In addition, an efficient wear‐leveling algorithm is proposed to prolong NVM device's lifespan. To evaluate the new architecture, we implement it as well as the wear‐leveling algorithm on real hardware and software platform. Experimental results show that MBSA can provide a big performance improvement, about 2.55X, and the wear‐leveling algorithm can efficiently balance write operations on NVM device with a negligible performance overhead (no more than 3%).
Wenzhi Chen, Zhongyong Lu, Yu Zhang 0036, Mohammad Mehedi Hassan, Abdulhameed Alelaiwi, Yang Xiang 0001
Concurr. Comput. Pract. Exp.2
2017 Precise contention-aware performance prediction on virtualized multicore system
Yuxia Cheng, Wenzhi Chen, Zonghui Wang, Yang Xiang 0001
J. Syst. Archit.2
2017 Adaptive Multimedia Data Forwarding for Privacy Preservation in Vehicular Ad-Hoc Networks
abstract
Vehicular ad-hoc networks (VANETs) have drawn much attention of researchers. The vehicles in VANETs frequently join and leave the networks, and therefore restructure the network dynamically and automatically. Forwarded messages in vehicular ad-hoc networks are primarily multimedia data, including structured data, plain text, sound, and video, which require access control with efficient privacy preservation. Ciphertext-policy attribute-based encryption (CP-ABE) is adopted to meet the requirements. However, solutions based on traditional CP-ABE suffer from challenges of the limited computational resources on-board units equipped in the vehicles, especially for the complex policies of encryption and decryption. In this paper, we propose a CP-ABE delegation scheme, which allows road side units (RSUs) to perform most of the computation, for the purpose of improving the decryption efficiency of the vehicles. By using decision tree to jointly optimize multiple factors, such as the distance from RSU, the communication and computational cost, the CP-ABE delegation scheme is adaptively activated based on the estimation of various vehicles decryption overhead. Experimental results thoroughly demonstrate that our scheme is effective and efficient for multimedia data forwarding in vehicular ad-hoc networks with privacy preservation.
Yingjie Xia, Wenzhi Chen, Xuejiao Liu 0002, Xuelong Li 0001, Yang Xiang 0001
IEEE Trans. Intell. Transp. Syst.2
2017 TC-Release++: An Efficient Timestamp-Based Coherence Protocol for Many-Core Architectures
abstract
As we enter the era of many-core, providing the shared memory abstraction through cache coherence has become progressively difficult. The standard directory-based coherence does not scale well with increasing core count. Timestamp-based hardware coherence protocols introduced recently offer an attractive alternative solution. This paper proposes a timestamp-based coherence protocol, called TC-Release++, that efficiently supports cache coherence in large-scale systems. Our approach is inspired by TC-Weak, a recently proposed timestamp-based coherence protocol targeting GPU architectures. We first design TC-Release in an attempt to straightforwardly port TC-Weak to general-purpose many-cores. But re-purposing TC-Weak for general-purpose many-core architectures is challenging due to significant differences both in architecture and the programming model. Indeed the performance of TC-Release turns out to be worse than conventional directory protocols. We overcome the limitations and overheads of TC-Release by exploiting simple hardware support to eliminate frequent memory stalls, and an optimized lifetime prediction mechanism to improve cache performance. The resulting optimized coherence protocol TC-Release++is highly scalable (storage scales logarithmically with core count) and shows better performance (3.0 percent) and comparable network traffic (within 1.3 percent) relative to the baseline MESI directory protocol. We use Murphi to formally verify that TC-Release++is error-free and imposes small verification cost.
Yuan Yao 0006, Wenzhi Chen, Tulika Mitra, Yang Xiang 0001
IEEE Trans. Parallel Distributed Syst.2
2016 Efficient Timestamp-Based Cache Coherence Protocol for Many-Core Architectures
abstract
As we enter the era of many-core, providing the shared memory abstraction through cache coherence has become progressively difficult. The de-facto standard directory-based cache coherence has been extensively studied; but it does not scale well with increasing core count. Timestamp-based hardware coherence protocols introduced recently offer an attractive alternative solution. In this paper, we propose a timestamp-based coherence protocol, called TC-Release++, that addresses the scalability issues of efficiently supporting cache coherence in large-scale systems.
Yuan Yao 0006, Zhiguo Ge, Tulika Mitra, Wenzhi Chen, Naxin Zhang
ICS5
2016 Efficient consolidation-aware VCPU scheduling on multicore virtualization platform
Yuxia Cheng, Wenzhi Chen, Qinming He, Yang Xiang 0001, Mohammad Mehedi Hassan, Abdulhameed Alelaiwi
Future Gener. Comput. Syst.3
2016 SEMD: Secure and efficient message dissemination with policy enforcement in VANET
Xuejiao Liu 0002, Yingjie Xia, Wenzhi Chen, Yang Xiang 0001, Mohammad Mehedi Hassan, Abdulhameed Alelaiwi
J. Comput. Syst. Sci.3
2016 A Robust Scheme for Feature-Preserving Mesh Denoising
abstract
In recent years researchers have made noticeable progresses in mesh denoising, that is, recovering high-quality 3D models from meshes corrupted with noise (raw or synthetic). Nevertheless, these state of the art approaches still fall short for robustly handling various noisy 3D models. The main technical challenge of robust mesh denoising is to remove noise while maximally preserving geometric features. In particular, this issue becomes more difficult for models with considerable amount of noise. In this paper we present a novel scheme for robust feature-preserving mesh denoising. Given a noisy mesh input, our method first estimates an initial mesh, then performs feature detection, identification and connection, and finally, iteratively updates vertex positions based on the constructed feature edges. Through many experiments, we show that our approach can robustly and effectively denoise various input mesh models with synthetic noise or raw scanned noise. The qualitative and quantitative comparisons between our method and the selected state of the art methods also show that our approach can noticeably outperform them in terms of both quality and robustness.
Xuequan Lu, Zhigang Deng 0001, Wenzhi Chen
IEEE Trans. Vis. Comput. Graph.3
2015 SelectDirectory: a selective directory for cache coherence in many-core architectures
Yuan Yao 0006, Zhiguo Ge, Tulika Mitra, Wenzhi Chen, Naxin Zhang
DATE5
2015 Fusion-Cache: A Refactored Content-Aware Host-Side SSD Cache
Wenzhi Chen, Zhongyong Lu
ICA3PP (2)2
2015 Affinity and Conflict-Aware Placement of Virtual Machines in Heterogeneous Data Centers
abstract
Virtual machine placement (VMP) problem has been a key issue in IaaS/PaaS cloud infrastructures. Many recent works on VMP prove that inter-VM relations such as memory share, traffic dependency and resource competition should be seriously considered to save energy, increase the performance of infrastructure, reduce service level agreement violation rates and provide better administrative capabilities to the cloud provider. However, most existing works consider the inter-VM relations without taking the heterogeneity of cloud data centers into account. In practice, heterogeneous physical machines (PM) in a heterogeneous data center are often partitioned into logical groups for load balancing and specific services, cloud users always assigned their VMs with specific PM requirements, which make the inter-VM relations far more complex. In this paper, we propose an efficient solution for VMP with inter-VM relation constraints in a heterogeneous data center. The experimental results prove that our solution can efficiently solve the complex problem with an acceptable runtime.
Kui Su, Lei Xu 0039, Wenzhi Chen, Zonghui Wang
ISADS4
2015 AMC: an adaptive multi-level cache algorithm in hybrid storage systems
abstract
Summary Hybrid storage systems that consist of flash‐based solid state drives (SSDs) and traditional disks are now widely used. In hybrid storage systems, there exists a two‐level cache hierarchy that regard dynamic random access memory (DRAM) as the first level cache and SSD as the second level cache for disk storage. However, this two‐level cache hierarchy typically uses independent cache replacement policies for each level, which makes cache resource management inefficient and reduces system performance. In this paper, we propose a novel adaptive multi‐level cache (AMC) replacement algorithm in hybrid storage systems. The AMC algorithm adaptively adjusts cache blocks between DRAM and SSD cache levels using an integrated solution. AMC uses combined selective promote and demote operations to dynamically determine the level in which the blocks are to be cached. In this manner, the AMC algorithm achieves multi‐level cache exclusiveness and makes cache resource management more efficient. By using real‐life storage traces, our evaluation shows the proposed algorithm improves hybrid multi‐level cache performance and also increases the SSD lifetime compared with traditional multi‐level cache replacement algorithms. Copyright © 2015 John Wiley & Sons, Ltd.
Yuxia Cheng, Wenzhi Chen, Zonghui Wang, Xinjie Yu, Yang Xiang 0001
Concurr. Comput. Pract. Exp.2
2015 Evaluation of semi-supervised learning method on action recognition
Haoquan Shen, Yan Yan 0002, Nicolas Ballas, Wenzhi Chen
Multim. Tools Appl.5
2015 A Lightweight Virtualization Solution for Android Devices
abstract
Mobile virtualization has emerged fairly recently and is considered a valuable way to mitigate security risks on Android devices. However, major challenges in mobile virtualization include runtime, hardware, resource overhead, and compatibility. In this paper, we propose a lightweight Android virtualization solution named Condroid, which is based on container technology. Condroid utilizes resource isolation based on namespaces feature and resource control based on cgroups feature. By leveraging them, Condroid can host multiple independent Android virtual machines on a single kernel to support mutilple Android containers. Furthermore, our implementation presents both a system service sharing mechanism to reduce memory utilization and a filesystem sharing mechanism to reduce storage usage. The evaluation results on Google Nexus 5 demonstrate that Condroid is feasible in terms of runtime, hardware resource overhead, and compatibility. Therefore, we find that Condroid has a higher performance than other virtualization solutions.
Wenzhi Chen, Lei Xu 0039, Guoxi Li, Yang Xiang 0001
IEEE Trans. Computers1
2014 DASH: A duplication-aware flash cache architecture in virtualization environment
abstract
With the rapid development of multi-core and multi-threading technologies, the performance gap between CPU and storage system is widening year by year, causing the storage system to be the bottleneck of the whole system performance. To alleviate this situation, flash memory has been used as the caching device of HDDs. On the other hand, cloud computing is becoming more and more popular and mature in industry field. As the key building block of it, virtualization technology allows several virtual machines (VMs) running on one single physical machine simultaneously, most of which usually run the same or similar operating systems and applications. In this scenario, flash cache will be occupied by many duplicate data blocks. However, existing flash cache architectures and replacement policies don't take this observation into consideration, which greatly limits the efficient use of the flash cache. In this paper, we propose a new duplication-aware flash cache architecture (DASH). In this architecture, flash cache is organized to cache only one copy of the duplicate data blocks, which can notably expand the effective cache capacity, making more I/O requests hit in the cache. Moreover, this architecture can reduce the amount of data written to flash cache, and thus the life span of flash device can be significantly prolonged. Experiments based on realistic applications show that, in some situations, our cache architecture can improve the cache hit ratio by 5 times, reduce the average I/O latency by 63% and eliminate flash cache writes by 81%.
Wenzhi Chen, Shuiqiao Yang, Zhongyong Lu, Zonghui Wang
ICPADS2
2014 AA-FVDM: An accident-avoidance full velocity difference model for animating realistic street-level traffic in rural scenes
abstract
ABSTRACT Most of existing traffic simulation efforts focus on urban regions with a coarse two‐dimensional representation; relatively few studies have been conducted to simulate realistic three‐dimensional traffic flows on a large, complex road web in rural scenes. In this paper, we present a novel agent‐based approach called accident‐avoidance full velocity difference model (abbreviated as AA‐FVDM) to simulate realistic street‐level rural traffics, on top of the existing FVDM. The main distinction between FVDM and AA‐FVDM is that FVDM cannot handle a critical real‐world traffic problem while AA‐FVDM settles this problem and retains the essence of FVDM. We also design a novel scheme to animate the lane‐changing maneuvering process (in particular, the execution course). Through numerous simulations, we demonstrate that besides addressing a previously unaddressed real‐world traffic problem, our AA‐FVDM method efficiently (in real time) simulates large‐scale traffic flows (tens of thousands of vehicles) with realistic, smooth effects. Furthermore, we validate our method using real‐world traffic data, and the validation results show that our method measurably outperforms state‐of‐the‐art traffic simulation methods.Copyright © 2013 John Wiley & Sons, Ltd.
Xuequan Lu, Wenzhi Chen, Zonghui Wang, Zhigang Deng 0001, Yangdong Ye
Comput. Animat. Virtual Worlds2
2014 A personality model for animating heterogeneous traffic behaviors
abstract
ABSTRACT How to automatically generate realistic and heterogeneous traffic behaviors has been a much needed yet challenging problem for numerous traffic simulation and urban planning applications. In this paper, we propose a novel approach to model heterogeneous traffic behaviors by adapting a well‐established personality trait model (i.e., Eysenck's PEN (psychoticism, extraversion and neuroticism) model) into widely used traffic simulation approaches. First, we collected a large amount of user feedback while users watch a variety of computer‐generated traffic simulation video clips. Then, we trained regression models to bridge low‐level traffic simulation parameters and high‐level perceived traffic behaviors (i.e., adjectives according to the PEN model and the three PEN traits). We also conducted an additional user study to validate the effectiveness and usefulness of our approach, in particular, high correlation coefficients and the Pearson values between users’ feedback and our model predictions prove the effectiveness of our approach. Furthermore, our approach can also produce interesting emergent traffic patterns including faster‐is‐slower effect and sticking‐in‐a‐pin‐wherever‐there‐is‐room effect. Copyright © 2014 John Wiley & Sons, Ltd.
Xuequan Lu, Zonghui Wang, Wenzhi Chen, Zhigang Deng 0001
Comput. Animat. Virtual Worlds4
2013 AAGA: Affinity-Aware Grouping for Allocation of Virtual Machines
abstract
Virtualization technology enables various application services to be distributed and encapsulated within virtual machines (VMs), which are dynamically allocated to physical machines (PMs) in cloud computing environments. However, in many existing virtualized systems, the limited network bandwidth often becomes a bottleneck resource, leading to the intensification of network competition and the performance degradation for communication or data intensive applications. Aiming at reducing communication overheads and improving the application performance, in this paper, we propose an Affinity-Aware Grouping method for Allocation of VMs (AAGA). Firstly, we identity and model the problem of affinity-aware grouping-based allocation for virtual machines, and propose a detailed grouping method based on which a heuristic bin packing algorithm is used to deploy VM groups into PMs. In order to demonstrate the effectiveness of AAGA, we create multiple real virtual clusters (multi-VCs) with 56 VMs running multi-VM applications and compare application performance with Non-Affinity-aware Grouping-based Allocation methods (NAGA). Experimental results show that AAGA achieves better performance than NAGA.
Jianhai Chen, Kevin Chiew, Deshi Ye, Liangwei Zhu, Wenzhi Chen
AINA5
2013 A User-Level NUMA-Aware Scheduler for Optimizing Virtual Machine Performance
Yuxia Cheng, Wenzhi Chen
APPT2
2013 FPGA based hardware-software co-designed dynamic binary translation system
abstract
Binary translation is used to allow applications of one instruction set architecture (ISA) to run on another, thereby maintaining the binary level compatibility across ISAs. Conventional software binary translation systems suffer performance loss because of architectural heterogeneity amongst ISAs, control flow translation and context switches. In this paper, we propose an FPGA based hardware-software co-designed dynamic binary translation (DBT) system, which moderates these issues at a low level of hardware cost. In our DBT system, we propose a MIPS condition code flags register and a modest ISA extension to bridge the architectural gap, a hardware address mapping mechanism to accelerate the handling of control flow instructions, and a scratchpad memory to reduce performance loss during context switches. We implement the system on Xilinx XC5VLX110T. Quantitative experiments reveal that the overall performance improvement is 56.1% over the baseline configuration, with only extra 1.4% of slices and 5.4% of BRAMs of Xilinx XC5VLX110T occupied.
Yuan Yao 0006, Zhongyong Lu, Qingsong Shi, Wenzhi Chen
FPL4
2012 Smart Ring: A Model of Node Failure Detection in High Available Cloud Data Center
Lei Xu 0039, Wenzhi Chen, Zonghui Wang, Huafei Ni
NPC2
2010 Improving host swapping using adaptive prefetching and paging notifier
abstract
In a virtualized system, the hypervisor may be forced to reclaim memory by swapping out pages of guest operating systems (OSes) when the regular memory balancing mechanisms, such as page sharing and ballooning, fail to revoke enough memory for reallocation purpose, which always leads to serious performance degradation. In this paper, we introduce Adaptive Swap Prefetcher (ASP) and Host Swapping Notifier (HSN), the effective and lightweight solutions to gracefully reduce the degradation in system performance when host swapping is triggered. ASP smartly prefetches more pages from the host swap file as long as the good spatial locality persists so as to reduce disk transfers. The guest OS will be notified by HSN when the hypervisor evicts pages, which then hides those pages from its memory reclamation routines to eliminate unnecessary guest swapping and to prevent the occurrence of double paging anomaly. Currently ASP and HSN are implemented in KVM, experimental results show that guest performance can be improved by a factory of 1.4x and 2x respectively using ASP and HSN.
Wenzhi Chen, Huijun Chen, Xiaoqin Chen, Dapeng Huang
HPDC1
2010 L4RW: Laziness-based Realistic Real-time Responsive Rebalance in Walking
abstract
Abstract We present a novel L4RW (Laziness‐based Realistic Real‐time Responsive Rebalance in Walking) technique to synthesize 4RW animations under unexpected external perturbations with minimal locomotion effort. We first devise a lazy dynamic rebalance model, which specifies the dynamic balance conditions, defines the rebalance effort, and selects the suitable rebalance strategy automatically using the laziness law after an unexpected perturbation. Based on the model, L4RW searches over a motion capture (mocap) database for an appropriate motion segment to follow, and the transition‐to motions is generated by interpolating the active response dynamic motion. A support vector machine (SVM) based training, classification, and predication algorithm is applied to reduce the search space, and it is trained offline only once. Our algorithm classifies the mocap database into many rebalance strategy‐specified subsets and then online predicts responsive motions in the subset according to the selected strategy. The rebalance effort, the ‘extrapolated center of mass’ (XCoM) and environment constraints are selected as feature attributes for the SVM feature vector. Furthermore, the subset's segments are sorted through the rebalance effort, then our algorithm searches for an acceptable segment starting from the least‐effort segment. Compared with previous methods, our search increases speed by over two orders of magnitude, and our algorithm creates more realistic and smooth 4RW animation.
Huansen Li, Pei Lv, Wenzhi Chen, Gengdai Liu
Comput. Graph. Forum4
2010 Handling occlusions in video-based augmented reality using depth information
abstract
Abstract Augmented Reality (AR) composes virtual objects with real scenes in a mixed environment where human–computer interaction has more semantic meanings. To seamlessly merge virtual objects with real scenes, correct occlusion handling is a significant challenge. We present an approach to separate occluded objects in multiple layers by utilizing depth, color, and neighborhood information. Scene depth is obtained by stereo cameras and two Gaussian local kernels are used to represent color, spatial smoothness. These three cues areintelligentlyfused in a probability framework, where the occlusion information can be safely estimated. We apply our method to handle occlusions in video‐based AR where virtual objects are simply overlapped on real scenes. Experiment results show the approach can correctly register virtual and real objects in different depth layers, and provide a spatial‐awareness interaction environment. Copyright © 2009 John Wiley & Sons, Ltd.
Jiejie Zhu, Wenzhi Chen
Comput. Animat. Virtual Worlds4
2009 Real time falling animation with active and protective responses
Wenzhi Chen, Gengdai Liu, Bing Tang
Vis. Comput.3
2008 SeVMM: VMM-Based Security Control Model
abstract
The security problem became more severe since the security requirement of different applications may conflict with the others in distributed application or grid computing. Virtualization technology can improvethe system's security, but did not satisfy the requirement of virtual resource sharing and inner-domain communication in distributed services and grid computing. By dividing the virtual resources into sharing virtual resources and normal ones, SeVMM provided secure mechanism for inter-domain communication control, which formed the base of multi-level security control model for virtual machine monitors, operating systems and applications. Case study and application showed that SeVMM improved the system's security without causing significant performance penalty.
Wenzhi Chen
CW1