EDBT 2026 Demo / reviewers in the wild / expert
Zheng Wang 0001
dblp:w/ZhengWang1
· DBLP profile ↗
133ranked-venue papers
8as first author
74since 2021 · last 2026
0000-0001-6157-0662ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 55 · 5 first-author · 29 since 2021Software engineering, systems software and programming languages · 24 · 1 first-author · 12 since 2021Artificial intelligence and machine learning · 20 · 10 since 2021Computer networks · 18 · 1 first-author · 12 since 2021Security and privacy · 12 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 since 2021Databases, data management, data science and information retrieval · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Lifting Optimized Binaries to Canonical Compiler IR via Structure-Aware Retrieval and Iterative VerificationabstractLifting stripped and highly optimized binaries to the canonical compiler intermediate representation (IR) enables program analysis when source code is unavailable.However, compiler optimizations severely distort controlflow and data-flow structure, making existing rule-based and LLM-based decompilation approaches brittle.We present BRIDGE, a system that reliably lifts optimized binaries to analysis-friendly compiler IR.BRIDGE combines control-flow-aware retrieval-augmented generation with feedback-driven verification.It uses pseudo-probe instrumentation to align optimized binary fragments with normalized IR semantics, and then employs an iterative refinement loop guided by static analysis and runtime feedback to improve executability and semantic consistency.We evaluate BRIDGE on HumanEval-Decompile and MBPP, lifting x86-64 and ARM64 binaries to LLVM IR.BRIDGE outperforms seven baselines, achieving an average of over 30% higher re-executability than the strongest general-purpose LLM baseline.Void func (){ PROBE(1); If else branch …… PROBE(2); PROBE(3); for Loop … PROBE(4);} Xiaoao Zhu, Jie Ren 0007, Zhiqiang Li 0003, Jie Zheng 0005, Zhanyong Tang, Zheng Wang 0001 |
ACL (1) | 6 |
| 2026 | Interpreter Memory Safety via Differential Fuzzing with a CHERI on TopabstractMemory safety is a critical issue in embedded systems. Although high-level languages like MicroPython simplify IoT development, their C-based runtimes remain vulnerable to memory errors triggered by Python code or native extensions. The CHERI (Capability Hardware Enhanced RISC Instructions) architecture offers hardware-enforced memory safety, but its effectiveness for exposing latent bugs in real-world interpreters has not yet been fully explored. We present diffCHERI:FruitFly, a novel differential testing framework for systematically uncovering memory defects in MicroPython across conventional (x86/ARM) and CHERI-enabled (Arm Morello) platforms. We mine historic vulnerabilities from diverse Python runtimes to extract recurring stress patterns, then use a large language model to generate new test programs, and apply Concrete Syntax Tree (CST) mutation to diversify inputs. Huanting Wang, Jeremy Singer, Zheng Wang 0001 |
ISMM | 4 |
| 2026 | KVSwap: Disk-aware KV Cache Offloading for Long-Context On-device InferenceabstractLanguage models (LMs) underpin emerging mobile and embedded AI applications like meeting and video summarization and document analysis, which often require processing multiple long-context inputs. Running an LM locally on-device improves privacy, enables offline use, and reduces cost, but long-context inference quickly hits a \emph{memory capacity wall} as the key-value (KV) cache grows linearly with context length and batch size. Existing KV-cache offloading schemes are designed to transfer cache data from GPU memory to CPU memory; however, they are not suitable for embedded and mobile systems, where the CPU and GPU (or NPU) typically share a unified memory and the non-volatile secondary storage (disk) offers limited I/O bandwidth. We present KVSwap, a software framework tailored for local devices that achieves high memory efficiency while effectively leveraging disk storage. KVSwap stores the full cache on disk, uses highly compact in-memory metadata to predict which entries to preload, overlaps computation with hardware-aware disk access, and orchestrates read patterns to match storage device characteristics. Our evaluation shows that across representative LMs and storage types, KVSwap delivers higher throughput under tight memory budgets while maintaining generation quality over existing KV cache offloading schemes. Chunwei Xia, Zheng Wang 0001 |
MobiSys | 3 |
| 2026 | SuperEar: Eavesdropping on Mobile Voice Calls via Stealthy Acoustic Metamaterials
Zhiyuan Ning 0003, Zhanyong Tang, Juan He 0007, Weizhi Meng 0001, Yuntian Chen, Jie Zhang 0028, Zheng Wang 0001 |
WWW | 7 |
| 2026 | LEGO-compiler: enhancing neural compilation through translation composability
Shuoming Zhang, Qiuchu Yu, Chunwei Xia, Zheng Wang 0001, Yunji Chen, Xiaobing Feng 0002, Huimin Cui |
CCF Trans. High Perform. Comput. | 5 |
| 2026 | The new compiler stack: a survey on the synergy of LLMs and compilers
Shuoming Zhang, Qiuchu Yu, Chunwei Xia, Zheng Wang 0001, Xiaobing Feng 0002, Huimin Cui |
CCF Trans. High Perform. Comput. | 5 |
| 2026 | Condition Number Analysis for MIMO-OTFS Communication SystemsabstractOrthogonal Time-Frequency Space (OTFS) modulation is an innovative modulation technique that operates in the two-dimensional delay-Doppler (DD) domain. It is specifically designed for high Doppler scenarios, where the channel can be transformed into an almost non-fading channel for DD domain symbol transmission. In multiple-input multiple-output (MIMO) OTFS systems, the DD domain input-output relation has considerable complexity, especially under fractional delays and Dopplers. The condition number is a key indicator of channel matrix quality as it reflects the sensitivity to noise and disturbance. In this paper, a novel Zak-OTFS modulation in the DD domain has been proposed as a competitive alternative to the conventional multi-carrier (MC) OTFS scheme, where MC-OTFS is implemented in two steps in the time-frequency (TF) domain. We focus on both MIMO-Zak-OTFS and MIMO-MC-OTFS schemes and study the condition numbers for the residual error covariance matrix and the channel matrix. We investigate the condition number under various system parameters, and the results indicate that OTFS outperforms MIMO-orthogonal frequency division multiplexing (MIMO-OFDM) in terms of condition number distributions. Meanwhile, the condition numbers of MIMO-Zak-OTFS matrices also outperform those of MIMO-MC-OTFS, while MIMO-Zak-OTFS shows a better BER performance compared to MIMO-MC-OTFS and MIMO-OFDM using both LMMSE and sphere decoders in the experimental results. Zheng Wang 0001, Bodong Shang |
IEEE Trans. Commun. | 1 |
| 2026 | DynEformer: A Unified Framework for Robust Workload Prediction Under Dynamic EnvironmentabstractWorkload prediction in multi-tenant edge cloud platforms (MT-ECP) is crucial for efficient application deployment and resource provisioning. However, the heterogeneous application patterns, variable infrastructure performance, and frequent deployments in MT-ECP pose significant challenges for accurate prediction. Existing clustering-based methods often incur excessive costs due to maintaining multiple data clusters and models, while end-to-end time-series prediction methods struggle with dynamic environments. To address these challenges, we perform a comprehensive analysis on a large-scale workload dataset in real-world MT-ECP and propose DynEformer, an end to-end framework with global pooling and static context aware ness, offering a unified workload prediction scheme for dynamic MT-ECP. Meticulously designed global pooling and information merging mechanisms can effectively identify and utilize global application patterns to drive local workload predictions. The integration of static content-aware mechanisms enhances model robustness in real-world scenarios. We also extend DynEformer's capabilities to Long-term workload forecasting (LTLF) and Long-period service (LPS) tasks. Experiments on six real-world datasets demonstrate that DynEformer achieves state-of-the-art performance, with a 32% relative improvement on nine baselines and a 52% improvement in application switching and new entity scenarios. Additional experiments on long-term prediction and online learning further confirm its effectiveness for LTLF and LPS tasks. Shaoyuan Huang, Zheng Wang 0001, Heng Zhang 0032, Xiaofei Wang 0001, Cheng Zhang 0019 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2026 | SynergyScale: Optimizing Offloading and Task Partitioning for Efficient Model Training
Jie Xu 0007, Zheng Wang 0001 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2026 | A Groupwise Add-Multiply-Shift-Accumulate Datapath for Efficient DNN AcceleratorsabstractMultiply–accumulate(MAC) units account for a large fraction of the power and area in modern deep neural network (DNN) accelerators. Although low-bitwidth quantization reduces hardware overhead, the high cost of multipliers remains a fundamental bottleneck in modern accelerator datapaths. This article proposes add–multiply–shift–accumulate (AMC), a groupwise arithmetic datapath that reduces multiplier count by sharing base multiplications across groups of neighboring weights and generating residual products using lightweight shift operation. To support efficient deployment, we design a compact residual encoding and buffer organization that allows AMC arrays to be constructed with minimal decoding and control overheads. While AMC can be directly applied to existing quantized models, we further introduce a lightweight residual-aware fine-tuning (RAF) procedure to increase AMC compatibility. We implement AMC-based accelerators in SystemVerilog and synthesize them in TSMC 28-nm CMOS technology across operating frequencies from 500MHz to 1GHz. At the compute unit level, AMC reduces arithmetic area by 39.5%–62.8% and dynamic power by 32.2%–60.3% compared with optimized baseline multipliers. When integrated into CNN and Vision Transformer accelerators, AMC achieves$1.34\times $–$18.90\times $higher area efficiency and up to$10.16\times $higher energy efficiency than prior designs while preserving baseline inference accuracy. Zhiwang Huo, Wenzhe Zhao 0001, Yuanchang Gong, Tian Xia 0008, Zheng Wang 0001, Pengju Ren |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2025 | Effects of Momentum in Implicit Bias of Gradient Flow for Diagonal Linear NetworksabstractThis paper targets on the regularization effect of momentum-based methods in regression settings and analyzes the popular diagonal linear networks to precisely characterize the implicit bias of continuous versions of heavy-ball (HB) and Nesterov's method of accelerated gradients (NAG). We show that, HB and NAG exhibit different implicit bias compared to GD for diagonal linear networks, which is different from the one for classic linear regression problem where momentum-based methods share the same implicit bias with GD. Specifically, the role of momentum in the implicit bias of GD is twofold: (a) HB and NAG induce extra initialization mitigation effects similar to SGD that are beneficial for generalization of sparse regression; (b) the implicit regularization effects of HB and NAG also depend on the initialization of gradients explicitly, which may not be benign for generalization. As a result, whether HB and NAG have better generalization properties than GD jointly depends on the aforementioned twofold effects determined by various parameters such as learning rate, momentum factor, and integral of gradients. Our findings highlight the potential beneficial role of momentum and can help understand its advantages in practice such as when it will lead to better generalization performance. Bochen Lyu, Zheng Wang 0001, Zhanxing Zhu |
AAAI | 3 |
| 2025 | Enhancing Deployment-Time Predictive Model Robustness for Code Analysis and OptimizationabstractSupervised machine learning techniques have shown promising results in code analysis and optimization problems. However, a learning-based solution can be brittle because minor changes in hardware or application workloads – such as facing a new CPU architecture or code pattern – may jeopardize decision accuracy, ultimately undermining model robustness. We introduce Prom, an open-source library to enhance the robustness and performance of predictive models against such changes during deployment. Prom achieves this by using statistical assessments to identify test samples prone to mispredictions and using feedback on these samples to improve a deployed model. We showcase Prom by applying it to 13 representative machine learning models across 5 code analysis and optimization tasks. Our extensive evaluation demonstrates that Prom can successfully identify an average of 96% (up to 100%) of mispredictions. By relabeling up to 5% of the Prom-identified samples through incremental learning, Prom can help a deployed model achieve a performance comparable to that attained during its model training phase. Huanting Wang, Patrick Lenihan, Zheng Wang 0001 |
CGO | 3 |
| 2025 | Dataflow-Guided Neuro-Symbolic Language Models for Type InferenceabstractLanguage Models (LMs) are increasingly used for type inference, aiding in error detection and software development.
Some real-world deployments of LMs require the model to run on local machines to safeguard the intellectual property of the source code. This setting often limits the size of the LMs that can be used. We present Nester, the first neuro-symbolic approach that enhances LMs for type inference by integrating symbolic learning without increasing model size. Nester breaks type inference into sub-tasks based on the data and control flow of the input code, encoding them as a modular high-level program. This program executes multi-step actions, such as evaluating expressions and analyzing conditional branches of the target code, combining static typing with LMs to infer potential types.
Evaluated on the ManyTypes4Py dataset in Python, Nester outperforms two state-of-the-art type inference methods (HiTyper and TypeGen), achieving 70.7\% Top-1 Exact Match, which is 18.3\% and 3.6\% higher than HiTyper and TypeGen, respectively. For complex type annotations like typing.Optional and typing.Union, Nester achieves 51.0\% and 16.7\%, surpassing TypeGen by 28.3\% and 5.8\%. Ge Li 0001, Yao Wan 0001, Hongyu Zhang 0002, Zhou Zhao 0001, Wenbin Jiang 0001, Xuanhua Shi, Hai Jin 0001, Zheng Wang 0001 |
ICML | 8 |
| 2025 | Optimizing Personalized Federated Learning Through Adaptive Layer-Wise LearningabstractReal-life deployment of federated Learning (FL) often faces non-IID data, which leads to poor accuracy and slow convergence. Personalized FL (pFL) tackles these issues by tailoring local models to individual data sources and using weighted aggregation methods for client-specific learning. However, existing pFL methods often fail to provide each local model with global knowledge on demand while maintaining low computational overhead. Additionally, local models tend to over-personalize their data during the training process, potentially dropping previously acquired global information. We propose FLAYER, a novel layer-wise learning method for pFL that optimizes local model personalization performance. FLAYER considers the different roles and learning abilities of neural network layers of individual local models. It incorporates global information for each local model as needed to initialize the local model cost-effectively. It then dynamically adjusts learning rates for each layer during local training, optimizing the personalized learning process for each local model while preserving global knowledge. Additionally, to enhance global representation in pFL, FLAYER selectively uploads parameters for global aggregation in a layer-wise manner. We evaluate FLAYER on four representative datasets in computer vision and natural language processing domains. Compared to eight state-of-the-art pFL methods, FLAYER improves the inference accuracy, on average, by 5.20% (up to 14.29%). Code is available at https://github.com/lancasterJie/FLAYER/. Weihang Chen, Jie Ren 0007, Zhiqiang Li 0003, Zheng Wang 0001 |
IJCAI | 5 |
| 2025 | Accelerating Tensor-Train Decomposition on Graph Neural NetworksabstractMemory footprint is a major concern when training graph neural networks (GNNs) on large graph data. Tensor-train decomposition (TTD) offers a potential solution by representing high-dimensional tensors with a set of smaller tensors, reducing memory overhead. However, existing TTD-based solutions for GNNs fail to reuse intermediate computation results and minimize memory data transfers to improve GNN performance. We introduce FALCON, a software framework to accelerate TTDbased GNN training. FALCON leverages the observation that a small subset of graph nodes with high edge degrees are frequently accessed, enabling the caching of intermediate results to reduce redundant computation and data transfers. Additionally, it incorporates multi-level graph partitioning and kernel optimization techniques to boost computational efficiency. We evaluated FALCON using three real-world datasets on three GPU platforms-NVIDIA 3090, 4090, and A100. Experimental results show that FALCON outperforms previous TTD-based frameworks, delivering a 1.3 to$8.17 \times$improvement in throughput while maintaining comparable or better efficiencies in memory footprint and model accuracy. Shenghao Qiu, Chunwei Xia, Zheng Wang 0001 |
IPDPS | 3 |
| 2025 | Leveraging Compilation Statistics for Compiler Phase OrderingabstractChoosing the optimal order and combination of compiler optimization passes - known as phase ordering - can enhance the performance of compiled binaries. However, existing approaches struggle to capture the subtle interaction between compiler passes and waste time on low-profitable pass sequences. We introduce CITROEN, a better approach for compiler phase ordering. CITROEN leverages pass-related compilation statistics to reject low-profitable compiler pass sequences to reduce the overhead of phase ordering search. It employs Bayesian optimization to navigate the search space, using compilation statistics instead of traditional tuning parameters to build an online cost model that provides both the performance prediction and the prediction uncertainty of compilation configurations. It dynamically allocates search iterations across source files to optimize search time in multi-file programs. We evaluate CITROEN by integrating it with the LLVM compiler and applying it to benchmarks from cBench and SPEC CPU 2017. CITROEN outperforms existing autotuning methods, discovering high-performing configurations quicker with fewer search iterations. Chunwei Xia, Zheng Wang 0001 |
IPDPS | 3 |
| 2025 | SecureMind: A Framework for Benchmarking Large Language Models in Memory Bug Detection and RepairabstractLarge language models (LLMs) hold great promise for automating software vulnerability detection and repair, but ensuring their correctness remains a challenge. While recent work has developed benchmarks for evaluating LLMs in bug detection and repair, existing studies rely on hand-crafted datasets that quickly become outdated. Moreover, systematic evaluation of advanced reasoning-based LLMs using chain-of-thought prompting for software security is lacking. We introduce SecureMind, an open-source framework for evaluating LLMs in vulnerability detection and repair, focusing on memory-related vulnerabilities. SecureMind provides a user-friendly Python interface for defining test plans, which automates data retrieval, preparation, and benchmarking across a wide range of metrics. Using SecureMind, we assess 10 representative LLMs, including 7 state-of-the-art reasoning models, on 16K test samples spanning 8 Common Weakness Enumeration (CWE) types related to memory safety violations. Our findings highlight the strengths and limitations of current LLMs in handling memory-related vulnerabilities. Huanting Wang, Dejice Jacob, David Kelly, Yehia El-khatib, Jeremy Singer, Zheng Wang 0001 |
ISMM | 6 |
| 2025 | Tuning LLM-based Code Optimization via Meta-Prompting: An Industrial PerspectiveabstractThere is a growing interest in leveraging multiple large language models (LLMs) for automated code optimization. However, industrial platforms deploying multiple LLMs face a critical challenge: prompts optimized for one LLM often fail with others, requiring expensive model-specific prompt engineering. This cross-model prompt engineering bottleneck severely limits the practical deployment of multi-LLM systems in production environments. We introduce Meta-Prompted Code Optimization (Mpco), a framework that automatically generates high-quality, task-specific prompts across diverse LLMs while maintaining industrial efficiency requirements. Mpco leverages meta-prompting to dynamically synthesize context-aware optimization prompts by integrating project metadata, task requirements, and LLM-specific contexts. It is an essential part of the ARTEMIS code optimization platform for automated validation and scaling.Our comprehensive evaluation on five real-world codebases with 366 hours of runtime benchmarking demonstrates Mpco’s effectiveness: it achieves overall performance improvements up to 19.06% with the best statistical rank across all systems compared to baseline methods. Analysis shows that 96% of the top-performing optimizations stem from meaningful edits. Through systematic ablation studies and meta-prompter sensitivity analysis, we identify that comprehensive context integration is essential for effective meta-prompting and that major LLMs can serve effectively as meta-prompters, providing actionable insights for industrial practitioners. Jingzhi Gong, Rafail Giavrimis, Paul Brookes, Vardan Voskanyan 0001, Fan Wu 0009, Mari Ashiga, Matthew Truscott, Michail Basios, Leslie Kanthan, Jie Xu 0007, Zheng Wang 0001 |
ASE | 11 |
| 2025 | MetaGuardian: Enhancing Voice Assistant Security through Advanced Acoustic MetamaterialsabstractVoice assistants (VAs) have become integral to daily life, yet their always-on microphones make them attractive targets for attacks that threaten user privacy and safety. We present MetaGuardian, the first system to leverage acoustic metamaterials to defend against three major classes of attacks for VAs - inaudible, adversarial, and laser-based - within a single, portable design. Unlike prior defenses, MetaGuardian can be seamlessly integrated into the enclosures of commercial smart devices, providing strong protection without requiring software modification, hardware redesign, or costly machine learning models. MetaGuardian leverages mutual impedance effects between metamaterial units to extend the protection range to 16–40 kHz, effectively blocking wideband inaudible attacks. It also employs a carefully designed coiled space structure to disrupt adversarial signals while preserving normal VA operations. Its universal design allows flexible adaptation to different devices, striking a balance between portability and protection effectiveness. In controlled evaluations, MetaGuardian achieves a high defense success rate across all attack types, offering a practical and reliable foundation for securing VAs on smart devices. Zhiyuan Ning 0003, Zheng Wang 0001, Zhanyong Tang |
MobiCom | 2 |
| 2025 | Scenario: User-Device Authentication on Smart IoTs Using Commodity RFIDabstractUser and device authentication are vital to the deployment of smart Internet of Things (IoT) devices. Unfortunately, achieving robust authentication on a diverse set of heterogeneous IoT devices remains an open problem. This paper presentsScenario, a generic authentication method to support user-device authentication on a wide range of IoT devices, using RFID-based wireless sensing.Scenarioonly requires attaching an RFID tag on the target device surface. It then uses the unique RFID signal characteristics introduced by the device material and user gestures to perform device and user authentication. We developed a prototype ofScenariousing commercial off-the-shelf devices and applied it to a multi-device smart environment. Experimental results show thatScenariois reliable, giving an average identification accuracy of 97.3% and 96.7% of the device and user authentication stages in diverse environments, respectively. Weiyuan Tong, Zhanyong Tang, Huanting Wang, Guixin Ye, Shuangjiao Zhai, Zheng Wang 0001 |
IEEE Trans. Dependable Secur. Comput. | 7 |
| 2025 | Accelerating Private Large Transformers Inference Through Fine-Grained Collaborative ComputationabstractHomomorphic encryption (HE) and secret sharing (SS) enable computations on encrypted data, providing significant privacy benefits for large transformer-based models (TBM) in sensitive sectors like medicine and finance. However, private TBM inference incurs significant costs due to the coarse-grained application of HE and SS. We present FASTLMPI, a new approach to accelerate private TBM inference through fine-grained computation optimization. Specifically, through the fine-grained co-design of homomorphic encryption and secret sharing, FASTLMPI achieves efficient protocols for matrix multiplication, SoftMax, LayerNorm, and GeLU. In addition, FASTLMPI introduces a precise segmented approximation technique for differentiable non-linear functions, improving its fitting accuracy while maintaining a low polynomial degree. Compared to solution BOLT (S&P’24), FASTLMPI shows a remarkable 25.1% to 55.3% decrease in runtime and an impressive 39.0% reduction in communication costs. Yuntian Chen, Zhanyong Tang, Tianpei Lu, Bingsheng Zhang, Zhiying Shi, Zheng Wang 0001 |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2025 | Trust Online Over-the-Air Computation for Wireless Federated LearningabstractUsing the wireless waveform superposition property, over-the-air computation (OAC) enables federated learning (FL) to achieve fast model aggregation. However, this computing paradigm is vulnerable to poisoning attacks due to the openness of a wireless channel over time, where malicious mobile devices can introduce cumulative errors for the global FL model in a time-varying wireless environment for each communication round. This article presents a trust online OAC (TO-OAC) scheme to minimize impacts on the global model introduced by malicious devices adjusting to dynamic attack and wireless channel fluctuations over time. TO-OAC achieves this by utilizing trustworthy security quantification of OAC for each FL training round. To optimize the cumulative training loss at the aggregation node with the long-term power and trust constraints of mobile devices, we propose a joint trust, power, and channel-aware algorithm to flexibly update local and global models in response to the dynamic changes in the wireless and secure environment. We analyze the performance limits for the aggregation of trust models, considering metrics for computation and communication through time. We then propose another trust online regularization over-the-air computation (TOR-OAC) as an improved version of the TO-OAC scheme to decrease convergence time while ensuring long-term trust and power limitation. Experimental results performed on real-life datasets show that the two proposed schemes (TO-OAC and TOR-OAC) outperform prior works, especially in noisy, time-varying wireless channels and malicious attacks. Mingjie Sun, Jie Zheng 0005, Hongyang Du 0001, Haijun Zhang 0001, Dusit Niyato, Jiawen Kang 0001, Jiacheng Wang 0001, Jie Ren 0007, Zheng Wang 0001 |
IEEE Trans. Mob. Comput. | 10 |
| 2024 | Optimizing Deep Learning Inference via Global Analysis and Tensor ExpressionsabstractOptimizing deep neural network (DNN) execution is important but becomes increasingly difficult as DNN complexity grows. Existing DNN compilers cannot effectively exploit optimization opportunities across operator boundaries, leaving room for improvement. To address this challenge, we present Souffle, an open-source compiler that optimizes DNN inference across operator boundaries. Souffle creates a global tensor dependency graph using tensor expressions, traces data flow and tensor information, and partitions the computation graph into subprograms based on dataflow analysis and resource constraints. Within a subprogram, Souffle performs local optimization via semantic-preserving transformations, finds an optimized program schedule, and improves instruction-level parallelism and data reuse. We evaluated Souffle using six representative DNN models on an NVIDIA A100 GPU. Experimental results show that Souffle consistently outperforms six state-of-the-art DNN optimizers by delivering a geometric mean speedup of up to 3.7× over TensorRT and 7.8× over Tensorflow XLA. Chunwei Xia, Qianqi Sun, Zheng Wang 0001, Yuan Wen, Xiaobing Feng 0002, Huimin Cui |
ASPLOS (1) | 4 |
| 2024 | Combining Structured Static Code Information and Dynamic Symbolic Traces for Software Vulnerability PredictionabstractDeep learning (DL) has emerged as a viable means for identifying software bugs and vulnerabilities. The success of DL relies on having a suitable representation of the problem domain. However, existing DL-based solutions for learning program representations have limitations - they either cannot capture the deep, precise program semantics or suffer from poor scalability. We present Concoction, the first DL system to learn program presentations by combining static source code information and dynamic program execution traces. Concoction employs unsupervised active learning techniques to determine a subset of important paths to collect dynamic symbolic execution traces. By implementing a focused symbolic execution solution, Concoction brings the benefits of static and dynamic code features while reducing the expensive symbolic execution overhead. We integrate Concoction with fuzzing techniques to detect function-level code vulnerabilities in C programs from 20 open-source projects. In 200 hours of automated concurrent test runs, Concoction has successfully uncovered vulnerabilities in all tested projects, identifying 54 unique vulnerabilities and yielding 37 new, unique CVE IDs. Concoction also significantly outperforms 16 prior methods by providing higher accuracy and lower false positive rates. Huanting Wang, Zhanyong Tang, Shin Hwei Tan, Jie Wang 0110, Hejun Fang, Chunwei Xia, Zheng Wang 0001 |
ICSE | 8 |
| 2024 | Seer: Proactive Revenue-Aware Scheduling for Live Streaming Services in Crowdsourced Cloud-Edge PlatformsabstractAs live streaming services skyrocket, Crowdsourced Cloud-edge service Platforms (CCPs) have surfaced as pivotal intermediaries catering to the mounting demand. Despite the role of stream scheduling to CCPs’ Quality of Service (QoS) and throughput, conventional optimization strategies struggle to enhancing CCPs’ revenue, primarily due to the intricate relationship between resource utilization and revenue. Additionally, the substantial scale of CCPs magnifies the difficulties of time-intensive scheduling. To tackle these challenges, we propose Seer, a proactive revenue-aware scheduling system for live streaming services in CCPs. The design of Seer is motivated by meticulous measurements of real-world CCPs environments, which allows us to achieve accurate revenue modeling and overcome three key obstacles that hinder the integration of prediction and optimal scheduling. Utilizing an innovative Preschedule-Execute-Re-schedule paradigm and flexible scheduling modes, Seer achieves efficient revenue-optimized scheduling in CCPs. Extensive evaluations demonstrate Seer’s superiority over competitors in terms of revenue, utilization, and anomaly penalty mitigation, boosting CCPs revenue by 147% and expediting scheduling 3.4× faster. Shaoyuan Huang, Zheng Wang 0001, Zhongtian Zhang, Heng Zhang 0032, Xiaofei Wang 0001 |
INFOCOM | 2 |
| 2024 | Optimizing General Matrix Multiplications on Modern Multi-core DSPsabstractGeneral Matrix Multiplication (GEMM) is a key subprogram in high-performance computing (HPC) and deep learning workloads. With the rising significance of power and energy consumption in HPC systems, accelerators based on Digital Signal Processors (DSPs) have been integrated into general-purpose HPC systems. Due to the architecture disparities, the GEMM optimization techniques used on conventional multi-core CPUs and GPGPUs are not always applicable to DSPs. This paper shares our experience in optimizing GEMM on multi-core GPDSPs, using a CPU-DSP processor as a case study. Our approach employs a range of techniques to optimize performance for DSP architectures. These include data partitioning, three-level pipelining, dedicated micro-kernel design, and improved vector reduction. These optimizations maximize the overlap between computation and communication while fully exploiting the capabilities of floating-point arithmetic units to achieve high performance. Our experimental results demonstrate that the performance attained by our optimization is up to 96% of the theoretical peak performance of the hardware. Kainan Yu, Xinxin Qi, Peng Zhang 0061, Jianbin Fang, Dezun Dong, Ruibo Wang, Tao Tang 0001, Chun Huang 0006, Yonggang Che, Zheng Wang 0001 |
IPDPS | 10 |
| 2024 | UPBEAT: Test Input Checks of Q# Quantum LibrariesabstractHigh-level programming models like Q# significantly simplify the complexity of programming for quantum computing. These models are supported by a set of foundation libraries for code development. However, errors can occur in the library implementation, and one common root cause is the lack of or incomplete checks on properties like values, length, and quantum states of inputs passed to user-facing subroutines. This paper presents Upbeat, a fuzzing tool to generate random test cases for bugs related to input checking in Q# libraries. Upbeat develops an automated process to extract constraints from the API documentation and the developer implemented input-checking statements. It leverages open-source Q# code samples to synthesize test programs. It frames the test case generation as a constraint satisfaction problem for classical computing and a quantum state model for quantum computing to produce carefully generated subroutine inputs to test if the input-checking mechanism is appropriately implemented. Under 100 hours of automated test runs, Upbeat has successfully identified 16 bugs in API implementations and 4 documentation errors. Of these, 14 have been confirmed, and 12 have been fixed by the library developers. Tianmin Hu, Guixin Ye, Zhanyong Tang, Shin Hwei Tan, Huanting Wang, Meng Li 0006, Zheng Wang 0001 |
ISSTA | 7 |
| 2024 | GraphCube: Interconnection Hierarchy-aware Graph ProcessingabstractProcessing large-scale graphs with billions to trillions of edges requires efficiently utilizing parallel systems. However, current graph processing engines do not scale well beyond a few tens of computing nodes because they are oblivious to the communication cost variations across the interconnection hierarchy. We introduce GraphCube, a better approach to optimizing graph processing on large-scale parallel systems with complex interconnections. GraphCube features a new graph partitioning approach to achieve better load balancing and minimize communication overhead across multiple levels of the interconnection hierarchy. We evaluate GraphCube by applying it to fundamental graph operations performed on synthetic and real-world graph datasets. Our evaluation used up to 79,024 computing nodes and 1.2+ million processor cores. Our large-scale experiments show that GraphCube outperforms state-of-the-art parallel graph processing methods in throughput and scalability. Furthermore, GraphCube outperformed the top-ranked systems on the Graph 500 list. Xinbiao Gan, Shenghao Qiu, Jiaqi Si, Jianbin Fang, Dezun Dong, Chunye Gong, Zheng Wang 0001 |
PPoPP | 10 |
| 2024 | Trust Management of Tiny Federated Learning in Internet of Unmanned Aerial VehiclesabstractLightweight training and distributed tiny data storage in local model will lead to the severe challenge of convergence for tiny federated learning (FL). Achieving fast convergence in tiny FL is crucial for many emerging applications in Internet of Unmanned Aerial Vehicles (IUAVs) networks. Excessive information exchange between UAVs and IoT devices could lead to security risks and data breaches, while insufficient information can slow down the learning process and negatively system performance experience due to significant computational and communication constraints in tiny FL hardware system. This paper proposes a trusting, low latency, and energy-efficient tiny wireless FL framework with blockchain (TBWFL) for IUAV systems. We develop a quantifiable model to determine the trustworthiness of IoT devices in IUAV networks. This model incorporates the time spent in communication, computation, and block production with a decay function in each round of FL at the UAVs. Then it combines the trust information from different UAVs, considering their credibility of trust recommendation. We formulate the TBWFL as an optimization problem that balances trustworthiness, learning speed, and energy consumption for IoT devices with diverse computing and energy capabilities. We decompose the complex optimization problem into three sub-problems for improved local accuracy, fast learning, trust verification, and energy efficiency of IoT devices. Our extensive experiments show that TBWFL offers higher trustworthiness, faster convergence, and lower energy consumption than the existing state-of-the-art FL scheme. Jie Zheng 0005, Jipeng Xu, Hongyang Du 0001, Dusit Niyato, Jiawen Kang 0001, Jiangtian Nie, Zheng Wang 0001 |
IEEE Internet Things J. | 7 |
| 2024 | Adaptive Modeling of Uncertainties for Traffic ForecastingabstractDeep neural networks (DNNs) have emerged as a dominant approach for developing traffic forecasting models. These models are typically trained to minimize error on averaged test cases and produce a single-point prediction, such as a scalar value for traffic speed or travel time. However, single-point predictions fail to account for prediction uncertainty that is critical for many transportation management scenarios, such as determining the best-or worst-case arrival time. We present, a generic framework to enhance the capability of an arbitrary DNN model for uncertainty modeling. requires little human involvement and does not change the base DNN architecture during deployment. Instead, it automatically learns a standard quantile function during the DNN model training to produce a prediction interval for the single-point prediction. The prediction interval defines a range where the true value of the traffic prediction is likely to fall. Furthermore, develops an adaptive scheme that dynamically adjusts the prediction interval based on the location and prediction window of the test input. We evaluated by applying it to five representative DNN models for traffic forecasting across seven public datasets. We then compared against six uncertainty quantification methods. Compared to the baseline uncertainty modeling techniques, with base DNN architectures delivers consistently better and more robust performance than the existing ones on the reported datasets. Yongchao Ye, Adnan Zeb, James Jian Qiao Yu, Zheng Wang 0001 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2024 | : Towards Collaborative and Cross-Domain Wi-Fi Sensing: A Case Study for Human Activity RecognitionabstractThe quality of a learning-based Wi-Fi sensing system is bounded by the quantity and quality of training data. However, obtaining sufficient and high-quality data across different domains is difficult due to extensive user involvement. We present CARING, a federated-learning-based framework to support collaborative and cross-domain Wi-Fi sensing. A key challenge of CARING is to allow the effective exchange and learning of knowledge across local models that are derived from heterogeneous data sources with uneven data distributions. We overcome this challenge by first extracting the activity-related representation to train local models. The shared global model aggregates received local model parameters and sends them back to individual devices for fine-tuning locally in the deployed environment. By leveraging the crowdsourced knowledge, CARING allows local models to quickly adapt to domain changes using just a few samples seen at test time. We demonstrate the benefit of CARING by applying it to activity recognition across three public datasets collected from 5 environments, 7 deployments, 31 users, and 29 activities. Experimental results show that CARING is highly effective and robust, improving the alternative approach for using single-sourced training data by up to 47%, giving an accuracy of over 80% (up to 100%) for various cross-domain scenarios. Xinyi Li 0005, Fengyi Song, Mina Luo, Kang Li 0005, Liqiong Chang, Xiaojiang Chen, Zheng Wang 0001 |
IEEE Trans. Mob. Comput. | 7 |
| 2024 | JavaScript Performance Tuning as a Crowdsourced ServiceabstractJavaScript (JS) is one of the most used programming languages for mobile applications. As JS is increasingly used in computation-intensive and latency-sensitive components, JS application performance can significantly impact user experience. While compilers play a crucial role in optimizing JS performance on mobile systems, their optimizations must be simple due to the computation and battery usage limitations of the underlying hardware platforms. We presentJSTuner, a machine-learning system to leverage compiler-based autotuning techniques to optimize JS performance by finding a good compiler optimization sequence.JSTuneris designed to reduce the cost of autotuning by using prior knowledge of JS programs collected through a crowdsourcing framework to bootstrap the search process. It allows the user to seamlessly utilize the computation resources of a cloud server to perform the heavy-lifting autotuning process for repeatedly running JS components. This enables aggressive search-based optimizations that are too expensive to run on the user's device. We evaluateJSTunerby applying it to 60 JS benchmarks across three distinct mobile devices and comparing it against four search-based techniques. Experimental results show that JSTuner consistently outperforms prior techniques and improves JS performance by 1.62x on average (up to 3.33x) over the default compiler setting used by the Chrome V8 JS engine. Jie Ren 0007, Zheng Wang 0001 |
IEEE Trans. Mob. Comput. | 3 |
| 2024 | Faster and Scalable MPI Applications LaunchingabstractDistributed parallel MPI applications are the dominant workload in many high-performance computing systems. While optimizing MPI application execution is a well-studied field, little work has considered optimizing the initial MPI application launching phase, which incurs extensive cross-machine communications and synchronization. The overhead of MPI application launching can be expensive, accounting for more than million core hours per 10K nodes annually on the production Tianhe-2A supercomputer, which will increase as the number of parallel machines used grows. Therefore, it is critical to optimize the MPI application launching process. This paper presents a novel approach to optimizing the MPI application launch. Our approach adopts a location-aware address generation rule to eliminate the need for address exchange and a topology-aware global communication scheme to optimize cross-machine synchronization. We then design a new application launch procedure to support the proposed optimizations to further reduce the pressure of the shared I/O system. Our techniques have been deployed to production in the Tianhe-2A supercomputer and the Next Generation Tianhe Supercomputer. Experimental results show that our approach scales well and outperforms alternative schemes, reducing the MPI application launching time by over 29% with 320K MPI processes. Yong Dong, Yiqin Dai, Kai Lu 0001, Ruibo Wang, Juan Chen 0001, Mingtian Shao, Zheng Wang 0001 |
IEEE Trans. Parallel Distributed Syst. | 8 |
| 2024 | DRLCAP: Runtime GPU Frequency Capping With Deep Reinforcement LearningabstractPower and energy consumption is the limiting factor of modern computing systems. As the GPU becomes a mainstream computing device, power management for GPUs becomes increasingly important. Current works focus on GPU kernel-level power management, with challenges in portability due to architecture-specific considerations. We presentDRLCap, a general runtime power management framework intended to support power management across various GPU architectures. It periodically monitors system-level information to dynamically detect program phase changes and model the workload and GPU system behavior. This elimination from kernel-specific constraints enhances adaptability and responsiveness. The framework leverages dynamic GPU frequency capping, which is the most widely used power knob, to control the power consumption.DRLCapemploys deep reinforcement learning (DRL) to adapt to the changing of program phases by automatically adjusting its power policy through online learning, aiming to reduce the GPU power consumption without significantly compromising the application performance. We evaluateDRLCapon three NVIDIA and one AMD GPU architectures. Experimental results show thatDRLCapimproves prior GPU power optimization strategies by a large margin. On average, it reduces the GPU energy consumption by 22% with less than 3% performance slowdown on NVIDIA GPUs. This translates to a 20% improvement in the energy efficiency measured by the energy-delay product (EDP) over the NVIDIA default GPU power management strategy. For the AMD GPU architecture,DRLCapsaves energy consumption by 10%, on average, with a 4% percentage loss, and improves energy efficiency by 8%. Yiming Wang 0010, Meng Hao 0002, Weizhe Zhang, Qiuyuan Tang, Zheng Wang 0001 |
IEEE Trans. Sustain. Comput. | 7 |
| 2024 | Model-Free GPU Online Energy OptimizationabstractGPUs play a central and indispensable role as accelerators in modern high-performance computing (HPC) platforms, enabling a wide range of tasks to be performed efficiently. However, the use of GPUs also results in significant energy consumption and carbon dioxide (CO2) emissions. This article presents MF-GPOEO, a model-free GPU online energy efficiency optimization framework. MF-GPOEO leverages a synthetic performance index and a PID controller to dynamically determine the optimal clock frequency configuration for GPUs. It profiles GPU kernel activity information under different frequency configurations and then compares GPU kernel execution time and gap duration between kernels to derive the synthetic performance index. With the performance index and measured average power, MF-GPOEO can use the PID controller to try different frequency configurations and find the optimal frequency configuration under the guidance of user-defined objective functions. We evaluate the MF-GPOEO by running it with 74 applications on an NVIDIA RTX3080Ti GPU. MF-GPOEO delivers a mean energy saving of 26.2% with a slight average execution time increase of 3.4% compared with NVIDIA's default clock scheduling strategy. Farui Wang, Meng Hao 0002, Weizhe Zhang, Zheng Wang 0001 |
IEEE Trans. Sustain. Comput. | 4 |
| 2023 | POINE2: Improving Poincaré Embeddings for Hierarchy-Aware Complex Query Reasoning over Knowledge GraphsabstractReasoning complex logical queries on incomplete and massive knowledge graphs (KGs) remains a significant challenge. The prevailing method for this problem is query embedding, which embeds KG units (i.e., entities and relations) and complex queries into low-dimensional space. Recent developments in the field show that embedding queries as geometric shapes is a viable means for modeling entity set and logical relationships between them. Despite being promising, current geometric-based methods face challenges in capturing hierarchical structures of complex queries, which leaves considerable room for improvement. This paper presents POINE2, a geometric-based query embedding framework based on hyperbolic geometry to handle complex queries on knowledge graphs. POINE2 maps entities and queries as geometric shapes on a Cartesian product space of Poincaré ball spaces. To capture the hierarchical structures of complex queries, we use the Poincaré radius to represent the different levels of the hierarchy, and we use the aperture of the shape to indicate semantic differences at the same level of the hierarchy. Additionally, POINE2 offers a flexible and expressive definition of logical operations. Experimental results show that POINE2 outperforms existing salient geometric-based embedding methods and significantly improves these methods on evaluation datasets. Qianren Mao, Jianxin Li 0002, Xingcheng Fu, Zheng Wang 0001 |
ECAI | 5 |
| 2023 | DisCo: Distilled Student Models Co-training for Semi-supervised Text MiningabstractMany text mining models are constructed by fine-tuning a large deep pre-trained language model (PLM) in downstream tasks.However, a significant challenge nowadays is maintaining performance when we use a lightweight model with limited labelled samples.We present DisCo, a semi-supervised learning (SSL) framework for fine-tuning a cohort of small student models generated from a large PLM using knowledge distillation.Our key insight is to share complementary knowledge among distilled student cohorts to promote their SSL effectiveness.DisCo employs a novel co-training technique to optimize a cohort of multiple small student models by promoting knowledge sharing among students under diversified views: model views produced by different distillation strategies and data views produced by various input augmentations.We evaluate DisCo on both semi-supervised text classification and extractive summarization tasks.Experimental results show that DisCo can produce student models that are 7.6× smaller and 4.8× faster in inference than the baseline PLMs while maintaining comparable performance.We also show that DisCo-generated student models outperform the similar-sized models elaborately tuned in distinct tasks. Weifeng Jiang, Qianren Mao, Chenghua Lin 0002, Jianxin Li 0002, Ting Deng, Zheng Wang 0001 |
EMNLP | 7 |
| 2023 | HighRPM: Combining Integrated Measurement and Sofware Power Modeling for High-Resolution Power MonitoringabstractIn an era where power and energy are the first-class constraints of computing systems, accurate power information is crucial for energy efficiency optimization in parallel computing systems. Existing power monitoring techniques rely on either software-centric power models that suffer from poor accuracy or integrated hardware measurement schemes that have a low reading update frequency and coarse granularity. These result in a low spatiotemporal resolution for power monitoring. This paper introduces HighRPM, a new method for accurately measuring power consumption on parallel computing systems. HighRPM combines coarse-grained power sensor readings and software power modeling techniques to improve temporal and spatial resolutions. To provide high-frequent power readings in the temporal domain, HighRPM employs statistical modeling and machine learning techniques to predict the long-term power trend and the short-term fluctuations in power consumption. To improve spatial coverage, HighRPM takes low-time resolution node-level power consumption and uses a neural network to distribute the power readings to lower-level computing components like CPUs and memory components. We evaluate HighRPM by applying it to both ARM-based and X86-based platforms. Experimental results show that HighRPM improves time resolution by 10 times, provides accurate readings for CPUs and memory, and reduces error by 7-24% compared to other power modeling methods. Xinxin Qi, Juan Chen 0001, Yong Dong, Yuan Yuan 0034, Tao Xu 0052, Rongyu Deng, Kexing Zhou, Zheng Wang 0001 |
ICPP | 9 |
| 2023 | Optimizing Multi-grid Computation and Parallelization on Multi-coresabstractMultigrid algorithms are widely used to solve large-scale sparse linear systems, which is essential for many high-performance workloads. The symmetric Gauss-Seidel (SYMGS) method is often responsible for the performance bottleneck of MG. This paper presents new methods to parallelize and enhance the computation and parallelization efficiency of the SYMGS and MG algorithms on multi-core CPUs. Our solution employs a matrix splitting strategy and a revised computation formula to decrease the computation operations and memory accesses in SYMGS. With this new SYMGS strategy, we can then merge the two most time-consuming components of MG. On top of these, we propose a new asynchronous parallelization scheme to reduce the synchronization overhead when parallelizing SYMGS. We demonstrate the benefit of our techniques by integrating them with the HPCG benchmark and two real-life applications. Evaluation conducted on four architectures, including three ARMv8 and one x86, shows that our techniques greatly surpass the performance of engineer- and vendor-tuned implementations across various workloads and platforms. Xiaojian Yang, Shengguo Li, Dezun Dong, Chun Huang 0006, Zheng Wang 0001 |
ICS | 6 |
| 2023 | Memory-aware Optimization for Sequences of Sparse Matrix-Vector MultiplicationsabstractThis paper presents a novel approach to optimize multiple invocations of a sparse matrix-vector multiplication (SpMV) kernel performed on the same sparse matrix A and dense vector x, like Ax, A2x, ⋯, Akx, and their linear combinations such as Ax + A2x. Such computations are frequently used in scientific applications for solving linear equations and in multi-grid methods. Existing SpMV optimization techniques typically focus on a single SpMV invocation and do not consider opportunities for optimization across a sequence of SpMV operations (SSpMV), leaving much room for performance improvement. Our work aims to bridge this performance gap. It achieve this by partitioning the sparse matrix into submatrices and devising a new computation pipeline that reduces memory access to the sparse matrix and exploits the data locality of the dense vector of SpMV. Additionally, we demonstrate how our approach can be integrated with parallelization schemes to further improve performance. We evaluate our approach on four distinct multi-core systems, including three ARM and one Intel platform. Experimental results show that our techniques improve the standard implementation and the highly-optimized Intel math kernel library (MKL) by a large margin. Shengguo Li, Dezun Dong, Xiaojian Yang, Zheng Wang 0001 |
IPDPS | 7 |
| 2023 | One for All: Unified Workload Prediction for Dynamic Multi-tenant Edge Cloud PlatformsabstractWorkload prediction in multi-tenant edge cloud platforms (MT-ECP) is vital for efficient application deployment and resource provisioning. However, the heterogeneous application patterns, variable infrastructure performance, and frequent deployments in MT-ECP pose significant challenges for accurate and efficient workload prediction. Clustering-based methods for dynamic MT-ECP modeling often incur excessive costs due to the need to maintain numerous data clusters and models, which leads to excessive costs. Existing end-to-end time series prediction methods are challenging to provide consistent prediction performance in dynamic MT-ECP. In this paper, we propose an end-to-end framework with global pooling and static content awareness, DynEformer, to provide a unified workload prediction scheme for dynamic MT-ECP. Meticulously designed global pooling and information merging mechanisms can effectively identify and utilize global application patterns to drive local workload predictions. The integration of static content-aware mechanisms enhances model robustness in real-world scenarios. Through experiments on five real-world datasets, DynEformer achieved state-of-the-art in the dynamic scene of MT-ECP and provided a unified end-to-end prediction scheme for MT-ECP. Shaoyuan Huang, Zheng Wang 0001, Heng Zhang 0032, Xiaofei Wang 0001, Cheng Zhang 0007 |
KDD | 2 |
| 2023 | Optimizing MPI Collectives on Shared Memory Multi-CoresabstractMessage Passing Interface (MPI) programs often experience performance slowdowns due to collective communication operations, like broadcasting and reductions. As modern CPUs integrate more processor cores, running multiple MPI processes on shared-memory machines to take advantage of hardware parallelism is becoming increasingly common. In this context, it is crucial to optimize MPI collective communications for shared-memory execution. However, existing MPI collective implementations on shared-memory systems have two primary drawbacks. The first is extensive redundant data movements when performing reduction collectives, and the second is the ineffective use of non-temporal instructions to optimize streamed data processing. To address these limitations, this paper proposes two optimization techniques that minimize data movements and enhance the use of non-temporal instructions. We evaluated our techniques by integrating them into the OpenMPI library and tested their performance using micro-benchmarks and real-world applications running on two multi-core clusters. Experimental results show that our approach significantly outperforms existing techniques, yielding a 1.2--6.4x performance improvement. Jintao Peng, Jianbin Fang, Jie Liu 0002, Bo Yang 0023, Shengguo Li, Zheng Wang 0001 |
SC | 8 |
| 2023 | Optimizing Direct Convolutions on ARM Multi-CoresabstractConvolution kernels are widely seen in deep learning workloads and are often responsible for performance bottlenecks. Recent research has demonstrated that a direct convolution approach can outperform the traditional convolution implementation based on tensor-to-matrix conversions. However, existing approaches for direct convolution still have room for performance improvement. We present nDirect, a new direct convolution approach that targets ARM-based multi-core CPUs commonly found in smartphones and HPC systems. nDirect is designed to be compatible with the data layout formats used by mainstream deep learning frameworks but offers new optimizations for the computational kernel, data packing, and parallelization. We evaluate nDirect by applying it to representative convolution kernels and demonstrating its performance on four distinct ARM multi-core CPU platforms. We compare nDirect against state-of-the-art convolution optimization techniques. Experimental results show that nDirect gives the best overall performance across evaluation scenarios and platforms. Weiling Yang, Jianbin Fang, Dezun Dong, Chun Huang 0006, Peng Zhang 0061, Tao Tang 0001, Zheng Wang 0001 |
SC | 8 |
| 2023 | A Generative and Mutational Approach for Synthesizing Bug-Exposing Test Cases to Guide Compiler FuzzingabstractRandom test case generation, or fuzzing, is a viable means for uncovering compiler bugs. Unfortunately, compiler fuzzing can be time-consuming and inefficient with purely randomly generated test cases due to the complexity of modern compilers. We present COMFUZZ, a focused compiler fuzzing framework. COMFUZZ aims to improve compiler fuzzing efficiency by focusing on testing components and language features that are likely to trigger compiler bugs. Our key insight is human developers tend to make common and repeat errors across compiler implementations; hence, we can leverage the previously reported buggy-exposing test cases of a programming language to test a new compiler implementation. To this end, COMFUZZ employs deep learning to learn a test program generator from open-source projects hosted on GitHub. With the machine-generated test programs in place, COMFUZZ then leverages a set of carefully designed mutation rules to improve the coverage and bug-exposing capabilities of the test cases. We evaluate COMFUZZ on 11 compilers for JS and Java programming languages. Within 260 hours of automated testing runs, we discovered 33 unique bugs across nine compilers, of which 29 have been confirmed and 22, including an API documentation defect, have already been fixed by the developers. We also compared COMFUZZ to eight prior fuzzers on four evaluation metrics. In a 24-hour comparative test, COMFUZZ uncovers at least 1.5× more bugs than the state-of-the-art baselines. Guixin Ye, Tianmin Hu, Zhanyong Tang, Zhenye Fan, Shin Hwei Tan, Wenxiang Qian, Zheng Wang 0001 |
ESEC/SIGSOFT FSE | 8 |
| 2023 | wrBench: Comparing Cache Architectures and Coherency Protocols on ARMv8 Many-Core Systems
Wanrong Gao, Jianbin Fang, Chun Huang 0006, Chuanfu Xu, Zheng Wang 0001 |
J. Comput. Sci. Technol. | 5 |
| 2023 | Programming bare-metal accelerators with heterogeneous threading models: a case study of Matrix-3000abstractAs the hardware industry moves toward using specialized heterogeneous many-core processors to avoid the effects of the power wall, software developers are finding it hard to deal with the complexity of these systems. In this paper, we share our experience of developing a programming model and its supporting compiler and libraries for Matrix-3000, which is designed for next-generation exascale supercomputers but has a complex memory hierarchy and processor organization. To assist its software development, we have developed a software stack from scratch that includes a low-level programming interface and a high-level OpenCL compiler. Our low-level programming model offers native programming support for using the bare-metal accelerators of Matrix-3000, while the high-level model allows programmers to use the OpenCL programming standard. We detail our design choices and highlight the lessons learned from developing system software to enable the programming of bare-metal accelerators. Our programming models have been deployed in the production environment of an exascale prototype system. Jianbin Fang, Peng Zhang 0061, Chun Huang 0006, Tao Tang 0001, Kai Lu 0001, Ruibo Wang, Zheng Wang 0001 |
Frontiers Inf. Technol. Electron. Eng. | 7 |
| 2023 | Lifelong Property Price Prediction: A Case Study for the Toronto Real Estate MarketabstractWe present LUCE, the first life-long predictive model for automated property valuation. LUCE addresses two critical issues of property valuation: the lack of recent sold prices and the sparsity of house data. It is designed to operate on limited volume of recent house transaction. As a departure from prior work, LUCE organizes the house data in a HIN where graph nodes are house entities and attributes that are important for house price valuation. We employ GCN to extract the spatial information from the HIN, and then use LSTM network to model the temporal dependencies over time. Unlike prior work, LUCE makes effective use of the limited house transactions in the past few months to update valuation information for all house entities. By providing a complete and up-to-date house valuation dataset, LUCE thus massively simplifies the downstream valuation task for the targeting properties. We demonstrate the benefit of LUCE by applying it to large, real-life datasets obtained from the Toronto real estate market. Extensive experimental results show that LUCE not only significantly outperforms prior property valuation methods but also often reaches and sometimes exceeds the valuation accuracy given by independent experts when using the actual realization price as the ground truth. Hao Peng 0001, Jianxin Li 0002, Zheng Wang 0001, Renyu Yang, Mingsheng Liu, Philip S. Yu, Lifang He 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2023 | Toward Wide-Area Contactless Wireless SensingabstractContactless wireless sensing without attaching a device to the target has achieved promising progress in recent years. However, one severe limitation is the small sensing range. This paper presents Widesee to realize wide-area sensing with only one transceiver pair. Widesee utilizes the LoRa signal to achieve a larger range of sensing and further incorporates drone’s mobility to broaden the sensing area. Widesee presents solutions across software and hardware to overcome two aspects of challenges for wide-range contactless sensing: (i) the interference brought by device mobility and LoRa’s high sensitivity; and (ii) the ambiguous target information such as location when employing just a single pair of transceivers for sensing. We have developed a working prototype of Widesee for human target detection and localization that are especially useful in emergency scenarios such as rescue search, and evaluated Widesee with both controlled experiments and the field study in a high-rise building. Extensive experiments demonstrate the great potential of Widesee for wide-area contactless sensing with a single LoRa transceiver pair hosted on a drone. Jie Xiong 0001, Sunghoon Ivan Lee, Zhanyong Tang, Zheng Wang 0001, Dingyi Fang, Xiaojiang Chen |
IEEE/ACM Trans. Netw. | 8 |
| 2023 | Janus: Latency-Aware Traffic Scheduling for IoT Data Streaming in Edge EnvironmentsabstractThis article focuses on a simple, yet fundamental question of distributed edge computing: “how to handle IoT traffic with different levels of sensitivity and criticality by satisfying the application-specific latency constraints?” This question arises in the practical deployment of edge computing, where user data can arrive at a much faster rate than that they can be processed by an edge node. Addressing this question is critical for meeting the latency requirement for latency-sensitive applications, but existing approaches are inadequate to the problem. We presentJanus, a multi-level traffic scheduling system for managing multiple data streams with various degrees of latency constraints. At the edge node level,Janususes multi-level queues to manage data streams with different latency constraints. It then allocates the output bandwidth of the edge node according to the requirements of applications in different priority queues, aiming to reduce the queuing and processing delay of latency-sensitive streams while maximizing the edge-node throughput. At the network level,Janusactively redirects incoming data streams to the less-loaded ones to achieve better network-wide load balance and improve the overall throughput. Experiments show thatJanusreduces the latency to only 16.6% of a non-priority based solution and improves the throughput by 1.7x of a state-of-the-art priority-aware data stream scheduling approach. Zhenyu Wen, Renyu Yang, Bin Qian 0002, Yubo Xuan, Lingling Lu, Zheng Wang 0001, Hao Peng 0001, Jie Xu 0007, Albert Y. Zomaya, Rajiv Ranjan 0001 |
IEEE Trans. Serv. Comput. | 6 |
| 2022 | Automating reinforcement learning architecture design for code optimizationabstractReinforcement learning (RL) is emerging as a powerful technique for solving complex code optimization tasks with an ample search space. While promising, existing solutions require a painstaking manual process to tune the right task-specific RL architecture, for which compiler developers need to determine the composition of the RL exploration algorithm, its supporting components like state, reward, and transition functions, and the hyperparameters of these models. This paper introduces SuperSonic, a new open-source framework to allow compiler developers to integrate RL into compilers easily, regardless of their RL expertise. SuperSonic supports customizable RL architecture compositions to target a wide range of optimization tasks. A key feature of SuperSonic is the use of deep RL and multi-task learning techniques to develop a meta-optimizer to automatically find and tune the right RL architecture from training benchmarks. The tuned RL can then be deployed to optimize new programs. We demonstrate the efficacy and generality of SuperSonic by applying it to four code optimization problems and comparing it against eight auto-tuning frameworks. Experimental results show that SuperSonic consistently improves hand-tuned methods by delivering better overall performance, accelerating the deployment-stage search by 1.75x on average (up to 100x). Huanting Wang, Zhanyong Tang, Cheng Zhang 0007, Chris Cummins, Hugh Leather, Zheng Wang 0001 |
CC | 7 |
| 2022 | Explicitly Modeling Importance and Coherence for Timeline SummarizationabstractTimeline summarization (TLS) identifies major events and generates short summaries on how the event evolves in a period of time. Existing timeline summarization methods generate summaries by considering the coverage and diversity of the content and temporized information but ignore the importance and coherence of sentences used in summary. However, ignoring such information often causes missing important facts in the generated TLS and confuses users. We propose a better approach for TLS by explicitly optimizing importance and coherence on top of coverage and diversity. We apply our approach to both direct and pipeline TLS frameworks. Experimental results show that our approach achieves better performance when compared with two state-of-the-art TLS methods. Qianren Mao, Jianxin Li 0002, JiaZheng Wang, Xi Li 0001, Zheng Wang 0001 |
ICASSP | 7 |
| 2022 | AIACC-Training: Optimizing Distributed Deep Learning Training through Multi-streamed and Concurrent Gradient CommunicationsabstractThere is a growing interest in training deep neural networks (DNNs) in a GPU cloud environment. This is typically achieved by running parallel training workers on multiple GPUs across computing nodes. Under such a setup, the communication overhead is often responsible for long training time and poor scalability. This paper presents AIACC-Training, a unified communication framework designed for the distributed training of DNNs in a GPU cloud environment. AIACC-Training permits a training worker to participate in multiple gradient communication operations simultaneously to improve network bandwidth utilization and reduce communication latency. It employs auto-tuning techniques to dynamically determine the right communication parameters based on the input DNN workloads and the underlying network infrastructure. AIACC-Training has been deployed to production at Alibaba GPU Cloud with 3000+ GPUs executing AIACC-Training optimized code at any time. Experiments performed on representative DNN workloads show that AIACC-Training outperforms existing solutions, improving the training throughput and scalability by a large margin. Lixiang Lin, Shenghao Qiu, Liang You, Long Xin, Jie Xu 0007, Zheng Wang 0001 |
ICDCS | 8 |
| 2022 | HiGIL: Hierarchical Graph Inference Learning for Fact CheckingabstractFact-checking is vital for countering fake news. This process requires verifying the truthfulness of a claim by reasoning about multiple pieces of evidence. The current dominant approach depends upon capturing the claim-evidence relations from a claim-evidence interaction graph. Existing solutions utilize phrase-level semantics on a single-granularity but ignore other hierarchical features, such as fact- and sentence-level textual semantics and their logical topology. Since the hierarchical features often provide hints to infer collaborative high-order clues that can be essential for fact-checking, they should not be overlooked. This paper proposes a better method to model the claim-evidence graph in a multi-granularity manner. Doing so allows one to exploit more textual semantics and logical topology between a claim and its evidence. To achieve the target, we first employ a graph inference learning framework to infer graph nodes on different granular semantic units within their hierarchical topology. Then, an inference learning procedure is designed to optimize the global textual similarity and local topological reachability from the claim-evidence graph. We evaluate our approach by applying it to fact-checking on an open dataset, and experimental results show that our technique outperforms existing graph-based techniques by a large margin. Qianren Mao, Yiming Wang 0010, Linfeng Du, Hao Peng 0001, Jia Wu 0001, Jianxin Li 0002, Zheng Wang 0001 |
ICDM | 8 |
| 2022 | Parallelizing and Balancing Coupled DSMC/PIC for Large-scale Particle SimulationsabstractIn high-performance and parallel computing, an important application class is particle simulation. Due to massive particle migration among distributed simulation workers across simulation iterations, achieving balanced runtime work distribution is vital for accelerating large-scale realistic particle simulations. This paper proposes a novel approach to enable dynamic load balance for distributed numerical particle simulations, specifically targeting the latest coupled DSMC/PI C method. Unlike prior work, our approach adopts a dual, nested unstructured grid organization to facilitate coupled DSMC/PIC computation and runtime grid distribution. Our implementation leverages both centralized and distributed communication strategies to dynamically migrate particles among arbitrary parallel processes. It then employs a load balancer - driven by a carefully designed analytical model and a grid remapping mechanism - to dynamically redistribute the simulation workloads among parallel simulation workers. By constantly monitoring and redis-tributing the simulation work across workers, our approach can adapt to the change of particle distribution across simulation iterations, avoiding a few workers becoming the performance bottleneck of the entire simulation process. We integrate our techniques into a coupled DSMC/PIC solver and apply them to simulate the plasma plume with hydrogen atoms and ions. Experimental results show that our approach can scale well up to 1500+ processes with billions of particles, exhibiting the state-of-the-art parallel simulation scalability and efficiency for plasma plume simulation. Haozhong Qiu, Chuanfu Xu, Dali Li, Zheng Wang 0001 |
IPDPS | 6 |
| 2022 | Towards Scalable Resource Management for SupercomputersabstractToday's supercomputers offer massive computation resources to execute a large number of user jobs. Effectively managing such large-scale hardware parallelism and workloads is essential for supercomputers. However, existing HPC resource management (RM) systems fail to capitalize on the hardware parallelism by following a centralized design used decades ago. They give poor scalability and inefficient performance on today's supercomputers, which will worsen in exascale computing. We present ESlurm, a better RM for supercomputers. As a departure from existing HPC RMs, ESlurm implements a distributed communication structure. It employs a new communication tree strategy and uses job runtime estimation to improve communications and job scheduling efficiency. ESlurm is deployed into production in a real supercomputer. We evaluate ESlurm on up to 20K nodes. Compared to state-of-the-art RM solutions, ESlurm exhibits better scalability, significantly reducing the resource usage of master nodes and improving data transfer and job scheduling efficiency by a large margin. Yiqin Dai, Yong Dong, Kai Lu 0001, Ruibo Wang, Wei Zhang 0027, Juan Chen 0001, Mingtian Shao, Zheng Wang 0001 |
SC | 8 |
| 2022 | STRONGHOLD: Fast and Affordable Billion-Scale Deep Learning Model TrainingabstractDeep neural networks (DNNs) with billion-scale parameters have demonstrated impressive performance in solving many tasks. Unfortunately, training a billion-scale DNN is out of the reach of many data scientists because it requires high-performance GPU servers that are too expensive to purchase and maintain. We present STRONGHOLD, a novel approach for enabling large DNN model training with no change to the user code. STRONGHOLD scales up the largest trainable model size by dynamically offloading data to the CPU RAM and enabling the use of secondary storage. It automatically determines the minimum amount of data to be kept in the GPU memory to minimize GPU memory usage. Compared to state-of-the-art offloading-based solutions, STRONGHOLD improves the trainable model size by 1.9x~6. Sx on a 32GB V100 GPU, with 1.2x~3.7x improvement on the training throughput. It has been deployed into production to successfully support large-scale DNN training. Wei Wang 0225, Shenghao Qiu, Renyu Yang, Songfang Huang, Jie Xu 0007, Zheng Wang 0001 |
SC | 7 |
| 2022 | MuchSUM: Multi-channel Graph Neural Network for Extractive SummarizationabstractRecent studies of extractive text summarization have leveraged BERT for document encoding with breakthrough performance. However, when using a pre-trained BERT-based encoder, existing approaches for selecting representative sentences for text summarization are inadequate since the encoder is not explicitly trained for representing sentences. Simply providing the BERT-initialized sentences to cross-sentential graph-based neural networks (GNNs) to encode semantic features of the sentences is not ideal because doing so fail to integrate other summary-worthy features like sentence importance and positions. This paper presents MuchSUM, a better approach for extractive text summarization. MuchSUM is a multi-channel graph convolutional network designed to explicitly incorporate multiple salient summary-worthy features. Specifically, we introduce three specific graph channels to encode the node textual features, node centrality features, and node position features, respectively, under bipartite word-sentence heterogeneous graphs. Then, a cross-channel convolution operation is designed to distill the common graph representations shared by different channels. Finally, the sentence representations of each channel are fused for extractive summarization. We also investigate three weighted graphs in each channel to infuse edge features for graph-based summarization modeling. Experimental results demonstrate our model can achieve considerable performance compared with some BERT-initialized graph-based extractive summarization systems. Qianren Mao, Hongdong Zhu, Cheng Ji 0001, Hao Peng 0001, Jianxin Li 0002, Zheng Wang 0001 |
SIGIR | 8 |
| 2022 | Detecting code vulnerabilities by learning from large-scale open source repositories
Rongze Xu, Zhanyong Tang, Guixin Ye, Huanting Wang, Xin Ke, Dingyi Fang, Zheng Wang 0001 |
J. Inf. Secur. Appl. | 7 |
| 2022 | FlowDNN: a physics-informed deep neural network for fast and accurate flow predictionabstractfor flow-related design optimization problems, e.g., aircraft and automobile aerodynamic design, computational fluid dynamics (CFD) simulations are commonly used to predict flow fields and analyze performance. While important, CFD simulations are a resource-demanding and time-consuming iterative process. The expensive simulation overhead limits the opportunities for large design space exploration and prevents interactive design. In this paper, we propose FlowDNN, a novel deep neural network (DNN) to efficiently learn flow representations from CFD results. FlowDNN saves computational time by directly predicting the expected flow fields based on given flow conditions and geometry shapes. FlowDNN is the first DNN that incorporates the underlying physical conservation laws of fluid dynamics with a carefully designed attention mechanism for steady flow prediction. This approach not only improves the prediction accuracy, but also preserves the physical consistency of the predicted flow fields, which is essential for CFD. Various metrics are derived to evaluate FlowDNN with respect to the whole flow fields or regions of interest (RoIs) (e.g., boundary layers where flow quantities change rapidly). Experiments show that FlowDNN significantly outperforms alternative methods with faster inference and more accurate results. It speeds up a graphics processing unit (GPU) accelerated CFD solver by more than 14 000×, while keeping the prediction error under 5%. Donglin Chen, Xiang Gao 0020, Chuanfu Xu, Siqi Wang 0001, Shizhao Chen, Jianbin Fang, Zheng Wang 0001 |
Frontiers Inf. Technol. Electron. Eng. | 7 |
| 2022 | Reinforcement Learning-Based Dialogue Guided Event Extraction to Exploit Argument RelationsabstractEvent extraction is a ftask for natural language processing. Finding the roles of event arguments like event participants is essential for event extraction. However, doing so for real-life event descriptions is challenging because an argument’s role often varies in different contexts. While the relationship and interactions between multiple arguments are useful for settling the argument roles, such information is largely ignored by existing approaches. This paper presents a better approach for event extraction by explicitly utilizing the relationships of event arguments. We achieve this through a carefully designed task-oriented dialogue system. To model the argument relation, we employ reinforcement learning and incremental learning to extract multiple arguments via a multi-turned, iterative process. Our approach leverages knowledge of the already extracted arguments of the same sentence to determine the role of arguments that would be difficult to decide individually. It then uses the newly obtained information to improve the decisions of previously extracted arguments. This two-way feedback process allows us to exploit the argument relations to effectively settle argument roles, leading to better sentence understanding and event extraction. Experimental results show that our approach consistently outperforms seven state-of-the-art event extraction methods for the classification of events and argument role and argument identification. Qian Li 0033, Hao Peng 0001, Jianxin Li 0002, Jia Wu 0001, Yuanxing Ning, Philip S. Yu, Zheng Wang 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 8 |
| 2022 | Fact-Driven Abstractive Summarization by Utilizing Multi-Granular Multi-Relational KnowledgeabstractAbstractive summarization generates a concise summary to capture the key ideas of the source text. This task underpins important applications like information retrieval, document comprehension, and event tracking. While much progress has been achieved, state-of-the-art summarization approaches often fail to generate high-quality summaries to reproduce factual details accurately. One of the key limitations of existing solutions is that they are primarily concerned about extracting facts from the source text but overlook other crucial factual information, such as the related time, locations, reasons, consequences, purposes, participants and involved parties. Furthermore, the current summarization frameworks are inadequate in modeling the complex semantic relations among facts and the corresponding factual information, leaving much room for improvement. This paper presentsFFSum, a novel summarization framework for exploiting multi-grained factual information to improve text summarization. To this end,FFSumconstructs an individual fine-grained factual graph with multiple relations among facts and the corresponding factual information. It employs a fact-driven graph attention network to integrate multi-granular factual representations at the encoding stage. It then uses a hybrid pointer network to retrieve factual pieces from the graph for the summary generation. We evaluate theFFSumby applying it to two real-world datasets. Experimental results show that theFFSumconsistently outperforms a state-of-the-art approach across evaluation datasets. Qianren Mao, Jianxin Li 0002, Hao Peng 0001, Shizhu He, Philip S. Yu, Zheng Wang 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 7 |
| 2022 | Lime: Low-Cost and Incremental Learning for Dynamic Heterogeneous Information NetworksabstractUnderstanding the interconnected relationships of large-scale information networks like social, scholar and Internet of Things networks is vital for tasks like recommendation and fraud detection. The vast majority of the real-world networks are inherently heterogeneous and dynamic, containing many different types of nodes and edges and can change drastically over time. The dynamicity and heterogeneity make it extremely challenging to reason about the network structure. Unfortunately, existing approaches are inadequate in modeling real-life dynamical networks as they either have strong assumption of a given stochastic process or fail to capture the heterogeneity of network structure, and they all require extensive computational resources. We introduceLime, a better approach for modeling dynamic and heterogeneous information networks.Limeis designed to extract high-quality network representation with significantly lower memory resources and computational time over the state-of-the-arts. Unlike prior work that uses a vector to encode each network node, we exploit the semantic relationships among network nodes to encode multiple nodes with similar semantics in shared vectors. By using many fewer node vectors, our approach significantly reduces the required memory space for encoding large-scale networks. To effectively trade information sharing for reduced memory footprint, we employ the recursive neural network (RsNN) with carefully designed optimization strategies to explore the node semantics in a novel cuboid space. We then go further by showing, for the first time, how an effective incremental learning approach can be developed – with the help of RsNN, our cuboid structure, and a set of novel optimization techniques – to allow a learning framework to quickly and efficiently adapt to a constantly evolving network. We evaluateLimeby applying it to three representative network-based tasks, node classification, node clustering and anomaly detection, performing on three large-scale datasets. We compareLimeagainst eleven prior state-of-the-art approaches for learning network representation. Our extensive experiments demonstrate thatLimenot only reduces the memory footprint by over 80 percent and the processing time over 2x when learning network representation but also delivers comparable performance for downstream processing tasks. We show that our incremental learning method can boost the learning time by up to 20x without compromising the quality of the learned network representation. Hao Peng 0001, Renyu Yang, Zheng Wang 0001, Jianxin Li 0002, Lifang He 0001, Philip S. Yu, Albert Y. Zomaya, Rajiv Ranjan 0001 |
IEEE Trans. Computers | 3 |
| 2022 | Optimizing Depthwise Separable Convolution Operations on GPUsabstractThe depthwise separable convolution is commonly seen in convolutional neural networks (CNNs), and is widely used to reduce the computation overhead of a standard multi-channel 2D convolution. Existing implementations of depthwise separable convolutions target accelerating model training with large batch sizes with a large number of samples to be processed at once. Such approaches are inadequate for small-batch-sized model training and the typical scenario of model inference where the model takes in a few samples at once. This article aims to bridge the gap of optimizing depthwise separable convolutions by targeting the GPU architecture. We achieve this by designing two novel algorithms to improve the column and row reuse of the convolution operation to reduce the number of memory operations performed on the width and the height dimensions of the 2D convolution. Our approach employs a dynamic tile size scheme to adaptively distribute the computational data across GPU threads to improve GPU utilization and to hide the memory access latency. We apply our approach on two GPU platforms: an NVIDIA RTX 2080Ti GPU and an embedded NVIDIA Jetson AGX Xavier GPU, and two data types: 32-bit floating point (FP32) and 8-bit integer (INT8). We compared our approach against cuDNN that is heavily tuned for the NVIDIA GPU architecture. Experimental results show that our approach delivers over 2× (up to 3×) performance improvement over cuDNN. We show that, when using a moderate batch size, our approach averagely reduces the end-to-end training time of MobileNet and EfficientNet by 9.7 and 7.3 percent respectively, and reduces the end-to-end inference time of MobileNet and EfficientNet by 12.2 and 11.6 percent respectively. Gangzhao Lu, Weizhe Zhang, Zheng Wang 0001 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2022 | Online Power Management for Multi-Cores: A Reinforcement Learning Based ApproachabstractPower and energy is the first-class design constraint for multi-core processors and is a limiting factor for future-generation supercomputers. While modern processor design provides a wide range of mechanisms for power and energy optimization, it remains unclear how software can make the best use of them. This article presents a novel approach for runtime power optimization on modern multi-core systems. Our policy combines power capping and uncore frequency scaling to match the hardware power profile to the dynamically changing program behavior at runtime. We achieve this by employing reinforcement learning (RL) to automatically explore the energy-performance optimization space from training programs, learning the subtle relationships between the hardware power profile, the program characteristics, power consumption and program running times. Our RL framework then uses the learned knowledge to adapt the chip's power budget and uncore frequency to match the changing program phases for any new, previously unseen program. We evaluate our approach on two computing clusters by applying our techniques to 11 parallel programs that were not seen by our RL framework at the training stage. Experimental results show that our approach can reduce the system-level energy consumption by 12 percent, on average, with less than 3 percent of slowdown on the application performance. By lowering the uncore frequency to leave more energy budget to allow the processor cores to run at a higher frequency, our approach can reduce the energy consumption by up to 17 percent while improving the application performance by 5 percent for specific workloads. Yiming Wang 0010, Weizhe Zhang, Meng Hao 0002, Zheng Wang 0001 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2022 | Dynamic GPU Energy Optimization for Machine Learning Training WorkloadsabstractGPUs are widely used to accelerate the training of machine learning workloads. As modern machine learning models become increasingly larger, they require a longer time to train, leading to higher GPU energy consumption. This paper presents GPOEO, an online GPU energy optimization framework for machine learning training workloads. GPOEO dynamically determines the optimal energy configuration by employing novel techniques for online measurement, multi-objective prediction modeling, and search optimization. To characterize the target workload behavior, GPOEO utilizes GPU performance counters. To reduce the performance counter profiling overhead, it uses an analytical model to detect the training iteration change and only collects performance counter data when an iteration shift is detected. GPOEO employs multi-objective models based on gradient boosting and a local search algorithm to find a trade-off between execution time and energy consumption. We evaluate the GPOEO by applying it to 71 machine learning workloads from two AI benchmark suites running on an NVIDIA RTX3080Ti GPU. Compared with the NVIDIA default scheduling strategy, GPOEO delivers a mean energy saving of 16.2% with a modest average execution time increase of 5.1%. Farui Wang, Weizhe Zhang, Shichao Lai, Meng Hao 0002, Zheng Wang 0001 |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2021 | Optimizing Barrier Synchronization on ARMv8 Many-Core ArchitecturesabstractSynchronization operations are commonly seen in OpenMP programs where a parallel construct often works with an explicit or implicit barrier operation. While OpenMP synchronization has been extensively studied on the traditional x86 CPU architectures, there is little work on understanding OpenMP barrier synchronization operations on ARMv8 high-performance many-cores. This paper presents the first comprehensive performance study on OpenMP barrier implementations on emerging ARMvS-based many-cores. We evaluate seven representative barrier algorithms on three distinct ARMv8 architectures: Phytium 2000+, ThunderX2, and Kunpeng920. We empirically show that the existing synchronization implementations exhibit poor scalability on ARMv8 architectures compared to the x86 counterpart. We then propose various optimization strategies for improving these widely used synchronization algorithms on each platform. We showcase that our optimizations yield 12.6x performance improvement over the GCC implementation and 4.7x improvement over the LLVM implementation, translating to 1.6x improvement over the state-of-the-art best-performing algorithm. We share our experience and practical insights on optimizing OpenMP synchronization operations on emerging ARMv8 multi-core CPU architectures. Wanrong Gao, Jianbin Fang, Chun Huang 0006, Chuanfu Xu, Zheng Wang 0001 |
CLUSTER | 5 |
| 2021 | HyFM: function merging for freeabstractFunction merging is an important optimization for reducing code size. It merges multiple functions into a single one, eliminating duplicate code among them. The existing state-of-the-art relies on a well-known sequence alignment algorithm to identify duplicate code across whole functions. However, this algorithm is quadratic in time and space on the number of instructions. This leads to very high time overheads and prohibitive levels of memory usage even for medium-sized benchmarks. For larger programs, it becomes impractical. Rodrigo Caetano Rocha, Pavlos Petoumenos, Zheng Wang 0001, Murray Cole, Kim M. Hazelwood, Hugh Leather |
LCTES | 3 |
| 2021 | RISE: robust wireless sensing using probabilistic and statistical assessmentsabstractWireless sensing builds upon machine learning shows encouraging results. However, adopting wireless sensing as a large-scale solution remains challenging as experiences from deployments have shown the performance of a machine-learned model to suffer when there are changes in the environment, e.g., when furniture is moved or when other objects are added or removed from the environment. We present Rise, a novel solution for enhancing the robustness and performance of learning-based wireless sensing techniques against such changes during a deployment. Rise combines probability and statistical assessments together with anomaly detection to identify samples that are likely to be misclassified and uses feedback on these samples to update a deployed wireless sensing model. We validate Rise through extensive empirical benchmarks by considering 11 representative sensing methods covering a broad range of wireless sensing tasks. Our results show that Rise can identify 92.3% of misclassifications on average. We showcase how Rise can be combined with incremental learning to help wireless sensing models retain their performance against dynamic changes in the operating environment to reduce the maintenance cost, paving the way for learning-based wireless sensing to become capable of supporting long-term monitoring in complex everyday environments. Shuangjiao Zhai, Zhanyong Tang, Petteri Nurmi, Dingyi Fang, Xiaojiang Chen, Zheng Wang 0001 |
MobiCom | 6 |
| 2021 | Automated conformance testing for JavaScript engines via deep compiler fuzzingabstractJavaScript (JS) is a popular, platform-independent programming language. To ensure the interoperability of JS programs across different platforms, the implementation of a JS engine should conform to the ECMAScript standard. However, doing so is challenging as there are many subtle definitions of API behaviors, and the definitions keep evolving. Guixin Ye, Zhanyong Tang, Shin Hwei Tan, Songfang Huang, Dingyi Fang, Lizhong Bian, Zheng Wang 0001 |
PLDI | 9 |
| 2021 | LIBSHALOM: optimizing small and irregular-shaped matrix multiplications on ARMv8 multi-coresabstractGeneral Matrix Multiplication (GEMM) is a key subroutine in highperformance computing. While the mainstream linear algebra libraries can deliver high performance on large and regular-shaped GEMM, they are inadequate for optimizing small and irregular-shaped GEMMs, which are commonly seen in new HPC applications. Some of the recent works in this direction have made promising progress on x86 architectures and GPUs but still leave much room for improvement on emerging HPC hardware built upon the ARMv8 architecture. We present LibShalom, an open-source library for optimizing small and irregular-shaped GEMMs, explicitly targeting the ARMv8 architecture. LibShalom builds upon the classical Goto algorithm but tailors it to minimize the expensive memory accessing overhead for data packing and processing small matrices. It uses analytic methods to determine GEMM kernel optimization parameters, enhancing the computation and parallelization efficiency of the GEMM kernels. We evaluate LibShalom by applying it to three ARMv8 multi-core architectures and comparing it against five mainstream linear algebra libraries. Experimental results show that LibShalom can consistently outperform existing solutions across GEMM workloads and hardware architectures. Weiling Yang, Jianbin Fang, Dezun Dong, Xing Su 0004, Zheng Wang 0001 |
SC | 5 |
| 2021 | Towards practical 3D ultrasound sensing on commercial-off-the-shelf mobile devices
Shuangjiao Zhai, Guixin Ye, Zhanyong Tang, Jie Ren 0007, Dingyi Fang, Baoying Liu, Zheng Wang 0001 |
Comput. Networks | 7 |
| 2021 | Combining Graph-Based Learning With Automated Data Collection for Code Vulnerability DetectionabstractThis paper presents FUNDED (Flow-sensitive vUl-Nerability coDE Detection), a novel learning framework for building vulnerability detection models. Funded leverages the advances in graph neural networks (GNNs) to develop a novel graph-based learning method to capture and reason about the program's control, data, and call dependencies. Unlike prior work that treats the program as a sequential sequence or an untyped graph, Funded learns and operates on a graph representation of the program source code, in which individual statements are connected to other statements through relational edges. By capturing the program syntax, semantics and flows, Funded finds better code representation for the downstream software vulnerability detection task. To provide sufficient training data to build an effective deep learning model, we combine probabilistic learning and statistical assessments to automatically gather high-quality training samples from open-source projects. This provides many real-life vulnerable code training samples to complement the limited vulnerable code samples available in standard vulnerability databases. We apply Funded to identify software vulnerabilities at the function level from program source code. We evaluate Funded on large real-world datasets with programs written in C, Java, Swift and Php, and compare it against six state-of-the-art code vulnerability detection models. Experimental results show that Funded significantly outperforms alternative approaches across evaluation settings. Huanting Wang, Guixin Ye, Zhanyong Tang, Shin Hwei Tan, Songfang Huang, Dingyi Fang, Yansong Feng 0002, Lizhong Bian, Zheng Wang 0001 |
IEEE Trans. Inf. Forensics Secur. | 9 |
| 2021 | Automatic translation of data parallel programs for heterogeneous parallelism through OpenMP offloading
Farui Wang, Weizhe Zhang, Meng Hao 0002, Gangzhao Lu, Zheng Wang 0001 |
J. Supercomput. | 6 |
| 2021 | eICIC Configuration of Downlink and Uplink Decoupling With SWIPT in 5G Dense IoT HetNetsabstractInterference management and power transfer can provide a significant improvement over the 5th generation mobile networks (5G) dense Internet of Things (IoT) heterogeneous networks (HetNets). In this paper, we present a novel approach to simultaneously manage inferences at the downlink (DL) and uplink (UL), and to identify opportunities for power transfer and additional UL transmissions integrated with existing protocols and infrastructures for enhanced inter-cell interference coordination (eICIC) protocol in dense IoT HetNets, while considering practical non-linear energy harvesting (EH) model. The design is formulated as the joint optimization of interference aware UL/DL decoupling, airtime resource allocation and energy transfer. The key insight of our algorithm is to translate the original, intractable joint-optimization problem into a problem space where a good approximate solution can be quickly found. We evaluate our scheme through theoretical analysis and simulation. The evaluation shows that our approach improves the system utility by over 20% compared to start-of-the-art in dense IoT HetNets. Compared to alternative schemes, our approach maintains the best user fairness and rate experience and can solve the problem in a fast and scalable way. Jie Zheng 0005, Haijun Zhang 0001, Dusit Niyato, Jie Ren 0007, Hai Wang 0010, Zheng Wang 0001 |
IEEE Trans. Wirel. Commun. | 8 |
| 2020 | Deep Program Structure Modeling Through Multi-Relational Graph-based LearningabstractDeep learning is emerging as a promising technique for building predictive models to support code-related tasks like performance optimization and code vulnerability detection. One of the critical aspects of building a successful predictive model is having the right representation to characterize the model input for the given task. Existing approaches in the area typically treat the program structure as a sequential sequence but fail to capitalize on the rich semantics of data and control flow information, for which graphs are a proven representation structure. Guixin Ye, Zhanyong Tang, Huanting Wang, Dingyi Fang, Jianbin Fang, Songfang Huang, Zheng Wang 0001 |
PACT | 7 |
| 2020 | Neighborhood Matching Network for Entity AlignmentabstractStructural heterogeneity between knowledge graphs is an outstanding challenge for entity alignment. This paper presents Neighborhood Matching Network (NMN), a novel entity alignment framework for tackling the structural heterogeneity challenge. NMN estimates the similarities between entities to capture both the topological structure and the neighborhood difference. It provides two innovative components for better learning representations for entity alignment. It first uses a novel graph sampling method to distill a discriminative neighborhood for each entity. It then adopts a cross-graph neighborhood matching module to jointly encode the neighborhood difference for a given entity pair. Such strategies allow NMN to effectively construct matching-oriented entity representations while ignoring noisy neighbors that have a negative impact on the alignment task. Extensive experiments performed on three entity alignment datasets show that NMN can well estimate the neighborhood similarity in more tough cases and significantly outperforms 12 previous state-of-the-art methods. Xiao Liu 0032, Yansong Feng 0002, Zheng Wang 0001, Dongyan Zhao 0001 |
ACL | 4 |
| 2020 | Vectorization-aware loop unrolling with seed forwardingabstractLoop unrolling is a widely adopted loop transformation, commonly used for enabling subsequent optimizations. Straight-line-code vectorization (SLP) is an optimization that benefits from unrolling. SLP converts isomorphic instruction sequences into vector code. Since unrolling generates repeatead isomorphic instruction sequences, it enables SLP to vectorize more code. However, most production compilers apply these optimizations independently and uncoordinated. Unrolling is commonly tuned to avoid code bloat, not maximizing the potential for vectorization, leading to missed vectorization opportunities. Rodrigo Caetano Rocha, Vasileios Porpodas, Pavlos Petoumenos, Fabrício Góes, Zheng Wang 0001, Murray Cole, Hugh Leather |
CC | 5 |
| 2020 | Optimizing GPU Memory Transactions for Convolution OperationsabstractConvolution computation is a common operation in deep neural networks (DNNs) and is often responsible for performance bottlenecks during training and inferencing. Existing approaches for accelerating convolution operations aim to reduce computational complexity. However, these strategies often increase the memory footprint with extra memory accesses, thereby leaving much room for performance improvement. This paper presents a novel approach to optimize memory access for convolution operations, specifically targeting GPU execution. Our approach leverages two optimization techniques to reduce the number of memory operations for convolution operations performed on the width and height dimensions. For convolution computations on the width dimension, we exploit shuffle instructions to exchange the overlapped columns of the input for reducing the number of memory transactions. For convolution operations on the height dimension, we multiply each overlapped row of the input with multiple rows of a filter to compute multiple output elements to improve the data locality of row elements. We apply our approach to 2D and multi-channel 2D convolutions on an NVIDIA 2080Ti GPU. For 2D convolution, our approach delivers over faster performance than the state-of-the-art image processing libraries. For multi-channel 2D convolutions, we obtain up to speedups over the quickest algorithm of cuDNN. We apply our approach to 2D and multi-channel 2D convolutions on an NVIDIA 2080Ti GPU. For 2D convolution, our approach delivers over 2× faster performance than the state-of-the-art image processing libraries. For multi-channel 2D convolutions, we obtain up to 1.3× speedups over the quickest algorithm of cuDNN. Gangzhao Lu, Weizhe Zhang, Zheng Wang 0001 |
CLUSTER | 3 |
| 2020 | FlowGAN: A Conditional Generative Adversarial Network for Flow Prediction in Various ConditionsabstractMany flow-related design optimization problems like aircraft and automobile aerodynamic design are solved via computational fluid dynamics (CFD) simulations. However, CFD simulations are known to be resource-demanding and time-consuming. Deep learning (DL) is emerging as a viable means to accelerate CFD simulations by directly predicting the outcomes of multiple simulation iterations. While promising, existing DL-based models have to be re-trained whenever the flow condition changes, which incurs significant training overhead for real-life scenarios with a wide range of flow conditions. This paper presents FLOWGAN, a novel conditional generative adversarial network for accurate prediction of flow fields in various conditions. FlowGAN is designed to directly obtain the generation of solutions to flow fields in various conditions based on observations rather than re-training. As FlowGAN does not rely on knowledge of the underlying governing equations, it can quickly adapt to various flow conditions and avoid the need for expensive re-training. We evaluate FlowGAN by applying it to scenarios of simulating both the whole flow field and selected regions of interest (RoI). Compared to the state-of-the-art DL based methods, FlowGAN significantly reduces the prediction errors by 2.27% while exhibiting a better generalization ability. Donglin Chen, Xiang Gao 0020, Chuanfu Xu, Shizhao Chen, Jianbin Fang, Zhenghua Wang, Zheng Wang 0001 |
ICTAI | 7 |
| 2020 | Camel: Smart, Adaptive Energy Optimization for Mobile Web InteractionsabstractWeb technology underpins many interactive mobile applications. However, energy-efficient mobile web interactions is an outstanding challenge. Given the increasing diversity and complexity of mobile hardware, any practical optimization scheme must work for a wide range of users, mobile platforms and web workloads. This paper presents CAMEL, a novel energy optimization system for mobile web interactions. CAMEL leverages machine learning techniques to develop a smart, adaptive scheme to judiciously trade performance for reduced power consumption. Unlike prior work, CAMEL directly models how a given web content affects the user expectation and uses this to guide energy optimization. It goes further by employing transfer learning and conformal predictions to tune a previously learned model in the end-user environment and improve it over time. We apply CAMEL to Chromium and evaluate it on four distinct mobile systems involving 1,000 testing webpages and 30 users. Compared to four state-of-the-art web-event optimizers, CAMEL delivers 22% more energy savings, but with 49% fewer violations on the quality of user experience, and exhibits orders of magnitudes less overhead when targeting a new computing environment. Jie Ren 0007, Petteri Nurmi, Miao Ma, Zhanyong Tang, Jie Zheng 0005, Zheng Wang 0001 |
INFOCOM | 9 |
| 2020 | Effective function merging in the SSA formabstractFunction merging is an important optimization for reducing code size. This technique eliminates redundant code across functions by merging them into a single function. While initially limited to identical or trivially similar functions, the most recent approach can identify all merging opportunities in arbitrary pairs of functions. However, this approach has a serious limitation which prevents it from reaching its full potential. Because it cannot handle phi-nodes, the state-of-the-art applies register demotion to eliminate them before applying its core algorithm. While a superficially minor workaround, this has a three-fold negative effect: by artificially lengthening the instruction sequences to be aligned, it hinders the identification of mergeable instruction; it prevents a vast number of functions from being profitably merged; it increases compilation overheads, both in terms of compile-time and memory usage. Rodrigo Caetano Rocha, Pavlos Petoumenos, Zheng Wang 0001, Murray Cole, Hugh Leather |
PLDI | 3 |
| 2020 | Parallel programming models for heterogeneous many-cores: a comprehensive survey
Jianbin Fang, Chun Huang 0006, Tao Tang 0001, Zheng Wang 0001 |
CCF Trans. High Perform. Comput. | 4 |
| 2020 | Compile-time code virtualization for android applications
Zhanyong Tang, Guixin Ye, Dongxu Peng, Dingyi Fang, Xiaojiang Chen, Zheng Wang 0001 |
Comput. Secur. | 7 |
| 2020 | Semantics-aware obfuscation scheme prediction for binary
Zhanyong Tang, Guixin Ye, Dongxu Peng, Dingyi Fang, Xiaojiang Chen, Zheng Wang 0001 |
Comput. Secur. | 7 |
| 2020 | Optimizing Deep Learning Inference on Embedded Systems Through Adaptive Model SelectionabstractDeep neural networks (DNNs) are becoming a key enabling technique for many application domains. However, on-device inference on battery-powered, resource-constrained embedding systems is often infeasible due to prohibitively long inferencing time and resource requirements of many DNNs. Offloading computation into the cloud is often unacceptable due to privacy concerns, high latency, or the lack of connectivity. Although compression algorithms often succeed in reducing inferencing times, they come at the cost of reduced accuracy. This article presents a new, alternative approach to enable efficient execution of DNNs on embedded devices. Our approach dynamically determines which DNN to use for a given input by considering the desired accuracy and inference time. It employs machine learning to develop a low-cost predictive model to quickly select a pre-trained DNN to use for a given input and the optimization constraint. We achieve this first by offline training a predictive model and then using the learned model to select a DNN model to use for new, unseen inputs. We apply our approach to two representative DNN domains: image classification and machine translation. We evaluate our approach on a Jetson TX2 embedded deep learning platform and consider a range of influential DNN models including convolutional and recurrent neural networks. For image classification, we achieve a 1.8x reduction in inference time with a 7.52% improvement in accuracy over the most capable single DNN model. For machine translation, we achieve a 1.34x reduction in inference time over the most capable single model with little impact on the quality of translation. Vicent Sanz Marco, Ben Taylor 0001, Zheng Wang 0001, Yehia El-khatib |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2020 | Using Generative Adversarial Networks to Break and Protect Text CaptchasabstractText-based CAPTCHAs remains a popular scheme for distinguishing between a legitimate human user and an automated program. This article presents a novel genetic text captcha solver based on the generative adversarial network. As a departure from prior text captcha solvers that require a labor-intensive and time-consuming process to construct, our scheme needs significantly fewer real captchas but yields better performance in solving captchas. Our approach works by first learning a synthesizer to automatically generate synthetic captchas to construct a base solver. It then improves and fine-tunes the base solver using a small number of labeled real captchas. As a result, our attack requires only a small set of manually labeled captchas, which reduces the cost of launching an attack on a captcha scheme. We evaluate our scheme by applying it to 33 captcha schemes, of which 11 are currently used by 32 of the top-50 popular websites. Experimental results demonstrate that our scheme significantly outperforms four prior captcha solvers and can solve captcha schemes where others fail. As a countermeasure, we propose to add imperceptible perturbations onto a captcha image. We demonstrate that our countermeasure can greatly reduce the success rate of the attack. Guixin Ye, Zhanyong Tang, Dingyi Fang, Zhanxing Zhu, Yansong Feng 0002, Pengfei Xu 0003, Xiaojiang Chen, Jungong Han, Zheng Wang 0001 |
ACM Trans. Priv. Secur. | 9 |
| 2020 | Optimizing Streaming Parallelism on Heterogeneous Many-Core ArchitecturesabstractAs many-core accelerators keep integrating more processing units, it becomes increasingly more difficult for a parallel application to make effective use of all available resources. An effective way of improving hardware utilization is to exploit spatial and temporal sharing of the heterogeneous processing units by multiplexing computation and communication tasks - a strategy known as heterogeneous streaming. Achieving effective heterogeneous streaming requires carefully partitioning hardware among tasks, and matching the granularity of task parallelism to the resource partition. However, finding the right resource partitioning and task granularity is extremely challenging, because there is a large number of possible solutions and the optimal solution varies across programs and datasets. This article presents an automatic approach to quickly derive a good solution for hardware resource partition and task granularity for task-based parallel applications on heterogeneous many-core architectures. Our approach employs a performance model to estimate the resulting performance of the target application under a given resource partition and task granularity configuration. The model is used as a utility to quickly search for a good configuration at runtime. Instead of hand-crafting an analytical model that requires expert insights into low-level hardware details, we employ machine learning techniques to automatically learn it. We achieve this by first learning a predictive model offline using training programs. The learned model can then be used to predict the performance of any unseen program at runtime. We apply our approach to 39 representative parallel applications and evaluate it on two representative heterogeneous many-core platforms: a CPU-XeonPhi platform and a CPU-GPU platform. Compared to the single-stream version, our approach achieves, on average, a 1.6x and 1.1x speedup on the XeonPhi and the GPU platform, respectively. These results translate to over 93 percent of the performance delivered by a theoretically perfect predictor. Peng Zhang 0061, Jianbin Fang, Canqun Yang, Chun Huang 0006, Tao Tang 0001, Zheng Wang 0001 |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2019 | Dual-View Ranking with Hardness Assessment for Zero-Shot LearningabstractZero-shot learning (ZSL) is to build recognition models for previously unseen target classes which have no labeled data for training by transferring knowledge from some other related auxiliary source classes with abundant labeled samples to the target ones with class attributes as the bridge. The key is to learn a similarity based ranking function between samples and class labels using the labeled source classes so that the proper (unseen) class label for a test sample can be identified by the function. In order to learn the function, single-view ranking based loss is widely used which aims to rank the true label prior to the other labels for a training sample. However, we argue that the ranking can be performed from the other view, which aims to place the images belonging to a label before the images from the other classes. Motivated by it, we propose a novel DuAl-view RanKing (DARK) loss for zeroshot learning simultaneously ranking labels for an image by point-to-point metric and ranking images for a label by pointto-set metric, which is capable of better modeling the relationship between images and classes. In addition, we also notice that previous ZSL approaches mostly fail to well exploit the hardness of training samples, either using only very hard ones or using all samples indiscriminately. In this work, we also introduce a sample hardness assessment method to ZSL which assigns different weights to training samples based on their hardness, which leads to a more accurate and robust ZSL model. Experiments on benchmarks demonstrate that DARK outperforms the state-of-the-arts for (generalized) ZSL. Guiguang Ding, Jungong Han, Xiaohan Ding, Sicheng Zhao, Zheng Wang 0001, Chenggang Yan 0001, Qionghai Dai |
AAAI | 6 |
| 2019 | Lattice CNNs for Matching Based Chinese Question AnsweringabstractShort text matching often faces the challenges that there are great word mismatch and expression diversity between the two texts, which would be further aggravated in languages like Chinese where there is no natural space to segment words explicitly. In this paper, we propose a novel lattice based CNN model (LCNs) to utilize multi-granularity information inherent in the word lattice while maintaining strong ability to deal with the introduced noisy information for matching based question answering in Chinese. We conduct extensive experiments on both document based question answering and knowledge based question answering tasks, and experimental results show that the LCNs models can significantly outperform the state-of-the-art matching models and strong baselines by taking advantages of better ability to distill rich but discriminative information from the word lattice input. Yuxuan Lai, Yansong Feng 0002, Xiaohan Yu 0005, Zheng Wang 0001, Kun Xu 0005, Dongyan Zhao 0001 |
AAAI | 4 |
| 2019 | Function Merging by Sequence AlignmentabstractResource-constrained devices for embedded systems are becoming increasingly important. In such systems, memory is highly restrictive, making code size in most cases even more important than performance. Compared to more traditional platforms, memory is a larger part of the cost and code occupies much of it. Despite that, compilers make little effort to reduce code size. One key technique attempts to merge the bodies of similar functions. However, production compilers only apply this optimization to identical functions, while research compilers improve on that by merging the few functions with identical control-flow graphs and signatures. Overall, existing solutions are insufficient and we end up having to either increase cost by adding more memory or remove functionality from programs. We introduce a novel technique that can merge arbitrary functions through sequence alignment, a bioinformatics algorithm for identifying regions of similarity between sequences. We combine this technique with an intelligent exploration mechanism to direct the search towards the most promising function pairs. Our approach is more than 2.4x better than the state-of-the-art, reducing code size by up to 25%, with an overall average of 6%, while introducing an average compilation-time overhead of only 15%. When aided by profiling information, this optimization can be deployed without any significant impact on the performance of the generated code. Rodrigo Caetano Rocha, Pavlos Petoumenos, Zheng Wang 0001, Murray Cole, Hugh Leather |
CGO | 3 |
| 2019 | Jointly Learning Entity and Relation Representations for Entity AlignmentabstractYuting Wu, Xiao Liu, Yansong Feng, Zheng Wang, Dongyan Zhao. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Xiao Liu 0032, Yansong Feng 0002, Zheng Wang 0001, Dongyan Zhao 0001 |
EMNLP/IJCNLP (1) | 4 |
| 2019 | Relation-Aware Entity Alignment for Heterogeneous Knowledge GraphsabstractEntity alignment is the task of linking entities with the same real-world identity from different knowledge graphs (KGs), which has been recently dominated by embedding-based methods. Such approaches work by learning KG representations so that entity alignment can be performed by measuring the similarities between entity embeddings. While promising, prior works in the field often fail to properly capture complex relation information that commonly exists in multi-relational KGs, leaving much room for improvement. In this paper, we propose a novel Relation-aware Dual-Graph Convolutional Network (RDGCN) to incorporate relation information via attentive interactions between the knowledge graph and its dual relation counterpart, and further capture neighboring structures to learn better entity representations. Experiments on three real-world cross-lingual datasets show that our approach delivers better and more robust results over the state-of-the-art alignment methods by learning better KG representations. Xiao Liu 0032, Yansong Feng 0002, Zheng Wang 0001, Rui Yan 0001, Dongyan Zhao 0001 |
IJCAI | 4 |
| 2019 | WideSee: towards wide-area contactless wireless sensingabstractContactless wireless sensing without attaching a device to the target has achieved promising progress in recent years. However, one severe limitation is the small sensing range. This paper presents WideSee to realize wide-area sensing with only one transceiver pair. WideSee utilizes the LoRa signal to achieve a larger range of sensing and further incorporates drone's mobility to broaden the sensing area. WideSee presents solutions across software and hardware to overcome two aspects of challenges for wide-range contactless sensing: (i) the interference brought by the device mobility and LoRa's high sensitivity; and (ii) the ambiguous target information such as location when employing just a single pair of transceivers. We have developed a working prototype of WideSee for human target detection and localization that are especially useful in emergency scenarios such as rescue search, and evaluated WideSee with both controlled experiments and the field study in a high-rise building. Extensive experiments demonstrate the great potential of WideSee for wide-area contactless sensing with a single LoRa transceiver pair hosted on a drone. Jie Xiong 0001, Xiaojiang Chen, Sunghoon Ivan Lee, Dianhe Han, Dingyi Fang, Zhanyong Tang, Zheng Wang 0001 |
SenSys | 9 |
| 2019 | Towards wide-area contactless human sensing: poster abstractabstractContactless wireless sensing without attaching a device to the target has achieved promising progress in recent years. However, one severe limitation in this field is the limited sensing range. This paper presents WideSee to realize wide-area sensing with only one transceiver pair. WideSee utilizes the LoRa signal to achieve a larger range of sensing and further incorporates drone's mobility to broaden the sensing area. We have developed a working prototype of WideSee for human target detection and localization that are especially useful in emergency scenarios like rescue and terrorist search. We also evaluated WideSee with field study in a high-rise building, which demonstrates the great potential of WideSee for supporting wide-area contactless sensing applications with a single LoRa transceiver pair hosted on a drone. Dianhe Han, Jie Xiong 0001, Sunghoon Ivan Lee, Xiaojiang Chen, Zhanyong Tang, Dingyi Fang, Zheng Wang 0001 |
SenSys | 10 |
| 2019 | Find me a safe zone: A countermeasure for channel state information based attacks
Jie Zhang 0028, Zhanyong Tang, Meng Li 0006, Dingyi Fang, Xiaojiang Chen, Zheng Wang 0001 |
Comput. Secur. | 6 |
| 2018 | Scale Up Event Extraction Learning via Automatic Training Data GenerationabstractThe task of event extraction has long been investigated in a supervised learning paradigm, which is bound by the number and the quality of the training instances. Existing training data must be manually generated through a combination of expert domain knowledge and extensive human involvement. However, due to drastic efforts required in annotating text, the resultant datasets are usually small, which severally affects the quality of the learned model, making it hard to generalize. Our work develops an automatic approach for generating training data for event extraction. Our approach allows us to scale up event extraction training instances from thousands to hundreds of thousands, and it does this at a much lower cost than a manual approach. We achieve this by employing distant supervision to automatically create event annotations from unlabelled text using existing structured knowledge bases or tables.We then develop a neural network model with post inference to transfer the knowledge extracted from structured knowledge bases to automatically annotate typed events with corresponding arguments in text.We evaluate our approach by using the knowledge extracted from Freebase to label texts from Wikipedia articles. Experimental results show that our approach can generate a large number of highquality training instances. We show that this large volume of training data not only leads to a better event extractor, but also allows us to detect multiple typed events. Yansong Feng 0002, Zheng Wang 0001, Rui Yan 0001, Chongde Shi, Dongyan Zhao 0001 |
AAAI | 4 |
| 2018 | Marrying Up Regular Expressions with Neural Networks: A Case Study for Spoken Language UnderstandingabstractThe success of many natural language processing (NLP) tasks is bound by the number and quality of annotated data, but there is often a shortage of such training data.In this paper, we ask the question: "Can we combine a neural network (NN) with regular expressions (RE) to improve supervised learning for NLP?".In answer, we develop novel methods to exploit the rich expressiveness of REs at different levels within a NN, showing that the combination significantly enhances the learning effectiveness when a small number of training examples are available.We evaluate our approach by applying it to spoken language understanding for intent detection and slot filling.Experimental results show that our approach is highly effective in exploiting the available training data, giving a clear boost to the RE-unaware NN. Bingfeng Luo, Yansong Feng 0002, Zheng Wang 0001, Songfang Huang, Rui Yan 0001, Dongyan Zhao 0001 |
ACL (1) | 3 |
| 2018 | Yet Another Text Captcha Solver: A Generative Adversarial Network Based ApproachabstractDespite several attacks have been proposed, text-based CAPTCHAs are still being widely used as a security mechanism. One of the reasons for the pervasive use of text captchas is that many of the prior attacks are scheme-specific and require a labor-intensive and time-consuming process to construct. This means that a change in the captcha security features like a noisier background can simply invalid an earlier attack. This paper presents a generic, yet effective text captcha solver based on the generative adversarial network. Unlike prior machine-learning-based approaches that need a large volume of manually-labeled real captchas to learn an effective solver, our approach requires significantly fewer real captchas but yields much better performance. This is achieved by first learning a captcha synthesizer to automatically generate synthetic captchas to learn a base solver, and then fine-tuning the base solver on a small set of real captchas using transfer learning. We evaluate our approach by applying it to 33 captcha schemes, including 11 schemes that are currently being used by 32 of the top-50 popular websites including Microsoft, Wikipedia, eBay and Google. Our approach is the most capable attack on text captchas seen to date. It outperforms four state-of-the-art text-captcha solvers by not only delivering a significant higher accuracy on all testing schemes, but also successfully attacking schemes where others have zero chance. We show that our approach is highly efficient as it can solve a captcha within 0.05 second using a desktop GPU. We demonstrate that our attack is generally applicable because it can bypass the advanced security features employed by most modern text captcha schemes. We hope the results of our work can encourage the community to revisit the design and practical use of text captchas. Guixin Ye, Zhanyong Tang, Dingyi Fang, Zhanxing Zhu, Yansong Feng 0002, Pengfei Xu 0003, Xiaojiang Chen, Zheng Wang 0001 |
CCS | 8 |
| 2018 | MOCL: an efficient openCL implementation for the matrix-2000 architectureabstractThis paper presents the design and implementation of an Open Computing Language (OpenCL) framework for the Matrix-2000 many-core architecture. This architecture is designed to replace the Intel XeonPhi accelerators of the TianHe-2 supercomputer. We share our experience and insights on how to design an effective OpenCL system for this new hardware accelerator. We propose a set of new analysis and optimizations to unlock the potential of the hardware. We extensively evaluate our approach using a wide range of OpenCL benchmarks on a single and multiple computing nodes. We present our design choices and provide guidance how to optimize code on the new Matrix-2000 architecture. Peng Zhang 0061, Tao Tang 0001, Jianbin Fang, Chun Huang 0006, Canqun Yang, Zheng Wang 0001 |
CF | 6 |
| 2018 | Proteus: network-aware web browsing on heterogeneous mobile systemsabstractWe present Proteus, a novel network-aware approach for optimizing web browsing on heterogeneous multi-core mobile systems. It employs machine learning techniques to predict which of the heterogeneous cores to use to render a given webpage and the operating frequencies of the processors. It achieves this by first learning offline a set of predictive models for a range of typical networking environments. A learnt model is then chosen at runtime to predict the optimal processor configuration, based on the web content, the network status and the optimization goal. We evaluate Proteus by implementing it into the open-source Chromium browser and testing it on two representative ARM big.LITTLE mobile multi-core platforms. We apply Proteus to the top 1,000 popular websites across seven typical network environments. Proteus achieves over 80% of best available performance. It obtains, on average, over 17% (up to 63%), 31% (up to 88%), and 30% (up to 91%) improvement respectively for load time, energy consumption and the energy delay product, when compared to two state-of-the-art approaches. Jie Ren 0007, Jianbin Fang, Yansong Feng 0002, Dongxiao Zhu, Zhunchen Luo, Jie Zheng 0005, Zheng Wang 0001 |
CoNEXT | 8 |
| 2018 | Exploiting Code Diversity to Enhance Code Virtualization ProtectionabstractCode virtualization built upon virtual machine (VM)technologies is emerging as a viable method for implementing code obfuscation to protect programs against unauthorized analysis. State-of-the-art VM-based protection approaches use a fixed set of virtual instructions and bytecode interpreters across programs. This, however, exposes a security vulnerability where an experienced attacker can use knowledge extracted from other programs to quickly uncover the mapping between virtual instructions and native code for applications protected under the same scheme. In this paper, we propose a novel VM-based code obfuscation system to address this problem. The core idea of our approach is to obfuscate the mapping between the opcodes of bytecode instructions and their semantics. We achieve this by partitioning each protected code region into multiple segments where the mapping of opcodes and their semantics is randomized in different ways in different segments. In this way, each bytecode instruction will be translated into different native code in different sections of the obfuscated code. This significantly increases the diversity of the program behavior. As a result, the knowledge of bytecode to native code mappings obtained from other programs will be less useful when targeting a new program. We evaluate our approach on a set of real-world applications and compare it against two state-of-the-art VM-based code obfuscation approaches. Experimental results show that our approach is effective, which provides stronger protection with comparable runtime overhead and code size. Zhanyong Tang, Guixin Ye, Xiaoqing Gong, Wei Wangg, Dingyi Fang, Zheng Wang 0001 |
ICPADS | 8 |
| 2018 | Evaluating Brush Movements for Chinese Calligraphy: A Computer Vision Based ApproachabstractChinese calligraphy is a popular, highly esteemed art form in the Chinese cultural sphere and worldwide. Ink brushes are the traditional writing tool for Chinese calligraphy and the subtle nuances of brush movements have a great impact on the aesthetics of the written characters. However, mastering the brush movement is a challenging task for many calligraphy learners as it requires many years’ practice and expert supervision. This paper presents a novel approach to help Chinese calligraphy learners to quantify the quality of brush movements without expert involvement. Our approach extracts the brush trajectories from a video stream; it then compares them with example templates of reputed calligraphers to produce a score for the writing quality. We achieve this by first developing a novel neural network to extract the spatial and temporal movement features from the video stream. We then employ methods developed in the computer vision and signal processing domains to track the brush movement trajectory and calculate the score. We conducted extensive experiments and user studies to evaluate our approach. Experimental results show that our approach is highly accurate in identifying brush movements, yielding an average accuracy of 90%, and the generated score is within 3% of errors when compared to the one given by human experts. Pengfei Xu 0003, Ziyu Guan, Xia Zheng, Xiaojiang Chen, Zhanyong Tang, Dingyi Fang, Xiaoqing Gong, Zheng Wang 0001 |
IJCAI | 9 |
| 2018 | Auto-tuning Streamed Applications on Intel Xeon PhiabstractMany-core accelerators, as represented by the XeonPhi coprocessors and GPGPUs, allow software to exploit spatial and temporal sharing of computing resources to improve the overall system performance. To unlock this performance potential requires software to effectively partition the hardware resource to maximize the overlap between host-device communication and accelerator computation, and to match the granularity of task parallelism to the resource partition. However, determining the right resource partition and task parallelism on a per program, per dataset basis is challenging. This is because the number of possible solutions is huge, and the benefit of choosing the right solution may be large, but mistakes can seriously hurt the performance. In this paper, we present an automatic approach to determine the hardware resource partition and the task granularity for any given streamed application, targeting the Intel XeonPhi architecture. Instead of hand-crafting the heuristic for which the process will have to repeat for each hardware generation, we employ machine learning techniques to automatically learn it. We achieve this by first learning a predictive model offline using training programs; we then use the learned model to predict the resource partition and task granularity for any unseen programs at runtime. We apply our approach to 23 representative parallel applications and evaluate it on a CPU-XeonPhi mixed heterogenous many-core platform. Our approach achieves, on average, a 1.6x (upto 5.6x) speedup, which translates to 94.5% of the performance delivered by a theoretically perfect predictor. Peng Zhang 0061, Jianbin Fang, Tao Tang 0001, Canqun Yang, Zheng Wang 0001 |
IPDPS | 5 |
| 2018 | Adaptive deep learning model selection on embedded systemsabstractThe recent ground-breaking advances in deep learning networks (DNNs) make them attractive for embedded systems. However, it can take a long time for DNNs to make an inference on resource-limited embedded devices. Offloading the computation into the cloud is often infeasible due to privacy concerns, high latency, or the lack of connectivity. As such, there is a critical need to find a way to effectively execute the DNN models locally on the devices. Ben Taylor 0001, Vicent Sanz Marco, Willy Wolff, Yehia El-khatib, Zheng Wang 0001 |
LCTES | 5 |
| 2018 | CrossSense: Towards Cross-Site and Large-Scale WiFi SensingabstractWe present CrossSense, a novel system for scaling up WiFi sensing to new environments and larger problems. To reduce the cost of sensing model training data collection, CrossSense employs machine learning to train, off-line, a roaming model that generates from one set of measurements synthetic training samples for each target environment. To scale up to a larger problem size, CrossSense adopts a mixture-of-experts approach where multiple specialized sensing models, or experts, are used to capture the mapping from diverse WiFi inputs to the desired outputs. The experts are trained offline and at runtime the appropriate expert for a given input is automatically chosen. We evaluate CrossSense by applying it to two representative WiFi sensing applications, gait identification and gesture recognition, in controlled single-link environments. We show that CrossSense boosts the accuracy of state-of-the-art WiFi sensing techniques from 20% to over 80% and 90% for gait identification and gesture recognition respectively, delivering consistently good performance - particularly when the problem size is significantly greater than that current approaches can effectively handle. Jie Zhang 0028, Zhanyong Tang, Meng Li 0006, Dingyi Fang, Petteri Nurmi, Zheng Wang 0001 |
MobiCom | 6 |
| 2018 | Towards Large-Scale RFID Positioning: A Low-cost, High-precision Solution Based on Compressive SensingabstractRFID-based positioning is emerging as a promising solution for inventory management in places like warehouses and libraries. However, existing solutions either are too sensitive to the environmental noise, or require deploying a large number of reference tags which incur expensive deployment cost and increase the chance of data collisions. This paper presents CSRP, a novel RFID based positioning system, which is highly accurate and robust to environmental noise, but relies on much less reference tags compared with the state-of-the-art. CSRP achieves this by employing an noise-resilient RFID fingerprint scheme and a compressive sensing based algorithm that can recover the target tag's position using a small number of signal measurements. This work provides a set of new analysis, algorithms and heuristics to guide the deployment of reference tags and to optimize the computational overhead. We evaluate CSRP in a deployment site with 270 commercial RFID tags. Experimental results show that CSRP can correctly identify 84.7% of the test items, achieving an accuracy that is comparable to the state-of-the-art, using an order of magnitude less reference tags. Liqiong Chang, Xinyi Li 0005, Ju Wang 0003, Haining Meng, Xiaojiang Chen, Dingyi Fang, Zhanyong Tang, Zheng Wang 0001 |
PerCom | 8 |
| 2018 | Enhance virtual-machine-based code obfuscation security through dynamic bytecode schedulingabstractCode virtualization built upon virtual machine (VM) technologies is emerging as a viable method for implementing code obfuscation to protect programs against unauthorized analysis. State-of-the-art VM-based protection approaches use a fixed scheduling structure where the program always follows a single, deterministic execution path for the same input. Such approaches, however, are vulnerable in certain scenarios where the attacker can reuse knowledge extracted from previously seen software to crack applications protected with the same obfuscation scheme. This paper presents Dsvmp, a novel VM-based code obfuscation approach for software protection. Dsvmp brings together two techniques to provide stronger code protection than prior VM-based approaches. Firstly, it uses a dynamic instruction scheduler to randomly direct the program to execute different paths without violating the correctness across different runs. By randomly choosing the program execution path, the application exposes diverse behavior, making it much more difficult for an attacker to reuse the knowledge collected from previous runs or similar applications to launch an attack. Secondly, it employs multiple VMs to further obfuscate the mapping from VM opcode to native machine instructions, so that the same opcode could be mapped to different native instructions at runtime, making code analysis even harder. We have implemented Dsvmp in a prototype system and evaluated it using a set of widely used applications. Experimental results show that Dsvmp provides stronger protection with comparable runtime overhead and code size, when it is compared to two commercial VM-based code obfuscation tools. Kaiyuan Kuang, Zhanyong Tang, Xiaoqing Gong, Dingyi Fang, Xiaojiang Chen, Zheng Wang 0001 |
Comput. Secur. | 6 |
| 2018 | Machine Learning in Compiler OptimizationabstractIn the last decade, machine-learning-based compilation has moved from an obscure research niche to a mainstream activity. In this paper, we describe the relationship between machine learning and compiler optimization and introduce the main concepts of features, models, training, and deployment. We then provide a comprehensive survey and provide a road map for the wide variety of different research areas. We conclude with a discussion on open issues in the area and potential research directions. This paper provides both an accessible introduction to the fast moving area of machine-learning-based compilation and a detailed bibliography of its main achievements. Zheng Wang 0001, Michael F. P. O'Boyle |
Proc. IEEE | 1 |
| 2018 | A Video-based Attack for Android Pattern LockabstractPattern lock is widely used for identification and authentication on Android devices. This article presents a novel video-based side channel attack that can reconstruct Android locking patterns from video footage filmed using a smartphone. As a departure from previous attacks on pattern lock, this new attack does not require the camera to capture any content displayed on the screen. Instead, it employs a computer vision algorithm to track the fingertip movement trajectory to infer the pattern. Using the geometry information extracted from the tracked fingertip motions, the method can accurately infer a small number of (often one) candidate patterns to be tested by an attacker. We conduct extensive experiments to evaluate our approach using 120 unique patterns collected from 215 independent users. Experimental results show that the proposed attack can reconstruct over 95% of the patterns in five attempts. We discovered that, in contrast to most people’s belief, complex patterns do not offer stronger protection under our attacking scenarios. This is demonstrated by the fact that we are able to break all but one complex patterns (with a 97.5% success rate) as opposed to 60% of the simple patterns in the first attempt. We demonstrate that this video-side channel is a serious concern for not only graphical locking patterns but also PIN-based passwords, as algorithms and analysis developed from the attack can be easily adapted to target PIN-based passwords. As a countermeasure, we propose to change the way the Android locking pattern is constructed and used. We show that our proposal can successfully defeat this video-based attack. We hope the results of this article can encourage the community to revisit the design and practical use of Android pattern lock. Guixin Ye, Zhanyong Tang, Dingyi Fang, Xiaojiang Chen, Willy Wolff, Adam J. Aviv, Zheng Wang 0001 |
ACM Trans. Priv. Secur. | 7 |
| 2017 | End-to-End Deep Learning of Optimization HeuristicsabstractAccurate automatic optimization heuristics are necessary for dealing with thecomplexity and diversity of modern hardware and software. Machine learning is aproven technique for learning such heuristics, but its success is bound by thequality of the features used. These features must be hand crafted by developersthrough a combination of expert domain knowledge and trial and error. This makesthe quality of the final model directly dependent on the skill and availabletime of the system architect. Our work introduces a better way for building heuristics. We develop a deepneural network that learns heuristics over raw code, entirely without using codefeatures. The neural network simultaneously constructs appropriaterepresentations of the code and learns how best to optimize, removing the needfor manual feature creation. Further, we show that our neural nets can transferlearning from one optimization problem to another, improving the accuracy of newmodels, without the help of human experts. We compare the effectiveness of our automatically generated heuristics againstones with features hand-picked by experts. We examine two challenging tasks:predicting optimal mapping for heterogeneous parallelism and GPU threadcoarsening factors. In 89% of the cases, the quality of our fully automaticheuristics matches or surpasses that of state-of-the-art predictive models usinghand-crafted features, providing on average 14% and 12% more performance withno human effort expended on designing features. Chris Cummins, Pavlos Petoumenos, Zheng Wang 0001, Hugh Leather |
PACT | 3 |
| 2017 | Learning with Noise: Enhance Distantly Supervised Relation Extraction with Dynamic Transition MatrixabstractBingfeng Luo, Yansong Feng, Zheng Wang, Zhanxing Zhu, Songfang Huang, Rui Yan, Dongyan Zhao. Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2017. Bingfeng Luo, Yansong Feng 0002, Zheng Wang 0001, Zhanxing Zhu, Songfang Huang, Rui Yan 0001, Dongyan Zhao 0001 |
ACL (1) | 3 |
| 2017 | Synthesizing benchmarks for predictive modeling
Chris Cummins, Pavlos Petoumenos, Zheng Wang 0001, Hugh Leather |
CGO | 3 |
| 2017 | Minimizing the cost of iterative compilation with active learning
William F. Ogilvie, Pavlos Petoumenos, Zheng Wang 0001, Hugh Leather |
CGO | 3 |
| 2017 | Real-Time Power Cycling in Video on Demand Data Centres Using Online Bayesian PredictionabstractEnergy usage in data centres continues to be a major and growing concern as an increasing number of everyday services depend on these facilities. Research in this area has examined topics including power smoothing using batteries and deep learning to control cooling systems, in addition to optimisation techniques for the software running inside data centres. We present a novel real-time power-cycling architecture, supported by a media distribution approach and online prediction model, to automatically determine when servers are needed based on demand. We demonstrate with experimental evaluation that this approach can save up to 31% of server energy in a cluster. Our evaluation is conducted on typical rack mount servers in a data centre testbed and uses a recent real-world workload trace from the BBC iPlayer, an extremely popular video on demand service in the UK. Vicent Sanz Marco, Zheng Wang 0001, Barry Porter |
ICDCS | 2 |
| 2017 | AppIS: Protect Android Apps Against Runtime Repackaging AttacksabstractApps repackaged through reverse engineering pose a significant security threat to the Android smart phone ecosystem. Previous solutions have mostly focused on the detection and identification of repackaged apps. Nevertheless, current app anti-repackaging services can only protect applications at a coarse level and get a significant performance overhead. These approaches can neither meet the performance requirements of Android nor achieve fine-grained protection against cumulative attack 1 at the same time. Specifically, these solutions rely on a fix-structure detecting engine and then will execute the same path at different times, which lead to the entire protection performs poorly when faced with dynamic cumulative attack, which is typical in real-world attack. This paper introduces AppIS, a reinforced anti-repackaging immune system, that is robust to app-repackaging attack scenarios. Unlike prior work, which mostly focuses on simple protection only from just one respect, our design exploits an interlocking guarding net with time diversity for the tamper-proofing of Android applications. The intuition underlying our design is that a dynamic and static combining method can provide a multi-level protection for the codes, core algorithm and sensitive data. We analyze and classify the existing threats on Android platform and furthermore abstract then model the repackaging attack scenarios. We then adopt a random controller used by the dispatcher to randomly construct guarding net with different structure every time. We have built a prototype of our design using Java Native Interface cross-layer calling mechanism for performance requirement. Results from a deployment of AppIS on several kinds of popular apps demonstrate that the new design can prevent our apps from cumulative attack without extra performance cost. Lina Song, Zhanyong Tang, Xiaoqing Gong, Xiaojiang Chen, Dingyi Fang, Zheng Wang 0001 |
ICPADS | 7 |
| 2017 | Optimise web browsing on heterogeneous mobile platforms: A machine learning based approachabstractWeb browsing is an activity that billions of mobile users perform on a daily basis. Battery life is a primary concern to many mobile users who often find their phone has died at most inconvenient times. The heterogeneous multi-core architecture is a solution for energy-efficient processing. However, the current mobile web browsers rely on the operating system to exploit the underlying hardware, which has no knowledge of individual web contents and often leads to poor energy efficiency. This paper describes an automatic approach to render mobile web workloads for performance and energy efficiency. It achieves this by developing a machine learning based approach to predict which processor to use to run the web rendering engine and at what frequencies the processors should operate. Our predictor learns offline from a set of training web workloads. The built predictor is then integrated into the browser to predict the optimal processor configuration at runtime, taking into account the web workload characteristics and the optimisation goal: whether it is load time, energy consumption or a trade-off between them. We evaluate our approach on a representative ARM big.LITTLE mobile architecture using the hottest 500 webpages. Our approach achieves 80% of the performance delivered by an ideal predictor. We obtain, on average, 45%, 63.5% and 81% improvement respectively for load time, energy consumption and the energy delay product, when compared to the Linux heterogeneous multi-processing scheduler. Jie Ren 0007, Hai Wang 0010, Zheng Wang 0001 |
INFOCOM | 4 |
| 2017 | Adaptive optimization for OpenCL programs on embedded heterogeneous systemsabstractHeterogeneous multi-core architectures consisting of CPUs and GPUs are commonplace in today’s embedded systems. These architectures offer potential for energy efficient computing if the application task is mapped to the right core. Realizing such potential is challenging due to the complex and evolving nature of hardware and applications. This paper presents an automatic approach to map OpenCL kernels onto heterogeneous multi-cores for a given optimization criterion – whether it is faster runtime, lower energy consumption or a trade-off between them. This is achieved by developing a machine learning based approach to predict which processor to use to run the OpenCL kernel and the host program, and at what frequency the processor should operate. Instead of hand-tuning a model for each optimization metric, we use machine learning to develop a unified framework that first automatically learns the optimization heuristic for each metric off-line, then uses the learned knowledge to schedule OpenCL kernels at runtime based on code and runtime information of the program. We apply our approach to a set of representative OpenCL benchmarks and evaluate it on an ARM big.LITTLE mobile platform. Our approach achieves over 93% of the performance delivered by a perfect predictor.We obtain, on average, 1.2x, 1.6x, and 1.8x improvement respectively for runtime, energy consumption and the energy delay product when compared to a comparative heterogeneous-aware OpenCL task mapping scheme. Ben Taylor 0001, Vicent Sanz Marco, Zheng Wang 0001 |
LCTES | 3 |
| 2017 | Improving spark application throughput via memory aware task co-location: a mixture of experts approachabstractData analytic applications built upon big data processing frameworks such as Apache Spark are an important class of applications. Many of these applications are not latency-sensitive and thus can run as batch jobs in data centers. By running multiple applications on a computing host, task co-location can significantly improve the server utilization and system throughput. However, effective task co-location is a non-trivial task, as it requires an understanding of the computing resource requirement of the co-running applications, in order to determine what tasks, and how many of them, can be co-located. State-of-the-art co-location schemes either require the user to supply the resource demands which are often far beyond what is needed; or use a one-size-fits-all function to estimate the requirement, which, unfortunately, is unlikely to capture the diverse behaviors of applications. Vicent Sanz Marco, Ben Taylor 0001, Barry Porter, Zheng Wang 0001 |
Middleware | 4 |
| 2017 | Cracking Android Pattern Lock in Five Attempts
Guixin Ye, Zhanyong Tang, Dingyi Fang, Xiaojiang Chen, Kwang In Kim, Ben Taylor 0001, Zheng Wang 0001 |
NDSS | 7 |
| 2017 | ALEA: A Fine-Grained Energy Profiling ToolabstractEnergy efficiency is becoming increasingly important, yet few developers understand how source code changes affect the energy and power consumption of their programs. To enable them to achieve energy savings, we must associate energy consumption with software structures, especially at the fine-grained level of functions and loops. Most research in the field relies on direct power/energy measurements taken from on-board sensors or performance counters. However, this coarse granularity does not directly provide the needed fine-grained measurements. This article presents ALEA, a novel fine-grained energy profiling tool based on probabilistic analysis for fine-grained energy accounting. ALEA overcomes the limitations of coarse-grained power-sensing instruments to associate energy information effectively with source code at a fine-grained level. We demonstrate and validate that ALEA can perform accurate energy profiling at various granularity levels on two different architectures: Intel Sandy Bridge and ARM big.LITTLE. ALEA achieves a worst-case error of only 2% for coarse-grained code structures and 6% for fine-grained ones, with less than 1% runtime overhead. Our use cases demonstrate that ALEA supports energy optimizations, with energy savings of up to 2.87 times for a latency-critical option pricing workload under a given power budget. Lev Mukhanov, Pavlos Petoumenos, Zheng Wang 0001, Konstantinos Parasyris, Dimitrios S. Nikolopoulos, Bronis R. de Supinski, Hugh Leather |
ACM Trans. Archit. Code Optim. | 3 |
| 2015 | Power Capping: What Works, What Does NotabstractPeak power consumption is the first order design constraint of data centers. Though peak power consumption is rarely, if ever, observed, the entire data center facility must prepare for it, leading to inefficient usage of its resources. The most prominent way for addressing this issue is to limit the power consumption of the data center IT facility far below its theoretical peak value. Many approaches have been proposed to achieve that, based on the same small set of enforcement mechanisms, but there has been no corresponding work on systematically examining the advantages and disadvantages of each such mechanism. In the absence of such a study, it is unclear what is the optimal mechanism for a given computing environment, which can lead to unnecessarily poor performance if an inappropriate scheme is used. This paper fills this gap by comparing for the first time five widely used power capping mechanisms under the same hardware/software setting. We also explore possible alternative power capping mechanisms beyond what has been previously proposed and evaluate them under the same setup. We systematically analyze the strengths and weaknesses of each mechanism, in terms of energy efficiency, overhead, and predictable behavior. We show how these mechanisms can be combined in order to implement an optimal power capping mechanism which reduces the slowdown compared to the most widely used mechanism by up to 88%. Our results provide interesting insights regarding the different trade-offs of power capping techniques, which will be useful for designing and implementing highly efficient power capping in the future. Pavlos Petoumenos, Lev Mukhanov, Zheng Wang 0001, Hugh Leather, Dimitrios S. Nikolopoulos |
ICPADS | 3 |
| 2014 | Active learning accelerated automatic heuristic construction for parallel program mappingabstractBuilding effective optimization heuristics is a challenging task which often takes developers several months if not years to complete. Predictive modelling has recently emerged as a promising solution, automatically constructing heuristics from training data, however, obtaining this data can take months per platform. This is becoming an ever more critical problem as the pace of change in architecture increases. Indeed, if no solution is found we shall be left with out of date heuristics which cannot extract the best performance from modern machines. William F. Ogilvie, Pavlos Petoumenos, Zheng Wang 0001, Hugh Leather |
PACT | 3 |
| 2014 | Exploitation of GPUs for the Parallelisation of Probably Parallel Legacy Code
Zheng Wang 0001, Daniel Christopher Powell, Björn Franke, Michael F. P. O'Boyle |
CC | 1 |
| 2014 | Smart multi-task scheduling for OpenCL programs on CPU/GPU heterogeneous platformsabstractHeterogeneous systems consisting of multiple CPUs and GPUs are increasingly attractive as platforms for high performance computing. Such platforms are usually programmed using OpenCL which provides program portability by allowing the same program to execute on different types of device. As such systems become more mainstream, they will move from application dedicated devices to platforms that need to support multiple concurrent user applications. Here there is a need to determine when and where to map different applications so as to best utilize the available heterogeneous hardware resources. In this paper, we present an efficient OpenCL task scheduling scheme which schedules multiple kernels from multiple programs on CPU/GPU heterogeneous platforms. It does this by determining at runtime which kernels are likely to best utilize a device. We show that speedup is a good scheduling priority function and develop a novel model that predicts a kernel's speedup based on its static code structure. Our scheduler uses this prediction and runtime input data size to prioritize and schedule tasks. This technique is applied to a large set of concurrent OpenCL kernels. We evaluated our approach for system throughput and average turn-around time against competitive techniques on two different platforms: a Core i7/Nvidia GTX590 and a Core i7/AMD Tahiti 7970 platforms. For system throughput, we achieve, on average, a 1.21x and 1.25x improvement over the best competitors on the NVIDIA and AMD platforms respectively. Our approach reduces the turnaround time, on average, by at least 1.5x and 1.2x on the NVIDIA and AMD platforms respectively, when compared to alternative approaches. Yuan Wen, Zheng Wang 0001, Michael F. P. O'Boyle |
HiPC | 2 |
| 2014 | Automatic and Portable Mapping of Data Parallel Programs to OpenCL for GPU-Based Heterogeneous SystemsabstractGeneral-purpose GPU-based systems are highly attractive, as they give potentially massive performance at little cost. Realizing such potential is challenging due to the complexity of programming. This article presents a compiler-based approach to automatically generate optimized OpenCL code from data parallel OpenMP programs for GPUs. A key feature of our scheme is that it leverages existing transformations, especially data transformations, to improve performance on GPU architectures and uses automatic machine learning to build a predictive model to determine if it is worthwhile running the OpenCL code on the GPU or OpenMP code on the multicore host. We applied our approach to the entire NAS parallel benchmark suite and evaluated it on distinct GPU-based systems. We achieved average (up to) speedups of 4.51× and 4.20× (143× and 67×) on Core i7/NVIDIA GeForce GTX580 and Core i7/AMD Radeon 7970 platforms, respectively, over a sequential baseline. Our approach achieves, on average, greater than 10× speedups over two state-of-the-art automatic GPU code generators. Zheng Wang 0001, Dominik Grewe, Michael F. P. O'Boyle |
ACM Trans. Archit. Code Optim. | 1 |
| 2014 | Integrating profile-driven parallelism detection and machine-learning-based mappingabstractCompiler-based auto-parallelization is a much-studied area but has yet to find widespread application. This is largely due to the poor identification and exploitation of application parallelism, resulting in disappointing performance far below that which a skilled expert programmer could achieve. We have identified two weaknesses in traditional parallelizing compilers and propose a novel, integrated approach resulting in significant performance improvements of the generated parallel code. Using profile-driven parallelism detection, we overcome the limitations of static analysis, enabling the identification of more application parallelism, and only rely on the user for final approval. We then replace the traditional target-specific and inflexible mapping heuristics with a machine-learning-based prediction mechanism, resulting in better mapping decisions while automating adaptation to different target architectures. We have evaluated our parallelization strategy on the NAS and SPEC CPU2000 benchmarks and two different multicore platforms (dual quad-core Intel Xeon SMP and dual-socket QS20 Cell blade). We demonstrate that our approach not only yields significant improvements when compared with state-of-the-art parallelizing compilers but also comes close to and sometimes exceeds the performance of manually parallelized codes. On average, our methodology achieves 96% of the performance of the hand-tuned OpenMP NAS and SPEC parallel benchmarks on the Intel Xeon platform and gains a significant speedup for the IBM Cell platform, demonstrating the potential of profile-guided and machine-learning- based parallelization for complex multicore platforms. Zheng Wang 0001, Georgios Tournavitis, Björn Franke, Michael F. P. O'Boyle |
ACM Trans. Archit. Code Optim. | 1 |
| 2013 | Smart, adaptive mapping of parallelism in the presence of external workloadabstractGiven the wide scale adoption of multi-cores in main stream computing, parallel programs rarely execute in isolation and have to share the platform with other applications that compete for resources. If the external workload is not considered when mapping a program, it leads to a significant drop in performance. This paper describes an automatic approach that combines compile-time knowledge of the program with dynamic runtime workload information to determine the best adaptive mapping of programs to available resources. This approach delivers increased performance for the target application without penalizing the existing workload. This approach is evaluated on NAS and SpecOMP parallel bench-mark programs across a wide range of workload scenarios. On average, our approach achieves performance gain of 1.5x over a state-of-art scheme on a 12 core machine. Murali Emani, Zheng Wang 0001, Michael F. P. O'Boyle |
CGO | 2 |
| 2013 | Portable mapping of data parallel programs to OpenCL for heterogeneous systemsabstractGeneral purpose GPU based systems are highly attractive as they give potentially massive performance at little cost. Re-alizing such potential is challenging due to the complexity of programming. This paper presents a compiler based approach to automatically generate optimized OpenCL code from data-parallel OpenMP programs for GPUs. Such an approach brings together the benefits of a clear high levellanguage (OpenMP) and an emerging standard (OpenCL) for heterogeneous multi-cores. A key feature of our scheme is that it leverages existing transformations, especially data transformations, to improve performance on GPU architectures and uses predictive modeling to automatically determine if it is worthwhile running the OpenCL code on the GPU or OpenMP code on the multi-core host. We applied our approach to the entire NAS parallel benchmark suite and evaluated it on two distinct GPU based systems: Core i7/NVIDIA GeForce GTX 580 and Core 17/AMD Radeon 7970. We achieved average (up to) speedups of 4.51x and 4.20x (143x and 67x) respectively over a sequential baseline. This is, on average, a factor 1.63 and 1.56 times faster than a hand-coded, GPU-specific OpenCL implementation developed by independent expert programmers. Dominik Grewe, Zheng Wang 0001, Michael F. P. O'Boyle |
CGO | 2 |
| 2013 | Using machine learning to partition streaming programsabstractStream-based parallel languages are a popular way to express parallelism in modern applications. The efficient mapping of streaming parallelism to today's multicore systems is, however, highly dependent on the program and underlying architecture. We address this by developing a portable and automatic compiler-based approach to partitioning streaming programs using machine learning. Our technique predicts the ideal partition structure for a given streaming application using prior knowledge learned offline. Using the predictor we rapidly search the program space (without executing any code) to generate and select a good partition. We applied this technique to standard StreamIt applications and compared against existing approaches. On a 4-core platform, our approach achieves 60% of the best performance found by iteratively compiling and executing over 3000 different partitions per program. We obtain, on average, a 1.90× speedup over the already tuned partitioning scheme of the StreamIt compiler. When compared against a state-of-the-art analytical, model-based approach, we achieve, on average, a 1.77× performance improvement. By porting our approach to an 8-core platform, we are able to obtain 1.8× improvement over the StreamIt default scheme, demonstrating the portability of our approach. Zheng Wang 0001, Michael F. P. O'Boyle |
ACM Trans. Archit. Code Optim. | 1 |
| 2011 | A workload-aware mapping approach for data-parallel programsabstractMuch compiler-orientated work in the area of mapping parallel programs to parallel architectures has ignored the issue of external workload. Given that the majority of platforms will not be dedicated to just one task at a time, the impact of other jobs needs to be addressed. As mapping is highly dependent on the underlying machine, a technique that is easily portable across platforms is also desirable. Dominik Grewe, Zheng Wang 0001, Michael F. P. O'Boyle |
HiPEAC | 2 |
| 2010 | Partitioning streaming parallelism for multi-cores: a machine learning based approachabstractStream based languages are a popular approach to expressing parallelism in modern applications. The efficient mapping of streaming parallelism to multi-core processors is, however, highly dependent on the program and underlying architecture. We address this by developing a portable and automatic compiler-based approach to partitioning streaming programs using machine learning. Our technique predicts the ideal partition structure for a given streaming application using prior knowledge learned off-line. Using the predictor we rapidly search the program space (without executing any code) to generate and select a good partition. We applied this technique to standard StreamIt applications and compared against existing approaches. On a 4-core platform, our approach achieves 60% of the best performance found by iteratively compiling and executing over 3000 different partitions per program. We obtain, on average, a 1.90x speedup over the already tuned partitioning scheme of the StreamIt compiler. When compared against a state-of-the-art analytical, model-based approach, we achieve, on average, a 1.77x performance improvement. By porting our approach to a 8-core platform, we are able to obtain 1.8x improvement over the StreamIt default scheme, demonstrating the portability of our approach. Zheng Wang 0001, Michael F. P. O'Boyle |
PACT | 1 |
| 2009 | Towards a holistic approach to auto-parallelization: integrating profile-driven parallelism detection and machine-learning based mappingabstractCompiler-based auto-parallelization is a much studied area, yet has still not found wide-spread application. This is largely due to the poor exploitation of application parallelism, subsequently resulting in performance levels far below those which a skilled expert programmer could achieve. We have identified two weaknesses in traditional parallelizing compilers and propose a novel, integrated approach, resulting in significant performance improvements of the generated parallel code. Using profile-driven parallelism detection we overcome the limitations of static analysis, enabling us to identify more application parallelism and only rely on the user for final approval. In addition, we replace the traditional target-specific and inflexible mapping heuristics with a machine-learning based prediction mechanism, resulting in better mapping decisions while providing more scope for adaptation to different target architectures. We have evaluated our parallelization strategy against the NAS and SPEC OMP benchmarks and two different multi-core platforms (dual quad-core Intel Xeon SMP and dual-socket QS20 Cell blade). We demonstrate that our approach not only yields significant improvements when compared with state-of-the-art parallelizing compilers, but comes close to and sometimes exceeds the performance of manually parallelized codes. On average, our methodology achieves 96% of the performance of the hand-tuned OpenMP NAS and SPEC parallel benchmarks on the Intel Xeon platform and gains a significant speedup for the IBM Cell platform, demonstrating the potential of profile-guided and machine-learning based parallelization for complex multi-core platforms. Georgios Tournavitis, Zheng Wang 0001, Björn Franke, Michael F. P. O'Boyle |
PLDI | 2 |
| 2009 | Mapping parallelism to multi-cores: a machine learning based approachabstractThe efficient mapping of program parallelism to multi-core processors is highly dependent on the underlying architecture. This paper proposes a portable and automatic compiler-based approach to mapping such parallelism using machine learning. It develops two predictors: a data sensitive and a data insensitive predictor to select the best mapping for parallel programs. They predict the number of threads and the scheduling policy for any given program using a model learnt off-line. By using low-cost profiling runs, they predict the mapping for a new unseen program across multiple input data sets. We evaluate our approach by selecting parallelism mapping configurations for OpenMP programs on two representative but different multi-core platforms (the Intel Xeon and the Cell processors). Performance of our technique is stable across programs and architectures. On average, it delivers above 96% performance of the maximum available on both platforms. It achieve, on average, a 37% (up to 17.5 times) performance improvement over the OpenMP runtime default scheme on the Cell platform. Compared to two recent prediction models, our predictors achieve better performance with a significant lower profiling cost. Zheng Wang 0001, Michael F. P. O'Boyle |
PPoPP | 1 |