Yuting Chen 0001

dblp:21/6510-1 · DBLP profile ↗
← Back
90ranked-venue papers
9as first author
28since 2021 · last 2026
0000-0002-4128-8966ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 69 · 7 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 1 first-author · 4 since 2021Systems, architecture and hardware · 10 · 7 since 2021Artificial intelligence and machine learning · 8 · 1 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Security and privacy · 1 · 1 since 2021Theory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2026 JOSer: Just-In-Time Object Serialization for Heavy Java Serialization Workloads
abstract
Object serialization is critical in Java, which preserves objects in memory and transfers them among software systems if needed. However, the serialization techniques of modern Java systems are usually inflexible and inefficient under heavy serialization workloads, as they rely on manually-defined object schemas or omni-functional serializers. To tackle the above problem, we reveal a novel, serialization-specific optimization opportunity in Java. Based on it, we develop JOSer (Just-in-time Object SERializer), an efficient, Just-in-Time (JIT) object serialization technique. At runtime, JOSer generates a set of class-specific, JIT-friendly object serializers (i.e., serialization code), and then continuously optimizes them with the JIT compiler of Java Virtual Machine (JVM). JOSer also shares the metadata of objects under serialization and the serializers under optimization. We evaluate JOSer against six Java serialization techniques including OpenJDK's built-in serialization technique. Overall, JOSer improves the throughput by up to 20~83× in serialization and 43~229× in deserialization. JOSer has been successfully deployed in real-world products, reducing serialization CPU usage of Flink by 35.32~41.14% and latency of search recommendations by 30+ ms. JOSer grounds Apache Fory#8482;, an open-source serialization framework available at https://fory.apache.org/.
Chaokun Yang, Pengbo Nie, Qianwei Yu, Chengcheng Wan 0001, He Jiang 0001, Yuting Chen 0001
ASPLOS (2)8
2026 Orchestrating optimization passes of machine learning compiler for reducing memory footprints of computation graphs
Qianwei Yu, Pengbo Nie, Chengcheng Wan 0001, He Jiang 0001, Jianjun Zhao 0001, Lei Qiao 0002, Yuting Chen 0001
J. Syst. Archit.10
2026 Synthetic Malware at Scale: Malicious Code Generation With Code Transplanting
abstract
Malicious code detection is one of the most essential tasks in safeguarding against security breaches, data compromise, and related threats. While machine learning has emerged as a predominant method for pattern detection, the training process is intricate due to the severe scarcity of malicious code samples. Consequently, machine learning detectors often encounter malicious patterns in limited and isolated scenarios, hindering their ability to generalize effectively across diverse threat landscapes. In this paper, we introduce MalCoder, a novel method for synthesizing malicious code samples. MalCoder enlarges the quantity and diversity of malicious instances by transplanting a set of malicious prototypes into a vast pool of benign code, thereby crafting a diverse array of malicious instances tailored to various application scenarios. For each malware prototype, MalCoder treats it as an incomplete code fragment and crafts its preceding and subsequent contexts through right-to-left and left-to-right code completion respectively. By leveraging GPTs with various sampling strategies, we can instantiate a large number of code samples bearing the malware prototype. Subsequently, MalCoder masks the original prototypes within the transplanted samples and fine-tunes an LLM code generator to reconstruct the original prototype. This process enables the model to seamlessly transplant malicious code fragments into benign code. During inference, MalCoder can automatically insert malicious fragments into benign samples at random positions, transforming benign code into malicious code. We apply MalCoder to a large pool of benign code in CodeSearchNet and craft over 50,000 malicious samples stemming from 39 malicious prototypes. Both qualitative and quantitative analyses show that the generated samples maintain key characteristics of malicious code while blending seamlessly with benign code, which helps in creating realistic and varied training data. Additionally, by using the generated samples as augmented training data, we witness a remarkable surge in malicious code detection capabilities. Specifically, the F1-score experiences a significant increase compared to utilizing only the original prototype samples.
Guangzhan Wang, Diwei Chen, Xiaodong Gu 0002, Yuting Chen 0001, Beijun Shen
IEEE Trans. Software Eng.4
2025 Conanj: Confidential Data Analysis for Confidential Computing of Java Programs
Xinyuan Miao, Yuting Chen 0001
APSEC2
2025 TRACED: A Temporal Graph Neural Networks-based Model for Data Prefetching
abstract
In modern microarchitectures, machine-learning-based prefetchers use past memory requests to learn access patterns and predict memory addresses, thereby prefetching data into the cache to mitigate the processor-memory speed gap. However, they face two key challenges in capturing irregular access patterns generated by complex data structures and algorithms. One is data dispersion: the disorderliness of memory addresses makes it difficult for prefetchers to extract meaningful data features. The other is temporal and spatial complexity: existing prefetchers fail to effectively learn temporal and spatial characteristics, and thus are unable to explore more complex access patterns. To resolve these challenges, we propose TRACED, a novel temporal graph neural network-based prefetcher aimed at learning access patterns of memory addresses. TRACED consists of two key components: a dynamic clustering component and a temporal graph neural network component. In the dynamic clustering component, we introduce a similarity function to quantify the similarity of memory addresses. Based on the quantified similarity, we dynamically group unordered memory addresses into different clusters. This ensures that the memory addresses in each cluster are ordered and change smoothly, thus resolving the first challenge. The temporal graph neural network component constructs a spatiotemporal graph to represent relationships among memory addresses. This helps capture temporal and spatial characteristics both across and within clusters, thus resolving the second challenge. This article demonstrates the effectiveness of the proposed prefetcher through experiments. Specifically, in terms of accuracy, TRACED outperforms BO, SPP, DOMINO, Delta-LSTM, and VOYAGER by 2.29%–40.83% on average. Furthermore, TRACED attains remarkable coverage of 55.67% and IPC of 43.75%, outperforming all competing approaches in both metrics.
He Jiang 0001, Liuwei Fu, Dong Liu 0025, Zhilei Ren, Yuting Chen 0001, Lei Qiao 0002
ACM Trans. Archit. Code Optim.5
2025 CodeMark: Contextual and Natural Watermarking for Tracing Code Snippet Provenance
abstract
Determining the origins of code snippets has gained increasing attention due to the popularity of large language models and the concern about their misuse in generating unlicensed or malicious code. Watermarking is considered a working solution for tracing code snippet provenance. However, source code watermarking requires more stringent and intricate rules than natural language or software watermarking, since one needs to ensure both readability and functionality of the watermarked code snippets. To this end, we propose a novel watermarking systemCodeMark, featured by variable renaming as the key. Surrounding variable renaming, several challenges emerge such as determining renaming candidates, defining the variable context, and providing diverse variable substitutes, etc. The design ofCodeMarkconquers these challenges by an end-to-end learning system, which interprets the code context through Graph Neural Networks (GNNs) and generates natural substitutes fitting the context by distilling from CodeBERT. Experiments illustrate thatCodeMarksurpasses the state-of-the-art watermarking systems in terms of watermarking requirements.
Wei Li 0254, Borui Yang, Yujie Sun 0001, Suyu Chen, Yuting Chen 0001, Liyao Xiang
IEEE Trans. Dependable Secur. Comput.5
2024 A Holistic Functionalization Approach to Optimizing Imperative Tensor Programs in Deep Learning
abstract
As deep learning empowers various fields, many domain-specific non-neural network operators have been proposed to improve the accuracy of deep learning models. Researchers often use the imperative programming diagram (PyTorch) to express these new operators, leaving the fusion optimization of these operators to deep learning compilers. Unfortunately, the inherent side effects introduced by imperative tensor programs, especially tensor-level mutations, often make optimization extremely difficult. Previous works either fail to eliminate the side effects of tensor-level mutations or require programmers to manually analyze and transform them. In this paper, we present a holistic functionalization approach (TensorSSA) to optimizing imperative tensor programs beyond control flow boundaries. We first introduce TensorSSA intermediate representation for removing tensor-level mutation and expanding the scope and ability of operator fusion. Based on TensorSSA IR, we propose a TensorSSA conversion algorithm that performs functionalization crossing the boundary of control flow. TensorSSA achieves a 1.79X (1.34X on average) speedup in representative deep learning tasks than state-of-the-art works.
Xingcheng Zhang, Shengen Yan, Yuting Chen 0001, Yueqian Zhang, Minxi Jin, Lijuan Jiang, Yun Liang 0001, Chao Yang 0002, Dahua Lin
DAC6
2024 OPASS: Orchestrating TVM's Passes for Lowering Memory Footprints of Computation Graphs
abstract
Deep learning (DL) compilers, such as TVM and TensorFlow, encompass a variety of passes for optimizing computation graphs (i.e., DL models). Despite the efforts on developing optimization passes, it remains a challenge in arranging these passes - most compilers employ fixed pass sequences that do not fit with computation graphs of diverse structures; on the other hand, optimization passes have cascade effects, making the structures of graphs under compilation volatile and as well making it difficult to generate optimal sequences for graphs. Inspired by recent progresses on static computing memory footprints (i.e., memory usages) of computation graphs, we introduce in this paper OPASS, a novel approach to orchestrating TVM's optimization passes for lowering memory footprints of computation graphs, and finally allowing the graphs to run on memory-constrained devices. The key idea is, given a computation graph$G$, to optimize the graph heuristically and iteratively: OPASS learns the effects of passes on the graph; it then optimizes$G$iteratively - each iteration picks up a pass by the reduction of the memory footprint of$G$and as well the implicit effects of the pass for further optimizations, letting the pass be applied. We evaluate OPASS on Rebench (a suite of computation graphs) and two real-world models (Transformer and ResNet). The results clearly show the strength of OPASS: it outperforms TVM's default sequence by$1.77\times$in reducing graphs' memory footprints, with affordable costs; it also offers extra memory reductions of$5\sim 12\%$by catching the implicit effects of passes. Furthermore, OPASS helps analyze positive/negative effects of passes to graphs' memory footprints, providing TVM developers with best practices for designing optimization pass sequences.
Pengbo Nie, Chengcheng Wan 0001, He Jiang 0001, Jianjun Zhao 0001, Yuting Chen 0001
ICSME7
2024 Few-shot code translation via task-adapted prompt learning
Xiaodong Gu 0002, Yuting Chen 0001, Beijun Shen
J. Syst. Softw.4
2024 Constructing exception handling chains for testing Java virtual machine implementations
abstract
Abstract The Java virtual machine (JVM) is the cornerstone of the Java platforms. A JVM's exception handling implementation interrupts, when the objective application encounters an exception (or an error), the normal execution of the application and performs specific handling tasks. However, little research has been done in systematically validating JVMs' exception handling implementations—test programs or even applications need to be carefully designed for throwing/catching exceptions at runtime; a JVM's exception handling implementation is also complicated, making it challenging to design tests for testing all of its functionalities. Inspired by the recent success of fuzz testing of compilers and JVM implementations, we introduce EHCBuilder, the first technique for fuzzing JVMs' exception handling implementations. The key idea is to construct exception handling chains, each of which abstracts a program's execution into a sequence of exception throwings, catchings, and/or handlings. A classfile seed can then be mutated into test programs with diverse exception handling chains, enabling (1) exceptions to be continuously thrown and caught at runtime, and (2) JVMs' exception handling implementations to be much more thoroughly tested. We have implemented EHCBuilder and evaluated EHCBuilder on popular JVM implementations including OpenJDK's HotSpot, Eclipse's OpenJ9, Azul's Zulu, and Oracle's GraalVM. Our results show that EHCBuilder can generate programs with very intricate exception handling chains and reveal differences among JVMs' exception handling implementations: Up to thousands of lines of source code in HotSpot's exception handling implementation are covered more than the original benchmarks; during 39 K iterations, EHCBuilder generates exception handling chains of different lengths, revealing 258 runtime differences. We classify the differences into four categories, and reveal a fast throw issue confirmed by HotSpot developers and another initCause issue confirmed by the OpenJ9 community.
Bochuan Chen, Yuting Chen 0001, Lei Bu
J. Softw. Evol. Process.3
2024 What's Wrong With Low-Code Development Platforms? An Empirical Study of Low-Code Development Platform Bugs
abstract
Low-code development platforms (LCDPs) are increasingly being introduced and leveraged by major IT enterprises to lower the threshold and promote the efficiency of software development. Like other software systems, LCDPs are also inevitable to have bugs. The bugs in LCDPs may cause unpredictable consequences as they pose risks to all the downstream software products. However, to the best of our knowledge, there exist no studies that ever consider the bugs caused by LCDPs. To handle the LCDP bugs better, in this article, we conduct an empirical study of the characteristics of LCDP bugs by examining 974 confirmed bugs of four dominant LCDPs (i.e., OutSystems, Mendix, Appsmith, and Budibase) from both commercial and open-source domains. These bugs are analyzed from three perspectives, including bug root causes, bug symptoms, and the affected stages of LCDPs. Based on the analysis, we obtain a series of valuable findings. For example, around 60% of the bugs reside in the stage of designing and specifying the developed applications. Over 37% of the bugs lead LCDPs to behave unexpectedly but without showing explicit signs. Moreover, the bugs relevant to the incorrect graphics of user interfaces are significant due to the characteristics of LCDPs. These findings point out the guidelines, challenges, and future directions to address LCDP bugs.
Dong Liu 0025, He Jiang 0001, Shikai Guo, Yuting Chen 0001, Lei Qiao 0002
IEEE Trans. Reliab.4
2023 Lejacon: A Lightweight and Efficient Approach to Java Confidential Computing on SGX
abstract
Intel's SGX is a confidential computing technique. It allows key functionalities of C/C++/native applications to be confidentially executed in hardware enclaves. However, numerous cloud applications are written in Java. For supporting their confidential computing, state-of-the-art approaches deploy Java Virtual Machines (JVMs) in enclaves and perform confidential computing on JVMs. Meanwhile, these JVM-in-enclave solutions still suffer from serious limitations, such as heavy overheads of running JVMs in enclaves, large attack surfaces, and deep computation stacks. To mitigate the above limitations, we for-malize a Secure Closed-World (SCW) principle and then propose Lejacon, a lightweight and efficient approach to Java confidential computing. The key idea is, given a Java application, to (1) separately compile its confidential computing tasks into a bundle of Native Confidential Computing (NCC) services; (2) run the NCC services in enclaves on the Trusted Execution Environment (TEE) side, and meanwhile run the non-confidential code on a JVM on the Rich Execution Environment (REE) side. The two sides interact with each other, protecting confidential computing tasks and as well keeping the Trusted Computing Base (TCB) size small. We implement Lejacon and evaluate it against OcclumJ (a state-of-the-art JVM-in-enclave solution) on a set of benchmarks using the BouncyCastle cryptography library. The evaluation results clearly show the strengths of Lejacon: it achieves compet-itive performance in running Java confidential code in enclaves; compared with OcclumJ, Lejacon achieves speedups by up to 16.2x in running confidential code and also reduces the TCB sizes by 90+% on average.
Xinyuan Miao, Sanhong Li, Pengbo Nie, Yuting Chen 0001, Beijun Shen, He Jiang 0001
ICSE8
2023 GenCoG: A DSL-Based Approach to Generating Computation Graphs for TVM Testing
abstract
TVM is a popular deep learning (DL) compiler. It is designed for compiling DL models, which are naturally computation graphs, and as well promoting the efficiency of DL computation. State-of-the-art methods, such as Muffin and NNSmith, allow developers to generate computation graphs for testing DL compilers. However, these techniques are inefficient — their generated computation graphs are either type-invalid or inexpressive, and hence not able to test the core functionalities of a DL compiler.
Pengbo Nie, Xinyuan Miao, Yuting Chen 0001, Chengcheng Wan 0001, Lei Bu, Jianjun Zhao 0001
ISSTA4
2023 InfeRE: Step-by-Step Regex Generation via Chain of Inference
abstract
Automatically generating regular expressions (abbrev. regexes) from natural language description (NL2RE) has been an emerging research area. Prior studies treat regex as a linear sequence of tokens and generate the final expressions autoregressively in a single pass. They did not take into account the step-by-step internal text-matching processes behind the final results. This significantly hinders the efficacy and interpretability of regex generation by neural language models. In this paper, we propose a new paradigm called InfeRE, which decomposes the generation of regexes into chains of step-bystep inference. To enhance the robustness, we introduce a self-consistency decoding mechanism that ensembles multiple outputs sampled from different models. We evaluate InfeRE on two publicly available datasets, NL-RX-Turk and KB13, and compare the results with state-of-the-art approaches and the popular tree-based generation approach TRANX. Experimental results show that InfeRE substantially outperforms previous baselines, yielding 16.3% and 14.7% improvement in DFA@5 accuracy on two datasets, respectively.
Xiaodong Gu 0002, Yuting Chen 0001, Beijun Shen
ASE3
2023 DuReSE: Rewriting Incomplete Utterances via Neural Sequence Editing
Xiaodong Gu 0002, Yuting Chen 0001, Beijun Shen
Neural Process. Lett.3
2023 Coverage-directed Differential Testing of X.509 Certificate Validation in SSL/TLS Implementations
abstract
Secure Sockets Layer (SSL) and Transport Security (TLS) are two secure protocols for creating secure connections over the Internet. X.509 certificate validation is important for security and needs to be performed before an SSL/TLS connection is established. Some advanced testing techniques, such as frankencert , have revealed, through randomly mutating Internet accessible certificates, that there exist unexpected, sometimes critical, validation differences among different SSL/TLS implementations. Despite these efforts, X.509 certificate validation still needs to be thoroughly tested as this work shows. This article tackles this challenge by proposing transcert , a coverage-directed technique to much more effectively test real-world certificate validation code. Our core insight is to (1) leverage easily accessible Internet certificates as seed certificates and (2) use code coverage to direct certificate mutation toward generating a set of diverse certificates. The generated certificates are then used to reveal discrepancies, thus potential flaws, among different certificate validation implementations. We implement transcert and evaluate it against frankencert , NEZHA , and RFCcert (three advanced fuzzing techniques) on five widely used SSL/TLS implementations. The evaluation results clearly show the strengths of transcert : During 10,000 iterations, transcert reveals 71 unique validation differences, 12×, 1.4×, and 7× as many as those revealed by frankencert , NEZHA , and RFCcert , respectively; it also supplements RFCcert in conformance testing of the SSL/TLS implementations against 120 validation rules, 85 of which are exclusively covered by transcert -generated certificates. We identify 17 root causes of validation differences, all of which have been confirmed and 11 have never been reported previously. The transcert -generated X.509 certificates also reveal that the primary goal of certificate chain validation is stated ambiguously in the widely adopted public key infrastructure standard RFC 5280.
Pengbo Nie, Chengcheng Wan 0001, Jiayu Zhu, Yuting Chen 0001, Zhendong Su 0001
ACM Trans. Softw. Eng. Methodol.5
2022 Hierarchical memory-constrained operator scheduling of neural architecture search networks
abstract
Neural Architecture Search (NAS) is widely used in industry, searching for neural networks meeting task requirements. Meanwhile, it faces a challenge in scheduling networks satisfying memory constraints. This paper proposes HMCOS that performs hierarchical memory-constrained operator scheduling of NAS networks: given a network, HMCOS constructs a hierarchical computation graph and employs an iterative scheduling algorithm to progressively reduce peak memory footprints. We evaluate HMCOS against RPO and Serenity (two popular scheduling techniques). The results show that HMCOS outperforms existing techniques in supporting more NAS networks, reducing 8.7~42.4% of peak memory footprints, and achieving 137--283x of speedups in scheduling.
Chengcheng Wan 0001, Yuting Chen 0001, He Jiang 0001, Lei Qiao 0002
DAC3
2022 Zebra: An Efficient, RDMA-Enabled Distributed Persistent Memory File System
Shengan Zheng, Yuting Chen 0001, Linpeng Huang
DASFAA (1)4
2022 Automatically repairing tensor shape faults in deep learning programs
Dangwei Wu, Beijun Shen, Yuting Chen 0001, He Jiang 0001, Lei Qiao 0002
Inf. Softw. Technol.3
2022 LocSeq: Automated Localization for Compiler Optimization Sequence Bugs of LLVM
abstract
Compiler bugs may be triggered when programs are optimized with optimization sequences. However, diagnosing compiler optimization sequence bugs is difficult due to limited debugging information. Although some techniques (e.g., DiWi and RecBi) have been proposed to automatically localize compiler bugs, no systematic work has been conducted to automatically localize compiler optimization sequence bugs. In this article, we propose LocSeq, a novel technique to automatically localize compiler optimization sequence bugs of LLVM. The core insight of LocSeq is based on the fact that the behaviors of optimizations may be influenced by each other, and thus, the innocent files may be excluded by constructing bug-free optimization sequences. First, given a buggy optimization sequence that triggers a compiler bug, in LocSeq, we transform the problem of the localization for a compiler optimization sequence bug to the problem of the construction for bug-free optimization sequences, which are helpful to localize buggy compiler files. Then, a constrained genetic algorithm is presented in LocSeq to generate a set of bug-free optimization sequences that share similar compiler execution traces with the buggy optimization sequence. Finally, LocSeq leverages a spectrum-based bug localization technique to localize the compiler optimization sequence bug by comparing the execution traces between bug-free optimization sequences and the buggy optimization sequence. To evaluate the effectiveness of LocSeq, we build a benchmark, including 60 optimization sequence bugs of LLVM, and compare LocSeq with the state-of-the-art techniques DiWi and RecBi. The experimental results show that LocSeq significantly outperforms DiWi and RecBi by up to 366.66%/72.27% and 250.00%/56.00% for localizing optimization sequence bugs within Top-1/5 files, respectively.
Zhide Zhou, He Jiang 0001, Zhilei Ren, Yuting Chen 0001, Lei Qiao 0002
IEEE Trans. Reliab.4
2021 Learning to Match Workers and Tasks via a Multi-View Graph Attention Network
abstract
The worker-task matching problem brings up unique characteristics that are not present in traditional matching scenarios, i.e., the huge flow of tasks with short lifespans, the importance of workers’ capabilities, and the quality of the completed tasks. These characteristics further pose significant challenges of data sparsity and comprehensive modeling.To address the two challenges, this paper proposes MvkGAN, a multi-view attention network on a bi-collaborative knowledge graph (BicKG). The core ideas of our work are 1) building BicKG from the data of workers, tasks, their interactions, and domain knowledge, and then leveraging it to reveal the latent interactions between workers and tasks to mitigate the data sparsity challenge; and 2) designing a multi-view knowledge graph attention network (MvkGAN) which learns to match workers and tasks, to meet the comprehensive modeling challenge. In this network, different features are organized as multiple views and these views are further connected by the attention mechanism.We have implemented MvkGAN and evaluated it against five state-of-the-art approaches (Wide&Deep, DeepFM, KGAT, Crow-dRex and PJFNN) on two real-world datasets. The evaluation results show that MvkGAN improves the accuracy by 6.90% and F1-score by 5.52% on average, and also has the ability of generating reasonable explanations.
Nan Cui, Chunqi Chen, Beijun Shen, Yuting Chen 0001
COMPSAC4
2021 ApproxiFuzzer: Fuzzing towards Deep Code Snippets in Java Programs
abstract
A real-world, complex software system can contain a number of code snippets. Many snippets are deep, surrounded by complicated triggering conditions and/or hidden in functions less frequently invoked. Fuzzing and symbolic execution are two mainstreams for exploring input spaces and increasing code coverage of complicated software systems. Meanwhile, it remains a challenge to determine whether a deep code snippet is reachable, and if it is reachable, which test(s) can reach it.This paper presents ApproxiFuzzer, an effective, demand-driven approach to fuzzing towards deep code snippets in Java programs. Given a program P, a target deep code snippet tcs, and a set of seeding test inputs, the key idea behind ApproxiFuzzer is to selectively mutate the test inputs and collect their execution traces such that the execution traces gradually approximate tcs; several measures are designed for measuring the distances between execution traces and the code snippet and directing the fuzzing process towards generating test inputs reaching tcs.We have implemented ApproxiFuzzer and evaluated it against Kelinci (an AFL-based fuzzer) and JDart (a concolic execution tool) on a set of real-world benchmarks. The evaluation clearly demonstrates the strengths of ApproxiFuzzer—ApproxiFuzzer outperforms Kelinci by 36× in efficiently generating test inputs, obtaining up to 18.2% higher code coverage; ApproxiFuzzer also outperforms JDart by 46.2∼96.2% in hitting deep code snippets.
Xintian Yu, Enze Ma, Pengbo Nie, Beijun Shen, Yuting Chen 0001
COMPSAC5
2021 JPDHeap: A JVM Heap Design for PM-DRAM Memories
abstract
Real-world e-commerce systems need large cache capacities. Persistent memory (PM) can be employed to enlarge JVMs’ cache capacities, meanwhile they incur heavy write slowdowns and garbage collection overheads. This paper proposes JPDheap, a JVM heap design for PM-DRAM memories. A JPDheap is composed of a standard Java heap on DRAM and another heap on PM. The core insight is to separate heap objects and store them on DRAM or PM, allowing objects to be accessed much more efficiently. Our evaluation shows that JPDheap outperforms state-of-the-art heap designs by up to 115.96% in increasing applications’ throughput and by up to 87.03% in decreasing the average latency.
Litong You, Tianxiao Gu, Shengan Zheng, Jianmei Guo, Sanhong Li, Yuting Chen 0001, Linpeng Huang
DAC6
2021 Reno: An RDMA-Enabled, Non-Volatile Memory-Optimized Key-Value Store
abstract
Remote direct memory access (RDMA) has been employed to boost remote data access for key-value stores, since it provides kernel-bypass, zero-copy and low-latency features. Meanwhile, existing RDMA-enabled key-value stores still have performance bottlenecks, since extra costs need to be spent on processing requests and keeping data consistency. This paper introduces Reno, an RDMA-enabled, NVM-optimized key-value store that supports fast remote access of persistent data. On the server-side, Reno is built atop a bucket-based hopscotch hash table, where the actual key-value items are stored in NVM (non-volatile memories), and metadata are managed by an in-DRAM index. On the client-side, Reno adopts a fully server-bypass paradigm for both remote read and write requests to achieve low latency and high throughput. We evaluate Reno on an Intel's Optane DC Persistent Memory platform with Infiniband network support. The results show the strengths of Reno. In particular, Reno outperforms its counterparts by 1.1~3.3x for remote reads and 1.9~4.8 x for remote writes in terms of latency; the speedups of concurrent throughput are up to 2.33 x, 2.36 x, 3.09 x and 5.08 x for read-only, read-heavy, write-heavy and write-only YCSB workloads, respectively.
Rulin Huang, Kaixin Huang, Yuting Chen 0001
ICPADS4
2021 Tensfa: Detecting and Repairing Tensor Shape Faults in Deep Learning Systems
abstract
Software developers frequently invoke deep learning (DL) APIs to incorporate learning solutions into software systems. However, misuses of these APIs can cause various DL faults, such as tensor shape faults. Tensor shape faults occur when restriction conditions of operations are not met; they are prevalent in practice, leading to many system crashes. Meanwhile, researchers and engineers still face a strong challenge in detecting tensor shape faults ─ static techniques incur heavy overheads in defining detection rules, and the only dynamic technique requires human engineers to rewrite APIs for tracking shape changes. To address the above challenge, we conduct a deep empirical study on crashing tensor shape faults (i.e., those causing programs to crash), categorizing them into four types and revealing twelve repair patterns. We then propose and implement Tensfa, an approach to detecting and repairing crashing tensor shape faults. Tensfa takes a machine learning method to learn from crash messages and employs decision trees in detecting tensor shape faults. Tensfa also provides the first automated solution to repairing the detected faults: it tracks shape properties by a customized Python debugger, analyzes their data dependences, and uses the twelve patterns to generate patches. We construct SFData, a set of 146 buggy programs with crashing tensor shape faults. Our Tensfa has been implemented and evaluated on SFData and IslamData (another dataset of tensor shape faults). The results clearly show the effectiveness of Tensfa. In particular, Tensfa achieves the state-of-the-art results: it reaches an F1-score of 96.88% in detecting the faults and repairs 80 out of 146 buggy programs in SFData.
Dangwei Wu, Beijun Shen, Yuting Chen 0001, He Jiang 0001, Lei Qiao 0002
ISSRE3
2021 ConLAR: Learning to Allocate Resources to Docker Containers under Time-Varying Workloads
abstract
Cloud platforms are increasingly using containers for lightweight virtualization. However, the mainstream operating systems are currently limited in their capabilities in customizing containers’ resource management. There remains two main challenges in resource allocations. First, the application workloads can be time-varying, leading to a problem of resource over- or under-allocations. Second, it becomes difficult to minimize the resource provisioning cost while guaranteeing the SLO (service level objective). To address these challenges, we propose ConLAR, a learning-to-allocate approach that predicts and allocates resources to Docker containers under time-varying workloads. ConLAR efficiently reduces over-provisioning cost with the SLO guarantees by taking an Observing-Predicting-Allocating-Executing paradigm: given a container, it observes the running of the online application and its environment, leverages the LSTM (long short term memory) model to predict its future workload, adaptively learns to construct resource allocation strategies with two objectives through RL (reinforcement learning), and executes them to scale container resources dynamically. We have evaluated ConLAR on two real-world workloads of ClarkNet and GoogleClusterData. The results clearly show the effectiveness of ConLAR. In particular, ConLAR achieves a resource over-provisioning cost of less than 16.5% and an SLO violations rate of 8.9%; it also shows good flexibility to learn different adaption policies.
Diwei Chen, Beijun Shen, Yuting Chen 0001
QRS3
2021 Context-Aware Conversational Recommendation of Trigger-Action Rules in IoT Programming
abstract
Trigger-action (TA) programming is a programming paradigm that allows end-users to automate and connect IoT devices and online services using if-trigger-then-action rules. Early studies have demonstrated this paradigms usability, but more recent work has also highlighted complexities that arise in realistic scenarios. To facilitate end-users in TA programming, we propose AutoTAR, a context-aware conversational recommendation technique for recommending TA rules. AutoTAR leverages a TA knowledge graph to encode semantic features and abstract functionalities of rules, and then takes a two-phase method to recommend TA rules to end-users: during the context-aware recommendation phase, it elicits user preferences from programming context and recommends the top-N rules using a mixed content and collaborative technique; during the conversational recommendation phase, it justifies recommendations by iteratively raising questions and collecting feedback from end-users. We evaluate AutoTAR on Mturk and real data collected from the IFTTT community. The results show that our method outperforms state-of-the-arts significantly — its context-aware recommendation outperforms RecRules by 26% on R@5 and 21% on NDCG@5; its conversational recommendation outperforms LARecommender (a conversational recommender with the LA model) by 67.64% on accuracy. In addition, AutoTAR is effective in solving three problems frequently occurring in TA rule recommendations, i.e., the cold-start problem, the repeat-consumption problem, and the incomplete-intent problem.
Mingxin Zhao, Qinyue Wu, Enze Ma, Beijun Shen, Yuting Chen 0001
Int. J. Softw. Eng. Knowl. Eng.5
2021 DeFiHap: Detecting and Fixing HiveQL Anti-Patterns
abstract
The emergence of Hive greatly facilitates the management of massive data stored in various places. Meanwhile, data scientists face challenges during HiveQL programming - they may not use correct and/or efficient HiveQL statements in their programs; developers may also introduce anti-patterns indeliberately into HiveQL programs, leading to poor performance, low maintainability, and/or program crashes. This paper presents an empirical study on HiveQL programming, in which 38 HiveQL anti-patterns are revealed. We then design and implement DeFiHap, the first tool for automatically detecting and fixing HiveQL anti-patterns. DeFiHap detects HiveQL anti-patterns via analyzing the abstract syntax trees of HiveQL statements and Hive configurations, and generates fix suggestions by rule-based rewriting and performance tuning techniques. The experimental results show that DeFiHap is effective. In particular, DeFiHap detects 25 anti-patterns and generates fix suggestions for 17 of them.
Yuetian Mao, Nan Cui, Tianjiao Du, Beijun Shen, Yuting Chen 0001
Proc. VLDB Endow.6
2020 Learning Code-Query Interaction for Enhancing Code Searches
abstract
Code search plays an important role in software development and maintenance. In recent years, deep learning (DL) has achieved a great success in this domain-several DL-based code search methods, such as DeepCS and UNIF, have been proposed for exploring deep, semantic correlations between code and queries; each method usually embeds source code and natural language queries into real vectors followed by computing their vector distances representing their semantic correlations. Meanwhile, deep learning-based code search still suffers from three main problems, i.e., the OOV (Out of Vocabulary) problem, the independent similarity matching problem, and the small training dataset problem. To tackle the above problems, we propose CQIL, a novel, deep learning-based code search method. CQIL learns code-query interactions and uses a CNN (Convolutional Neural Network) to compute semantic correlations between queries and code snippets. In particular, CQIL employs a hybrid representation to model code-query correlations, which solves the OOV problem. CQIL also deeply learns the code-query interaction for enhancing code searches, which solves the independent similarity matching and the small training dataset problems. We evaluate CQIL on two datasets (CODEnn and CosBench). The evaluation results show the strengths of CQIL-it achieves the MAP@1 values, 0.694 and 0.574, on CODEnn and CosBench, respectively. In particular, it outperforms DeepCS and UNIF, two state-of-the-art code search methods, by 13.6% and 18.1% in MRR, respectively, when the training dataset is insufficient.
Wei Li 0254, Haozhe Qin, Shuhan Yan, Beijun Shen, Yuting Chen 0001
ICSME5
2020 Guided, Deep Testing of X.509 Certificate Validation via Coverage Transfer Graphs
abstract
SSL and TLS are two secure protocols for creating secure connections over the Internet. X.509 certificate validation is important for security and needs to be performed before an SSL/TLS connection is established. However, state-of-the-art testing techniques, such as frankencert and mucert, have revealed, through randomly mutating Internet accessible certificates, that there exist unexpected, sometimes critical, validation differences among different SSL/TLS implementations. Despite these strong efforts, certificate validation is still not thoroughly tested and more effective techniques are needed as this work shows. To this end, this paper introduces transcert, a novel approach for effectively guiding fuzzing to perform deep testing of X.509 certificate validation. The goal of transcert is to generate certificates that trigger diverse executions; it achieves this goal by introducing the concept of a coverage transfer graph to efficiently, precisely abstract program executions. In particular, it records the execution of how a given certificate is validated by a reference SSL/TLS implementation. It then constructs a coverage transfer graph to model the coverage transfer from a test certificate (seed) to its mutated certificates (mutants), and explores the coverage transfer graph by iteratively sampling and mutating certificates. We have implemented transcert and evaluated it against frankencert and mucert on four state-of-the-art SSL/TLS implementations. The evaluation results clearly show the strengths of transcert- during 10,000 iterations, transcert has revealed 3,469 validation differences, 8× as many as those revealed by frankencert and mucert. We have identified 11 root causes of validation differences, all of which have been confirmed and five have never been reported previously. We also found that the primary goal of certificate chain validation is stated ambiguously in the widely-adopted PKI standard RFC 5280.
Jiayu Zhu, Chengcheng Wan 0001, Pengbo Nie, Yuting Chen 0001, Zhendong Su 0001
ICSME4
2020 Learning to Recommend Trigger-Action Rules for End-User Development - A Knowledge Graph Based Approach
Qinyue Wu, Beijun Shen, Yuting Chen 0001
ICSR3
2020 Are the Code Snippets What We Are Searching for? A Benchmark and an Empirical Study on Code Search with Natural-Language Queries
abstract
Code search methods, especially those that allow programmers to raise queries in a natural language, plays an important role in software development. It helps to improve programmers' productivity by returning sample code snippets from the Internet and/or source-code repositories for their natural-language queries. Meanwhile, there are many code search methods in the literature that support natural-language queries. Difficulties exist in recognizing the strengths and weaknesses of each method and choosing the right one for different usage scenarios, because (1) the implementations of those methods and the datasets for evaluating them are usually not publicly available, and (2) some methods leverage different training datasets or auxiliary data sources and thus their effectiveness cannot be fairly measured and may be negatively affected in practical uses. To build a common ground for measuring code search methods, this paper builds CosBench, a dataset that consists of 1000 projects, 52 code-independent natural-language queries with ground truths, and a set of scripts for calculating four metrics on code research results. We have evaluated four IR (Information Retrieval)-based and two DL (Deep Learning)-based code search methods on CosBench. The empirical evaluation results clearly show the usefulness of the CosBench dataset and various strengths of each code search method. We found that DL-based methods are more suitable for queries on reusing code, and IR-based ones for queries on resolving bugs and learning API uses.
Shuhan Yan, Yuting Chen 0001, Beijun Shen, Lingxiao Jiang
SANER3
2020 Semantic Service Search in IT Crowdsourcing Platform: A Knowledge Graph-Based Approach
abstract
Understanding user’s search intent in vertical websites like IT service crowdsourcing platform relies heavily on domain knowledge. Meanwhile, searching for services accurately on crowdsourcing platforms is still difficult, because these platforms do not contain enough information to support high-performance search. To solve these problems, we build and leverage a knowledge graph named ITServiceKG to enhance search performance of crowdsourcing IT services. The main ideas are to (1) build an IT service knowledge graph from Wikipedia, Baidupedia, CN-DBpedia, StuQ and data in IT service crowdsourcing platforms, (2) use properties and relations of entities in the knowledge graph to expand user query and service information, and (3) apply a listwise approach with relevance features and topic features to re-rank the search results. The results of our experiments indicate that our approach outperforms the traditional search approaches.
Qinyue Wu, Duankang Fu, Beijun Shen, Yuting Chen 0001
Int. J. Softw. Eng. Knowl. Eng.4
2020 Feedback2Code: A Deep Learning Approach to Identifying User-Feedback-Related Source Code Files
abstract
Users frequently raise feedback when using software products. Feedback from users regarding their experiences and expectations and software defects they found adds values to software maintenance and evolution — software managers collect user feedback and then dispatch feedback issues that developers (and/or maintainers) need to track and process. Feedback tracking is often supported by open source platforms and collaborative software systems. Meanwhile, there still exists a gap between feedback issues and source code: since user feedback is usually informal and arbitrary, engineers have to spend much effort on comprehending issues and identifying which source code files need to be improved or fixed. This paper introduces a deep learning approach, Feedback2Code , which facilitates identification of user-feedback-related source code files. The core idea is to (1) explore latent semantics of user feedback and source code using several deep learning techniques such as Multi-Layer Perceptron (MLP), Convolutional Neutral Network (CNN) and skip-gram and (2) establish a multi-correlation model to explore linkages between feedback issues and source code files. Given a feedback issue, the linkages then allow engineers to identify source code files that are highly relevant to the issue. We have implemented Feedback2Code and evaluated it against ChangeAdvisor (a state-of-the-art approach) on 24 open source projects. The evaluation results clearly show the strength of Feedback2Code : for 103793 feedback issues, Feedback2Code successfully established 101190 feedback-code linkages and achieved a precision that is [Formula: see text] higher than that of ChangeAdvisor . Feedback2Code also achieved an MRR and an MAP that are [Formula: see text] and [Formula: see text] higher than those of ChangeAdvisor , respectively. Furthermore, we also found that a Feedback2Code -trained model can be easily transferred, allowing feedback-code linkages to be established in new projects with a little history data.
Shuhan Yan, Tianjiao Du, Beijun Shen, Yuting Chen 0001, Zhilei Ren
Int. J. Softw. Eng. Knowl. Eng.4
2019 Reinforcement Learning of Code Search Sessions
abstract
Searching and reusing online code is a common activity in software development. Meanwhile, like many general-purposed searches, code search also faces the session search problem: in a code search session, the user needs to iteratively search for code snippets, exploring new code snippets that meet his/her needs and/or making some results highly ranked. This paper presents Cosoch, a reinforcement learning approach to session search of code documents (code snippets with textual explanations). Cosoch is aimed at generating a session that reveals user intentions, and correspondingly searching and reranking the resulting documents. More specifically, Cosoch casts a code search session into a Markov decision process, in which rewards measuring the relevances between the queries and the resulting code documents guide the whole session search. We have built a dataset, say CosoBe, from StackOverflow, containing 103 code search sessions with 378 pieces of user feedback. We have also evaluated Cosoch on CosoBe. The evaluation results show that Cosoch achieves an average NDCG@3 score of 0.7379, outperforming StackOverflow by 21.3%.
Wei Li 0254, Shuhan Yan, Beijun Shen, Yuting Chen 0001
APSEC4
2019 An Adaptive Approach to Recommending Obfuscation Rules for Java Bytecode Obfuscators
abstract
Bytecode obfuscation is an essential technique for protecting intellectual property and defending against Man-AtThe-End (MATE) attacks to Java/Android applications. Several bytecode obfuscators have been developed for modifying or refactoring Java bytecode (.class) so that it becomes hard to understand but remains fully functional. These obfuscators usually integrate a variety of obfuscation rules, allowing obfuscation algorithms to be combined and enforced on the applications. Meanwhile, it still remains a difficulty: Given a bytecode file f, which obfuscation rule(s) need to be applied such that f can get obfuscated sufficiently? This paper presents ORChooser (Obfuscation Rule Chooser), an adaptive approach to recommending a small number obfuscation rules for Java bytecode obfuscators. The key idea of ORChooser is, given a bytecode obfuscator, to (1)randomly select/unselect obfuscation rules for the obfuscator, and (2)calculate the obfuscation distance between the bytecode before and after obfuscation. Furthermore, ORChooser takes an iterative process to adaptively obfuscate the bytecode file f such that the obfuscated code is far away from f. We have implemented ORChooser and evaluated it on a state-of-the-art bytecode obfuscators: Android R8. The evaluation results clearly show the strength of ORChooser. In particular, within 5 iterations, ORChooser chose about 25% of obfuscation rules for R8, reducing more than 29% of the bytecode size. The similarity between the bytecode files before and after obfuscation is less than 27%, indicating that the ORChooser-supported obfuscators have obfuscated bytecode sufficiently and reduce its comprehensibility significantly.
Yanru Peng, Yuting Chen 0001, Beijun Shen
COMPSAC (1)2
2019 A Frequency-Aware Spatio-Temporal Network for Traffic Flow Prediction
Shunfeng Peng, Yanyan Shen, Yanmin Zhu 0006, Yuting Chen 0001
DASFAA (2)4
2019 Deep differential testing of JVM implementations
abstract
The Java Virtual Machine (JVM) is the cornerstone of the widely-used Java platform. Thus, it is critical to ensure the reliability and robustness of popular JVM implementations. However, little research exists on validating production JVMs. One notable effort is classfuzz, which mutates Java bytecode syntactically to stress-test different JVMs. It is shown that classfuzz mainly produces illegal bytecode files and uncovers defects in JVMs' startup processes. It remains a challenge to effectively test JVMs' bytecode verifiers and execution engines to expose deeper bugs. This paper tackles this challenge by introducing classming, a novel, effective approach to performing deep, differential JVM testing. The key of classming is a technique, live bytecode mutation, to generate, from a seed bytecode file f, likely valid, executable (live) bytecode files: (1) capture the seed f's live bytecode, the sequence of its executed bytecode instructions; (2) repeatedly manipulate the control- and data-flow in f's live bytecode to generate semantically different mutants; and (3) selectively accept the generated mutants to steer the mutation process toward live, diverse mutants. The generated mutants are then employed to differentially test JVMs. We have evaluated classming on mainstream JVM implementations, including OpenJDK's HotSpot and IBM's J9, by mutating the DaCapo benchmarks. Our results show that classming is very effective in uncovering deep JVM differences. More than 1,800 of the generated classes exposed JVM differences, and more than 30 triggered JVM crashes. We analyzed and reported the JVM runtime differences and crashes, of which 14 have already been confirmed/fixed, including a highly critical security vulnerability in J9 that allowed untrusted code to disable the security manager and elevate its privileges (CVE-2017-1376).
Yuting Chen 0001, Ting Su 0001, Zhendong Su 0001
ICSE1
2019 CocoQa: Question Answering for Coding Conventions Over Knowledge Graphs
abstract
Coding convention plays an important role in guaranteeing software quality. However, coding conventions are usually informally presented and inconvenient for programmers to use. In this paper, we present CocoQa, a system that answers programmer's questions about coding conventions. CocoQa answers questions by querying a knowledge graph for coding conventions. It employs 1) a subgraph matching algorithm that parses the question into a SPARQL query, and 2) a machine comprehension algorithm that uses an end-to-end neural network to detect answers from searched paragraphs. We have implemented CocoQa, and evaluated it on a coding convention QA dataset. The results show that CocoQa can answer questions about coding conventions precisely. In particular, CocoQa can achieve a precision of 82.92% and a recall of 91.10%. Repository: https://github.com/14dtj/CocoQa/ Video: https://youtu.be/VQaXi1WydAU.
Tianjiao Du, Junming Cao, Qinyue Wu, Wei Li 0254, Beijun Shen, Yuting Chen 0001
ASE6
2019 Constructing a Knowledge Base of Coding Conventions from Online Resources
abstract
Coding conventions are a set of coding guidelines used by software developers to improve the readability of source code, increase software maintainability, and promote the reuse of coding patterns.In this paper, we introduce CCBase, a knowledge base of coding conventions, that was constructed from online resources.Specifically, CCBase was constructed as follows.We designed the ontology of the coding convention domain, crawled data related to coding conventions from a variety of online resources, and then extracted entities and relations using an NLP-enabled rule matching method.To uncover the latent relations, we further proposed a similarity metric to reveal the similar-to and relate-to relations, and developed a RCE algorithm to establish a unified type hierarchy of coding conventions.The resulting knowledge base contains 3139 coding conventions for Java and C++, with 3761 entities and 767 relations.Furthermore, we have extended the usability of CCBase by developing a question answering system on the base.We have conducted experiments to evaluate CCBase.The experimental results show that CCBase has a wide coverage on entities and relations in coding conventions domain, and the QA system achieves an F1 score of 84.5% on 214 questions raised in StackOverflow.
Junming Cao, Tianjiao Du, Beijun Shen, Wei Li 0254, Qinyue Wu, Yuting Chen 0001
SEKE6
2019 Generating SQL Statements from Natural Language Queries: A Multitask Learning Approach (S)
abstract
NL2SQL advocates an idea of helping engineers and/or end users generate SQL statements from natural language queries.However, it still remains a strong challenge in improving its precision and scalability.This paper introduces MultiSQL, a multitask deep learning approach to performing NL2SQL.MultiSQL unifies the task representations and trains a model in parallel on multiple tasks, including NL2SQL, machine translation, etc.It employs a multitask question-answering network for jointly learning all tasks and transferring knowledge among tasks.We have evaluated MultiSQL on two query datasets: WikiSQL (an open sourced dataset) and CnSQL (a Chinese dataset we created).The evaluation results clearly show the effectiveness of MultiSQL.In particular, the accuracies achieved by MultiSQL approximate those achieved by the state-of-the-art NL2SQL methods on WikiSQL, and its accuracy is 78%, which is 17% higher than the "Chinese2English + NL2SQL" method on CnSQL.
Chunqi Chen, Yunxiang Xiong, Beijun Shen, Yuting Chen 0001
SEKE4
2019 Enhancing Semantic Search of Crowdsourcing IT Services using Knowledge Graph
abstract
Mining search intents in vertical websites like IT service crowdsourcing platform relies heavily on domain knowledge.Meanwhile, it still remains a difficulty of searching services in crowdsourcing platforms, as these platforms do contain much insufficient information, for example, users tend to use images describing IT services for the purpose of advertisements.To solve these problems, we build and leverage a knowledge graph to enhance searching of crowdsourcing IT services.The key idea is to (1) build an IT service knowledge graph from StackOverflow tag synonym system, Wikipedia, StuQ and data in IT service crowdsourcing platforms, (2) plug two activities into the basic search processterm expansion and service re-ranking, (3) use superordinates, hypernyms, synonyms, descriptions and relations of entities in the knowledge graph to expand user query and service information, and (4) apply a learning-to-rank model with four features to re-rank the search results, enforcing those more relevant services have the higher-ranking position.We have conducted several experiments to evaluate our approach.The results show that our approach achieves an MRR 34.9% higher and a Recall@15 11% higher than those of a basic search approach.
Duankang Fu, Shufan Zhou, Beijun Shen, Yuting Chen 0001
SEKE4
2019 CrowDevBot: A Task-Oriented Conversational Bot for Software Crowdsourcing Platform (S)
abstract
With the trends of developing software on the Internet, many software crowdsourcing platforms are emerging.They attract a lot of developers to bid for crowdsourced projects and develop software systems collaboratively.In this paper, we present CrowDevBot, a task-oriented conversational bot for software crowdsourcing platform, that aims to assist online users in completing crowdsourcing-related tasks in a more natural manner.The key idea of CrowDevBot is to: (1) combine a rulebased method and an SVM-NaiveBayes-C4.5 integrated learning method to discover users' intention; (2) employ an integrated CRF (conditional random field) method with novel features to improve the performance of slot filling; and (3) leverage a software service knowledge base to unify entity names and predefine the key slots of user query.We implement CrowDevBot and integrate it into JointForce, an IT software crowdsourcing platform in China.To the best of our knowledge, this is the first time that a task-oriented conversational bot is practically used in software crowdsourcing platform(s).We evaluated our approach on real data set from JointForce.The results show that our intention detecting method achieves F1-score of 87% on the limited training data.For the slot filling, the F1-score of our integrated CRF model reaches 82%, 8% higher than that of the normal CRF model.
Zeyu Ni, Beijun Shen, Yuting Chen 0001, Zhangyuan Meng, Junming Cao
SEKE3
2019 API recommendation for event-driven Android application development
Weizhao Yuan, Lingxiao Jiang, Yuting Chen 0001, Jianjun Zhao 0001, Haibo Yu 0001
Inf. Softw. Technol.4
2019 JDap: Supporting in-memory data persistence in javascript using Intel's PMDK
Litong You, Qipeng Zhang, Tianyou Li, Chen Li 0009, Yuting Chen 0001, Linpeng Huang
J. Syst. Archit.6
2018 SPMP: A JavaScript Support for Shared Persistent Memory on Node.js
Qipeng Zhang, Tianyou Li, Yuting Chen 0001, Linpeng Huang, Andy Rudoff
ICA3PP (2)4
2018 JSNVM: Supporting Data Persistence in JavaScript Using Non-Volatile Memory
abstract
Data persistence refers to an characteristic that data needs to outlive its creator. In recent years, NVM (Non-volatile Memory), a new type of computer memory supporting memory-like byte-addressable access and disk-like persistence, became a promising technique for facilitating data persistence in memory and its management. So far NVM has been supported in many programming languages such as C, C++ and Java. Meanwhile, it has not yet been well supported in JavaScript, a popular scripting language for software development. This paper presents JSNVM, an extension to JavaScript's runtime and execution engine, to support data persistence in JavaScript using NVM. JSNVM consists of (1) a JavaScript persistent object pool that serves as a persistent heap in which persistent objects are created and managed. This pool is also enhanced for guaranteeing data safety and consistency; and (2) a set of JavaScript persistent APIs that provide programmers with a support in creating, managing, and accessing persistent data in an easy-to-use and safe manner. We have implemented JSNVM and evaluated it against two database-supported data persistence styles (i.e., MongoDB's object store and indexedDB's binary store) on micro-benchmarks and real-world applications. The evaluation results show that compared with MongoDB's object store, JSNVM can achieve a 1.6× speedup; when applied to access binary data, JSNVM can achieve a 24.6× speedup over indexedDB's binary store, denoting that JSNVM can enhance performance of JavaScript applications in practice. JSNVM is planned to be publicly available by the end of 2018.
Yanmin Zhu 0006, Yuting Chen 0001, Linpeng Huang, Tianyou Li
ICPADS3
2018 The role of model checking in software engineering
Anil Kumar Karna, Yuting Chen 0001, Haibo Yu 0001, Hao Zhong 0001, Jianjun Zhao 0001
Frontiers Comput. Sci.2
2017 CRSearcher: Searching Code Database for Repairing Bugs
abstract
With the exponentially rising of software development in the past decades, millions of software products have been created. Existing empirical studies show that many code snippets are similar. Although there exist many difficulties in maintaining these similar code snippets, we believe that it is feasible to leverage the similarity to enhance program repairs, as bugs may have already been repaired in many other similar code snippets.
Yingyi Wang, Yuting Chen 0001, Beijun Shen, Hao Zhong 0001
Internetware2
2017 Cold-Start Developer Recommendation in Software Crowdsourcing: A Topic Sampling Approach
abstract
Recently, software crowdsourcing platforms, which provide paid tasks for developers, become attractive to both employers and developers.Developers expect to find tasks that match their interests and capabilities via crowdsourcing platforms, and thus recommender systems play important roles in these platforms.However, we still face several challenges when building a recommender system for a crowdsourcing platform.A major challenge is how to recommend tasks to cold-start developers whose task interaction data is not available.This paper presents a novel, topic sampling approach to tackling with the cold-start developer recommendation problem.First, it employs a general method for modeling developers and tasks, which solves the data heterogeneous issue across different platforms.After that, it casts the cold-start developer recommendation problem into a multi-optimization problem, and takes a topic-sampling based genetic algorithm to recommend tasks.More specifically, our approach is different from traditional solutions in that it leverages task descriptions and popularity-to-be, allowing new tasks to be recommended to cold-start developers.To evaluate the effectiveness of the proposed approach, we have conducted experiments on a large dataset crawled from three real-world software crowdsourcing platforms.Compared with other state-ofthe-art recommendation solutions, the experimental results show that the proposed approach improves 75% of precision and recall on average.
Wenkai Mo, Beijun Shen, Yuting Chen 0001
SEKE4
2017 Guided, stochastic model-based GUI testing of Android apps
abstract
Mobile apps are ubiquitous, operate in complex environments and are developed under the time-to-market pressure. Ensuring their correctness and reliability thus becomes an important challenge. This paper introduces Stoat, a novel guided approach to perform stochastic model-based testing on Android apps. Stoat operates in two phases: (1) Given an app as input, it uses dynamic analysis enhanced by a weighted UI exploration strategy and static analysis to reverse engineer a stochastic model of the app's GUI interactions; and (2) it adapts Gibbs sampling to iteratively mutate/refine the stochastic model and guides test generation from the mutated models toward achieving high code and model coverage and exhibiting diverse sequences. During testing, system-level events are randomly injected to further enhance the testing effectiveness.
Ting Su 0001, Guozhu Meng, Yuting Chen 0001, Geguang Pu, Yang Liu 0003, Zhendong Su 0001
ESEC/SIGSOFT FSE3
2016 Heterogeneous Cross-Company Effort Estimation through Transfer Learning
abstract
Software effort estimation is vital but challenging activity during software development. In many small or medium-sized companies, such challenges are stemmed from historical data shortage. The problem can be solved by leveraging cross-company data for effort estimation. While in practice, cross-company effort estimation may not be easy to take because the cross-company data for effort estimation can be heterogenous. In this paper, we propose a novel approach named Mixture of Canonical Correlation Analysis and Restricted Boltzmann Machines (MCR) to address data heterogeneity issue in cross-company effort estimation. The essential ideas in MCR are (1) to present a unified metric representing heterogenous effort estimation data; and (2) to combine Canonical Correlation Analysis and Restricted Boltzmann Machines method to estimate effort in heterogenous cross-company effort estimation. The MCR approach is evaluated on 5 public datasets in PROMISE repository. The evaluation results show that: (1) for estimations with partially different metrics, the MCR approach outperforms within-company effort estimator KNN with a decrease in MMRE by 0.60, an increase in PRED(25) by 0.16, and a decrease in MdMRE by 0.19; (2) for estimations with totally different metrics, the MCR approach outperforms within-company effort estimator KNN with a decrease in MMRE by 0.49, an increase in PRED(25) by 0.08, and a decrease in MdMRE by 0.10.
Shensi Tong, Yuting Chen 0001, Beijun Shen
APSEC3
2016 GRETA: Graph-Based Tag Assignment for GitHub Repositories
abstract
GitHub is a well-known software community where a large number of software repositories are hosted. Since large amounts of documents and code in GitHub repositories are in a mess, users cannot search or understand them efficiently. One solution is to employ a tag system, which annotates each repository with several tags. Thus, the GitHub repositories can be more efficiently accessed and understood. However, GitHub does not provide any automated tools of tagging repositories. In order to tackle this problem, we propose GRETA, a novel graph-based approach to assigning tags for repositories on GitHub. The core insight of GRETA is (1) to construct an Entity-Tag Graph (ETG) for GitHub using the domain knowledge from StackOverflow, and (2) to assign tags for repositories by taking a random walk algorithm. We have implemented GRETA and also developed a repository search engine for GitHub using the tag assignment results of GRETA. We have evaluated GRETA against several baseline methods to investigate its effectiveness of tagging GitHub repositories. The results show GRETA achieves up to 35% of F-Measure, outperforming the baseline methods. Besides, the GRETA-based search engine gains a higher NDCG value than the search engine provided by GitHub, indicating that it significantly enhances the search ability on GitHub with the tagged repositories.
Xuyang Cai, Jiangang Zhu, Beijun Shen, Yuting Chen 0001
COMPSAC4
2016 Software Defect Prediction Using Semi-Supervised Learning with Change Burst Information
abstract
Software defect prediction is an important software quality assurance technique. It utilizes historical project data and previously discovered defects to predict potential defects. However, most of existing methods assume that large amounts of labeled historical data are available for prediction, while in the early stage of the life cycle, projects may lack the data needed for building such predictors. In addition, most of existing techniques use static code metrics as predictors, while they omit change information that may introduce risks into software development. In this paper, we take these two issues into consideration, and propose a semi-supervised based defect prediction approach - extRF. extRF extends the classical supervised Random Forest algorithm by self-training paradigm. It also employs change burst information for improving accuracy of software defect prediction. We also conduct an experiment to evaluate extRF against three other supervised machine learners (i.e. Logistic Regression, Naive Bayes, Random Forest) and compare the effectiveness of code metrics, change burst metrics, and a combination of them. Experimental results show that extRF trained with a small size of labeled dataset achieves comparable performance to some supervised learning approaches trained with a larger size of labeled dataset. When only 2% of Eclipse 2.0 data are used for training, extRF can achieve F-measure about 0.562, approximate to that of LR (a supervised learning approach) at labeled sampling rate of 50%. Besides, change burst metrics outperform code metrics in that F-measure rises to a peak value of 0.75 for Eclipse 3.0 and JDT.Core.
Beijun Shen, Yuting Chen 0001
COMPSAC3
2016 SatiIndicator: Leveraging User Reviews to Evaluate User Satisfaction of SourceForge Projects
abstract
Quality of software (QoS) is important for users, as it may lead to high cost when a user or a company happens to pick up a software project with low quality. In recent years, many software quality assessment models take user satisfaction as an important metric for measuring software quality. However, user satisfaction on a software project is usually not precisely evaluated. In this paper, we propose a novel, automated approach called SatiIndicator to evaluate user satisfaction of a software project by analyzing user reviews with user opinions and emotions. The essential idea of SatiIndicator is to (1) use a topic model to cluster all aspects in a software genre into different topics and compute the weight of each topic, (2) take sentiment analysis and calculate the sentiment strength of every aspect and pure attitude reviews, and (3) evaluate the user satisfaction score for a software project. Wilson Interval is applied to punish the software projects with insufficient reviews in order to keep fairness. We have evaluated SatiIndicator on ten software genres in SourceForge. The evaluation results show that when software projects have sufficient reviews, SatiIndicator performs 35% higher than baselines at p@3, 15% higher than baselines at p@15 and over 85% Spearman Coefficient with the ground truth. When software projects have insufficient reviews, SatiIndicator performs 30% higher than baselines at p@3, 15% higher than baselines at p@15 and over 60% Spearman Coefficient with the ground truth.
Zhenzheng Qian, Beijun Shen, Wenkai Mo, Yuting Chen 0001
COMPSAC4
2016 Multi-perspective change impact analysis using linked data of software engineering
abstract
Change impact analysis plays an important role in software maintenance and evolution. However existing researches mostly focus on one single artifact. Software development is usually accompanied by various types of software artifacts, such as requirement documents, software architectures, test cases, source code, etc., requiring a much more comprehensive change impact analysis. This paper presents a novel approach to multi-perspective change impact analysis that is able to address heterogeneous software artifacts. The essential idea of the novel approach is (1) to adopt semantic web to construct automatically ontology based software engineering linked data, which links requirements, classes, code, bug reports, commits, developers, test cases and others, (2) to build a weighted change impact matrix/graph using the dependency features extracted from linked data, and (3) to follow a change impact propagation algorithm to analyze the overall change impacts. We have conducted experiments on two open source projects (HtmlUnit and OpenRocket) to evaluate our approach. The experimental results show that our approach achieves better F-measure and stability than existing multi-perspective change impact analysis approaches.
Chengcheng Wan 0001, Zece Zhu, Yuting Chen 0001
Internetware4
2016 Multicast routing tree for sequenced packet transmission in software-defined networks
abstract
Multicast denotes an idea of sending data to numbers of receivers from one source in one transmission. It has been widely applied in group communication (e.g., media streaming, multi-point video conferencing). Multicast routing tree (MRT) is usually built to keep the right paths to transmit data, where data copies are created in parent nodes and then forwarded to child nodes. However, constructing an MRT is usually difficult for a given network topology; finding an optimal multicast routing tree with the minimal cost is a proven NP-complete problem. Moreover, multicast applications usually run in local or small networks due to the limitations in flexibility, scalability, and security.
Renke Wu, Haojie Zhou, Haibo Yu 0001, Yuting Chen 0001, Hao Zhong 0001
Internetware5
2016 Rule-directed code clone synchronization
abstract
Code clones are prevalent in software systems due to many factors in software development. Detecting code clones and managing consistency between them along code evolution can be very useful for reducing clone-related bugs and maintenance costs. Despite some early attempts at detecting code clones and managing the consistency between them, the state-of-the-art tool can only handle simple code clones whose structures are identical or quite similar. However, existing empirical studies show that clones can have quite different structures with their evolution, which can easily go beyond the capability of the state-of-the-art tool. In this paper, we propose CCSync, a novel, rule-directed approach, which paves the structure differences between the code clones and synchronizes them even when code clones become quite different in their structures. The key steps of this approach are, given two code clones, to (1) extract a synchronization rule from the relationship between the clones, and (2) once one code fragment is updated, propagate the modifications to the other following the synchronization rule. We have implemented a tool for CCSync and evaluated its effectiveness on five Java projects. Our results shows that there are many code clones suitable for synchronization, and our tool achieves precisions of up to 92% and recalls of up to 84%. In particular, more than 76% of our generated revisions are identical with manual revisions.
Xiao Cheng 0006, Hao Zhong 0001, Yuting Chen 0001, Zhenjiang Hu 0002, Jianjun Zhao 0001
ICPC3
2016 LockPeeker: detecting latent locks in Java APIs
abstract
Detecting lock-related defects has long been a hot research topic in software engineering. Many efforts have been spent on detecting such deadlocks in concurrent software systems. However, latent locks may be hidden in application programming interface (API) methods whose source code may not be accessible to developers. Many APIs have latent locks. For example, our study has shown that J2SE alone can have 2,000+ latent locks. As latent locks are less known by developers, they can cause deadlocks that are hard to perceive or diagnose. Meanwhile, the state-of-the-art tools mostly handle API methods as black boxes, and cannot detect deadlocks that involve such latent locks. In this paper, we propose a novel black-box testing approach, called LockPeeker, that reveals latent locks in Java APIs. The essential idea of LockPeeker is that latent locks of a given API method can be revealed by testing the method and summarizing the locking effects during testing execution. We have evaluated LockPeeker on ten real-world Java projects. Our evaluation results show that (1) LockPeeker detects 74.9% of latent locks in API methods, and (2) it enables state-of-the-art tools to detect deadlocks that otherwise cannot be detected.
Hao Zhong 0001, Yuting Chen 0001, Jianjun Zhao 0001
ASE3
2016 Coverage-directed differential testing of JVM implementations
abstract
Java virtual machine (JVM) is a core technology, whose reliability is critical. Testing JVM implementations requires painstaking effort in designing test classfiles (*.class) along with their test oracles. An alternative is to employ binary fuzzing to differentially test JVMs by blindly mutating seeding classfiles and then executing the resulting mutants on different JVM binaries for revealing inconsistent behaviors. However, this blind approach is not cost effective in practice because most of the mutants are invalid and redundant. This paper tackles this challenge by introducing classfuzz, a coverage-directed fuzzing approach that focuses on representative classfiles for differential testing of JVMs’ startup processes. Our core insight is to (1) mutate seeding classfiles using a set of predefined mutation operators (mutators) and employ Markov Chain Monte Carlo (MCMC) sampling to guide mutator selection, and (2) execute the mutants on a reference JVM implementation and use coverage uniqueness as a discipline for accepting representative ones. The accepted classfiles are used as inputs to differentially test different JVM implementations and find defects. We have implemented classfuzz and conducted an extensive evaluation of it against existing fuzz testing algorithms. Our evaluation results show that classfuzz can enhance the ratio of discrepancy-triggering classfiles from 1.7% to 11.9%. We have also reported 62 JVM discrepancies, along with the test classfiles, to JVM developers. Many of our reported issues have already been confirmed as JVM defects, and some even match recent clarifications and changes to the Java SE 8 edition of the JVM specification.
Yuting Chen 0001, Ting Su 0001, Chengnian Sun, Zhendong Su 0001, Jianjun Zhao 0001
PLDI1
2016 Evaluating quality-in-use of FLOSS through analyzing user reviews
abstract
Quality-in-use (QU) is an important measure for evaluating the quality of a software system from user respective. Several approaches have been proposed to evaluate QU of FLOSS. Meanwhile, they are usually less effective, as they usually assume that sufficient, precise usage statistics can be collected, which may be impractical for evaluating many real-world FLOSS systems. This paper presents QUIndicator, a novel, fine-grained approach to evaluating QU of FLOSS using user reviews. The key idea of QUIndicator is to, for a specific FLOSS system, (1) use a topic model to cluster its user reviews into different topics, transform topics into characteristics of a QU model and compute the weight of each characteristic, (2) take review aspect as the minimum analysis unit, and also apply sentiment analysis to analyze the sentiment strength of each review aspect, and (3) match review aspects with their corresponding characteristics in QU model and evaluate the QU of the system. Wilson interval is adopted to keep fairness by punishing FLOSS systems with insufficient reviews. We have evaluated QUIndicator on ten FLOSS genre datasets. The evaluation results show that when a FLOSS system has sufficient reviews, QUIndicator can achieve a p@3 value up to 30% higher than those of baselines, and also achieves over 75% Spearman Coefficient with the ground truth. When the system has insufficient reviews, QUIndicator can achieve a p@3 value up to 25% higher than those of baselines and over 55% Spearman Coefficient with the ground truth.
Zhenzheng Qian, Chengcheng Wan 0001, Yuting Chen 0001
SNPD3
2016 Supporting Selective Undo for Refactoring
abstract
Due to various considerations, programmers often need to backtrack their code. Furthermore, as the most recent edit may not be the wrong edit, programmers sometimes have to backtrack their code for arbitrary edits, which is referred as selective undo in this paper. To meet the needs, researchers have proposed various approaches to support selective undo. However, to the best of our knowledge, these approaches can support only simple edits, and cannot handle refactoring, although most code editors already provide various refactoring actions. Indeed, it is challenging to support selective undo for refactoring, since multiple code elements and complicated actions can be involved. In this paper, we present a novel approach that leverages Bidirectional Transformation (BX) to support selective undo for refactoring. We evaluate our approach on a recent refactoring tool that transfers enhanced for loops to lambda expressions. Our results show that our approach achieves an accuracy of up to 89%.
Xiao Cheng 0006, Yuting Chen 0001, Zhenjiang Hu 0002, Tao Zan, Hao Zhong 0001, Jianjun Zhao 0001
SANER2
2015 TBIL: A Tagging-Based Approach to Identity Linkage Across Software Communities
abstract
Nowadays, developers can be involved in several software developer communities like StackOverflow and Github. Meanwhile, accounts from different communities are usually less connected. Linking these accounts, which is called identity linkage, is a prerequisite of many interesting studies such as investigating activities of one developer in two or more communities. Many researches have been performed on social networks, but very few of them can be adapted to software communities, as information of users provided in these communities has a huge difference to that in social networks. We tackle with the problem by introducing TBIL, a novel tagging-based approach to identity linkage among software communities. The essential idea of this approach is to employ skills (measured by tags), usernames and concerned topics of developers as hints, and to use a decision tree-based algorithm and another heuristic greedy matching algorithm to link user identities. We measure the effectiveness of TBIL on two well-known software communities, i.e., StackOverflow and Github. The results show that our method is feasible and practical in linking developer identities. In particular, the F-Score of our method is 0.15 higher than previous identity linkage methods in software communities.
Wenkai Mo, Beijun Shen, Yuting Chen 0001, Jiangang Zhu
APSEC3
2015 JaConTeBe: A Benchmark Suite of Real-World Java Concurrency Bugs (T)
abstract
Researchers have proposed various approaches to detect concurrency bugs and improve multi-threaded programs, but performing evaluations of the effectiveness of these approaches still remains a substantial challenge. We survey the existing evaluations and find out that they often use code or bugs not representative of real world. To improve representativeness, we have prepared JaConTeBe, a benchmark suite of 47 confirmed concurrency bugs from 8 popular open-source projects, supplemented with test cases for reproducing buggy behaviors. Running three approaches on JaConTeBe shows that our benchmark suite confirms some limitations of the three approaches. We submitted JaConTeBe to the SIR repository (a software-artifact repository for rigorous controlled experiments), and it was included as a part of SIR.
Darko Marinov, Hao Zhong 0001, Yuting Chen 0001, Jianjun Zhao 0001
ASE4
2015 Towards Effective Developer Recommendation in Software Crowdsourcing
abstract
Crowdsourcing has attracted increasing attention from both industry and academia since it was proposed.Now a lot of work is finished by crowdsourcing, such as logo design, website promotion, industrial design, copywriting, software development, translation and image annotation.Although software crowdsourcing achieves positive results in practice, we still face a challenge of assigning suitable developers to specific tasks.In this paper, we propose a novel approach that recommends developers.In particular, our approach supports: comprehensively measuring the tasks and developers in software crowdsourcing, and recommending developers on the basis of the developer-task competence, task-task similarity, and soft power.
Shixiong Zhao, Beijun Shen, Yuting Chen 0001, Hao Zhong 0001
SEKE3
2015 Guided differential testing of certificate validation in SSL/TLS implementations
abstract
Certificate validation in SSL/TLS implementations is critical for Internet security. There is recent strong effort, namely frankencert, in automatically synthesizing certificates for stress-testing certificate validation. Despite its early promise, it remains a significant challenge to generate effective test certificates as they are structurally complex with intricate syntactic and semantic constraints. This paper tackles this challenge by introducing mucert, a novel, guided technique to much more effectively test real-world certificate validation code. Our core insight is to (1) leverage easily accessible Internet certificates as seed certificates, and (2) diversify them by adapting Markov Chain Monte Carlo (MCMC) sampling. The diversified certificates are then used to reveal discrepancies, thus potential flaws, among different certificate validation implementations. We have implemented mucert and extensively evaluated it against frankencert. Our experimental results show that mucert is significantly more cost-effective than frankencert. Indeed, 1K mucerts (i.e., mucert-mutated certificates) yield three times as many distinct discrepancies as 8M frankencerts (i.e., frankencert-synthesized certificates), and 200 mucerts can achieve higher code coverage than 100,000 frankencerts. This improvement is significant as it incurs much cost to test each generated certificate. We have analyzed and reported 20+ latent discrepancies (presumably missed by frankencert), and reported an additional 357 discrepancy-triggering certificates to SSL/TLS developers, who have already confirmed some of our reported issues and are investigating causes of all the reported discrepancies. In particular, our reports have led to bug fixes, active discussions in the community, and proposed changes to relevant IETF’s RFCs. We believe that mucert is practical and effective for helping improve the robustness of SSL/TLS implementations.
Yuting Chen 0001, Zhendong Su 0001
ESEC/SIGSOFT FSE1
2014 Mining Developer Mailing List to Predict Software Defects
abstract
It has been studied that the communication among software stakeholders can be used to predict potential software defects. Yet researchers have rarely studied the relations between the software and the mailing lists of the developers. In this paper, we research on how to predict software defects by mining the mailing lists of the software developers. First, we extract both the structural and the unstructured information from mailing lists as metrics. The structural information is calculated through analyzing the social network hidden in the mailing lists, and the unstructured information is obtained through taking topical and textual analysis of the lists. Second, we design a mailing list-based approach to predicting software defects. We have also analyzed the software repository of several open source projects by linking their bug tracking data-bases to the mailing list archives. The experimental results provide empirical evidence that the mailing list metrics are related to software quality and can be used as predictors of defect-proneness. Furthermore, we found that (1) messages having certain structures may indicate some defect related files, (2) the sentiment and some topic-specific mailing models are of strong correlations with the software defects.
Beijun Shen, Yuting Chen 0001
APSEC (1)3
2014 A Scenario-Based Approach to Predicting Software Defects Using Compressed C4.5 Model
abstract
Defect prediction approaches use software metrics and fault data to learn which software properties are associated with what kinds of software faults in programs. One trend of existing techniques is to predict the software defects in a program construct (file, class, method, and so on) rather than in a specific function scenario, while the latter is important for assessing software quality and tracking the defects in software functionalities. However, it still remains a challenge in that how a functional scenario is derived and how a defect prediction technique should be applied to a scenario. In this paper, we propose a scenario-based approach to defect prediction using compressed C4.5 model. The essential idea of this approach is to use a k-medoids algorithm to cluster functions followed by deriving functional scenarios, and then to use the C4.5 model to predict the fault in the scenarios. We have also conducted an experiment to evaluate the scenario-based approach and compared it with a file-based prediction approach. The experimental results show that the scenario-based approach provides with high performance by reducing the size of the decision tree by 52.65% on average and also slightly increasing the accuracy.
Biwen Li, Beijun Shen, Yuting Chen 0001, Jinshuang Wang
COMPSAC4
2014 AspectBreeze: integrating trustworthiness aspects into graph grammar supported architecture description language
abstract
Aspect-oriented software development (AOSD) has been developed for supporting a long-standing idea of Separation of Concerns (SoC) and enhancing software non-functional attributes including modularity, reusability, and maintainability. Many architectural description languages (ADLs) also provide software developers with support in specifying aspects in software architectures. However, when applied to describe trustworthy software systems, these ADLs face difficulties in using aspects to specify trustworthy attributes and maintaining consistency between the architectures before and after weaving of aspects. In this paper, we extend Breeze, a graph grammar supported ADL, to AspectBreeze. AspectBreeze allows trustworthiness aspects to be easily defined and seamlessly woven into base architectures. Further, architectures are defined along with graph grammars that allow any change to a base architecture to be reflected in the corresponding architecture in which aspects are woven. This paper also presents a case study of an online auction system to show how AspectBreeze is used and how graph grammars can maintain the consistency between the architectures before and after weaving of trustworthiness aspects.
Jingzhou Liu, Yuting Chen 0001, Chen Li 0009, Jianjun Zhao 0001
Internetware2
2014 A constraint-weaving approach to points-to analysis for AspectJ
Yuting Chen 0001, Jianjun Zhao 0001
Frontiers Comput. Sci.2
2013 Platform independent analysis of probabilities on execution paths of multithreaded programs
abstract
A concurrent program is intuitively associated with probability. In this paper we propose a platform independent approach, called ProbPP, to analyzing the probabilities on the execution paths of the multithreaded programs. The main idea of ProbPP is to calculate the probabilities on the basis of two kinds of probabilities: Primitive Dependent Probabilities (PDPs) representing the control dependent probabilities among the program statements and Thread Execution Probabilities (TEPs) representing the probabilities of threads being scheduled to execute. We have also conducted two preliminary experiments to evaluate the effectiveness and performance of ProbPP, and the experimental results show that ProbPP can provide engineers with acceptable accuracy.
Yuting Chen 0001
ICIS1
2013 Constraint-based locality analysis for X10 programs
abstract
X10 is a HPC (High Performance Computing) programming language proposed by IBM for supporting a PGAS (Partitioned Global Address Space) programming model offering a shared address space. The address space can be further partitioned into several logical locations where objects and activities (or threads) will be dynamically created. An analysis of locations can help to check the safety of object accesses through exploring which objects and activities may reside in which locations, while in practice the objects and activities are usually designated at runtime and their locations may also vary under different environments. In this paper, we propose a constraint-based locality analysis method called Leopard for X10. Leopard calculates the points-to relations for analyzing the objects and activities in a program and uses a place constraint graph to analyze their locations.We have developed a tool to support Leopard, and conducted an experiment to evaluate its effectiveness and efficiency. The experimental results show that Leopard can calculate the locations of objects and activities precisely.
Yuting Chen 0001, Jianjun Zhao 0001
PEPM2
2013 Extracting URLs from JavaScript via program analysis
abstract
With the extensive use of client-side JavaScript in web applications, web contents are becoming more dynamic than ever before. This poses significant challenges for search engines, because more web URLs are now embedded or hidden inside JavaScript code and most web crawlers are script-agnostic, significantly reducing the coverage of search engines. We present a hybrid approach that combines static analysis with dynamic execution, overcoming the weakness of a purely static or dynamic approach that either lacks accuracy or suffers from huge execution cost. We also propose to integrate program analysis techniques such as statement coverage and program slicing to improve the performance of URL mining.
Qi Wang 0017, Jingyu Zhou, Yuting Chen 0001, Jianjun Zhao 0001
ESEC/SIGSOFT FSE3
2012 Formal Specification-Based Inspection for Verification of Programs
abstract
Software inspection is a static analysis technique that is widely used for defect detection, but which suffers from a lack of rigor. In this paper, we address this problem by taking advantage of formal specification and analysis to support a systematic and rigorous inspection method. The aim of the method is to use inspection to determine whether every functional scenario defined in the specification is implemented correctly by a set of program paths and whether every program path of the program contributes to the implementation of some functional scenario in the specification. The method is comprised of five steps: deriving functional scenarios from the specification, deriving paths from the program, linking scenarios to paths, analyzing paths against the corresponding scenarios, and producing an inspection report, and allows for a systematic and automatic generation of a checklist for inspection. We present an example to show how the method can be used, and describe an experiment to evaluate its performance by comparing it to perspective-based reading (PBR). The result shows that our method may be more effective in detecting function-related defects than PBR but slightly less effective in detecting implementation-related defects. We also describe a prototype tool to demonstrate the supportability of the method, and draw some conclusions about our work.
Shaoying Liu, Yuting Chen 0001, Fumiko Nagoya, John A. McDermid
IEEE Trans. Software Eng.2
2011 Probabilistic Points-to Analysis for Java
Jianjun Zhao 0001, Yuting Chen 0001
CC3
2011 Frequency Estimation of Virtual Call Targets for Object-Oriented Programs
Sai Zhang 0001, Jianjun Zhao 0001, Yuting Chen 0001
ECOOP5
2010 BPGen: an automated breakpoint generator for debugging
abstract
During debugging processes, breakpoints are frequently used to inspect and understand runtime behaviors of programs. Although most development environments offer convenient breakpoint facilities, the use of these environments usually requires considerable human efforts in order to generate useful breakpoints. Before setting breakpoints or typing breakpoint conditions, developers usually have to make some judgements and hypotheses on the basis of their observations and experience. To reduce this kind of efforts we present a tool, named BPGen, to automatically generate breakpoints for debugging. BPGen uses three well-known dynamic fault localization techniques in tandem to identify suspicious program statements and states, through which both conditional and unconditional breakpoints are generated. BPGen is implemented as an Eclipse plugin for supplementing the existing Eclipse JDT debugger.
Dacong Yan, Jianjun Zhao 0001, Yuting Chen 0001, Shengqian Yang
ICSE (2)4
2010 A Rigorous Method for Inspection of Model-Based Formal Specifications
abstract
Writing formal specifications can help developers understand users' requirements, and build a solid foundation for implementation. But like other activities in software development, it is error-prone, especially for large-scale systems. In practice, effective detection of specification errors still remains a challenge. In this paper, we put forward a rigorous, systematic method for the inspection of model-based formal specifications. The method makes good use of the well-defined consistency properties of a specification to provide precise rules and guidelines for inspection. The inspection process utilizes both well-defined expressions derived from the specification and human inspectors' judgments to find errors. We present a case study of the method by describing how it is applied to inspect an Automated Teller Machine (ATM) software specification to investigate the method's feasibility, and explore potential challenges in using it. We also describe a prototype software tool including its functions and distinct features to demonstrate the tool supportability of the method.
Shaoying Liu, John A. McDermid, Yuting Chen 0001
IEEE Trans. Reliab.3
2009 A Divergence-Oriented Approach to Adaptive Random Testing of Java Programs
abstract
Adaptive Random Testing (ART) is a testing technique which is based on an observation that a test input usually has the same potential as its neighbors in detection of a specific program defect. ART helps to improve the efficiency of random testing in that test inputs are selected evenly across the input spaces. However, the application of ART to object-oriented programs (e.g., C++ and Java) still faces a strong challenge in that the input spaces of object-oriented programs are usually high dimensional, and therefore an even distribution of test inputs in a space as such is difficult to achieve. In this paper, we propose a divergence-oriented approach to adaptive random testing of Java programs to address this challenge. The essential idea of this approach is to prepare for the tested program a pool of test inputs each of which is of significant difference from the others, and then to use the ART technique to select test inputs from the pool for the tested program. We also develop a tool called ARTGen to support this testing approach, and conduct experiment to test several popular open-source Java packages to assess the effectiveness of the approach. The experimental result shows that our approach can generate test cases with high quality.
Xucheng Tang, Yuting Chen 0001, Jianjun Zhao 0001
ASE3
2008 A Systematic Approach for Integrating Fault Trees into System Statecharts
abstract
As software systems are encompassing a wide range of fields and applications, software reliability becomes a crucial step. The need for safety analysis and test cases that have high probability to uncover plausible faults are necessities in proving software quality. System models that represent only the operational behavioral of a system are incomplete sources for deriving test cases and performing safety analysis before the implementation process. Therefore, a system model that encompasses faults is required. This paper presents a technique that formalizes a safety model through the incorporation of faults with system specifications. The technique focuses on introducing semantic faults through the integration of fault trees with system specifications or statechart. The method uses a set of systematic transformation rules that tries to maintain the semantics of both fault trees and statechart representations during the transformation of fault trees into statechart notations.
Omar el Ariss, Dianxiang Xu, W. Eric Wong, Yuting Chen 0001, Yann-Hang Lee
COMPSAC4
2008 A Review Approach to Detecting Violations of Consistency between Specification and Program Structures
abstract
The application of specification-based program verification techniques (e.g., black-box testing, formal proof) faces strong challenges in practice when the gap between the structure of a specification and that of its program is large. This paper describes a view-based program review approach to addressing these challenges. The essential idea of the approach is first to derive comparable views from the specification and program, and then detect and eliminate the violations of structural consistency in the program views on the basis of a set of criteria. We also developed a prototype tool to support the review approach, and conducted a case study to assess the effectiveness of the approach.
Yuting Chen 0001, Shaoying Liu, W. Eric Wong
Int. J. Softw. Eng. Knowl. Eng.1
2008 A relation-based method combining functional and structural testing for test case generation
Shaoying Liu, Yuting Chen 0001
J. Syst. Softw.2
2006 A Tool-Supported Review Approach to Detecting Structural Consistency Violations
Yuting Chen 0001, Shaoying Liu, Fumiko Nagoya
ICECCS1
2005 A Tool and Case Study for Specification-Based Program Review
abstract
Effective tool support is crucial for successfully applying software review techniques in practice. In this paper, we describe the design and implementation of a software tool to support an approach to reviewing programs on the basis of their formal specifications. The approach was initially proposed in our previous publication to improve the rigor, repeatability, and effectiveness of existing code review methods. We also present a case study in which we reviewed an ATM system to assess the performance of the review approach when used with the software tool. The results of the case study show that the approach is effective in detecting errors in programs and the tool is helpful in enhancing the efficiency of the review process.
Fumiko Nagoya, Shaoying Liu, Yuting Chen 0001
COMPSAC (1)3
2005 A Framework for SOFL-Based Program Review
abstract
Program review is a practical and cost-effective method for detecting errors in program code. This paper describes our recent work aiming to provide support for revealing errors which usually arise from inappropriate implementations of desired specifications. In our approach, the SOFL specification language is employed for specifying software systems. We provide a framework that guides reviewers to compare a code with its specification for effective detection of potential defects.
Yuting Chen 0001, Shaoying Liu, Fumiko Nagoya
ICECCS1
2005 Design of a Tool for Specification-Based Program Review
abstract
Program review is an effective means for enhancing software quality. In this paper we describe the design of a software tool to support our proposed "function-path" approach to reviewing programs based on SOFL specifications. The approach includes four steps: (1) deriving all the functional scenarios from a formal specification, (2) generating all the necessary program paths in a program, (3) establishing the mapping between the functional scenarios in the specification and the program paths as implemented functions in the program, and (4) reviewing all the program paths against their functional scenarios in the specification.
Fumiko Nagoya, Shaoying Liu, Yuting Chen 0001
ICECCS3
2005 An Automated Approach to Specification-Based Program Inspection
Shaoying Liu, Fumiko Nagoya, Yuting Chen 0001, Masashi Goya, John A. McDermid
ICFEM3
2004 An Approach to Detecting Domain Errors Using Formal Specification-Based Testing
abstract
Domain testing, a technique for testing software or portions of software dominated by numerical processing, is intended to detect domain errors that usually arise from incorrect implementations of desired domains. This paper describes our recent work aiming to provide support for revealing domain errors using formal specifications. In our approach, formal specifications serve as a means for domain modeling. We describe a strong domain testing strategy that guide testers to select a set of test points so that the potential domain errors can be effectively detected, and apply our approach in two case studies for test cases generation.
Yuting Chen 0001, Shaoying Liu
APSEC1
2004 An Investigation of the Approach to Specification-Based Program Review through Case Studies
abstract
Software review is an effective means to enhance the quality of software systems. However, traditional review methods emphasize the importance of the way to organize reviews and rely on the quality of the reviewers' experience and personal skills. In this paper we propose a new approach to rigorously reviewing programs based on their formal specifications. The fundamental idea of the approach is to use a formal specification as a standard to check whether all the required functions and properties in the specification are correctly implemented by its program. To help investigate the effectiveness and the weakness of the approach, we conduct two case studies of reviewing two program systems that implement the same formal specification of "A Research Management Policy" using different strategies, and present the evaluation of the case studies. The results show that the review approach is effective in detecting faults when the reviewer is different from the programmer, but less effective when the reviewer is the same as the programmer.
Fumiko Nagoya, Shaoying Liu, Yuting Chen 0001
ICECCS3
2004 An Approach to Integration Testing Based on Data Flow Specifications
Yuting Chen 0001, Shaoying Liu, Fumiko Nagoya
ICTAC1