Chengcheng Wan 0001

dblp:183/4488 · DBLP profile ↗
← Back
34ranked-venue papers
10as first author
26since 2021 · last 2026
0000-0001-9162-9688ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 23 · 6 first-author · 18 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 1 since 2021Systems, architecture and hardware · 4 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 3 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 JOSer: Just-In-Time Object Serialization for Heavy Java Serialization Workloads
abstract
Object serialization is critical in Java, which preserves objects in memory and transfers them among software systems if needed. However, the serialization techniques of modern Java systems are usually inflexible and inefficient under heavy serialization workloads, as they rely on manually-defined object schemas or omni-functional serializers. To tackle the above problem, we reveal a novel, serialization-specific optimization opportunity in Java. Based on it, we develop JOSer (Just-in-time Object SERializer), an efficient, Just-in-Time (JIT) object serialization technique. At runtime, JOSer generates a set of class-specific, JIT-friendly object serializers (i.e., serialization code), and then continuously optimizes them with the JIT compiler of Java Virtual Machine (JVM). JOSer also shares the metadata of objects under serialization and the serializers under optimization. We evaluate JOSer against six Java serialization techniques including OpenJDK's built-in serialization technique. Overall, JOSer improves the throughput by up to 20~83× in serialization and 43~229× in deserialization. JOSer has been successfully deployed in real-world products, reducing serialization CPU usage of Flink by 35.32~41.14% and latency of search recommendations by 30+ ms. JOSer grounds Apache Fory#8482;, an open-source serialization framework available at https://fory.apache.org/.
Chaokun Yang, Pengbo Nie, Qianwei Yu, Chengcheng Wan 0001, He Jiang 0001, Yuting Chen 0001
ASPLOS (2)6
2026 CRONUS: Counterexample-Guided Constraint Learning for Network Update Synthesis
Jianshuo Xu, Hongtai Zhu, Jincheng Ding, Runxuan Fang, Yechuan Xia, Haiqin Wu, Chengcheng Wan 0001, Geguang Pu
INFOCOM7
2026 Understanding the Effectiveness of Mutators in Mutation-Based Protocol Fuzzing
Jiayi Jiang, Yiutak Choi, Ting Su 0001, Haiying Sun, Chengcheng Wan 0001, Geguang Pu
SANER6
2026 PADD: Prefix-based Attention Divergence Detector for LLM Jailbreaks
abstract
Large Language Models (LLMs) have achieved remarkable progress in understanding and generation tasks, yet they remain highly susceptible to adversarial prompt attacks that bypass safety safeguards and induce the generation of harmful content. Existing pre-generation and post-generation defense methods typically rely on intent recognition, surface-level features, or external classifiers, rendering them vulnerable to evasion via subtle prompt perturbations while incurring substantial computational overhead.
Ziqun Bao, Jiaqiang Niu, Yuchen Shao, Chengcheng Wan 0001
WWW4
2026 Orchestrating optimization passes of machine learning compiler for reducing memory footprints of computation graphs
Qianwei Yu, Pengbo Nie, Chengcheng Wan 0001, He Jiang 0001, Jianjun Zhao 0001, Lei Qiao 0002, Yuting Chen 0001
J. Syst. Archit.4
2026 Keeper: Automated Testing and Fixing of Machine Learning Software - RCR Report
abstract
This artifact aims to provide source code, benchmark suite, results, and materials used in our study “Keeper: Automated Testing and Fixing of Machine Learning Software” [ 3 ]. We developed an automated testing and fixing tool Keeper and its IDE plugin for ML software. It automatically detects software defects and attempts to change how ML APIs are used to alleviate software misbehavior. This artifact provides guidelines to set up and execute Keeper and also guidelines to interpret our evaluation results. We hope this artifact can motivate and help future research to further tackle ML API misuses. All related data are available online.
Chengcheng Wan 0001, Shicheng Liu, Sophie Xie, Yuhan Liu 0004, Michael Maire, Henry Hoffmann, Shan Lu 0001
ACM Trans. Softw. Eng. Methodol.1
2025 Are LLMs Correctly Integrated into Software Systems?
abstract
Large language models (LLMs) provide effective solutions in various application scenarios, with the support of retrieval-augmented generation (RAG). However, developers face challenges in integrating LLM and RAG into software systems, due to lacking interface specifications, various requirements from software context, and complicated system management. In this paper, we have conducted a comprehensive study of 100 open-source applications that incorporate LLMs with RAG support, and identified 18 defect patterns. Our study reveals that 77 % of these applications contain more than three types of integration defects that degrade software functionality, efficiency, and security. Guided by our study, we propose systematic guidelines for resolving these defects in software life cycle. We also construct an open-source defect library HYDRANGEA [1].
Yuchen Shao, Yuheng Huang 0004, Lei Ma 0003, Ting Su 0001, Chengcheng Wan 0001
ICSE6
2025 Between Lines of Code: Unraveling the Distinct Patterns of Machine and Human Programmers
abstract
Large language models have catalyzed an unprece-dented wave in code generation. While achieving significant advances, they blur the distinctions between machine- and human-authored source code, causing integrity and authenticity issues of software artifacts. Previous methods such as DetectGPthave proven effective in discerning machine-generated texts, but they do not identify and harness the unique patterns of machine-generated code. Thus, its applicability falters when applied to code. In this paper, we carefully study the specific patterns that characterize machine- and human-authored code. Through a rigorous analysis of code attributes such as lexical diversity, conciseness, and naturalness, we expose unique patterns inherent to each source. We particularly notice that the syntactic segmentation of code is a critical factor in identifying its provenance. Based on our findings, we propose DetectCodeGPT, a novel method for detecting machine-generated code, which improves DetectGPT by capturing the distinct stylized patterns of code. Diverging from conventional techniques that depend on external LLMs for perturbations, DetectCodeGPT perturbs the code corpus by strategically inserting spaces and newlines, ensuring both efficacy and efficiency. Experiment results show that our approach significantly outperforms state-of-the-art techniques in detecting machine-generated code.1.
Yuling Shi, Hongyu Zhang 0002, Chengcheng Wan 0001, Xiaodong Gu 0002
ICSE3
2025 On the Effectiveness of Large Language Models in Domain-Specific Code Generation
abstract
Large language models (LLMs) such as ChatGPT have shown remarkable capabilities in code generation. Despite significant achievements, they rely on enormous training data to acquire a broad spectrum of open-domain knowledge. Besides, their evaluation revolves around open-domain benchmarks like HumanEval, which primarily consist of programming contests. Therefore, it is hard to fully characterize the intricacies and challenges associated with particular domains (e.g., Web, game, and math). In this article, we conduct an in-depth study of the LLMs in domain-specific code generation. Our results demonstrate that LLMs exhibit sub-optimal performance in generating domain-specific code, due to their limited proficiency in utilizing domain-specific libraries. We further observe that incorporating API knowledge as prompts can empower LLMs to generate more professional code. Based on these findings, we further investigate how to effectively incorporate API knowledge into the code generation process. We experiment with three strategies for incorporating domain knowledge, namely, external knowledge inquirer, chain-of-thought prompting, and chain-of-thought fine-tuning. We refer to these strategies as a new code generation approach called DomCoder . Experimental results show that all strategies of DomCoder improve the effectiveness of domain-specific code generation under certain settings.
Xiaodong Gu 0002, Yalan Lin, Hongyu Zhang 0002, Chengcheng Wan 0001, Zhao Wei, Yong Xu 0010, Juhong Wang
ACM Trans. Softw. Eng. Methodol.6
2024 BinPRE: Enhancing Field Inference in Binary Analysis Based Protocol Reverse Engineering
abstract
Protocol reverse engineering (PRE) aims to infer the specification of network protocols when the source code is not available. Specifically, field inference is one crucial step in PRE to infer the field formats and semantics. To perform field inference, binary analysis based PRE techniques are one major approach category. However, such techniques face two key challenges --- (1) the format inference is fragile when the logics of processing input messages may vary among different protocol implementations, and (2) the semantic inference is limited by inadequate and inaccurate inference rules.
Jiayi Jiang, Chengcheng Wan 0001, Haoyi Chen, Haiying Sun, Ting Su 0001
CCS3
2024 CFP: A Reinforcement Learning Framework for Comprehensive Fairness-Performance Trade-Off in Machine Learning
Simiao Zhang, Jitao Bai, Menghong Guan, Yueling Zhang, Jun Sun 0001, Yihao Huang 0001, Jiaping Wang, Chengcheng Wan 0001, Ting Su 0001, Geguang Pu
ICANN (1)8
2024 OPASS: Orchestrating TVM's Passes for Lowering Memory Footprints of Computation Graphs
abstract
Deep learning (DL) compilers, such as TVM and TensorFlow, encompass a variety of passes for optimizing computation graphs (i.e., DL models). Despite the efforts on developing optimization passes, it remains a challenge in arranging these passes - most compilers employ fixed pass sequences that do not fit with computation graphs of diverse structures; on the other hand, optimization passes have cascade effects, making the structures of graphs under compilation volatile and as well making it difficult to generate optimal sequences for graphs. Inspired by recent progresses on static computing memory footprints (i.e., memory usages) of computation graphs, we introduce in this paper OPASS, a novel approach to orchestrating TVM's optimization passes for lowering memory footprints of computation graphs, and finally allowing the graphs to run on memory-constrained devices. The key idea is, given a computation graph$G$, to optimize the graph heuristically and iteratively: OPASS learns the effects of passes on the graph; it then optimizes$G$iteratively - each iteration picks up a pass by the reduction of the memory footprint of$G$and as well the implicit effects of the pass for further optimizations, letting the pass be applied. We evaluate OPASS on Rebench (a suite of computation graphs) and two real-world models (Transformer and ResNet). The results clearly show the strength of OPASS: it outperforms TVM's default sequence by$1.77\times$in reducing graphs' memory footprints, with affordable costs; it also offers extra memory reductions of$5\sim 12\%$by catching the implicit effects of passes. Furthermore, OPASS helps analyze positive/negative effects of passes to graphs' memory footprints, providing TVM developers with best practices for designing optimization pass sequences.
Pengbo Nie, Chengcheng Wan 0001, He Jiang 0001, Jianjun Zhao 0001, Yuting Chen 0001
ICSME3
2024 FIPSER: Improving Fairness Testing of DNN by Seed Prioritization
abstract
As a rapidly evolving AI technology, deep neural networks are becoming increasingly integrated into human society, yet raising concerns about fairness issues. Previous studies have proposed a metric called causal fairness to measure the fairness of machine learning models and proposed some search algorithms to mine individual discrimination instance pairs (IDIPs). Fairness issues can be alleviated by retraining models with corrected IDIPs. However, the number of samples that are used as seeds for these methods is often limited due to the pursuit of efficiency. In addition, the quantity of IDIPs generated on different seeds varies, so it makes sense to select appropriate samples as seeds, which has not been sufficiently considered in past studies. In this paper, we study the imbalance in IDIP quantities for various datasets and sensitive attributes, highlighting the need for selecting and ranking seed samples. Then, we proposed FIPSER, a feature importance and perturbation potential-based seed prioritization method. Our experimental results show that, on average, when applied to the current state-of-the-art method of IDIP mining, FIPSER can improve its effectiveness by 45% and efficiency by 11%.
Yueling Zhang, Min Zhang 0007, Chengcheng Wan 0001, Ting Su 0001, Geguang Pu
ASE5
2024 ChameleonAPI: Automatic and Efficient Customization of Neural Networks for ML Applications
Yuhan Liu 0004, Chengcheng Wan 0001, Kuntai Du, Henry Hoffmann, Junchen Jiang, Shan Lu 0001, Michael Maire
OSDI2
2024 Keeper: Automated Testing and Fixing of Machine Learning Software
abstract
The increasing number of software applications incorporating machine learning (ML) solutions has led to the need for testing techniques. However, testing ML software requires tremendous human effort to design realistic and relevant test inputs and to judge software output correctness according to human common sense. Even when misbehavior is exposed, it is often unclear whether the defect is inside ML API or the surrounding code and how to fix the implementation. This article tackles these challenges by proposing Keeper, an automated testing and fixing tool for ML software. The core idea of Keeper is designing pseudo-inverse functions that semantically reverse the corresponding ML task in an empirical way and proxy common human judgment of real-world data. It incorporates these functions into a symbolic execution engine to generate tests. Keeper also detects code smells that degrade software performance. Once misbehavior is exposed, Keeper attempts to change how ML APIs are used to alleviate the misbehavior. Our evaluation on a variety of applications shows that Keeper greatly improves branch coverage, while identifying 74 previously unknown failures and 19 code smells from 56 out of 104 applications. Our user studies show that 78% of end-users and 95% of developers agree with Keeper’s detection and fixing results.
Chengcheng Wan 0001, Shicheng Liu, Sophie Xie, Yuhan Liu 0004, Henry Hoffmann, Michael Maire, Shan Lu 0001
ACM Trans. Softw. Eng. Methodol.1
2024 VarGAN: Adversarial Learning of Variable Semantic Representations
abstract
Variable names are of critical importance in code representation learning. However, due to diverse naming conventions, variables often receive arbitrary names, leading to long-tail, out-of-vocabulary (OOV), and other well-known problems. While the Byte-Pair Encoding (BPE) tokenizer has addressed the surface-level recognition of low-frequency tokens, it has not noticed the inadequate training of low-frequency identifiers by code representation models, resulting in an imbalanced distribution of rare and common identifiers. Consequently, code representation models struggle to effectively capture the semantics of low-frequency variable names. In this paper, we propose VarGAN, a novel method for variable name representations. VarGAN strengthens the training of low-frequency variables through adversarial training. Specifically, we regard the code representation model as a generator responsible for producing vectors from source code. Additionally, we employ a discriminator that detects whether the code input to the generator contains low-frequency variables. This adversarial setup regularizes the distribution of rare variables, making them overlap with their corresponding high-frequency counterparts in the vector space. Experimental results demonstrate that VarGAN empowers CodeBERT to generate code vectors that exhibit more uniform distribution for both low- and high-frequency identifiers. There is an improvement of 8% in similarity and relatedness scores compared to VarCLR in the IdBench benchmark. VarGAN is also validated in downstream tasks, where it exhibits enhanced capabilities in capturing token- and code-level semantics.
Yalan Lin, Chengcheng Wan 0001, Shuwen Bai, Xiaodong Gu 0002
IEEE Trans. Software Eng.2
2023 Stitcher: Learned Workload Synthesis from Historical Performance Footprints
Chengcheng Wan 0001, Joyce Cahoon, Wenjing Wang 0005, Katherine Lin, Sean Liu, Raymond Truong, Alexandra M. Ciortea, Konstantinos Karanasos, Subru Krishnan
EDBT1
2023 HotGPT: How to Make Software Documentation More Useful with a Large Language Model?
abstract
It is well known that valuable information is contained in the natural language components of software systems, like comments and manual, and such information can be used to improve system performance and reliability. Past research has attempted to extract such information through task-specific machine learning models and tool chains. Here, we investigate a general, one-model-fit-all solution through a state-of-the-art large language model (e.g., the GPT series). Our investigation covers three representative tasks: extracting locking rules from comments, synthesizing exception predicates from comments, and identifying performance-related configurations; it reveals challenges and opportunities in applying large language models to system maintenance tasks.
Yiming Su, Chengcheng Wan 0001, Utsav Sethi, Shan Lu 0001, Madan Musuvathi, Suman Nath
HotOS2
2023 GenCoG: A DSL-Based Approach to Generating Computation Graphs for TVM Testing
abstract
TVM is a popular deep learning (DL) compiler. It is designed for compiling DL models, which are naturally computation graphs, and as well promoting the efficiency of DL computation. State-of-the-art methods, such as Muffin and NNSmith, allow developers to generate computation graphs for testing DL compilers. However, these techniques are inefficient — their generated computation graphs are either type-invalid or inexpressive, and hence not able to test the core functionalities of a DL compiler.
Pengbo Nie, Xinyuan Miao, Yuting Chen 0001, Chengcheng Wan 0001, Lei Bu, Jianjun Zhao 0001
ISSTA5
2023 Self-Supervised Query Reformulation for Code Search
abstract
Automatic query reformulation is a widely utilized technology for enriching user requirements and enhancing the outcomes of code search. It can be conceptualized as a machine translation task, wherein the objective is to rephrase a given query into a more comprehensive alternative. While showing promising results, training such a model typically requires a large parallel corpus of query pairs (i.e., the original query and a reformulated query) that are confidential and unpublished by online code search engines. This restricts its practicality in software development processes. In this paper, we propose SSQR, a self-supervised query reformulation method that does not rely on any parallel query corpus. Inspired by pre-trained models, SSQR treats query reformulation as a masked language modeling task conducted on an extensive unannotated corpus of queries. SSQR extends T5 (a sequence-to-sequence model based on Transformer) with a new pre-training objective named corrupted query completion (CQC), which randomly masks words within a complete query and trains T5 to predict the masked content. Subsequently, for a given query to be reformulated, SSQR identifies potential locations for expansion and leverages the pre-trained T5 model to generate appropriate content to fill these gaps. The selection of expansions is then based on the information gain associated with each candidate. Evaluation results demonstrate that SSQR outperforms unsupervised baselines significantly and achieves competitive performance compared to supervised methods.
Yuetian Mao, Chengcheng Wan 0001, Yuze Jiang, Xiaodong Gu 0002
ESEC/SIGSOFT FSE2
2023 Run-Time Prevention of Software Integration Failures of Machine Learning APIs
abstract
Due to the under-specified interfaces, developers face challenges in correctly integrating machine learning (ML) APIs in software. Even when the ML API and the software are well designed on their own, the resulting application misbehaves when the API output is incompatible with the software. It is desirable to have an adapter that converts ML API output at runtime to better fit the software need and prevent integration failures. In this paper, we conduct an empirical study to understand ML API integration problems in real-world applications. Guided by this study, we present SmartGear, a tool that automatically detects and converts mismatching or incorrect ML API output at run time, serving as a middle layer between ML API and software. Our evaluation on a variety of open-source applications shows that SmartGear detects 70% incompatible API outputs and prevents 67% potential integration failures, outperforming alternative solutions.
Chengcheng Wan 0001, Yuhan Liu 0004, Kuntai Du, Henry Hoffmann, Junchen Jiang, Michael Maire, Shan Lu 0001
Proc. ACM Program. Lang.1
2023 Coverage-directed Differential Testing of X.509 Certificate Validation in SSL/TLS Implementations
abstract
Secure Sockets Layer (SSL) and Transport Security (TLS) are two secure protocols for creating secure connections over the Internet. X.509 certificate validation is important for security and needs to be performed before an SSL/TLS connection is established. Some advanced testing techniques, such as frankencert , have revealed, through randomly mutating Internet accessible certificates, that there exist unexpected, sometimes critical, validation differences among different SSL/TLS implementations. Despite these efforts, X.509 certificate validation still needs to be thoroughly tested as this work shows. This article tackles this challenge by proposing transcert , a coverage-directed technique to much more effectively test real-world certificate validation code. Our core insight is to (1) leverage easily accessible Internet certificates as seed certificates and (2) use code coverage to direct certificate mutation toward generating a set of diverse certificates. The generated certificates are then used to reveal discrepancies, thus potential flaws, among different certificate validation implementations. We implement transcert and evaluate it against frankencert , NEZHA , and RFCcert (three advanced fuzzing techniques) on five widely used SSL/TLS implementations. The evaluation results clearly show the strengths of transcert : During 10,000 iterations, transcert reveals 71 unique validation differences, 12×, 1.4×, and 7× as many as those revealed by frankencert , NEZHA , and RFCcert , respectively; it also supplements RFCcert in conformance testing of the SSL/TLS implementations against 120 validation rules, 85 of which are exclusively covered by transcert -generated certificates. We identify 17 root causes of validation differences, all of which have been confirmed and 11 have never been reported previously. The transcert -generated X.509 certificates also reveal that the primary goal of certificate chain validation is stated ambiguously in the widely adopted public key infrastructure standard RFC 5280.
Pengbo Nie, Chengcheng Wan 0001, Jiayu Zhu, Yuting Chen 0001, Zhendong Su 0001
ACM Trans. Softw. Eng. Methodol.2
2022 Hierarchical memory-constrained operator scheduling of neural architecture search networks
abstract
Neural Architecture Search (NAS) is widely used in industry, searching for neural networks meeting task requirements. Meanwhile, it faces a challenge in scheduling networks satisfying memory constraints. This paper proposes HMCOS that performs hierarchical memory-constrained operator scheduling of NAS networks: given a network, HMCOS constructs a hierarchical computation graph and employs an iterative scheduling algorithm to progressively reduce peak memory footprints. We evaluate HMCOS against RPO and Serenity (two popular scheduling techniques). The results show that HMCOS outperforms existing techniques in supporting more NAS networks, reducing 8.7~42.4% of peak memory footprints, and achieving 137--283x of speedups in scheduling.
Chengcheng Wan 0001, Yuting Chen 0001, He Jiang 0001, Lei Qiao 0002
DAC2
2022 Automated Testing of Software that Uses Machine Learning APIs
abstract
An increasing number of software applications incorporate machine learning (ML) solutions for cognitive tasks that statistically mimic human behaviors. To test such software, tremendous human effort is needed to design image/text/audio inputs that are relevant to the software, and to judge whether the software is processing these inputs as most human beings do. Even when misbehavior is exposed, it is often unclear whether the culprit is inside the cognitive ML API or the code using the API.
Chengcheng Wan 0001, Shicheng Liu, Sophie Xie, Henry Hoffmann, Michael Maire, Shan Lu 0001
ICSE1
2022 Doppler: Automated SKU Recommendation in Migrating SQL Workloads to the Cloud
abstract
Selecting the optimal cloud target to migrate SQL estates from on-premises to the cloud remains a challenge. Current solutions are not only time-consuming and error-prone, requiring significant user input, but also fail to provide appropriate recommendations. We present Doppler, a scalable recommendation engine that provides right-sized Azure SQL Platform-as-a-Service (PaaS) recommendations without requiring access to sensitive customer data and queries. Doppler introduces a novel price-performance methodology that allows customers to get a personalized rank of relevant cloud targets solely based on low-level resource statistics, such as latency and memory usage. Doppler supplements this rank with internal knowledge of Azure customer behavior to help guide new migration customers towards one optimal target. Experimental results over a 9-month period from prospective and existing customers indicate that Doppler can identify optimal targets and adapt to changes in customer workloads. It has also found cost-saving opportunities among over-provisioned cloud customers, without compromising on capacity or other requirements. Doppler has been integrated and released in the Azure Data Migration Assistant v5.5, which receives hundreds of assessment requests daily.
Joyce Cahoon, Wenjing Wang 0005, Katherine Lin, Sean Liu, Raymond Truong, Chengcheng Wan 0001, Alexandra M. Ciortea, Sreraman Narasimhan, Subru Krishnan
Proc. VLDB Endow.8
2021 Are Machine Learning Cloud APIs Used Correctly?
abstract
Machine learning (ML) cloud APIs enable developers to easily incorporate learning solutions into software systems. Unfortunately, ML APIs are challenging to use correctly and efficiently, given their unique semantics, data requirements, and accuracy-performance tradeoffs. Much prior work has studied how to develop ML APIs or ML cloud services, but not how open-source applications are using ML APIs. In this paper, we manually studied 360 representative open-source applications that use Google or AWS cloud-based ML APIs, and found 70% of these applications contain API misuses in their latest versions that degrade functional, performance, or economical quality of the software. We have generalized 8 anti-patterns based on our manual study and developed automated checkers that identify hundreds of more applications that contain ML API misuses.
Chengcheng Wan 0001, Shicheng Liu, Henry Hoffmann, Michael Maire, Shan Lu 0001
ICSE1
2020 Orthogonalized SGD and Nested Architectures for Anytime Neural Networks
abstract
We propose a novel variant of SGD customized for training network architectures that support anytime behavior: such networks produce a series of increasingly accurate outputs over time. Efficient architectural designs for these networks focus on re-using internal state; subnetworks must produce representations relevant for both imme- diate prediction as well as refinement by subse- quent network stages. We consider traditional branched networks as well as a new class of re- cursively nested networks. Our new optimizer, Orthogonalized SGD, dynamically re-balances task-specific gradients when training a multitask network. In the context of anytime architectures, this optimizer projects gradients from later out- puts onto a parameter subspace that does not in- terfere with those from earlier outputs. Experi- ments demonstrate that training with Orthogonal- ized SGD significantly improves generalization accuracy of anytime networks.
Chengcheng Wan 0001, Henry Hoffmann, Shan Lu 0001, Michael Maire
ICML1
2020 Guided, Deep Testing of X.509 Certificate Validation via Coverage Transfer Graphs
abstract
SSL and TLS are two secure protocols for creating secure connections over the Internet. X.509 certificate validation is important for security and needs to be performed before an SSL/TLS connection is established. However, state-of-the-art testing techniques, such as frankencert and mucert, have revealed, through randomly mutating Internet accessible certificates, that there exist unexpected, sometimes critical, validation differences among different SSL/TLS implementations. Despite these strong efforts, certificate validation is still not thoroughly tested and more effective techniques are needed as this work shows. To this end, this paper introduces transcert, a novel approach for effectively guiding fuzzing to perform deep testing of X.509 certificate validation. The goal of transcert is to generate certificates that trigger diverse executions; it achieves this goal by introducing the concept of a coverage transfer graph to efficiently, precisely abstract program executions. In particular, it records the execution of how a given certificate is validated by a reference SSL/TLS implementation. It then constructs a coverage transfer graph to model the coverage transfer from a test certificate (seed) to its mutated certificates (mutants), and explores the coverage transfer graph by iteratively sampling and mutating certificates. We have implemented transcert and evaluated it against frankencert and mucert on four state-of-the-art SSL/TLS implementations. The evaluation results clearly show the strengths of transcert- during 10,000 iterations, transcert has revealed 3,469 validation differences, 8× as many as those revealed by frankencert and mucert. We have identified 11 root causes of validation differences, all of which have been confirmed and five have never been reported previously. We also found that the primary goal of certificate chain validation is stated ambiguously in the widely-adopted PKI standard RFC 5280.
Jiayu Zhu, Chengcheng Wan 0001, Pengbo Nie, Yuting Chen 0001, Zhendong Su 0001
ICSME2
2020 ALERT: Accurate Learning for Energy and Timeliness
Chengcheng Wan 0001, Muhammad Husni Santriaji, Eri Rogers, Henry Hoffmann, Michael Maire, Shan Lu 0001
USENIX ATC1
2019 View-centric performance optimization for database-backed web applications
abstract
Web developers face the stringent task of designing informative web pages while keeping the page-load time low. This task has become increasingly challenging as most web contents are now generated by processing ever-growing amount of user data stored in back-end databases. It is difficult for developers to understand the cost of generating every web-page element, not to mention explore and pick the web design with the best trade-off between performance and functionality. In this paper, we present Panorama, a view-centric and database-aware development environment for web developers. Using database-aware program analysis and novel IDE design, Panorama provides developers with intuitive information about the cost and the performance-enhancing opportunities behind every HTML element, as well as suggesting various global code refactorings that enable developers to easily explore a wide spectrum of performance and functionality trade-offs.
Cong Yan, Chengcheng Wan 0001, Shan Lu 0001, Alvin Cheung
ICSE3
2018 SMOPAT: Mining semantic mobility patterns from trajectories of private vehicles
Chengcheng Wan 0001, Yanmin Zhu 0006, Jiadi Yu, Yanyan Shen
Inf. Sci.1
2016 Multi-perspective change impact analysis using linked data of software engineering
abstract
Change impact analysis plays an important role in software maintenance and evolution. However existing researches mostly focus on one single artifact. Software development is usually accompanied by various types of software artifacts, such as requirement documents, software architectures, test cases, source code, etc., requiring a much more comprehensive change impact analysis. This paper presents a novel approach to multi-perspective change impact analysis that is able to address heterogeneous software artifacts. The essential idea of the novel approach is (1) to adopt semantic web to construct automatically ontology based software engineering linked data, which links requirements, classes, code, bug reports, commits, developers, test cases and others, (2) to build a weighted change impact matrix/graph using the dependency features extracted from linked data, and (3) to follow a change impact propagation algorithm to analyze the overall change impacts. We have conducted experiments on two open source projects (HtmlUnit and OpenRocket) to evaluate our approach. The experimental results show that our approach achieves better F-measure and stability than existing multi-perspective change impact analysis approaches.
Chengcheng Wan 0001, Zece Zhu, Yuting Chen 0001
Internetware1
2016 Evaluating quality-in-use of FLOSS through analyzing user reviews
abstract
Quality-in-use (QU) is an important measure for evaluating the quality of a software system from user respective. Several approaches have been proposed to evaluate QU of FLOSS. Meanwhile, they are usually less effective, as they usually assume that sufficient, precise usage statistics can be collected, which may be impractical for evaluating many real-world FLOSS systems. This paper presents QUIndicator, a novel, fine-grained approach to evaluating QU of FLOSS using user reviews. The key idea of QUIndicator is to, for a specific FLOSS system, (1) use a topic model to cluster its user reviews into different topics, transform topics into characteristics of a QU model and compute the weight of each characteristic, (2) take review aspect as the minimum analysis unit, and also apply sentiment analysis to analyze the sentiment strength of each review aspect, and (3) match review aspects with their corresponding characteristics in QU model and evaluate the QU of the system. Wilson interval is adopted to keep fairness by punishing FLOSS systems with insufficient reviews. We have evaluated QUIndicator on ten FLOSS genre datasets. The evaluation results show that when a FLOSS system has sufficient reviews, QUIndicator can achieve a p@3 value up to 30% higher than those of baselines, and also achieves over 75% Spearman Coefficient with the ground truth. When the system has insufficient reviews, QUIndicator can achieve a p@3 value up to 25% higher than those of baselines and over 55% Spearman Coefficient with the ground truth.
Zhenzheng Qian, Chengcheng Wan 0001, Yuting Chen 0001
SNPD2
2016 An empirical study on recovering requirement-to-code links
abstract
Requirements traceability provides support for critical software engineering activities such as change impact analysis and requirements validation. Unfortunately many organizations have ineffective traceability practices in place, largely because of poor communication and time pressure problems. Therefore researchers have proposed various approaches to automatically recover requirement-to-code links. Typically, these approaches are based on Information Retrieval techniques, and use various features such as synonyms, verb-object phrases, and structural information. Although many links are thus recovered, the effectiveness of individual features is not fully evaluated, and it is rather difficult to combine different features to produce better results. In this paper, we implement a tool, called R2C, that combines various features to recover requirement-to-code links. With the support of R2C, we conduct an empirical study to understand the effectiveness of these features in recovering requirement-to-code links. Our results show that verb-object phrase is the most effective feature in recovering such links. A preliminary case study indicates that our tuning combines different features to produce better results than IR-based technique using a single feature.
Chengcheng Wan 0001
SNPD2