EDBT 2026 Demo / reviewers in the wild / expert
Mengting Yuan 0001
dblp:06/10596-1
· DBLP profile ↗
32ranked-venue papers
1as first author
21since 2021 · last 2026
0000-0001-8758-8668ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 18 · 10 since 2021Systems, architecture and hardware · 11 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 4 · 1 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Pyls: Enabling Python Hardware Synthesis with Dynamic Polymorphism via LCRS EncodingabstractPython dominates AI development and is the most widely used dynamic programming language, but synthesizing its polymorphic functions into hardware remains challenging. Existing HLS solutions support only static subsets of Python, forcing CPU offload with costly communication overhead. We present Pyls, the first framework that synthesizes dynamically polymorphic Python into monolithic hardware via Left-Child Right-Sibling (LCRS) encoding. Key to our approach is representing all Python objects as LCRS trees, enabling uniform hardware handling of dynamic types. Pyls automatically converts objects to fixed-width formats, generates XLS IR designs, and implements a tree memory architecture for efficient runtime type resolution. On FPGA platforms, Pyls demonstrates speedups of 5.19× and 3.98× over two ASIC CPUs, 303.29× over a soft-core processor, and 282.66× over a heterogeneous SoC design. Bolei Tong, Yongyan Fang, Chaorui Wang, Qing'an Li, Jingling Xue, Mengting Yuan 0001 |
CGO | 6 |
| 2026 | CMakeSonar: A Static Approach to Detecting CMake Bugs with a Fine-Grained Type SystemabstractAs build systems and their scripts grow in size and complexity, detecting bugs in build configurations becomes increasingly challenging due to the rich functionality and weak typing of build scripting languages. This paper introduces CM ake S onar , the first static approach to precisely identifying semantic bugs in CMake scripts. CM ake S onar addresses this challenge by (1) designing a fine-grained type system that captures the runtime semantics of CMake values, and (2) performing a flow-sensitive analysis that detects inconsistent and ill-typed value usages by solving type constraints. Our approach identifies configuration and usage errors that can silently affect build correctness, portability, and deployment safety. In our evaluation, CM ake S onar identifies 155 bugs across 36 real-world CMake projects on GitHub, of which 23 have been accepted and fixed by developers. With a false positive rate of 4.32 % and a recall of 97.48 % , CM ake S onar demonstrates that precise static analysis can effectively uncover high-impact bugs in untyped build systems. Haotian Han, Zihang Zhong, Qing'an Li, Jingling Xue, Mengting Yuan 0001 |
Proc. ACM Program. Lang. | 5 |
| 2025 | Calibro: Compilation-Assisted Linking-Time Binary Code Outlining for Code Size Reduction in Android ApplicationsabstractRecent Android systems have employed pre-compilation technology to boost app launch speed and runtime performance. However, this generates large OAT files that over-consume scarce memory and storage resources in mobile devices. This paper conducts an evaluation of code redundancy in popular production android applications and observes that the code redundancy is up to 25%. To reduce the code size via redundancy elimination, this paper proposes Calibro, a compilation-assisted linking-time binary code outlining method. Calibro consists of two parts, the Compilation-Time code Outlining (CTO) and the Linking-Time Binary code Outlining (LTBO) with information collected at compilation- time. Additionally, a paralleled suffix tree method is proposed to reduce the building time overhead, and a hot function filtering method is proposed to effectively mitigate run-time performance degradation caused by code outlining. Experimental results show that compared to the baseline (the original AOSP version with all available code size optimization enabled), the proposed approach reduces code size in Android applications by more than 15.19% on average, with negligible runtime performance degradation and tolerable building time overhead. Hence the proposed code outlining approach is promising for production deployment. Hanming Sun, Wenhan Shang, Mengting Yuan 0001, Jingqin Fu, Jiang Ma, Chun Jason Xue, Qing'an Li |
CGO | 4 |
| 2025 | Accelerating graph substitutions in DNN optimization by heuristic algorithmsabstractAbstract Graph substitution is a key optimization technique used in deep learning frameworks. Traditional search-based methods are one way to address the problem of graph substitution. However, with the ongoing expansion of deep neural networks (DNNs), the exploration of their vast equivalent graph search space becomes increasingly time-consuming. In this paper, we propose two heuristic methods to accelerate the search process in graph substitution, offering a relatively novel direction compared to existing methods. The first method employs a Memory-Augmented heuristic to optimize computation graphs. To further enhance the efficiency of computation graph optimization, the second method uses the simulated annealing method. This method adds computation graphs with degraded performance into the candidate set with a certain probability. The experimental results show that without significant compromise of inference performance, these two methods can find graph substitutions delivering similar DNN computing performance compared to existing searching methods, while the overall searching time can be reduced from hours to seconds. The source code is available at https://github.com/hudevictor/MAS-SAS . Chun Hu, Yufan Huang, Mengting Yuan 0001, Qing'an Li |
Neural Process. Lett. | 5 |
| 2024 | Unsupervised and Supervised Co-learning for Comment-based Codebase Refining and its Application in Code SearchabstractBackground: Code pre-training and large language models are heavily dependent on data quality. These models require a vast, high-quality corpus matching text descriptions with codes to establish semantic correlations between natural and programming languages. Unlike NLP tasks, code comment heavily relies on specialized programming knowledge and is often limited in quantity and variety. Thus, most widely available open-source datasets are established with compromise and noise from platforms, such as StackOverflow, where code snippets are often incomplete. This may lead to significant errors when deploying the trained models in real-world applications. Aims: Comments as a substitute for queries are used to build code search datasets from GitHub. While comments describe code functionality and details, they often contain noise and differ from queries. Thus, our research focuses on improving the syntactic and semantic quality of code comments. Method: We propose a comment-based data refinement framework CoCoRF 1 via an unsupervised and supervised co-learning technique. It applies manually defined rules for syntax filtering and constructs a bootstrap query corpus via the WTFF algorithm for training the TVAE model for further semantic filtering. Results: Our study shows that CoCoRF achieves high efficiency with less computational resource, and outperforms comparison models in DeepCS code search task. Conclusions: Our findings indicate that the CoCoRF framework significantly improves the performance of code search tasks by enhancing the quality of code datasets. Gang Hu 0003, Xiaoqin Zeng, Wanlong Yu, Min Peng 0002, Mengting Yuan 0001, Liang Duan |
ESEM | 5 |
| 2024 | IVE: Accelerating Enumeration-Based Subgraph Matching via Exploring Isolated VerticesabstractThe performance of the enumeration-based sub-graph matching, which searches all isomorphic subgraphs in the data graph, is crucial to various applications. The upper bound of the complexity for the enumeration-based method is exponential to the number of query graph vertices, denoted as$n$. We propose a novel subgraph matching algorithm called the Isolated Vertices Exploration (IVE). The IVE leverages isolated vertices during the reordering and enumeration phases, thereby significantly accelerating the subgraph matching process. During the enumeration, the isolated vertices can be matched by using a quick bipartite graph matching algorithm. Consequently, the complexity of matching the remaining non-isolated vertices is exponential to the number of non-isolated vertices, denoted as$n^{\prime}$. For the reordering, we designed the Maximum Deleted Edges (MDE) to minimize$n^{\prime}$. MDE iteratively selects the query vertex with the maximum edges. According to the experimental results,$n^{\prime}$is less than$0.8n$for 99.8% of arbitrary graphs. Moreover, IVE outperforms the state-of-the-art algorithms in various scenarios with different sizes, sparsities and fields, achieving a performance speedup of up to 80.3x. Zite Jiang, Shuai Zhang 0040, Xingzhong Hou, Mengting Yuan 0001, Haihang You |
ICDE | 4 |
| 2024 | Transformer-based code search for software Q&A sitesabstractAbstract In software Q&A sites, there are many code‐solving examples of individual program problems, and these codes with explanatory natural language descriptions are easy to understand and reuse. Code search in software Q&A sites increases the productivity of developers. However, previous approaches to code search fail to capture structural code information and the interactivity between source codes and natural queries. In other words, most of them focus on specific code structures only. This paper proposes TCS (Transformer‐based code search), a novel neural network, to catch structural information for searching valid source codes from the query, which is vital for code search. The multi‐head attention mechanism in Transformer helps TCS learn enough information about the underlying semantic vector representation of codes and queries. An aligned attention matrix is also employed to catch relationships between codes and queries. Experimental results show that the proposed TCS can learn more structural information and has better performance than existing models. Yaohui Peng, Gang Hu 0003, Mengting Yuan 0001 |
J. Softw. Evol. Process. | 4 |
| 2024 | Fast Subgraph Matching by Dynamic Graph EditingabstractSubgraph matching is a challenging NP-complete problem that involves finding identical subgraphs of a query graph$q$in a larger data graph$G$. It has numerous applications in diverse fields, including social and biological networks. However, existing subgraph matching algorithms assume that the graph structure is fixed, which limits their performance in solving more difficult matching cases. To address this issue, we propose a novel approach called Dynamic Graph Editing (DGE), which dynamically edits the query graph to optimize the subgraph matching algorithm. Based on this approach, we introduce an efficient enumeration method called Dynamic Graph Editing Enumeration, which significantly improves the performance of the algorithm. Our experimental results show that DGE outperforms current state-of-the-art algorithms in terms of computational efficiency and ability to solve more complex subgraph matching cases. Zite Jiang, Shuai Zhang 0040, Boxiao Liu, Xingzhong Hou, Mengting Yuan 0001, Haihang You |
IEEE Trans. Serv. Comput. | 5 |
| 2024 | iTCRL: Causal-Intervention-Based Trace Contrastive Representation Learning for Microservice SystemsabstractNowadays, microservice architecture has become mainstream way of cloud applications delivery. Distributed tracing is crucial to preserve the observability of microservice systems. However, existing trace representation approaches only concentrate on operations, relationships and metrics related to service invocations. They ignore service events that denotes meaningful, singular point in time during the service's duration. In this paper, we propose iTCRL, a novel trace contrastive representation learning approach based on causal intervention. This approach first constructs a unified graph representation for each trace to describe the runtime status of service events in traces and the complex relationships between them. Then, Causal-intervention-based Trace Contrastive Learning is proposed, which learns trace representations from causal perspective based on the unified graph representations of traces. It uses causal intervention to generate contrastive views, heterogeneous graph neural network-based trace encoder to learn trace representations, and direct causal effect to guide the training of trace encoder. Experimental results on three datasets show that iTCRL outperforms all baselines in terms of trace classification, trace anomaly detection, trace sampling and noise robustness, and also validate the contribution of Causal-intervention-based Trace Contrastive Learning. Xiangbo Tian, Shi Ying 0001, Tiangang Li, Mengting Yuan 0001, Ruijin Wang, Yishi Zhao, Jianga Shang |
IEEE Trans. Software Eng. | 4 |
| 2023 | Statistical Type Inference for Incomplete ProgramsabstractWe propose a novel two-stage approach, Stir, for inferring types in incomplete programs that may be ill-formed, where whole-program syntactic analysis often fails. In the first stage, Stir predicts a type tag for each token by using neural networks, and consequently, infers all the simple types in the program. In the second stage, Stir refines the complex types for the tokens with predicted complex type tags. Unlike existing machine-learning-based approaches, which solve type inference as a classification problem, Stir reduces it to a sequence-to-graph parsing problem. According to our experimental results, Stir achieves an accuracy of 97.37 % for simple types. By representing complex types as directed graphs (type graphs), Stir achieves a type similarity score of 77.36 % and 59.61 % for complex types and zero-shot complex types, respectively. Yaohui Peng, Qiongling Yang, Hanwen Guo, Qing'an Li, Jingling Xue, Mengting Yuan 0001 |
ESEC/SIGSOFT FSE | 7 |
| 2023 | A hybrid-order local search algorithm for set k-cover problem in wireless sensor networks
Boxiao Liu, Mengting Yuan 0001, Haihang You |
Frontiers Comput. Sci. | 2 |
| 2023 | SLF: A passive parallelization of subgraph isomorphism
Wenle Liang, Wenyong Dong, Mengting Yuan 0001 |
Inf. Sci. | 3 |
| 2023 | A heterogeneous processing-in-memory approach to accelerate quantum chemistry simulation
Zeshi Liu, Wenqian Dong, Mengting Yuan 0001, Haihang You, Dong Li 0001 |
Parallel Comput. | 4 |
| 2023 | Effective Stack Wear Leveling for NVMabstractWith the rapid growth of data processed by computer systems, nonvolatile memory (NVM), represented by phase change memory (PCM), is regarded as a promising next-generation storage technology as it offers superior advantages over DRAM. However, PCM suffers from a severe write durability problem, leading to an extremely short lifespan under the uneven write patterns of real-world programs. We observe that loops are one of the primary causes of uneven writes on the stack. To alleviate this problem, we present Loop2Recursion, a compiler-assisted stack wear leveling technique that automatically transforms loops into recursive functions. In addition, we propose several optimizations to reduce the stack sizes and instruction counts of the generated recursive functions, two schemes to limit recursion depth, and selective loop transformation for cache-enabled architectures. Experimental results demonstrate that Loop2Recursion outperforms state-of-the-art methods by significantly improving stack wear leveling with a greatly reduced performance overhead. Jifeng Wu, Wei Li 0241, Mengting Yuan 0001, Chun Jason Xue, Jingling Xue, Qing'an Li |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2023 | DRONE: An Efficient Distributed Subgraph-Centric Framework for Processing Large-Scale Power-law GraphsabstractNowadays, the ever-increasing volume of graph-structured data such as social networks, graph databases and knowledge graphs requires to be processed efficiently and scalably. These natural graphs commonly found in the real world have highly skewed power-law degree distribution and are called power-law graphs. The subgraph-centric programming model is a promising approach applied in many state-of-the-art distributed graph computing frameworks. However, the performance of subgraph-centric frameworks is limited when processing large-scale power-law graphs. When deployed to the subgraph-centric framework, existing graph partitioning algorithms are not suitable for power-law graphs. In this paper, we present a novel distributed graph computing framework, DRONE (Distributed gRaph cOmputiNg Engine), which leverages the subgraph-centric model and the vertex-cut graph partitioning strategy. DRONE also supports the fault tolerance mechanism to accommodate the increasing scale of machines with negligible overhead (6.48% on average). We further study the execution workflow of DRONE and propose an efficient and balanced graph partition algorithm (EBV) for DRONE. Experiments show that DRONE reduces the running time on real-world graphs by 25.6%, on average, compared to the state-of-the-art distributed graph computing frameworks. In addition, the EBV graph partition algorithm reduces the replication factor by at least 21.8% than other self-based partition algorithms. Our results indicate that DRONE has excellent potential in processing large-scale power-law graphs. Shuai Zhang 0040, Zite Jiang, Xingzhong Hou, Mengting Yuan 0001, Haihang You |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2022 | Recovering Container Class Types in C++ BinariesabstractWe present TIARA, a novel approach to recovering container classes in c++ binaries. Given a variable address in a c++ binary, TIARA first applies a new type-relevant slicing algorithm incorporated with a decay function, TSLICE, to obtain an inter-procedural forward slice of instructions expressed as a CFG to summarize how the variable is used in the binary (as our primary contribution). TIARA then makes use of a GCN (Graph Convolutional Network) to learn and predict the container type for the variable (as our secondary contribution). According to our evaluation, TIARA can advance the state of the art in inferring commonly used container types in a set of eight large real-world COTS c++ binaries efficiently (in terms of the overall analysis time) and effectively (in terms of precision, recall and F1 score). Xuezheng Xu, Qing'an Li, Mengting Yuan 0001, Jingling Xue |
CGO | 4 |
| 2022 | GLite: a fast and efficient automatic graph-level optimizer for large-scale DNNsabstractWe propose a scalable graph-level optimizer named GLite to speed up search-based optimizations on large neural networks. GLite leverages a potential-based partitioning strategy to partition large computation graphs into small subgraphs without losing profitable substitution patterns. To avoid redundant subgraph matching, we propose a dynamic programming algorithm to reuse explored matching patterns. The experimental results show that GLite reduces the running time of search-based optimizations from hours to milliseconds, without compromising in inference performance. Jiaqi Li 0006, Min Peng 0002, Qing'an Li, Meizheng Peng, Mengting Yuan 0001 |
DAC | 5 |
| 2022 | An empirical study of the effectiveness of IR-based bug localization for large-scale industrial projects
Wei Li 0241, Qing'an Li, Yunlong Ming, Weijiao Dai, Shi Ying 0001, Mengting Yuan 0001 |
Empir. Softw. Eng. | 6 |
| 2022 | Fast and efficient parallel breadth-first search with power-law graph transformation
Zite Jiang, Shuai Zhang 0040, Mengting Yuan 0001, Haihang You |
Frontiers Comput. Sci. | 4 |
| 2021 | An Efficient and Balanced Graph Partition Algorithm for the Subgraph-Centric Programming Model on Large-scale Power-law GraphsabstractNowadays, the parallel processing of power-law graphs is one of the biggest challenges in the field of graph computation. The subgraph-centric programming model is a promising approach and has been applied in many state-of-the-art distributed graph computing frameworks. The graph partition algorithm plays an important role in the overall performance of subgraph-centric frameworks. However, traditional graph partition algorithms have significant difficulties in processing large-scale power-law graphs. The major problem is the communication bottleneck found in many subgraph-centric frameworks. Detailed analysis indicates that the communication bottleneck is caused by the huge communication volume or the extreme message imbalance among partitioned subgraphs. The traditional partition algorithms do not consider both factors at the same time, especially on power-law graphs. In this paper, we propose a novel efficient and balanced vertex-cut graph partition algorithm (EBV) which grants appropriate weights to the overall communication cost and communication balance. We observe that the number of replicated vertices and the balance of edge and vertex assignment have a great influence on communication patterns of distributed subgraph-centric frameworks, which further affect the overall performance. Based on this insight, We design an evaluation function that quantifies the proportion of replicated vertices and the balance of edges and vertices assignments as important parameters. Besides, we sort the order of edge processing by the sum of end-vertices' degrees from small to large. Experiments show that EBV reduces replication factor and communication by at least 21.8% and 23.7% respectively than other self-based partition algorithms. When deployed in the subgraph-centric framework, it reduces the running time on power-law graphs by an average of 16.8% compared with the state-of-the-art partition algorithm. Our results indicate that EBV has a great potential in improving the performance of subgraph-centric frameworks for the parallel large-scale power-law graph processing. Shuai Zhang 0040, Zite Jiang, Xingzhong Hou, Zhen Guan, Mengting Yuan 0001, Haihang You |
ICDCS | 5 |
| 2021 | A novel webpage layout aesthetic evaluation model for quantifying webpage layout design
Hongyan Wan, Wanting Ji, Guoqing Wu 0004, Xiaoyun Jia, Xue Zhan, Mengting Yuan 0001, Ruili Wang 0001 |
Inf. Sci. | 6 |
| 2020 | Loop2Recursion: Compiler-Assisted Wear Leveling for Non-Volatile MemoryabstractNon-Volatile Memory (NVM) technologies, such as Phase Change Memory (PCM), herald the next generation of main memory as they offer superior features compared with DRAM. Unfortunately, NVM's limited write endurance hinders its adoption as its lifetime can be extremely short under skew writes. This paper observes that the loops in programs are one of the primary causes of uneven writes as they introduce the hot data and cause a large number of stack frames to be allocated to the same locations. To alleviate this problem, we present Loop2Recursion, a compile-time wear leveling technique for transforming loops into recursions automatically. Our approach is flexible as it can avoid a substantial memory overhead by limiting the depth of recursion. Experimental results demonstrate that Loop2Recursion can significantly improve the wear leveling over stack area compared to the state-of-the-art methods, while incurring only negligible performance overhead. Wei Li 0241, Mengting Yuan 0001, Chun Jason Xue, Jingling Xue, Qing'an Li |
ICCD | 3 |
| 2020 | Neural joint attention code search over structure embeddings for software Q&A sites
Gang Hu 0003, Min Peng 0002, Yihan Zhang 0005, Qianqian Xie, Mengting Yuan 0001 |
J. Syst. Softw. | 5 |
| 2020 | Unsupervised software repositories mining and its application to code searchabstractSummary Software repositories are crucial resources for many software tasks, including code retrieval and annotation. Programming forums provide questions and answers (Q&A) from software developers, containing abundant code‐description posts for exchanging knowledge about programming issues. However, most posts provide personal opinions of users that are often not adequately confirmed or outdated. Mining software repositories in such open and unrestricted forums is challenging. Since the posts can be arbitrary and noisy, it is difficult to get unified labels for supervised noise elimination. Different from existing mining approaches, this paper proposes Code‐Description Mining Framework (CodeMF), an unsupervised framework to eliminate noisy posts and extract high quality software repositories from programming forums. CodeMF treats all social features of the posts as discrete‐time signals for kernel principal component analysis and further performs wavelet transform feature fusion to find the delicate changes (noises in temporal signals). We conduct comprehensive experiments on StackOverflow. Experimental results demonstrate that CodeMF can effectively reduce running time and improve precision via mining high‐quality software repositories for various programming languages, especially for the large‐scale codebases. To further illustrate the effect of CodeMF applied in software tasks, we introduce it to improve the performance of query‐expansion code search. Meanwhile, for SQL and C# programs, compared to the state‐of‐the‐art query‐expansion method QECK, the improvement of QECK CodeMF is 2% and 6% on Recall@10, and 4% and 14% on mean reciprocal rank, respectively. Gang Hu 0003, Min Peng 0002, Yihan Zhang 0005, Qianqian Xie, Wang Gao 0002, Mengting Yuan 0001 |
Softw. Pract. Exp. | 6 |
| 2019 | A Wear Leveling Aware Memory Allocator for Both Stack and Heap Management in PCM-based Main Memory SystemsabstractPhase change memory (PCM) has been considered as a replacement of DRAM, due to its potentials in high storage density and low leakage power. However, the limited write endurance presents critical challenges. Various wear leveling techniques have been proposed to mitigate this issue from different perspectives, including both hardware and software levels. This paper proposes a wear leveling aware memory allocator, which (1) always prefers allocating memory blocks with less writes upon memory requests, and (2) leaves blocks allocated more than a threshold value unallocable temporarily. Furthermore, for the first time, this allocator provides a uniform management scheme for both stack and heap areas, thus could better balance writes in stack and heap areas. Experimental evaluations show that, compared to state-of-the-art memory allocators (i.e., glibc malloc, NVMalloc and Walloc), the proposed memory allocator improves the PCM wear leveling, in terms of CoV (a wear leveling indicator) by 41.9%, 30.3%, and 35.8%, respectively. Wei Li 0241, Ziqi Shuai, Chun Jason Xue, Mengting Yuan 0001, Qing'an Li |
DATE | 4 |
| 2019 | Software Defect Prediction Based on Cost-Sensitive Dictionary LearningabstractSoftware defect prediction technology has been widely used in improving the quality of software system. Most real software defect datasets tend to have fewer defective modules than defective-free modules. Highly class-imbalanced data typically make accurate predictions difficult. The imbalanced nature of software defect datasets makes the prediction model classifying a defective module as a defective-free one easily. As there exists the similarity during the different software modules, one module can be represented by the sparse representation coefficients over the pre-defined dictionary which consists of historical software defect datasets. In this study, we make use of dictionary learning method to predict software defect. We optimize the classifier parameters and the dictionary atoms iteratively, to ensure that the extracted features (sparse representation) are optimal for the trained classifier. We prove the optimal condition of the elastic net which is used to solve the sparse coding coefficients and the regularity of the elastic net solution. Due to the reason that the misclassification of defective modules generally incurs much higher cost risk than the misclassification of defective-free ones, we take the different misclassification costs into account, increasing the punishment on misclassification defective modules in the procedure of dictionary learning, making the classification inclining to classify a module as a defective one. Thus, we propose a cost-sensitive software defect prediction method using dictionary learning (CSDL). Experimental results on the 10 class-imbalance datasets of NASA show that our method is more effective than several typical state-of-the-art defect prediction methods. Hongyan Wan, Guoqing Wu 0004, Mali Yu, Mengting Yuan 0001 |
Int. J. Softw. Eng. Knowl. Eng. | 4 |
| 2017 | Software Defect Prediction Using Dictionary LearningabstractWith the popularization of software version control system and defect tracking tools, large amounts of software development data is recorded.How to effectively use these data to improve the quality of software development, has become a hot topic in recent years.Software defect prediction technology can take full advantage of the historical data to build predictive models and automatically detect defective modules for efficient software test to improve the quality of a software system.But the class-imbalanced data makes the prediction model classifying a modules as a defective-free one easily, while the misclassification of defective modules generally incurs much higher cost risk than the misclassification of defective-free ones.To resolve this problem, we propose a cost-sensitive software defect prediction method using dictionary learning.It iteratively optimizes the classifier parameters and the dictionary atoms, to ensure that the extracted features (sparse representation) are optimal for the trained classifier; Moreover, we take the different misclassification costs into account, increasing the punishment on misclassification defective modules in the procedure of dictionary learning, making classification inclining to classify a module as a defective one.Experimental results on the 10 classimbalanced data sets of NASA show that our method is more effective than other methods. Hongyan Wan, Guoqing Wu 0004, Rui Wang 0036, Mengting Yuan 0001 |
SEKE | 6 |
| 2016 | Heterogeneous Defect Prediction via Exploiting Correlation SubspaceabstractSoftware defect prediction generally builds models from intra-project data.Lack of training data at the early stage of software testing limits the efficiency of prediction in practice.Thereby researchers proposed cross-project defect prediction using the data from other projects.Most previous efforts assumed the cross-project defect data have the same metrics set which means the metrics used and size of metrics set are same in the data of projects.However, in real scenarios, this assumption may not hold.In addition, software defect datasets have the class imbalance problem increasing the difficulty for the learner to predict defects.In this paper, we advance canonical correlation analysis for deriving a joint feature space for associating crossproject data and propose a novel support vector machine algorithm which incorporates the correlation transfer information into classifier design for cross-project prediction.Moreover, we take different misclassification costs into consideration to make the classification inclining to classify a module as a defective one, alleviating the impact of imbalanced data.Experiments on public heterogeneous datasets from different projects show that our method is more effective, compared to state-of-the-art methods. Guoqing Wu 0004, Min Jiang 0005, Hongyan Wan, Guoan You, Mengting Yuan 0001 |
SEKE | 6 |
| 2016 | High quality information extraction and query-oriented summarization for automatic query-reply in social network
Min Peng 0002, Binlong Gao, Mengting Yuan 0001, Fei Li 0007 |
Expert Syst. Appl. | 5 |
| 2016 | Exploiting Correlation Subspace to Predict Heterogeneous Cross-Project DefectsabstractCross-project defect prediction trains a prediction model using historical data from source projects and applies the model to target projects. Most previous efforts assumed the cross-project data have the same metrics set, which means the metrics used and the size of metrics set are the same. However, this assumption may not hold in practical scenarios. In addition, software defect datasets have the class-imbalance problem which increases the difficulty for the learner to predict defects. In this paper, we advance canonical correlation analysis by deriving a joint feature space for associating cross-project data. We also propose a novel support vector machine algorithm which incorporates the correlation transfer information into classifier design for cross-project prediction. Moreover, we take different misclassification costs into consideration to make the classification inclining to classify a module as a defective one, alleviating the impact of imbalanced data. The experimental results show that our method is more effective compared to state-of-the-art methods. Guoqing Wu 0004, Hongyan Wan, Guoan You, Mengting Yuan 0001, Min Jiang 0005 |
Int. J. Softw. Eng. Knowl. Eng. | 5 |
| 2014 | Using Genetic Algorithms to Repair JUnit Test CasesabstractJUnit test repair has been proposed as a way to alleviate the burden of maintaining the broken tests caused by evolving software. Existing techniques for JUnit test repair either focus on repairing failing assertions or fixing test case compilation errors rendered by evolving method declarations. The empirical work suggests that the synthesis of new method calls is often needed when repairing test cases in practice. In this work, we propose Test Fix, an approach to fix broken JUnit test cases by synthesizing new method calls. Test Fix reduces the synthesis to a search problem and uses a genetic algorithm to solve it. Evaluated on real world applications, preliminary experimental results show that Test Fix can repair broken tests with adding or deleting method calls. Evaluations on the performance of JUnit test repair technique using a genetic algorithm against a random search algorithm are also conducted. Experimental results indicate the clear superiority of genetic algorithms over random search algorithm. Guoqing Wu 0004, Mengting Yuan 0001 |
APSEC (1) | 4 |
| 2013 | Minimizing code size via page selection optimization on partitioned memory architecturesabstractFor 8-bit microcontrollers, bank-switching is commonly used to increase memory capacity. The disadvantage of this technique is that bank (page) selection instructions are introduced when switching active data (program) bank. The page selection problem is to minimize the number of page selection instructions inserted. While previous efforts work on optimizing bank selection instructions for the data segment, our work focuses on minimizing page selection instructions for the program segment. Minimizing page selection instructions is a more challenging problem as the size of each procedure being allocated is affected by the number of inserted page selection instructions. In this paper, we first give a formal definition of the page selection problem, and then we formulate the problem as an Integer Linear Programming (ILP) to find the optimal solution. We introduce a tabu search heuristic algorithm, TMSEARCH, to solve the problem efficiently. The experimental results show that ILP can find optimal solutions for small-scale problems, and TMSEARCH is able to find good solutions for all benchmarks within reasonable time. Com-pared to a commercial compiler, TMSEARCH reduces total code size between 0.04% and 19.3%, and reduces page selection instructions between 24.3% and 78.7%. Mengting Yuan 0001, Chun Jason Xue, Qing'an Li, Yingchao Zhao 0001 |
CASES | 1 |