VLDB 2026 Research / reviewers in the wild / expert
Songqiang Chen
dblp:257/0002
· DBLP profile ↗
21ranked-venue papers
3as first author
20since 2021 · last 2026
0000-0002-1220-8728ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 18 · 3 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Subgraph-Oriented Testing for Deep Learning LibrariesabstractDeep Learning (DL) libraries, such as PyTorch, are widely used for building and deploying DL models on various hardware platforms. Meanwhile, they are found to contain bugs that lead to incorrect calculation results and cause issues like non-convergence training and inaccurate prediction of DL models. Thus, many efforts have been made to test DL libraries and reveal bugs. However, existing DL library testing methods manifest limitations: model-level testing methods cause complexity in fault localization. Meanwhile, API-level testing methods often generate invalid inputs or primarily focus on extreme inputs that lead to crash failures; they also ignore testing realistic API interactions. These limitations may lead to missing detection of bugs, even in the frequently used APIs. To address these limitations, we propose SORT (Subgraph-Oriented Realistic Testing) to differential test DL libraries on different hardware platforms. SORT takes popular API interaction patterns, represented as frequent subgraphs of model computation graphs, as test subjects. In this way, it introduces realistic API interaction sequences while maintaining efficiency in locating faulty APIs for observed errors. Besides, SORT prepares test inputs by referring to extensive features of runtime inputs for each API in executing real-life benchmark data. The generated inputs are expected to better simulate such valid real inputs and reveal bugs that are more likely to happen in real-life usage. Evaluation on 728 frequent subgraphs of 49 popular PyTorch models demonstrates that SORT achieves a 100% valid input generation rate, detects more precision bugs than existing methods, and reveals interaction-related bugs missed by single-API testing. 18 precision bugs in PyTorch are identified and reported to PyTorch developers. Xiaoyuan Xie, Songqiang Chen, Jinfu Chen 0002 |
IEEE Trans. Software Eng. | 3 |
| 2025 | Testing Deep Learning Libraries with Semantic Equivalent API PatternsabstractTesting deep learning (DL) libraries has garnered significant research attention since bugs within DL libraries can lead to incorrect predictions of neural models and mislead downstream applications. While some current testing methods for DL libraries focus on detecting crashes or inconsistencies in outputs across different devices or libraries, they often overlook non-crash bugs that may arise in all devices or reside in APIs that lack easily identifiable functionally equivalent counterpart APIs in other DL libraries. To address this gap, researchers proposed performing testing of semantic equivalent API sequences within a DL library. Such API sequences are identified with manually crafted rules. However, existing rules are formed by an ad hoc review of the documentation, resulting in their failure to cover and test some valuable APIs. In this paper, we propose a method for systematically constructing semantic equivalent API rules based on several semantic equivalent patterns. These patterns are derived by considering API function characteristics that often influence output consistency. Based on these patterns, we instantiated 11 concrete rules that help validate four critical yet previously not well-tested aspects of DL libraries, i.e., consistency across distributed and centralized executions, equivalent implementations, input variations, and invertible operations. Testing results of TensorFlow and PyTorch using our newly formulated rules show that our rules cover an additional 167 APIs in PyTorch and 451 APIs in TensorFlow compared to baselines and reveal 47 previously unknown bugs. Songqiang Chen, Haoyu Peng, Xiaoyuan Xie |
APSEC | 2 |
| 2025 | CodeCleaner: Mitigating Data Contamination for LLM BenchmarkingabstractData contamination presents a critical barrier preventing widespread industrial adoption of advanced software engineering techniques that leverage large language models (LLMs).This phenomenon occurs when evaluation data inadvertently overlaps with the public code repositories used to train LLMs, severely undermining the credibility of performance evaluations.Code refactoring, which comprises code restructuring and variable renaming, has emerged as a promising measure to mitigate data contamination.However, the lack of automated code refactoring tools and scientifically validated refactoring techniques has hampered widespread industrial implementation.To bridge the gap, this paper presents the first systematic study to examine the efficacy of code refactoring operators at multiple scales (method-level, class-level, and crossclass level) and in different programming languages.We develop CodeCleaner, including 11 operators for Python in multiple scales and 4 for Java.We elaborate on the rationale for why these operators could work to resolve data contamination and use both data-wise (e.g., N-gram matching overlap ratio) and model-wise metrics (e.g., perplexity) to quantify the efficacy after operators are applied.A drop of 75% overlap ratio is found when applying all operators in CodeCleaner, demonstrating their effectiveness in addressing data contamination.Besides, we migrate four operators to Java, showing their generalizability to another language.We also observed an average of 19% decrease in LLMs' performance after applying our operators.We make CodeCleaner online available at https://github.com/ArabelaTso/CodeCleaner-v1 to facilitate further studies on mitigating LLM data contamination. Jialun Cao, Songqiang Chen, Wuqi Zhang, Hau Ching Lo, Yeting Li, Shing-Chi Cheung |
Internetware | 2 |
| 2025 | LspFuzz: Hunting Bugs in Language ServersabstractThe Language Server Protocol (LSP) has revolutionized the integration of code intelligence in modern software development. There are approximately 300 LSP server implementations for various languages and 50 editors offering LSP integration. However, the reliability of LSP servers is a growing concern, as crashes can disable all code intelligence features and significantly impact productivity, while vulnerabilities can put developers at risk even when editing untrusted source code. Despite the widespread adoption of LSP, no existing techniques specifically target LSP server testing. To bridge this gap, we present LspFuzz, a grey-box hybrid fuzzer for systematic LSP server testing. Our key insight is that effective LSP server testing requires holistic mutation of source code and editor operations, as bugs often manifest from their combinations. To satisfy the sophisticated constraints of LSP and effectively explore the input space, we employ a two-stage mutation pipeline: syntax-aware mutations to source code, followed by context-aware dispatching of editor operations. We evaluated LspFuzz on four widely used LSP servers. LspFuzz demonstrated superior performance compared to baseline fuzzers, and uncovered previously unknown bugs in real-world LSP servers. Of the 51 bugs we reported, 42 have been confirmed, 26 have been fixed by developers, and two have been assigned CVE numbers. Our work advances the quality assurance of LSP servers, providing both a practical tool and foundational insights for future research in this domain. Hengcheng Zhu 0001, Songqiang Chen, Valerio Terragni, Lili Wei 0001, Yepang Liu 0001, Shing-Chi Cheung |
ASE | 2 |
| 2024 | FastLog: An End-to-End Method to Efficiently Generate and Insert Logging StatementsabstractLogs play a crucial role in modern software systems, serving as a means for developers to record essential information for future software maintenance. As the performance of these log-based maintenance tasks heavily relies on the quality of logging statements, various works have been proposed to assist developers in writing appropriate logging statements. However, these works either only support developers in partial sub-tasks of this whole activity; or perform with a relatively high time cost and may introduce unwanted modifications. To address their limitations, we propose FastLog, which can support the complete logging statement generation and insertion activity, in a very speedy manner. Specifically, given a program method, FastLog first predicts the insertion position in the finest token level, and then generates a complete logging statement to insert. We further use text splitting for long input texts to improve the accuracy of predicting where to insert logging statements. A comprehensive empirical analysis shows that our method outperforms the state-of-the-art approach in both efficiency and output quality, which reveals its great potential and practicality in current real-time intelligent development environments. Xiaoyuan Xie, Songqiang Chen, Jifeng Xuan |
ISSTA | 3 |
| 2024 | MR-Adopt: Automatic Deduction of Input Transformation Function for Metamorphic TestingabstractWhile a recent study reveals that many developer-written test cases can encode a reusable Metamorphic Relation (MR), over 70% of them directly hard-code the source input and follow-up input in the encoded relation. Such encoded MRs, which do not contain an explicit input transformation to transform the source inputs to corresponding follow-up inputs, cannot be reused with new source inputs to enhance test adequacy. Congying Xu, Songqiang Chen, Shing-Chi Cheung, Valerio Terragni, Hengcheng Zhu 0001, Jialun Cao |
ASE | 2 |
| 2024 | SURE: A Visualized Failure Indexing Approach Using Program Memory SpectrumabstractFailure indexing is a longstanding crux in software debugging, the goal of which is to automatically divide failures (e.g., failed test cases) into distinct groups according to the culprit root causes, as such multiple faults residing in a faulty program can be handled independently and simultaneously. The community of failure indexing has long been plagued by two challenges: (1) The effectiveness of division is still far from promising. Specifically, existing failure indexing techniques only employ a limited source of software runtime data, for example, code coverage, to be failure proximity and further divide them, which typically delivers unsatisfactory results. (2) The outcome can be hardly comprehensible. Specifically, a developer who receives the division result is just aware of how all failures are divided, without knowing why they should be divided the way they are. This leads to difficulties for developers to be convinced by the division result, which in turn affects the adoption of the results. To tackle these two problems, in this article, we propose SURE , a vi SU alized failu R e ind E xing approach using the program memory spectrum (PMS). We first collect the runtime memory information (i.e., variables’ names and values, as well as the depth of the stack frame) at several preset breakpoints during the execution of a failed test case, and transform the gathered memory information into a human-friendly image (called PMS). Then, any pair of PMS images that serve as proxies for two failures is fed to a trained Siamese convolutional neural network, to predict the likelihood of them being triggered by the same fault. Last, a clustering algorithm is adopted to divide all failures based on the mentioned likelihood. In the experiments, we use 30% of the simulated faults to train the neural network, and use 70% of the simulated faults as well as real-world faults to test. Results demonstrate the effectiveness of SURE: It achieves 101.20% and 41.38% improvements in faults number estimation, as well as 105.20% and 35.53% improvements in clustering, compared with the state-of-the-art technique in this field, in simulated and real-world environments, respectively. Moreover, we carry out a human study to quantitatively evaluate the comprehensibility of PMS, revealing that this novel type of representation can help developers better comprehend failure indexing results. Xihao Zhang, Xiaoyuan Xie, Songqiang Chen, Quanming Liu, Ruizhi Gao |
ACM Trans. Softw. Eng. Methodol. | 4 |
| 2024 | Word Closure-Based Metamorphic Testing for Machine TranslationabstractWith the wide application of machine translation, the testing of Machine Translation Systems (MTSs) has attracted much attention. Recent works apply Metamorphic Testing (MT) to address the oracle problem in MTS testing. Existing MT methods for MTS generally follow the workflow of input transformation and output relation comparison, which generates a follow-up input sentence by mutating the source input and compares the source and follow-up output translations to detect translation errors, respectively. These methods use various input transformations to generate the test case pairs and have successfully triggered numerous translation errors. However, they have limitations in performing fine-grained and rigorous output relation comparison and thus may report many false alarms and miss many true errors. In this article, we propose a word closure-based output comparison method to address the limitations of the existing MTS MT methods. We first propose word closure as a new comparison unit, where each closure includes a group of correlated input and output words in the test case pair. Word closures suggest the linkages between the appropriate fragment in the source output translation and its counterpart in the follow-up output for comparison. Next, we compare the semantics on the level of word closure to identify the translation errors. In this way, we perform a fine-grained and rigorous semantic comparison for the outputs and thus realize more effective violation identification. We evaluate our method with the test cases generated by five existing input transformations and the translation outputs from three popular MTSs. Results show that our method significantly outperforms the existing works in violation identification by improving the precision and recall and achieving an average increase of 29.9% in F1 score. It also helps to increase the F1 score of translation error localization by 35.9%. Xiaoyuan Xie, Songqiang Chen, Shing-Chi Cheung |
ACM Trans. Softw. Eng. Methodol. | 3 |
| 2024 | Metamorphic Testing of Image Captioning Systems via Image-Level ReductionabstractThe Image Captioning (IC) technique is widely used to describe images in natural language. However, even state-of-the-art IC systems can still produce incorrect captions and lead to misunderstandings. Recently, some IC system testing methods have been proposed. However, these methods still rely on pre-annotated information and hence cannot really alleviate the difficulty in identifying the test oracle. Furthermore, their methods artificially manipulate objects, which may generate unreal images as test cases and thus lead to less meaningful testing results. Thirdly, existing methods have various requirements on the eligibility of source test cases, and hence cannot fully utilize the given images to perform testing. To tackle these issues, in this paper, we proposeReICto perform metamorphic testing for the IC systems with some image-level reduction transformations like image cropping and stretching. Instead of relying on the pre-annotated information,ReICuses a localization method to align objects in the caption with corresponding objects in the image, and checks whether each object is correctly described or deleted in the caption after transformation. With the image-level reduction transformations,ReICdoes not artificially manipulate any objects and hence can avoid generating unreal follow-up images. Additionally, it eliminates the requirement on the eligibility of source test cases during the metamorphic transformation process, as well as decreases the ambiguity and boosts the diversity among the follow-up test cases, which consequently enables testing to be performed on any test image and reveals more distinct valid violations. We employReICto test five popular IC systems. The results demonstrate thatReICcan sufficiently leverage the provided test images to generate follow-up cases of good realism, and effectively detect a great number of distinct violations, without the need for any pre-annotated information. Xiaoyuan Xie, Xingpeng Li, Songqiang Chen |
IEEE Trans. Software Eng. | 3 |
| 2023 | Properly Offer Options to Improve the Practicality of Software Document Completion ToolsabstractWith the great progress in deep learning and natural language processing, many completion tools are proposed to help practitioners efficiently fill in various fields in software document. However, most of these tools offer their users only one option and this option generally requires much revision to meet a satisfactory quality, which hurts much practicality of the completion tools. By finding that the beam search model of such tools often generates a much better output at relatively high confidence and considering the interactive use of such tools, we advise such tools to offer multiple high-confidence model outputs for more chances of offering a good option. And we further suggest these tools offer dissimilar outputs to expand the chance of including a better output in a few options. To evaluate our whole idea, we design a clustering-based initial method to help these tools properly offer some dissimilar model outputs as options. We adopt this method to improve nine completion tools for three software document fields. Results show it can help all the nine tools offer an option that needs less revision from users and thus effectively improve the practicality of tools. Songqiang Chen, Xiaoyuan Xie |
ICPC | 2 |
| 2023 | qaAskeR+: a novel testing method for question answering software via asking recursive questions
Xiaoyuan Xie, Songqiang Chen |
Autom. Softw. Eng. | 3 |
| 2022 | Towards the Robustness of Multiple Object Tracking SystemsabstractDue to the wide use of visual perception techniques in safety-critical fields, existing studies have tested the robustness of the essential object detection systems in scenarios with different image content. However, the applications that perceive one video with multiple image frames, such as autonomous driving, usually further require the trajectories of objects. This is mainly realized by combining detecting objects and associating detected objects in frames using multiple object tracking (MOT) systems. Thus, it is also essential to test the robustness of MOT systems, particularly in their exclusive scenarios that involve variety beyond the static image content. In this paper, we propose a novel testing method with five new Metamorphic Relations to realize the robustness test for MOT systems in two typical categories of scenarios, i.e., the speed variety of tracked objects and temporary camera failures. Our method also properly addresses the oracle problem and the lack of test cases for some rare scenarios to make the test efficient and diverse. Finally, we use our method to test three typical MOT systems and effectively reveal numerous and diverse MOT errors. We also extensively discuss the performance of tested systems and summarize two typical scenes where they often misbehave. Xiaoyuan Xie, Ying Duan, Songqiang Chen, Jifeng Xuan |
ISSRE | 3 |
| 2022 | Boosting the Revealing of Detected Violations in Deep Learning Testing: A Diversity-Guided MethodabstractDue to the ability to bypass the oracle problem, Metamorphic Testing (MT) has been a popular technique to test deep learning (DL) software. However, no work has taken notice of the prioritization for Metamorphic test case Pairs (MPs), which is quite essential and beneficial to the effectiveness of MT in DL testing. When the fault-sensitive MPs apt to trigger violations and expose defects are not prioritized, the revealing of some detected violations can be greatly delayed or even missed to conceal critical defects. In this paper, we propose the first method to prioritize the MPs for DL software, so as to boost the revealing of detected violations in DL testing. Specifically, we devise a new type of metric to measure the execution diversity of DL software on MPs based on the distribution discrepancy of the neuron outputs. The fault-sensitive MPs are next prioritized based on the devised diversity metric. Comprehensive evaluation results show that the proposed prioritization method and diversity metric can effectively prioritize the fault-sensitive MPs, boost the revealing of detected violations, and even facilitate the selection and design of the effective Metamorphic Relations for the image classification DL software. Xiaoyuan Xie, Pengbo Yin, Songqiang Chen |
ASE | 3 |
| 2022 | Generating Multiscale Maps From Satellite Images via Series Generative Adversarial NetworksabstractConsidering the success of generative adversarial networks (GANs) for image-to-image translation, researchers have attempted to translate satellite images to maps (si2map) through GAN for cartography. However, these studies involved limited scales, which hinders multiscale map creation. By extending their method, high-resolution satellite images can be trivially translated to multiscale maps through scale-wise si2map generators trained for certain scales. However, this strategy has two theoretical limitations. First, inconsistency between high-resolution satellite images and object generalization on multiscale maps (SI-M inconsistency) increasingly complicates the extraction of geographical information from satellite images for generators with decreasing scale. Second, as si2map translation is cross-domain, generators incur high computation costs to transform the pixel distribution on satellite images to that on maps. Thus, we designed a series strategy of generators for multiscale si2map translation to address these limitations. In this strategy, high-resolution satellite images are inputted to an si2map generator to output large-scale maps, which are translated to multiscale maps through series multiscale map generators. The series strategy avoids SI-M inconsistency as high-resolution satellite images are only translated to large-scale maps and transforms cross-domain translation to approximately intradomain translation when generating multiscale maps. Our experimental results showed better quality multiscale map generation with the series strategy, as shown by average increases of 11.69%, 53.78%, 55.42%, and 72.34% in the structural similarity index (SSIM), edge structural similarity index (ESSI), intersection over union (road), and intersection over union (water) for data from Mexico City and Tokyo at zoom levels 17–13. Xu Chen 0042, Bangguo Yin, Songqiang Chen, Haifeng Li 0007 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | MULA: A Just-In-Time Multi-labeling System for Issue ReportsabstractA very important function of an issue tracking system is to assign labels to issue reports, such as bug, feature, enhancement, etc., in order to categorize issues to facilitate various development activities. In practice, it is very common that an issue has multiple labels. However, current works are mainly based on single-label prediction, which are not suitable for just-in-time multi-labeling services, due to the low efficiency. Therefore, in this paper, we propose MULA, a just-in-time MUlti-LAbeling system, which learns and automatically assigns multiple labels to issue reports. We have built a dataset with 81,601 entries and 11 labels, as the first benchmark for this task, and implemented a GitHub app. To the best of our knowledge, this is the first work and tool for online multi-labeling GitHub issues based on their categories. We conduct a comprehensive empirical study, including comparisons with five commonly adopted labeling models that show the superiority of MULA, as well as an evaluation that shows high consistency between MULA’s suggestions and developers’ opinions. Xiaoyuan Xie, Yuhui Su, Songqiang Chen, Lin Chen 0015, Jifeng Xuan, Baowen Xu |
IEEE Trans. Reliab. | 3 |
| 2021 | Where to Handle an Exception? Recommending Exception Handling Locations from a Global PerspectiveabstractException handling is an effective mechanism to guarantee software reliability in modern programming languages. An exception interrupts the program execution and propagates backwards along the call chain until the exception is caught by an exception handler. In software development practices, developers may be confused in determining where to place the exception handler in the call chain. The reason is that exception handling requires a developer to take a comprehensive consideration from a global perspective of the software project. In this paper, we propose an automatic approach EHAdvisor, which recommends exception handling locations from the global perspective of the project. EHAdvisor first trains a binary classification model based on four types of features, including architectural features, project features, functional features, and exception features. Then, for a new code snippet with exceptions, EHAdvisor predicts the exception catching probability for each method in the call chain based on the classification model and recommends Top-K exception handling locations based on the probability ranking. We conducted experiments on a dataset from 29 high-quality open source projects. Experimental results show that EHAdvisor achieves an average Top-1 recommendation success rate of 70.83% for across-project location recommendation and an average Top-1 accuracy of 86.21% for intra-project recommendation. Experiments on the importance scores show that global features, such as project features and architectural features, are evidently important to the recommendation of exception handling locations. Xiangyang Jia, Songqiang Chen, Xingqi Zhou, Run Yu 0002, Xu Chen 0042, Jifeng Xuan |
ICPC | 2 |
| 2021 | Testing Your Question Answering Software via Asking RecursivelyabstractQuestion Answering (QA) is an attractive and challenging area in NLP community. There are diverse algorithms being proposed and various benchmark datasets with different topics and task formats being constructed. QA software has also been widely used in daily human life now. However, current QA software is mainly tested in a reference-based paradigm, in which the expected outputs (labels) of test cases need to be annotated with much human effort before testing. As a result, neither the just-in-time test during usage nor the extensible test on massive unlabeled real-life data is feasible, which keeps the current testing of QA software from being flexible and sufficient. In this paper, we propose a method, qaAskeR, with three novel Metamorphic Relations for testing QA software. qaAskeR does not require the annotated labels but tests QA software by checking its behaviors on multiple recursively asked questions that are related to the same knowledge. Experimental results show that qaAskeR can reveal violations at over 80% of valid cases without using any preannotated labels. Diverse answering issues, especially the limited generalization on question types across datasets, are revealed on a state-of-the-art QA algorithm. Songqiang Chen, Xiaoyuan Xie |
ASE | 1 |
| 2021 | Property-based Test for Part-of-Speech Tagging ToolabstractPart-of-Speech (POS) tagging for sentences is a basic and widely-used Natural Language Processing (NLP) technique. People rely heavily on it to predict POS tags that serve as the base for many advanced NLP tasks, such as sentiment analysis, word sense disambiguation, and information retrieval. However, POS tagging tools could make wrong predictions, which bring consequent error propagation to the advanced tasks and even cause serious threats in critical application domains. In this paper, we propose to test POS tagging tools with Metamorphic Testing against some properties that they should follow. The preliminary exploration with two groups of Metamorphic Relations shows that our method can effectively reveal defects of three common POS tagging tools (i.e., spaCy, NLTK, and Flair) on handling fairly simple intra- and inter-sentence transformation regarding adverbial clause and sentence appending. This demonstrates the great potential of our method to deliver a systematic test and reveal the unaware issues, which may benefit the validation, repair, and improvement, for POS tagging tools. Songqiang Chen, Xiaoyuan Xie |
ASE | 2 |
| 2021 | Validation on machine reading comprehension software without annotated labels: a property-based methodabstractMachine Reading Comprehension (MRC) in Natural Language Processing has seen great progress recently. But almost all the current MRC software is validated with a reference-based method, which requires well-annotated labels for test cases and tests the software by checking the consistency between the labels and the outputs. However, labeling test cases of MRC could be very costly due to their complexity, which makes reference-based validation hard to be extensible and sufficient. Furthermore, solely checking the consistency and measuring the overall score may not be sensible and flexible for assessing the language understanding capability. In this paper, we propose a property-based validation method for MRC software with Metamorphic Testing to supplement the reference-based validation. It does not refer to the labels and hence can make much data available for testing. Besides, it validates MRC software against various linguistic properties to give a specific and in-depth picture on linguistic capabilities of MRC software. Comprehensive experimental results show that our method can successfully reveal violations to the target linguistic properties without the labels. Moreover, it can reveal problems that have been concealed by the traditional validation. Comparison according to the properties provides deeper and more concrete ideas about different language understanding capabilities of the MRC software. Songqiang Chen, Xiaoyuan Xie |
ESEC/SIGSOFT FSE | 1 |
| 2021 | SMAPGAN: Generative Adversarial Network-Based Semisupervised Styled Map Tile Generation MethodabstractTraditional online map tiles, which are widely used on the Internet, such as by Google Maps and Baidu Maps, are rendered from vector data. The timely updating of online map tiles from vector data, for which generation is time-consuming, is a difficult mission. Generating map tiles over time from remote sensing images is relatively simple and can be performed quickly without vector data. However, this approach used to be challenging or even impossible. Inspired by image-to-image translation (img2img) techniques based on generative adversarial networks (GANs), we proposed a semisupervised generation of styled map tiles based on the GANs (SMAPGAN) model to generate styled map tiles directly from remote sensing images. In this model, we designed a semisupervised learning strategy to pretrain SMAPGAN on rich unpaired samples and fine-tune it on limited paired samples in reality. We also designed the image gradient L1 loss and the image gradient structure loss to generate a styled map tile with global topological relationships and detailed edge curves for objects, which are important in cartography. Moreover, we proposed the edge structural similarity index (ESSI) as a metric to evaluate the quality of the topological consistency between the generated map tiles and ground truth. The experimental results show that SMAPGAN outperforms state-of-the-art (SOTA) works according to the mean squared error, the structural similarity index, and the ESSI. Also, SMAPGAN gained higher approval than SOTA in a human perceptual test on the visual realism of cartography. Our work shows that SMAPGAN is a new tool with excellent potential for producing styled map tiles. Our implementation of SMAPGAN is available at https://github.com/imcsq/SMAPGAN. Xu Chen 0042, Songqiang Chen, Bangguo Yin, Jian Peng 0009, Xiaoming Mei, Haifeng Li 0007 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2020 | Stay Professional and Efficient: Automatically Generate Titles for Your Bug ReportsabstractBug reports in a repository are generally organized line by line in a list-view, with their titles and other meta-data displayed. In this list-view, a concise and precise title plays an important role that enables project practitioners to quickly and correctly digest the core idea of the bug, without carefully reading the corresponding details. However, the quality of bug report titles varies in open-source communities, which may be due to the limited time and unprofessionalism of authors. To help report authors efficiently draft good-quality titles, we propose a method, named iTAPE, to automatically generate titles for their bug reports. iTAPE formulates title generation into a one-sentence summarization task. By properly tackling two domain-specific challenges (i.e. lacking off-the-shelf dataset and handling the low-frequency human-named tokens), iTAPE then generates titles using a Seq2Seq-based model. A comprehensive experimental study shows that iTAPE can obtain fairly satisfactory results, in terms of the comparison with three latest one-sentence summarization works, as well as the feedback from human evaluation. Songqiang Chen, Xiaoyuan Xie, Bangguo Yin, Yuanxiang Ji, Lin Chen 0015, Baowen Xu |
ASE | 1 |