Pin Ji

dblp:223/1135 · DBLP profile ↗
← Back
5ranked-venue papers
4as first author
5since 2021 · last 2025
0009-0002-3154-4409ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 5 · 4 first-author · 5 since 2021
YearPublicationVenuePosition
2025 SPlice: Automated Testing for Speech Translation via Syntactic Analysis
abstract
With the advancement of Deep Learning, the performance of speech translation systems has made remarkable progress. However, similar to traditional software, speech translation systems can still suffer from software defects that can lead to incorrect translations with potentially serious consequences. These systems are also vulnerable to real-world environmental interference, making their behavior unpredictable. The black-box nature of deep neural networks renders traditional testing methods ineffective, while the lack of diverse test cases and the challenge of constructing test oracles further hinder the implementation of their testing. To address this, we introduce syntactic structure invariance, a linguistically inspired concept that captures the structural containment between a pair of derivationally related sentences and their corresponding translations. Based on this concept, we propose a novel speech translation testing method, SPlice. SPlice simulates environmental disturbances on a seed speech, disassembles it into a template and speech blocks, and then generates multiple derivational speech pairs by inserting the blocks back into the template. SPlice detects translation errors by checking whether the syntactic structure invariance relation is violated in the translation results corresponding to the speech pairs. To validate SPlice, we experiment with three industrial speech translation systems: Google Translate, Youdao Translator, and Iflytek Translator. With 600 speeches crawled from the BBC as seed tests, SPlice detects 1,640, 1,101, and 1,305 translation errors with around 90.7% precision. The experimental results show that SPlice can effectively detect errors in the speech translation results with high precision, providing valuable information for developers to improve system performance.
Ji Qi 0005, Pin Ji, Jia Liu 0008, Yang Feng 0003
QRS3
2025 NLPLego: Assembling Test Generation for Natural Language Processing Applications
abstract
With the development of Deep Learning, Natural Language Processing (NLP) applications have reached or even exceeded human-level capabilities in certain tasks. Although NLP applications have shown good performance, they can still have bugs like traditional software and even lead to serious consequences. Inspired by Lego blocks and syntax structure analysis, we propose an assembling test generation method for NLP applications or models and implement it in NLPLego . The key idea of NLPLego is to assemble the sentence skeleton and adjuncts in order by simulating the building of Lego blocks to generate multiple grammatically and semantically correct sentences based on one seed sentence. The sentences generated by NLPLego have derivation relations and different degrees of variation. These characteristics make it well-suited for integration with metamorphic testing theory, addressing the challenge of test oracle absence in NLP application testing. To validate NLPLego , we conduct experiments on three commonly used NLP tasks (i.e., machine reading comprehension, sentiment analysis, and semantic similarity measures), focusing on the efficiency of test generation and the quality and effectiveness of generated tests. We select five advanced NLP models and one popular industrial NLP software as the tested subjects. Given seed tests from SQuAD 2.0, SST, and QQP, NLPLego successfully detects 1,732, 3,140, and 261,879 incorrect behaviors with around 93.1% precision in three tasks, respectively. The experiment results show that NLPLego can efficiently generate high-quality tests for multiple NLP tasks to detect erroneous behaviors effectively. In the case study, we analyze the testing results provided by NLPLego to obtain intuitive representations of the different NLP capabilities of the tested subjects. The case study confirms that NLPLego can provide developers with clarity on the direction to improve NLP models or applications, laying the foundation for enhancing performance.
Pin Ji, Yang Feng 0003, Ruohao Zhang, Ruichen Xue, Weitao Huang, Jia Liu 0015
ACM Trans. Softw. Eng. Methodol.1
2025 MoCo: Fuzzing Deep Learning Libraries via Assembling Code
abstract
The rapidly developing Deep Learning (DL) techniques have been applied in software systems of various types. However, they can also pose new safety threats with potentially serious consequences, especially in safety-critical domains. DL libraries serve as the underlying foundation for DL systems, and bugs in them can have unpredictable impacts that directly affect the behaviors of DL systems. Previous research on fuzzing DL libraries still has limitations in generating tests corresponding to crucial testing scenarios and constructing test oracles. In this paper, we proposeMoCo, a novel fuzzing testing method for DL libraries via assembling code. The seed tests used byMoCoare code files that implement DL models, covering both model construction and training in the most common real-world application scenarios for DL libraries.MoCofirst disassembles the seed code files to extract templates and code blocks, then applies code block mutation operators (e.g., API replacement, random generation, and boundary checking) to generate new code blocks that fit the template. To ensure the correctness of the code block mutation, we employ the Large Language Model to parse the official documents of DL libraries for information about the parameters and the constraints between them. By inserting context-appropriate code blocks into the template,MoCocan generate a tree of code files with intergenerational relations. According to the derivation relations in this tree, we construct the test oracle based on the execution state consistency and the calculation result consistency. Since the granularity of code assembly is controlled rather than randomly divergent, we can quickly pinpoint the lines of code where the bugs are located and the corresponding triggering conditions. We conduct a comprehensive experiment to evaluate the efficiency and effectiveness ofMoCousing three widely-used DL libraries (i.e., TensorFlow, PyTorch, and Jittor). During the experiments,MoCodetects 77 new bugs of four types in three DL libraries, where 55 bugs have been confirmed, and 39 bugs have been fixed by developers. The experimental results demonstrate thatMoCocan generate high-quality tests that cover crucial testing scenarios and detect different types of bugs, which helps developers improve the reliability of DL libraries.
Pin Ji, Yang Feng 0003, Duo Wu, Lingyue Yan, Penglin Chen, Jia Liu 0015
IEEE Trans. Software Eng.1
2022 ASRTest: automated testing for deep-neural-network-driven speech recognition systems
abstract
With the rapid development of deep neural networks and end-to-end learning techniques, automatic speech recognition (ASR) systems have been deployed into our daily and assist in various tasks. However, despite their tremendous progress, ASR systems could also suffer from software defects and exhibit incorrect behaviors. While the nature of DNN makes conventional software testing techniques inapplicable for ASR systems, lacking diverse tests and oracle information further hinders their testing. In this paper, we propose and implement a testing approach, namely ASR, specifically for the DNN-driven ASR systems. ASRTest is built upon the theory of metamorphic testing. We first design the metamorphic relation for ASR systems and then implement three families of transformation operators that can simulate practical application scenarios to generate speeches. Furthermore, we adopt Gini impurity to guide the generation process and improve the testing efficiency. To validate the effectiveness of ASRTest, we apply ASRTest to four ASR models with four widely-used datasets. The results show that ASRTest can detect erroneous behaviors under different realistic application conditions efficiently and improve 19.1% recognition performance on average via retraining with the generated data. Also, we conduct a case study on an industrial ASR system to investigate the performance of ASRTest under the real usage scenario. The study shows that ASRTest can detect errors and improve the performance of DNN-driven ASR systems effectively.
Pin Ji, Yang Feng 0003, Jia Liu 0008, Zhenyu Chen 0001
ISSTA1
2021 Automated Testing for Machine Translation via Constituency Invariance
abstract
With the development of deep neural networks, machine translation has achieved significant progress and integrated with people’s daily lives to assist in various tasks. However, machine translators, which are essentially one kind of software, also suffer from software defects. Translation errors might cause misunderstanding or even lead to marketing blunders, and political crisis. Thus, almost all translation service providers have feedback channels of incorrect translations to collect training data and improve product performance. Inspired by the syntax structure analysis, we introduce the constituency invariance, which reflects the structural similarity between a simple sentence and sentences derived from it, to test machine translators. We implement it into an automated tool CIT to detect translation errors by checking the constituency invariance relation between the translation results. CIT adopts constituency parse trees to represent the syntactic structures of sentences and employs an efficient data augmentation method to derive multiple new sentences based on one sentence. To validate CIT, we experiment with three widely-used machine translators, i.e., Bing Microsoft Translator, Google Translate, and Youdao Translator. With 600 seed sentences as input, CIT detects 2212, 1910, and 1590 translation errors with around 77% precision. We have submitted detected errors to the development teams. Until we submit this paper, Google, Bing, and Youdao have fixed 15.4%, 32.0%, 14.3% of reported errors, respectively.
Pin Ji, Yang Feng 0003, Jia Liu 0008, Baowen Xu
ASE1