VLDB 2026 Research / reviewers in the wild / expert
Guoliang Dong
dblp:232/2133
· DBLP profile ↗
9ranked-venue papers
4as first author
6since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 7 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Evaluating and Mitigating Linguistic Discrimination in Large Language Models: Perspectives on Safety Equity and Knowledge EquityabstractLarge language models (LLMs) typically provide multilingual support and demonstrate remarkable capabilities in solving tasks described in different languages. However, LLMs can exhibit linguistic discrimination due to the uneven distribution of training data across languages. That is, LLMs struggle to maintain consistency when handling the same task in different languages, compromising both safety equity and knowledge equity. In this paper, we first systematically evaluate the linguistic discrimination of LLMs from two aspects: safety and quality, using a form of metamorphic testing. The metamorphic relationship we examine is that LLMs are expected to deliver outputs with similar semantics when prompted with inputs that have the same meaning. We conduct this evaluation with two datasets based on four representative LLMs. The results show that LLMs exhibit stronger human alignment capabilities with queries in English, French, Russian, and Spanish compared to queries in Bengali, Georgian, Nepali and Maithili. Moreover, for queries in English, Danish, Czech and Slovenian, LLMs tend to produce responses with a higher quality compared to the other languages. Upon these findings, we propose LDFighter, a similarity-based voting method, to mitigate the linguistic discrimination in LLMs. We comprehensively evaluate LDFighter against a spectrum of queries including benign, harmful, and adversarial prompts. The results show that LDFighter significantly reduces jailbreak success rates and improves response quality. All code, data, and the technical appendix are publicly available at: \url{https://github.com/dgl-prc/ldfighter}. Guoliang Dong, Haoyu Wang 0017, Jun Sun 0001, Xinyu Wang 0001 |
IJCAI | 1 |
| 2025 | A Comprehensive Study of OOP-Related Bugs in C++ CompilersabstractModern C++, a programming language characterized by its extensive use of object-oriented programming (OOP) features, is widely used for system programming. However, C++ compilers often struggle to correctly handle these sophisticated OOP features, resulting in numerous high-profile compiler bugs that can lead to crashes or miscompilation. Despite the significance of OOP-related bugs, existing studies largely overlook OOP features, hindering their ability to discover such bugs. To assist both compiler fuzzer designers and compiler developers, we conduct a comprehensive study of the compiler bugs caused by incorrectly handling C++ OOP-related features. First, we systematically extract 788 OOP-related C++ compiler bugs from GCC and LLVM. Second, derived from the core concepts of OOP and C++, we manually identified a two-level taxonomy of the OOP-related features leading to compiler bugs, which consists of 6 primary categories (e.g.,Abstraction & Encapsulation,Inheritance, andRuntime Polymorphism), along with 17 secondary categories (e.g.,Constructors & DestructorsandMultiple Inheritance). Third, we systematically analyze the root causes, symptoms, fixes, options, and C++ standard versions of these bugs. Our analysis yields 13 key findings, highlighting that features related to the construction and destruction of objects lead to the highest number of bugs, crashes are the most frequent symptom, and while the average time from bug introduction to discovery is 1856 days, fixing the bug once discovered takes only 174 days on average. Additionally, more than half of the bugs can be triggered without any compiler options. These findings offer valuable insights not only for developing new compiler testing approaches but also for improving language design and compiler engineering. Inspired by these findings, we developed a proof-of-concept compiler fuzzer OOPFuzz, specifically targeting OOP-related bugs in C++ compilers. We applied it against the newest release versions of GCC and LLVM. In about 3 hours, it detected 9 bugs, of which 3 have been confirmed by the developers, including a bug of LLVM that had persisted for 13 years. The results indicate our taxonomy and analysis provide valuable insights for future research in compiler testing. Bo Wang 0050, Chong Chen 0002, Junjie Chen 0003, Youfang Lin, Guoliang Dong, Jun Sun 0001 |
IEEE Trans. Software Eng. | 7 |
| 2022 | Repairing Adversarial Texts Through Perturbation
Guoliang Dong, Jingyi Wang 0004, Jun Sun 0001, Sudipta Chattopadhyay 0001, Xinyu Wang 0001, Jie Shi 0013, Jin Song Dong 0001 |
TASE | 1 |
| 2022 | Automatic Fairness Testing of Neural Classifiers Through Adversarial SamplingabstractAlthough deep learning has demonstrated astonishing performance in many applications, there are still concerns about its dependability. One desirable property of deep learning applications with societal impact is fairness (i.e., non-discrimination). Unfortunately, discrimination might be intrinsically embedded into the models due to the discrimination in the training data. As a countermeasure, fairness testing systemically identifies discriminatory samples, which can be used to retrain the model and improve the model’s fairness. Existing fairness testing approaches however have two major limitations. First, they only work well on traditional machine learning models and have poor performance (e.g., effectiveness and efficiency) on deep learning models. Second, they only work on simple structured (e.g., tabular) data and are not applicable for domains such as text. In this work, we bridge the gap by proposing a scalable and effective approach for systematically searching for discriminatory samples while extending existing fairness testing approaches to address a more challenging domain, i.e., text classification. Compared with state-of-the-art methods, our approach only employs lightweight procedures like gradient computation and clustering, which is significantly more scalable and effective. Experimental results show that on average, our approach explores the search space much more effectively (9.62 and 2.38 times more than the state-of-the-art methods respectively on tabular and text datasets) and generates much more discriminatory samples (24.95 and 2.68 times) within a same reasonable time. Moreover, the retrained models reduce discrimination by 57.2 and 60.2 percent respectively on average. Peixin Zhang 0001, Jingyi Wang 0004, Jun Sun 0001, Xinyu Wang 0001, Guoliang Dong, Xingen Wang, Jin Song Dong 0001 |
IEEE Trans. Software Eng. | 5 |
| 2021 | Towards Repairing Neural Networks CorrectlyabstractNeural networks are increasingly applied to support decision-making in safety-critical applications (like autonomous cars, unmanned aerial vehicles, and face recognition-based authentication). While many impressive static verification techniques have been proposed to tackle the correctness problem of neural networks, existing static verification techniques still do not answer the natural question: what is the subsequent measure that one should take if the DNN is not verified? In this work, we propose a runtime repairing method to ensure the correctness of neural networks within certain input regions. Given a neural network and a safety property, we first adopt state-of-the-art static verification techniques to verify the neural networks. In the case that the verification fails, we strategically identify locations to introduce additional gates which “correct” neural network behaviors at runtime whilst keeping the modifications small. Experiment results show that our approach effectively generates neural networks which are guaranteed to satisfy the properties, whilst being consistent with the original neural network most of the time. Guoliang Dong, Jun Sun 0001, Xingen Wang, Xinyu Wang 0001 |
QRS | 1 |
| 2021 | A novel fuzzy clustering algorithm based on rough set and inhibitive factorabstractSummary As an important data analysis method, rough set theory can be used to analyze data with uncertainties. Rough set is able to acquire knowledge by the indistinguishable relationship among data objects without any prior knowledge. Rough set theory provides a new theoretical means for solving soft computing problems and has a wide application space in data mining. Meanwhile, the fuzzy C‐means clustering algorithm is sensitive to noisy points and has low convergence speed. In order to deal with the above problems, a novel fuzzy clustering algorithm based on rough set and inhibitive factor is proposed. According to the related concepts of rough set theory, the membership model of fuzzy C‐means algorithm is redefined. Besides, an inhibitive factor is set to improve the convergence speed of the algorithm under the premise of guaranteeing the clustering effect. Lei-Yu Tang, Guoliang Dong, Jiancong Fan |
Concurr. Comput. Pract. Exp. | 4 |
| 2020 | White-box fairness testing through adversarial samplingabstractAlthough deep neural networks (DNNs) have demonstrated astonishing performance in many applications, there are still concerns on their dependability. One desirable property of DNN for applications with societal impact is fairness (i.e., non-discrimination). In this work, we propose a scalable approach for searching individual discriminatory instances of DNN. Compared with state-of-the-art methods, our approach only employs lightweight procedures like gradient computation and clustering, which makes it significantly more scalable than existing methods. Experimental results show that our approach explores the search space more effectively (9 times) and generates much more individual discriminatory instances (25 times) using much less time (half to 1/7). Peixin Zhang 0001, Jingyi Wang 0004, Jun Sun 0001, Guoliang Dong, Xinyu Wang 0001, Xingen Wang, Jin Song Dong 0001 |
ICSE | 4 |
| 2020 | Towards Interpreting Recurrent Neural Networks through Probabilistic AbstractionabstractNeural networks are becoming a popular tool for solving many real-world problems such as object recognition and machine translation, thanks to its exceptional performance as an end-to-end solution. However, neural networks are complex black-box models, which hinders humans from interpreting and consequently trusting them in making critical decisions. Towards interpreting neural networks, several approaches have been proposed to extract simple deterministic models from neural networks. The results are not encouraging (e.g., low accuracy and limited scalability), fundamentally due to the limited expressiveness of such simple models. Guoliang Dong, Jingyi Wang 0004, Jun Sun 0001, Yang Zhang 0016, Xinyu Wang 0001, Jin Song Dong 0001, Xingen Wang |
ASE | 1 |
| 2019 | Adversarial sample detection for deep neural network through model mutation testingabstractDeep neural networks (DNN) have been shown to be useful in a wide range of applications. However, they are also known to be vulnerable to adversarial samples. By transforming a normal sample with some carefully crafted human imperceptible perturbations, even highly accurate DNN make wrong decisions. Multiple defense mechanisms have been proposed which aim to hinder the generation of such adversarial samples. However, a recent work show that most of them are ineffective. In this work, we propose an alternative approach to detect adversarial samples at runtime. Our main observation is that adversarial samples are much more sensitive than normal samples if we impose random mutations on the DNN. We thus first propose a measure of 'sensitivity' and show empirically that normal samples and adversarial samples have distinguishable sensitivity. We then integrate statistical hypothesis testing and model mutation testing to check whether an input sample is likely to be normal or adversarial at runtime by measuring its sensitivity. We evaluated our approach on the MNIST and CIFAR10 datasets. The results show that our approach detects adversarial samples generated by state-of-the-art attacking methods efficiently and accurately. Jingyi Wang 0004, Guoliang Dong, Jun Sun 0001, Xinyu Wang 0001, Peixin Zhang 0001 |
ICSE | 2 |