Shuguang Chen

dblp:15/7102 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
7since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Systems, architecture and hardware · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Human Cognitive Pattern Simulation for Crowdsourced Test Report Consistency Detection
abstract
Crowdsourced testing has emerged as a prominent paradigm in software testing by leveraging the diversity of crowdworkers. In this paradigm, crowd-workers are required to submit a test report for each identified bug, which typically contains a textual description and a bug screenshot. However, due to varying worker expertise, many reports exhibit inconsistencies between the textual description and the bug screenshot, which hinder the report review process. Existing methods address this issue by automatically detecting report consistency, typically through matching the UI widgets referenced in the textual description with those visible in the bug screenshot. However, such methods focus only on surface-level element correspondence and fail to capture the abstract bug semantics, such as the functional meaning and bug-triggering context. Consequently, they lack the ability to detect more subtle but realistic inconsistencies. To bridge this gap, we propose INCONHUNTER, a novel method for crowdsourced test report consistency detection that explicitly simulates human cognitive pattern. In this pattern, humans typically adopt two complementary reasoning strategies. If the textual description allows them to form an expectation about the visual bug features, they assess consistency by verifying if the expected features appear in the bug screenshot. Otherwise, they shift to reasoning about if the bug-triggering context described in the report aligns with the app state shown in the bug screenshot. INCONHUNTERinstantiates this cognitive pattern through two LLM-powered modules, each dedicated to one reasoning strategy. We evaluate INCONHUNTERthrough experiments on our dataset with 2,310 labeled crowdsourced test reports, and results show that INCONHUNTERoutperforms baselines by 14.00%–19.28%, demonstrating superior effectiveness, monetary-based cost efficiency, and alignment with human cognitive pattern.
Yuchen Ling, Shengcheng Yu, Shuguang Chen, Liuming Wang, Chunrong Fang, Jia Liu 0015, Zhenyu Chen 0001
IEEE Trans. Software Eng.3
2024 ERD-CQC : Enhanced Rule and Dependency Code Quality Check for Java
abstract
In the field of software development, the application of code quality check tools has become a key factor in improving product quality and development efficiency. While many existing tools are effective at detecting common problems in code, there are still some limitations. Firstly, these tools rely on predefined rules that may not fully encompass real-world coding challenges. Secondly, a lack of consideration of dependencies leads to failure to report violations occurring across files or modules. Third, the metrics used by these tools primarily focus on object-oriented programming, limiting their ability to assess software quality from the perspective of nationalized standards. To address these issues, this work proposes a dependency-enhanced method namely ERD-CQC for code quality detection and measurement. ERD-CQC provides 88 detection rules and 45 metrics, supplementing checking rules in categories such as Circuit Breaking, Serializable, and Security. ERD-CQC constructs an infused graph by integrating abstract syntax trees (ASTs), entities, and dependencies for violation detection. Based on the detection results, ERD-CQC provides a code quality measurement system with 4 nationalized standard dimensions for the purpose of measuring code quality from multiple perspectives. To validate the effectiveness of ERD-CQC, we manually examined 647 compliant and 528 non-compliant code snippets. ERD-CQC achieves the recall and F1 score exceeding 98%. We also collected open-source projects and closed-source projects in the real world, containing a total of 4,319 non-compliant code snippets. On this real-world benchmark, the average F1 score of ERD-CQC is 11.44% higher than the advanced tool SonarQube. Finally, we visualized the quality measurement results based on metrics and found that open-source and closed-source projects have certain patterns in metric performance. Our work will benefit developers in checking, evaluating, and monitoring their software quality comprehensively.
Wuxia Jin, Liuming Wang, Shuguang Chen, Yihan Wang 0020, Haijun Wang 0002, Ting Liu 0002
Internetware5
2024 Fault-tolerant deep learning inference on CPU-GPU integrated edge devices with TEEs
Hongjian Xu, Longlong Liao, Xinqi Liu, Shuguang Chen, Zhixuan Liang, Yuanlong Yu 0001
Future Gener. Comput. Syst.4
2024 Dual-attentive cascade clustering learning for visible-infrared person re-identification
Xianju Wang, Cuiqun Chen, Shuguang Chen
Multim. Tools Appl.4
2023 Deploying User-space TCP at Cloud Scale with LUNA
Lingjun Zhu, Erci Xu, Shuguang Chen, Xingyu Liao, Zhendan Yang, Zhongqing Chen, Yijun Hou, Jiaji Zhu, Jiesheng Wu
USENIX ATC7
2022 Style Transfer as Data Augmentation: A Case Study on Named Entity Recognition
abstract
In this work, we take the named entity recognition task in the English language as a case study and explore style transfer as a data augmentation method to increase the size and diversity of training data in low-resource scenarios.We propose a new method to effectively transform the text from a high-resource domain to a low-resource domain by changing its style-related attributes to generate synthetic data for training.Moreover, we design a constrained decoding algorithm along with a set of key ingredients for data selection to guarantee the generation of valid and coherent data.Experiments and analysis on five different domain pairs under different data regimes demonstrate that our approach can significantly improve results compared to current state-of-the-art data augmentation methods.Our approach is a practical solution to data scarcity, and we expect it to be applicable to other NLP tasks. 1
Shuguang Chen, Leonardo Neves, Thamar Solorio
EMNLP1
2021 Data Augmentation for Cross-Domain Named Entity Recognition
abstract
Current work in named entity recognition (NER) shows that data augmentation techniques can produce more robust models.However, most existing techniques focus on augmenting in-domain data in low-resource scenarios where annotated data is quite limited.In contrast, we study cross-domain data augmentation for the NER task.We investigate the possibility of leveraging data from highresource domains by projecting it into the lowresource domains.Specifically, we propose a novel neural architecture to transform the data representation from a high-resource to a low-resource domain by learning the patterns (e.g.style, noise, abbreviations, etc.) in the text that differentiate them and a shared feature space where both domains are aligned.We experiment with diverse datasets and show that transforming the data to the low-resource domain representation achieves significant improvements over only using data from highresource domains. 1
Shuguang Chen, Gustavo Aguilar, Leonardo Neves, Thamar Solorio
EMNLP (1)1
2007 antiCODE: a natural sense-antisense transcripts database
abstract
BACKGROUND: Natural antisense transcripts (NATs) are endogenous RNA molecules that exhibit partial or complete complementarity to other RNAs, and that may contribute to the regulation of molecular functions at various levels. In recent years, large-scale NAT screens in several model organisms have produced much data, but there is no database to assemble all these data. AntiCODE intends to function as an integrated NAT database for this purpose. RESULTS: This release of antiCODE contains more than 30,000 non-redundant natural sense-antisense transcript pairs from 12 eukaryotic model organisms. In order to provide an integrated NAT research platform, efficient browser, search and Blast functions have been included to enable users to easily access information through parameters such as species, accession number, overlapping patterns, coding potential etc. In addition to the collected information, antiCODE also introduces a simple classification system to facilitate the study of natural antisense transcripts. CONCLUSION: Though a few similar databases also dealing with NATs have appeared lately, antiCODE is the most comprehensive among these, comprising almost all currently detected NAT pairs.
Yifei Yin, Yi Zhao 0013, Changning Liu, Shuguang Chen, Runsheng Chen
BMC Bioinform.5