VLDB 2026 Research / reviewers in the wild / expert
Sen Fang
dblp:290/0510
· DBLP profile ↗
16ranked-venue papers
9as first author
16since 2021 · last 2026
0000-0002-9918-7180ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 9 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 3 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Stable Signer: Hierarchical Sign Language Generative ModelabstractSign Language Production (SLP) is the process of converting the complex input text into a real video. Most previous works focused on the Text2Gloss, Gloss2Pose, Pose2Vid stages, and some concentrated on Prompt2Gloss and Text2Avatar stages. However, this field has made slow progress due to the inaccuracy of text conversion, pose generation, and the rendering of poses into real human videos in these stages, resulting in gradually accumulating errors. Therefore, in this paper, we streamline the traditional redundant structure, simplify and optimize the task objective, and design a new sign language generative model called Stable Signer. It redefines the SLP task as a hierarchical generation end-to-end task that only includes text understanding (Prompt2Gloss, Text2Gloss) and Pose2Vid, and executes text understanding through our proposed new Sign Language Understanding Linker called SLUL, and generates hand gestures through the named SLP-MoE hand gesture rendering expert block to end-to-end generate high-quality and multi-style sign language videos. SLUL is trained using the newly developed Semantic-Aware Gloss Masking Loss (SAGM Loss). Its performance has improved by 48.6% compared to the current SOTA generation methods. Sen Fang, Yalin Feng, Hongbin Zhong, Dimitris N. Metaxas |
ACL (1) | 1 |
| 2025 | SignDiff: Diffusion Model for American Sign Language ProductionabstractIn this paper, we propose a dual-condition diffusion pre-training model named SignDIFF that can generate human sign language speakers from a skeleton pose. SignDiff has a novel Frame Reinforcement Network called FR-Net, similar to dense human pose estimation work, which enhances the correspondence between text lexical symbols and sign language dense pose frames, reduces the occurrence of multiple fingers in the diffusion model. In addition, we propose a new method for American Sign Language Production (ASLP), which can generate ASL skeletal pose videos from text input, integrating two new improved modules and a new loss function to improve the accuracy and quality of sign language skeletal posture and enhance the ability of the model to train on largescale data. We propose a simple baseline for ASL production and report the scores of 17.19 and 12.85 on BLEU-4 on the How2Sign dev/test sets. We evaluated our model on the previous mainstream dataset PHOENIX14T, and our method achieved the SOTA results. In addition, our image quality far exceeds all previous results by 10 percentage points in terms of SSIM. Sen Fang, Chunyu Sui, Yanghao Zhou, Hongbin Zhong, Yapeng Tian, Chen Chen 0001 |
FG | 1 |
| 2025 | Deep Smart Contract Intent DetectionabstractIn recent years, research in software security has concentrated on identifying vulnerabilities in smart contracts to prevent significant losses of crypto assets on blockchains. Despite early successes in this area, detecting developers' intents in smart contracts has become a more pressing issue, as malicious intents have caused substantial financial losses. Unfortunately, existing research lacks effective methods for detecting development intents in smart contracts. To address this gap, we propose SMARTINTENTNN (Smart Contract Intent Neural Network), a deep learning model designed to automatically detect development intents in smart contracts. SMARTINTENTNN leverages a pre-trained sentence encoder to generate contextual representations of smart contracts, employs a K-means clustering model to identify and highlight prominent intent features, and utilizes a bidirectional LSTM-based deep neural network for multi-label classification. We trained and evaluated SMARTINTENTNN on a dataset containing over 40,000 real-world smart contracts, employing self-comparison baselines in our experimental setup. The results show that SMARTINTENTNN achieves an F1-score of 0.8633 in identifying intents across 10 distinct categories, outperforming all baselines and addressing the gap in smart contract detection by incorporating intent analysis. Youwei Huang, Sen Fang, Jiachun Tao, Tao Zhang 0001 |
SANER | 2 |
| 2025 | Enhancing Android Malware Detection: The Influence of ChatGPT on Decision-centric TaskabstractWith the rise of large language models, such as ChatGPT, non-decisional models have been applied to various tasks. Moreover, ChatGPT has drawn attention to the traditional decision-centric task of Android malware detection. Despite effective detection methods proposed by scholars, they face low interpretability issues. Specifically, while these methods excel in classifying applications as benign or malicious and can detect malicious behavior, they often fail to provide detailed explanations for the decisions they make. This challenge raises concerns about the reliability of existing detection schemes and questions their true ability to understand complex data. In this study, we investigate the influence of the non-decisional model, ChatGPT, on the traditional decision-centric task of Android malware detection. We choose three state-of-the-art solutions, Drebin , \(XM_{AL}\) , and MaMaDroid , conduct a series of experiments on publicly available datasets, and carry out a comprehensive comparison and analysis. Our findings indicate that these decision-driven solutions primarily rely on statistical patterns within datasets to make decisions, rather than genuinely understanding the underlying data. In contrast, ChatGPT, as a non-decisional model, excels in providing comprehensive analysis reports, substantially enhancing interpretability. Furthermore, we conduct surveys among experienced developers. The result highlights developers’ preference for ChatGPT, as it offers in-depth insights and enhances efficiency and understanding of challenges. Meanwhile, these studies and analyses offer profound insights, presenting developers with a novel perspective on Android malware detection—enhancing the reliability of detection results from a non-decisional perspective. Yao Li 0017, Sen Fang, Tao Zhang 0001, Haipeng Cai |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2025 | RepairLLaMA: Efficient Representations and Fine-Tuned Adapters for Program RepairabstractAutomated Program Repair (APR) has evolved significantly with the advent of Large Language Models (LLMs). Fine-tuning LLMs for program repair is a recent avenue of research, with many dimensions which have not been explored. Existing work mostly fine-tune LLMs with naive code representations and does not scale to frontier models. To address this problem, we propose RepairLLaMA, a novel program repair approach that 1) identifies optimal code representations for APR with fine-tuned models, and 2) pioneers state-of-the-art parameter-efficient fine-tuning technique (PEFT) for program repair. This results in RepairLLaMA producing a highly effective ‘program repair adapter’ for fixing bugs with AI. Our experiments demonstrate the validity of both concepts. First, fine-tuning adapters with program repair specific code representations enables the model to use meaningful repair signals and produce better patches. Second, parameter-efficient fine-tuning helps fine-tuning to converge and clearly contributes to the effectiveness of RepairLLaMA in fixing bugs outside the fine-tuning data distribution. Overall, RepairLLaMA correctly fixes 144 Defects4J v2, 109 HumanEval-Java, and 20 GitBug-Java bugs, outperforming all baselines. André Silva 0002, Sen Fang, Martin Monperrus |
IEEE Trans. Software Eng. | 2 |
| 2024 | Erratum to "Effective Prediction of Bug-Fixing Priority via Weighted Graph Convolutional Networks"abstractIn [1], the affiliation of the primary authors should be as follows: Sen Fang, Youshuai Tan, Tao Zhang 0001, Zhou Xu 0003, Hui Liu 0003 |
IEEE Trans. Reliab. | 1 |
| 2024 | Supersonic: Learning to Generate Source Code Optimizations in C/C++abstractSoftware optimization refines programs for resource efficiency while preserving functionality. Traditionally, it is a process done by developers and compilers. This paper introduces a third option, automated optimization at the source code level. We present SUPERSONIC, a neural approach targeting minor source code modifications for optimization. Using a seq2seq model, SUPERSONIC is trained on C/C++ program pairs (xt, xt+1), wherext+1is an optimized version ofxt, and outputs a diff. SUPERSONIC’s performance is benchmarked against OpenAI’s GPT-3.5-Turbo and GPT-4 on competitive programming tasks. The experiments show that SUPERSONIC not only outperforms both models on the code optimization task but also minimizes the extent of the change with a model more than 600x smaller than GPT-3.5-Turbo and 3700x smaller than GPT-4. Zimin Chen, Sen Fang, Martin Monperrus |
IEEE Trans. Software Eng. | 2 |
| 2023 | Exploring Efficient-Tuned Learning Audio Representation Method from BriVL
Sen Fang, Yangjian Wu, Teik Toe Teoh |
ICONIP (15) | 1 |
| 2023 | RepresentThemAll: A Universal Learning Representation of Bug ReportsabstractDeep learning techniques have shown promising performance in automated software maintenance tasks associated with bug reports. Currently, all existing studies learn the customized representation of bug reports for a specific downstream task. Despite early success, training multiple models for multiple downstream tasks faces three issues: complexity, cost, and compatibility, due to the customization, disparity, and uniqueness of these automated approaches. To resolve the above challenges, we propose RepresentThemAll, a pre-trained approach that can learn the universal representation of bug reports and handle multiple downstream tasks. Specifically, RepresentThemAll is a universal bug report framework that is pre-trained with two carefully designed learning objectives: one is the dynamic masked language model and another one is a contrastive learning objective, “find yourself”. We evaluate the performance of RepresentThemAll on four downstream tasks, including duplicate bug report detection, bug report summarization, bug priority prediction, and bug severity prediction. Our experimental results show that RepresentThemAll outperforms all baseline approaches on all considered downstream tasks after well-designed fine-tuning. Sen Fang, Tao Zhang 0001, Youshuai Tan, He Jiang 0001, Xin Xia 0001, Xiaobing Sun 0001 |
ICSE | 1 |
| 2023 | Java Code Clone Detection by Exploiting Semantic and Syntax Information From Intermediate Code-Based GraphabstractCode clone detection plays a critical role in the field of software engineering. To achieve this goal, developers are required to have rich development experience for finding the “functional” clone code. However, this is unfriendly to novice developers. Although many approaches were proposed to automatically detect code clones, the results are not satisfactory. A major reason is that it is difficult to extract syntax and semantic information from the source code. To resolve this problem, in this article, we develop a novel graph representation approach based on intermediate code to detect the functional code clones. This graph representation is built based on intermediate code compiled from the source code. By using it, we can easily utilize graph embedding techniques to extract syntactic and semantic features from abstract syntax tree, control flow graph, and DFG generated from intermediate code. After that, we use the Softmax classifier to detect functional code clone pairs. We evaluate the performance of the proposed graph representation approach based on intermediate code for the code clone detection task on the BigCloneBench dataset. In order to improve performance, the embedded representation of intermediate code is initialized based on pretrained vectors learned from the collected LLVM IR dataset in advance. The experimental results show that our proposed intermediate code-based graph approach performs better than existing functional code clone detection approaches. Especially for the type-4 code clone detection, our approach outperforms the baseline approaches by an average of 33.49% in the term ofF1 score. Dawei Yuan, Sen Fang, Tao Zhang 0001, Zhou Xu 0003, Xiapu Luo |
IEEE Trans. Reliab. | 2 |
| 2023 | Do Pretrained Language Models Indeed Understand Software Engineering Tasks?abstractArtificial intelligence (AI) for software engineering (SE) tasks has recently achieved promising performance. In this article, we investigate to what extent the pre-trained language model truly understands those SE tasks such as code search, code summarization, etc. We conduct a comprehensive empirical study on a board set of AI for SE (AI4SE) tasks by feeding them with variant inputs: 1) with various masking rates and 2) with sufficient input subset method. Then, the trained models are evaluated on different SE tasks, including code search, code summarization, and duplicate bug report detection. Our experimental results show that pre-trained language models are insensitive to the given input, thus they achieve similar performance in these three SE tasks. We refer to this phenomenon asoverinterpretation, where a model confidently makes a decision without salient features, or where a model finds some irrelevant relationships between the final decision and the dataset. Our study investigates two approaches to mitigate the overinterpretation phenomenon: whole word mask strategy and ensembling. To the best of our knowledge, we are thefirstto reveal this overinterpretation phenomenon to the AI4SE community, which is an important reminder for researchers to design the input for the models and calls for necessary future work in understanding and implementing AI4SE tasks. Yao Li 0017, Tao Zhang 0001, Xiapu Luo, Haipeng Cai, Sen Fang, Dawei Yuan |
IEEE Trans. Software Eng. | 5 |
| 2023 | STRE: An Automated Approach to Suggesting App Developers When to Stop Reading ReviewsabstractIt is well known that user feedback (i.e., reviews) plays an essential role in mobile app maintenance. Users upload their troubles, app issues, or praises, to help developers refine their apps. However, reading tremendous amounts of reviews to retrieve useful information is a challenging job. According to our manual studies, reviews are full of repetitive opinions, thus developers could stop reading reviews when no more new helpful information appears. Developers can extract useful information from partial reviews to ameliorate their app and then develop a new version. However, it is tough to have a good trade-off between getting enough useful feedback and saving more time. In this paper, we propose a novel approach, named STRE, which utilizes historical reviews to suggest the time when most of the useful information appears in reviews of a certain version. We evaluate STRE on 62 recent versions of five apps from Apple's App Store. Study results demonstrate that our approach can help developers save their time by up to 98.33% and reserve enough useful reviews before stopping to read reviews such that developers do not spend additional time in reading redundant reviews over the suggested stopping time. At the same time, STRE can complement existing review categorization approaches that categorize reviews to further assist developers. In addition, we find that the missed top-word-related reviews appearing after the suggested stopping time contain limited useful information for developers. Finally, we find that 12 out of 13 of the emerging bugs from the studied versions appear before the suggested stopping time. Our approach demonstrates the value of automatically refining information from reviews. Youshuai Tan, Jinfu Chen 0002, Weiyi Shang, Tao Zhang 0001, Sen Fang, Xiapu Luo, Zijie Chen 0005, Shuhao Qi |
IEEE Trans. Software Eng. | 5 |
| 2022 | PRHAN: Automated Pull Request Description Generation Based on Hybrid Attention Network
Sen Fang, Tao Zhang 0001, Youshuai Tan, Zhou Xu 0003, Zhi-Xin Yuan, Ling-Ze Meng |
J. Syst. Softw. | 1 |
| 2022 | Facial makeup transfer with GAN for different aging faces
Sen Fang, Mingxing Duan, Kenli Li 0001, Keqin Li 0001 |
J. Vis. Commun. Image Represent. | 1 |
| 2021 | Self-Attention Networks for Code Search
Sen Fang, Youshuai Tan, Tao Zhang 0001, Yepang Liu 0001 |
Inf. Softw. Technol. | 1 |
| 2021 | Effective Prediction of Bug-Fixing Priority via Weighted Graph Convolutional NetworksabstractWith the increasing number of software bugs, bug fixing plays an important role in software development and maintenance. To improve the efficiency of bug resolution, developers utilize bug reports to resolve given bugs. Especially, bug triagers usually depend on bugs' descriptions to suggest priority levels for reported bugs. However, manual priority assignment is a time-consuming and cumbersome task. To resolve this problem, recent studies have proposed many approaches to automatically predict the priority levels for the reported bugs. Unfortunately, these approaches still face two challenges that include words' nonconsecutive semantics in bug reports and the imbalanced data. In this article, we propose a novel approach that graph convolutional networks (GCN) based on weighted loss function to perform the priority prediction for bug reports. For the first challenge, we build a heterogeneous text graph for bug reports and apply GCN to extract words' semantics in bug reports. For the second challenge, we construct a weighted loss function in the training phase. We conduct the priority prediction on four open-source projects, including Mozilla, Eclipse, Netbeans, and GNU compiler collection. Experimental results show that our method outperforms two baseline approaches in terms of the F-measure by weighted average of 13.22%. Sen Fang, Youshuai Tan, Tao Zhang 0001, Zhou Xu 0003, Hui Liu 0003 |
IEEE Trans. Reliab. | 1 |