Yihong Dong

dblp:64/11335 · DBLP profile ↗
← Back
56ranked-venue papers
10as first author
53since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 32 · 6 first-author · 31 since 2021Software engineering, systems software and programming languages · 10 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Systems, architecture and hardware · 3 · 2 since 2021Databases, data management, data science and information retrieval · 2Computer networks · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Large Language Model Unlearning for Source Code
abstract
While Large Language Models (LLMs) excel at code generation, their inherent tendency toward verbatim memorization of training data introduces critical risks like copyright infringement, insecurity emission, and deprecated API utilization, etc. A straightforward yet promising defense is unlearning, i.e., erasing or down-weighting the offending snippets through post-training. However, we find its application to source code often tends to spill over, damaging the basic knowledge of programming languages learned by the LLM and degrading the overall capability. To ease this challenge, we propose PROD for precise source code unlearning. PROD surgically zeroes out the prediction probability of the prohibited tokens, and renormalizes the remaining distribution so that the generated code stays correct. By excising only the targeted snippets, PROD achieves precise forgetting without much degradation of the LLM's overall capability. To facilitate in-depth evaluation against PROD, we establish an unlearning benchmark consisting of three downstream tasks (i.e., unlearning of copyrighted code, insecure code, and deprecated APIs), and introduce Pareto Dominance Ratio (PDR) metric, which indicates both the forget quality and the LLM utility. Our comprehensive evaluation demonstrates that PROD achieves superior overall performance between forget quality and model utility compared to existing unlearning approaches across three downstream tasks, while consistently exhibiting improvements when applied to LLMs of varying series. PROD also exhibits superior robustness against adversarial attacks without generating or exposing the data to be forgotten. These results underscore that our approach not only successfully extends the application boundary of unlearning techniques to source code, but also holds significant implications for advancing reliable code generation.
Yihong Dong, Huangzhao Zhang, Tangxinyu Wang, Yingwei Ma, Rongyu Cao, Binhua Li, Zhi Jin 0001, Wenpin Jiao, Yongbin Li 0001, Ge Li 0001
AAAI2
2026 RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization
abstract
Yihong Dong, Xue Jiang, Yongding Tao, Huanyu Liu, Kechi Zhang, Lili Mou, Rongyu Cao, Yingwei MA, Jue Chen, Binhua Li, Zhi Jin, Fei Huang, Yongbin Li, Ge Li. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yihong Dong, Yongding Tao, Huanyu Liu 0001, Kechi Zhang, Lili Mou, Rongyu Cao, Yingwei Ma, Jue Chen 0003, Binhua Li, Zhi Jin 0001, Fei Huang 0002, Yongbin Li 0001, Ge Li 0001
ACL (1)1
2026 Saber: Efficient Sampling with Adaptive Acceleration and Backtracking Enhanced Remasking for Diffusion Language Model in Code Generation
abstract
Diffusion language models (DLMs) are emerging as a compelling alternative to the dominant autoregressive paradigm, offering inherent advantages in parallel generation and bidirectional context modeling. However, for the tasks with strict structural constraints such as code generation, DLMs face a critical trade-off between inference speed and output quality, where accelerating generation by reducing sampling steps often leads to catastrophic performance collapse.We find that the fundamental reasons are: 1) the generation difficulty is uneven in the structured sequence decoding steps, making DLM’s static acceleration strategy suboptimal; 2) the context of tokens generated by DLM evolves continuously, causing early high-confidence predictions to turn into irreversible errors.In this paper, we introduce efficient Sampling with Adaptive acceleration and Backtracking Enhanced Remasking (i.e., Saber), a novel training-free sampling algorithm for DLMs that the first to improve both inference speed and output quality in code generation. Saber dynamically adjusts the number of tokens unmasked per step based on the model’s evolving confidence, and utilizes a backtracking mechanism to revert tokens whose confidence drops as new context emerges, with its effectiveness supported by theoretical analysis.Extensive experiments on multiple mainstream code generation benchmarks show that Saber boosts Pass@1 accuracy by an average of 1.9% over mainstream DLM sampling methods, while achieving an average 251.4% inference speedup. By leveraging the inherent advantages of DLMs, our work significantly narrows the performance gap with autoregressive models in code generation.
Yihong Dong, Zhaoyu Ma, Zhiyuan Fan, Jiaru Qian, Yongmin Li 0004, Jianha Xiao, Zhi Jin 0001, Ge Li 0001
ACL (1)1
2026 CODERL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment
abstract
Xue Jiang, Yihong Dong, Mengyang Liu, Deng Hongyi, Tian Wang, Yongding Tao, Zhi Jin, Wenpin Jiao, Ge Li. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yihong Dong, Mengyang Liu, Hongyi Deng, Yongding Tao, Zhi Jin 0001, Wenpin Jiao, Ge Li 0001
ACL (1)2
2026 KoCo-Bench: Can Large Language Models Leverage Domain Knowledge in Software Development?
abstract
Xue Jiang, Ge Li, Jiaru Qian, Xianjie Shi, Chenjie Li, Hao Zhu, Ziyu Wang, Jielun Zhang, Zeyu Zhao, Kechi Zhang, Jia Li, Wenpin Jiao, Zhi Jin, Yihong Dong. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Ge Li 0001, Jiaru Qian, Xianjie Shi, Chenjie Li, Jielun Zhang, Kechi Zhang, Jia Li 0012, Wenpin Jiao, Zhi Jin 0001, Yihong Dong
ACL (1)14
2026 Attributed Network Representation Learning Based on Graph Neural Network: A Comprehensive Survey
abstract
ABSTRACT An attributed network encodes richer information through node and edge attributes. Attributed network representation learning (ANRL) seeks to obtain low‐dimensional node embeddings by jointly modeling structural topology and attribute semantics. Graph neural network (GNN)‐based methods, which leverage recursive message passing, have become the mainstream approach in this area. However, existing reviews provide limited systematic categorization and comparative analysis. In this paper, we classify existing GNN‐based attributed network embedding methods into six categories: graph convolution network (GCN)‐based methods, heterogeneous graph neural network‐based methods, graph autoencoder‐based methods, bidirectional encoder representations from transformers (BERT)‐based methods, hyper‐graph neural network (HGNN)‐based methods, and Bayesian graph neural network‐based methods. We not only summarize a large number of attributed net‐work embedding methods but also analyze and compare these methods. Additionally, we introduce some typical application scenarios in this field. Finally, we discuss the challenges and highlight several future research directions.
Jiangbo Qian, Yihong Dong
Concurr. Comput. Pract. Exp.5
2026 Regularized spatio-temporal weighted graph convolution network based on multimodal fusion for mild cognitive impairment detection
Yihong Dong, Yuehan Wu
Expert Syst. Appl.2
2025 Rethinking Repetition Problems of LLMs in Code Generation
abstract
With the advent of neural language models, the performance of code generation has been significantly boosted.However, the problem of repetitions during the generation process continues to linger.Previous work has primarily focused on content repetition, which is merely a fraction of the broader repetition problem in code generation.A more prevalent and challenging problem is structural repetition.In structural repetition, the repeated code appears in various patterns but possesses a fixed structure, which can be inherently reflected in grammar.In this paper, we formally define structural repetition and propose an efficient decoding approach called RPG, which stands for Repetition Penalization based on Grammar, to alleviate the repetition problems in code generation for LLMs.Specifically, RPG first leverages grammar rules to identify repetition problems during code generation, and then strategically decays the likelihood of critical tokens that contribute to repetitions, thereby mitigating them in code generation.To facilitate this study, we construct a new dataset CodeRepetEval to comprehensively evaluate approaches for mitigating the repetition problems in code generation.Extensive experimental results demonstrate that RPG substantially outperforms the best-performing baselines on CodeRepetEval dataset as well as HumanEval and MBPP benchmarks, effectively reducing repetitions and enhancing the quality of generated code. 1
Yihong Dong, Bin Gu 0006, Zhi Jin 0001, Ge Li 0001
ACL (1)1
2025 LLM-Powered Test Case Generation for Detecting Bugs in Plausible Programs
abstract
Kaibo Liu, Zhenpeng Chen, Yiyang Liu, Jie M. Zhang, Mark Harman, Yudong Han, Yun Ma, Yihong Dong, Ge Li, Gang Huang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Kaibo Liu, Zhenpeng Chen 0001, Jie Zhang 0050, Mark Harman, Yudong Han 0001, Yun Ma 0002, Yihong Dong, Ge Li 0001, Gang Huang 0001
ACL (1)8
2025 CodeDPO: Aligning Code Models with Self Generated and Verified Source Code
abstract
Kechi Zhang, Ge Li, Yihong Dong, Jingjing Xu, Jun Zhang, Jing Su, Yongfei Liu, Zhi Jin. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Kechi Zhang, Ge Li 0001, Yihong Dong, Yongfei Liu, Zhi Jin 0001
ACL (1)3
2025 BANER: Boundary-Aware LLMs for Few-Shot Named Entity Recognition
abstract
Despite the recent success of two-stage prototypical networks in few-shot named entity recognition (NER), challenges such as over/under-detected false spans in the span detection stage and unaligned entity prototypes in the type classification stage persist. Additionally, LLMs have not proven to be effective few-shot information extractors in general. In this paper, we propose an approach called Boundary-Aware LLMs for Few-Shot Named Entity Recognition to address these issues. We introduce a boundary-aware contrastive learning strategy to enhance the LLM’s ability to perceive entity boundaries for generalized entity spans. Additionally, we utilize LoRAHub to align information from the target domain to the source domain, thereby enhancing adaptive cross-domain classification capabilities. Extensive experiments across various benchmarks demonstrate that our framework outperforms prior methods, validating its effectiveness. In particular, the proposed strategies demonstrate effectiveness across a range of LLM architectures. The code and data are released on https://github.com/UESTC-GQJ/BANER.
Quanjiang Guo, Yihong Dong, Ling Tian, Zhao Kang 0001, Yu Zhang 0193
COLING2
2025 A Novel Audio-Visual Multimodal Semi-Supervised Model Based on Graph Neural Networks for Depression Detection
abstract
There is a significant correlation between depression, verbal behavior, and facial expressions. By analyzing patients’ audio and facial visuals, depression assessments can be conducted. However, existing work is predominantly based on single modalities. Additionally, acquiring a sufficient amount of labeled data in clinical settings is challenging and costly. To leverage multimodal audio-visual data while addressing the issue of lacking trainable labeled data, we propose an audiovisual multimodal semi-supervised depression detection model based on Graph Neural Networks (AVS-GNN). This model first extracts dual-modality temporal information from audio features and facial visual features of patients and obtains modality-specific high-level embedding representations through graph representation learning. Subsequently, it utilizes graph-based contrastive unsupervised learning to capture consistency information between pairs of unlabeled samples across different modalities and to facilitate cross-modal interactions. We specifically designed a hybrid weighted pseudo-labeling strategy to assign high-confidence pseudo-labels to unlabeled data and further retraining the model. Experiments on two depression datasets show that this model outperforms baseline methods across all evaluation metrics.
Chenjian Sun, Yihong Dong
ICASSP3
2025 Adaptive Cross-Variable Spectral Filtering for Time Series Forecasting
Zhi Cheng, Yihong Dong
ICIC (7)2
2025 UnCert-CoT: Uncertainty-Aware Chain-of-Thought for Code Generation with Large Language Model
Ge Li 0001, Jia Li 0012, Hong Mei 0001, Zhi Jin 0001, Yihong Dong, Qibin Zheng
ICIC (23)7
2025 ROCODE: Integrating Backtracking Mechanism and Program Analysis in Large Language Models for Code Generation
abstract
Large language models (LLMs) have achieved impressive performance in code generation recently, offering programmers revolutionary assistance in software development. However, due to the auto-regressive nature of LLMs, they are susceptible to error accumulation during code generation. Once an error is produced, LLMs can merely continue to generate the subsequent code conditioned on it, given their inability to adjust previous outputs. Existing LLM-based approaches typically consider post-revising after code generation, leading to the challenging resolution of accumulated errors and the significant wastage of resources. Ideally, LLMs should rollback and resolve the occurred error in time during code generation, rather than proceed on the basis of the error and wait for post-revising after generation. In this paper, we propose Rocode,which integrates the backtracking mechanism and program analysis into LLMs for code generation. Specifically, we employ program analysis to perform incremental error detection during the generation process. When an error is detected, the backtracking mechanism is triggered to priming rollback strategies and constraint regeneration, thereby eliminating the error early and ensuring continued generation on the correct basis. Experiments on multiple code generation benchmarks show that ROCODE can significantly reduce the errors generated by LLMs, with a compilation pass rate of 99.1 %. The test pass rate is relatively improved by up to 23.8% compared to the best baseline approach. Compared to the post-revising baseline, the token cost is reduced by 19.3%. Moreover, our approach is model-agnostic and achieves consistent improvements across nine representative LLMs.
Yihong Dong, Yongding Tao, Huanyu Liu 0001, Zhi Jin 0001, Ge Li 0001
ICSE2
2025 Line-level Semantic Structure Learning for Code Vulnerability Detection
abstract
Unlike the flow structure of natural languages, programming languages have an inherent rigidity in structure and grammar.However, existing detection methods based on pre-trained models typically treat code as a natural language sequence, ignoring its unique structural information.This hinders the models from understanding the code's semantic and structual information.To address this problem, we introduce the Code Structure-Aware Network through Line-level Semantic Learning (CSLS), which comprises four components: code preprocessing, global semantics awareness, line semantic awareness, and line semantic structure awareness.The preprocessing step transforms the code into two types of text: global code text and line-level code text.Unlike typical preprocessing methods, CSLS preserves structural elements such as line breaks and indentation characters while processing the global text.While preserving global code semantics, the CSLS network emphasizes capturing structural relationships between line semantics.By modeling each line's semantics, CSLS treats line-level semantics as the smallest structural unit to learn nonlinear structural relationships, thereby improving code vulnerability detection accuracy.We conducted extensive experiments on vulnerability detection datasets from real projects.Results show that our preprocessing method significantly enhances the performance of existing baseline models.Additionally, the CSLS model outperforms the state-of-the-art baselines in code vulnerability detection, achieving 70.57% accuracy on the Devign dataset and a 49.59% F1 score on the Reveal dataset.These results demonstrate the importance of preserving and utilizing code structure information to improve the performance of code vulnerability detection models.
Ge Li 0001, Jia Li 0011, Yihong Dong, Yingfei Xiong 0001, Zhi Jin 0001
Internetware4
2025 Aligning LLMs to Fully Utilize the Cross-file Context in Repository-level Code Completion
abstract
Large Language Models (LLMs) have shown promising results in repository-level code completion, which completes code based on the in-file and cross-file context of a repository. The cross-file context typically contains different types of information (e.g., relevant APIs and similar code) and is lengthy. In this paper, we found that LLMs struggle to fully utilize the information in the cross-file context. We hypothesize that one of the root causes of the limitation is the misalignment between pre-training (i.e., relying on nearby context) and repo-level code completion (i.e., frequently attending to long-range cross-file context).To address the above misalignment, we propose Code Long-context Alignment - CoLA, a purely data-driven approach to explicitly teach LLMs to focus on the cross-file context. Specifically, CoLA constructs a large-scale repo-level code completion dataset - CoLA-132K, where each sample contains the long cross-file context (up to 128K tokens) and requires generating context-aware code (i.e., cross-file API invocations and code spans similar to cross-file context). Through a two-stage training pipeline upon CoLA-132K, LLMs learn the capability of finding relevant information in the cross-file context, thus aligning LLMs with repo-level code completion. We apply CoLA to multiple popular LLMs (e.g., aiXcoder-7B) and extensive experiments on CoLA-132K and a public benchmark - CrossCodeEval. Our experiments yield the following results. ❶ Effectiveness. CoLA substantially improves the performance of multiple LLMs in repo-level code completion. For example, it improves aiXcoder-7B by up to 19.7% in exact match. ❷ Generalizability. The capability learned by CoLA can generalize to new languages (i.e., languages not in training data). ❸ Enhanced Context Utilization Capability. We design two probing experiments, which show CoLA improves the capability of LLMs in utilizing the information (i.e., relevant APIs and similar code) in cross-file context. Our datasets and model weights are released in [1].
Jia Li 0011, Huanyu Liu 0001, Xianjie Shi, He Zong, Yihong Dong, Kechi Zhang, Siyuan Jiang, Zhi Jin 0001, Ge Li 0001
ASE6
2025 Thinking Longer, Not Larger: Enhancing Software Engineering Agents via Scaling Test-Time Compute
abstract
Recent advancements in software engineering agents have demonstrated promising capabilities in automating program improvements. However, their reliance on closed-source or resource-intensive models introduces significant deployment challenges in private environments, prompting a critical question: How can personally deployable open-source LLMs (e.g., 32B models running on a single GPU) achieve comparable code reasoning performanceƒ To this end, we propose a unified Test-Time Compute (TTC) scaling framework that leverages increased inference-time computation instead of larger models. Our framework incorporates two complementary strategies: internal TTC and external TTC. Internally, we introduce a development-contextualized trajectory synthesis method leveraging real-world software repositories to bootstrap multi-stage reasoning processes, such as fault localization and patch generation. We further enhance trajectory quality through rejection sampling, rigorously evaluating trajectories along accuracy and complexity. Externally, we propose a novel development-process-based search strategy guided by reward models and execution verification. This approach enables targeted computational allocation at critical development decision points, overcoming limitations of existing "end-point only" verification methods.Evaluations on SWE-bench Verified demonstrate our 32B model achieves a 46% issue resolution rate, surpassing significantly larger models such as DeepSeek R1 671B and OpenAI o1. Additionally, we provide the empirical validation of the test-time scaling phenomenon within SWE agents, revealing that models dynamically allocate more tokens to increasingly challenging problems, effectively enhancing reasoning capabilities. We publicly release all training data, models, and code to facilitate future research.1. In fact, our method has been deployed in Tongyi Lingma, an IDE-based coding assistant developed by Alibaba Cloud, where it helps developers solve real-world programming problems.
Yingwei Ma, Yongbin Li 0001, Yihong Dong, Yanhao Li, Rongyu Cao, Jue Chen 0003, Fei Huang 0002, Binhua Li
ASE3
2025 Alzheimer's Disease Recognition Based on Adaptive Graph Normalization Flow for Incomplete Multimodal Data Fusion
Yihong Dong, Haihao Yan, Linlin Gao
MICCAI (8)2
2025 Reasoning is Periodicity? Improving Large Language Models Through Effective Periodicity Modeling
abstract
Periodicity, as one of the most important basic characteristics, lays the foundation for facilitating structured knowledge acquisition and systematic cognitive processes within human learning paradigms. However, the potential flaws of periodicity modeling in Transformer affect the learning efficiency and establishment of underlying principles from data for large language models (LLMs) built upon it. In this paper, we demonstrate that integrating effective periodicity modeling can improve the learning efficiency and performance of LLMs. We introduce FANformer, which adapts Fourier Analysis Network (FAN) into attention mechanism to achieve efficient periodicity modeling, by modifying the feature projection process of attention mechanism. Extensive experimental results on language modeling show that FANformer consistently outperforms Transformer when scaling up model size and training tokens, underscoring its superior learning efficiency. Our pretrained FANformer-1B exhibits marked improvements on downstream tasks compared to open-source LLMs with similar model parameters or training tokens. Moreover, we reveal that FANformer exhibits superior ability to learn and apply rules for reasoning compared to Transformer. The results position FANformer as an effective and promising architecture for advancing LLMs.
Yihong Dong, Ge Li 0001, Yongding Tao, Kechi Zhang, Lecheng Wang, Huanyu Liu 0001, Jiazheng Ding, Jia Li 0011, Jinliang Deng, Hong Mei 0001
NeurIPS1
2025 FAN: Fourier Analysis Networks
abstract
Despite the remarkable successes of general-purpose neural networks, such as MLPs and Transformers, we find that they exhibit notable shortcomings in modeling and reasoning about periodic phenomena, achieving only marginal performance within the training domain and failing to generalize effectively to out-of-domain (OOD) scenarios. Periodicity is ubiquitous throughout nature and science. Therefore, neural networks should be equipped with the essential ability to model and handle periodicity. In this work, we propose FAN, a novel neural network that effectively addresses periodicity modeling challenges while offering broad applicability similar to MLP with fewer parameters and FLOPs. Periodicity is naturally integrated into FAN's structure and computational processes by introducing the Fourier Principle. Unlike existing Fourier-based networks, which possess particular periodicity modeling abilities but face challenges in scaling to deeper networks and are typically designed for specific tasks, our approach overcomes this challenge to enable scaling to large-scale models and maintains the capability to be applied to more types of tasks. Through extensive experiments, we demonstrate the superiority of FAN in periodicity modeling tasks and the effectiveness and generalizability of FAN across a range of real-world tasks. Moreover, we reveal that compared to existing Fourier-based networks, FAN accommodates both periodicity modeling and general-purpose modeling well.
Yihong Dong, Ge Li 0001, Yongding Tao, Kechi Zhang, Jia Li 0011, Jinliang Deng
NeurIPS1
2025 SATURN: SAT-based Reinforcement Learning to Unleash LLMs Reasoning
abstract
How to design reinforcement learning (RL) tasks that effectively unleash the reasoning capability of large language models (LLMs) remains an open question. Existing RL tasks (e.g., math, programming, and constructing reasoning tasks) suffer from three key limitations: (1) Scalability. They rely heavily on human annotation or expensive LLM synthesis to generate sufficient training data. (2) Verifiability. LLMs' outputs are hard to verify automatically and reliably. (3) Controllable Difficulty. Most tasks lack fine-grained difficulty control, making it hard to train LLMs to develop reasoning ability from easy to hard. To address these limitations, we propose Saturn, a SAT-based RL framework that uses Boolean Satisfiability (SAT) problems to train and evaluate LLMs reasoning. Saturn enables scalable task construction, rule-based verification, and precise difficulty control. Saturn designs a curriculum learning pipeline that continuously improves LLMs' reasoning capability by constructing SAT tasks of increasing difficulty and training LLMs from easy to hard. To ensure stable training, we design a principled mechanism to control difficulty transitions. We introduce Saturn-2.6k, a dataset of 2,660 SAT problems with varying difficulty. It supports the evaluation of how LLM reasoning changes with problem difficulty. We apply Saturn to DeepSeek-R1-Distill-Qwen and obtain Saturn-1.5B and Saturn-7B. We achieve several notable results: (1) On SAT problems, Saturn-1.5B and Saturn-7B achieve average pass@3 improvements of +14.0 and +28.1, respectively. (2) On math and programming tasks, Saturn-1.5B and Saturn-7B improve average scores by +4.9 and +1.8 on benchmarks (e.g., AIME, LiveCodeBench). (3) Compared to the state-of-the-art (SOTA) approach in constructing RL tasks, Saturn achieves further improvements of +8.8\%. We release the source code, data, and models to support future research.
Huanyu Liu 0001, Ge Li 0001, Jia Li 0011, Kechi Zhang, Yihong Dong
NeurIPS6
2025 Recursive Transformer: Boosting Reasoning Ability with State Stack
abstract
The Transformer architecture has emerged as a landmark advancement within the broad field of artificial intelligence, effectively catalyzing the advent of large language models (LLMs). However, despite its remarkable capabilities and the substantial progress it has facilitated, the Transformer architecture still has some limitations. One such intrinsic limitation is its inability to effectively recognize regular expressions or deterministic context-free grammars. Standard Transformers lack an explicit mechanism for recursion and structured state transitions, which can hinder systematic generalization on nested and hierarchical patterns. Drawing inspiration from pushdown automata, which efficiently resolve deterministic context-free grammars using stacks, we equip layers with a differentiable stack and propose StackTrans with recursion to address the aforementioned issue within LLMs. Unlike previous approaches that modify the attention computation, StackTrans explicitly incorporates hidden state stacks between Transformer layers. This design maintains compatibility with existing frameworks like flash-attention. Specifically, our design features stack operations -- such as pushing and popping hidden states -- that are differentiable and can be learned in an end-to-end manner. Our comprehensive evaluation spans benchmarks for both Chomsky hierarchy and large-scale natural languages. Across these diverse tasks, StackTrans consistently outperforms standard Transformer models and other baselines. We have successfully scaled StackTrans up from 360M to 7B parameters. In particular, our from-scratch pretrained model StackTrans-360M outperforms several larger open-source LLMs with 2–3x more parameters, showcasing its superior efficiency and reasoning capability.
Kechi Zhang, Ge Li 0001, Jia Li 0012, Huangzhao Zhang, Yihong Dong, Jia Li 0011, Zhi Jin 0001
NeurIPS5
2025 A Novel Adaptive Graph Neural Network Based on Meta-Learning for Cross-Domain Few-Shot Learning
abstract
ABSTRACT In recent years, Graph Neural Network (GNN)‐based Few‐Shot Learning (FSL) methods have achieved significant performance. Such methods typically assume that base and novel classes share the same underlying distribution, which limits their practical applicability. Cross‐domain FSL has emerged as a promising direction to tackle real‐world distribution shifts. However, existing methods still suffer from significant domain shift, leading to severe degradation in generalization performance. To address this challenge, we introduce an enhanced Adversarial Feature Augmentation (AFA) module that, unlike prior methods which align only global feature distributions, explicitly accounts for class decision boundaries. Additionally, we propose a novel Meta‐Learning Adaptive Graph Neural Network (MLA‐GNN) framework to alleviate the limitations of scenario training in cross‐domain problems. This model consists of two key components: a meta network that adaptively generates task‐level parameters and a target network that receives the generated parameters, namely, the Task‐specific Graph Neural Network. The meta network encodes task contextual information and generates specific parameters for the target network, enabling the GNN to flexibly adjust classification boundaries across different tasks, thereby maintaining decision effectiveness on data from different domains. We train our model on the MiniImageNet dataset and test it on five few‐shot datasets. Extensive experiments demonstrate the superiority of the proposed method in cross‐domain few‐shot classification tasks.
Yuehan Wu, Jieyi Yang, Shiliang Lai, Yihong Dong
Concurr. Comput. Pract. Exp.4
2025 Multi-Temporal Granularity Concept Induction for semantically driven video summarization
Junren Huang, Jiangbo Qian, Yihong Dong
Expert Syst. Appl.4
2025 S2CA: Shared Concept Prototypes and Concept-level Alignment for text-video retrieval
Jiangbo Qian, Yihong Dong
Neurocomputing4
2025 Multiscale Spectral Augmentation for Graph Contrastive Learning for fMRI analysis to diagnose psychiatric disease
Yihong Dong, Shoubo Peng
Knowl. Based Syst.2
2025 An Adaptive Dual-channel Multi-modal graph neural network for few-shot learning
Jieyi Yang, Yihong Dong
Knowl. Based Syst.2
2025 CodeScore: Evaluating Code Generation by Learning Code Execution
abstract
A proper code evaluation metric (CEM) profoundly impacts the evolution of code generation, which is an important research field in NLP and software engineering. Prevailing match-based CEMs (e.g., BLEU, Accuracy, and CodeBLEU) suffer from two significant drawbacks. 1. They primarily measure the surface differences between codes without considering their functional equivalence. However, functional equivalence is pivotal in evaluating the effectiveness of code generation, as different codes can perform identical operations. 2. They are predominantly designed for the Ref-only input format. However, code evaluation necessitates versatility in input formats. Aside from Ref-only, there are NL-only and Ref and NL formats, which existing match-based CEMs cannot effectively accommodate. In this article, we propose CodeScore, a large language model (LLM)-based CEM, which estimates the functional correctness of generated code on three input types. To acquire CodeScore, we present UniCE, a unified code generation learning framework, for LLMs to learn code execution (i.e., learning PassRatio and Executability of generated code) with unified input. Extensive experimental results on multiple code evaluation datasets demonstrate that CodeScore absolutely improves up to 58.87% correlation with functional correctness compared to other CEMs, achieves state-of-the-art performance, and effectively handles three input formats.
Yihong Dong, Jiazheng Ding, Ge Li 0001, Zhuo Li 0013, Zhi Jin 0001
ACM Trans. Softw. Eng. Methodol.1
2024 Signal Transformer: Complex-Valued Attention and Meta-Learning for Signal Recognition
abstract
Deep neural networks have been shown as a class of useful tools for addressing signal recognition issues in recent years, especially for identifying the nonlinear feature structures of signals. However, this power of most deep learning techniques heavily relies on an abundant amount of training data, so the performance of classic neural nets decreases sharply when the number of training data samples is small or unseen data are presented in the testing phase. This calls for an advanced strategy, i.e., model-agnostic meta-learning (MAML), which can capture the invariant representation of the data samples or signals. In this paper, inspired by the special structure of the signal, i.e., real and imaginary parts consisted in practical time-series signals, we propose a Complex-valued Attentional MEta Learner (CAMEL) for few-shot signal recognition in the complex domain by leveraging attention and meta-learning. Experimental results showcase the superiority of the proposed CAMEL compared with the state-of-the-art methods.
Yihong Dong, Muqiao Yang, Songtao Lu, Qingjiang Shi
ICASSP2
2024 EvoCodeBench: An Evolving Code Generation Benchmark with Domain-Specific Evaluations
abstract
How to evaluate Large Language Models (LLMs) in code generation remains an open question. Many benchmarks have been proposed, but they have two limitations, i.e., data leakage and lack of domain-specific evaluation.The former hurts the fairness of benchmarks, and the latter hinders practitioners from selecting superior LLMs for specific programming domains.To address these two limitations, we propose a new benchmark - EvoCodeBench, which has the following advances: (1) Evolving data. EvoCodeBench will be dynamically updated every period (e.g., 6 months) to avoid data leakage. This paper releases the first version - EvoCodeBench-2403, containing 275 samples from 25 repositories.(2) A domain taxonomy and domain labels. Based on the statistics of open-source communities, we design a programming domain taxonomy consisting of 10 popular domains. Based on the taxonomy, we annotate each sample in EvoCodeBench with a domain label. EvoCodeBench provides a broad platform for domain-specific evaluations.(3) Domain-specific evaluations. Besides the Pass@k, we compute the Domain-Specific Improvement (DSI) and define LLMs' comfort and strange domains. These evaluations help practitioners select superior LLMs in specific domains and discover the shortcomings of existing LLMs.Besides, EvoCodeBench is collected by a rigorous pipeline and aligns with real-world repositories in multiple aspects (e.g., code distributions).We evaluate 8 popular LLMs (e.g., gpt-4, DeepSeek Coder, StarCoder 2) on EvoCodeBench and summarize some insights. EvoCodeBench reveals the actual abilities of these LLMs in real-world repositories. For example, the highest Pass@1 of gpt-4 on EvoCodeBench-2403 is only 20.74%. Besides, we evaluate LLMs in different domains and discover their comfort and strange domains. For example, gpt-4 performs best in most domains but falls behind others in the Internet domain. StarCoder 2-15B unexpectedly performs well in the Database domain and even outperforms 33B LLMs. We release EvoCodeBench, all prompts, and LLMs' completions for further community analysis.
Jia Li 0011, Ge Li 0001, Xuanming Zhang, Yunfei Zhao 0003, Yihong Dong, Zhi Jin 0001, Binhua Li, Fei Huang 0002, Yongbin Li 0001
NeurIPS5
2024 Swin transformer-based traffic video text tracking
Jinyao Yu, Jiangbo Qian, Chong Wang 0001, Yihong Dong
Appl. Intell.5
2024 Animation line art colorization based on the optical flow method
abstract
Abstract Coloring an animation sketch sequence is a challenging task in computer vision since the information contained in line sketches is too sparse, and the colors need to be uniform between continuous frames. Many the existing colorization algorithms can only be applied to one image and can be considered color filling algorithms. Such algorithms only provide a color result that fits within a reasonable range and can not be applied to the coloring of frame sequences. This paper proposes an end‐to‐end two‐stage optical flow colorization network to solve the animation frame sequence colorization problem. The first stage of the network finds the direction of the color pixel flow from the detail change between a given reference frame and the next frame of line artwork and then completes the initial coloring process. The second stage of the network performs color correction and clarifies the output of the first stage. Since our algorithm does not directly colorize the image but finds the path of the color change to colorize it, it ensures a consistent color space for the sequence frames after colorization. We conduct experiments on an animation dataset, and the results show that our algorithm is effective. The code is available at https://github.com/silenye/Colorization .
Jiangbo Qian, Chong Wang 0001, Yihong Dong, Baisong Liu
Comput. Animat. Virtual Worlds4
2024 A Graph Contrastive Learning Model Based on Structural and Semantic View for HIN Recommendation
abstract
Abstract With the rapid growth of information in the Internet era, people are in great need of recommendation methods to filter information. At present, recommendation methods which based on heterogeneous information network (HIN) have attracted wide attention. Recently, HIN-based recommendation methods need to be modeled from two aspects: node structural association and semantic association. To this end, we propose a graph contrastive learning model based on structural and semantic view for HIN recommendation (GCL-SS). GCL-SS utilizes U-I interactive view to obtain node structural embeddings, and utilizes U-I semantic view to obtain node semantic embeddings. Based on these two kinds of embeddings, we establish a self-supervised contrastive learning mechanism to effectively integrate structural information and semantic information of user (item) nodes in HIN, and finally learn a more discriminative user (item) embedding. In addition, in order to strengthen the semantic association between nodes, we innovatively utilize time sequence encoder (TSE), such as LSTM, to encode semantic homogeneous network decomposed by HIN in U-I semantic view. At last, based on the user and item embeddings, we adopt bilinear decoder to model the potential association between user and item, so as to realize rating prediction of user to item. The experimental results on three real datasets confirm that our GCL-SS model performs better than state-of-the-art recommendation methods in rating prediction task. In addition, the results of four ablation experiments indicate that our GCL-SS model can effectively improve the performance of rating prediction in recommendation.
Ruowang Yu, Yihong Dong, Jiangbo Qian
Neural Process. Lett.3
2024 Self-Collaboration Code Generation via ChatGPT
abstract
Although large language models (LLMs) have demonstrated remarkable code-generation ability, they still struggle with complex tasks. In real-world software development, humans usually tackle complex tasks through collaborative teamwork, a strategy that significantly controls development complexity and enhances software quality. Inspired by this, we present a self-collaboration framework for code generation employing LLMs, exemplified by ChatGPT. Specifically, through role instructions, (1) Multiple LLM agents act as distinct “experts,” each responsible for a specific subtask within a complex task; (2) Specify the way to collaborate and interact, so that different roles form a virtual team to facilitate each other’s work, ultimately the virtual team addresses code generation tasks collaboratively without the need for human intervention. To effectively organize and manage this virtual team, we incorporate software-development methodology into the framework. Thus, we assemble an elementary team consisting of three LLM roles (i.e., analyst, coder, and tester) responsible for software development’s analysis, coding, and testing stages. We conduct comprehensive experiments on various code-generation benchmarks. Experimental results indicate that self-collaboration code generation relatively improves 29.9–47.1% Pass@1 compared to the base LLM agent. Moreover, we showcase that self-collaboration could potentially enable LLMs to efficiently handle complex repository-level tasks that are not readily solved by the single LLM agent.
Yihong Dong, Zhi Jin 0001, Ge Li 0001
ACM Trans. Softw. Eng. Methodol.1
2024 Self-Planning Code Generation with Large Language Models
abstract
Although large language models (LLMs) have demonstrated impressive ability in code generation, they are still struggling to address the complicated intent provided by humans. It is widely acknowledged that humans typically employ planning to decompose complex problems and schedule solution steps prior to implementation. To this end, we introduce planning into code generation to help the model understand complex intent and reduce the difficulty of problem-solving. This paper proposes a self-planning code generation approach with large language models, which consists of two phases, namely planning phase and implementation phase. Specifically, in the planning phase, LLM plans out concise solution steps from the intent combined with few-shot prompting. Subsequently, in the implementation phase, the model generates code step by step, guided by the preceding solution steps. We conduct extensive experiments on various code-generation benchmarks across multiple programming languages. Experimental results show that self-planning code generation achieves a relative improvement of up to 25.4% in Pass@1 compared to direct code generation, and up to 11.9% compared to Chain-of-Thought of code generation. Moreover, our self-planning approach also enhances the quality of the generated code with respect to correctness, readability, and robustness, as assessed by humans.
Yihong Dong, Lecheng Wang, Qiwei Shang, Ge Li 0001, Zhi Jin 0001, Wenpin Jiao
ACM Trans. Softw. Eng. Methodol.2
2023 Antecedent Predictions Are More Important Than You Think: An Effective Method for Tree-Based Code Generation
abstract
Code generation focuses on automatically converting natural language (NL) utterances into code snippets. Sequence-to-tree (Seq2Tree) approaches are proposed for code generation with the aim of ensuring grammatical correctness of the generated code. These approaches generate subsequent Abstract Syntax Tree (AST) nodes based on the preceding predictions of AST nodes. However, existing Seq2Tree approaches tend to treat both antecedent predictions and subsequent predictions equally, which poses a challenge for models to produce accurate subsequent predictions if the antecedent predictions are incorrect under the constraints of the AST. Given this challenge, it is necessary to pay more attention to antecedent predictions compared to subsequent predictions. To this end, this paper proposes a novel and effective method, named Antecedent Prioritized (AP) Loss, which prioritizes antecedent predictions by leveraging the position information of the generated AST nodes. We design an AST-to-Vector (AST2Vec) method that maps AST node positions to two-dimensional vectors, thereby modeling the position information of AST nodes. To evaluate the effectiveness of our proposed loss, we implement and train an Antecedent Prioritized Tree-based code generation model called APT. Experiments on four benchmark datasets demonstrate that with better antecedent predictions and accompanying subsequent predictions, APT achieves significant improvements, indicating the superiority and generality of our proposed method.
Yihong Dong, Ge Li 0001, Zhi Jin 0001
ECAI1
2023 Seq2Seq or Seq2Tree: Generating Code Using Both Paradigms via Mutual Learning
abstract
Code generation aims to automatically generate the source code based on given natural language (NL) descriptions, which is of great significance for automated software development. Some code generation models follow a language model-based paradigm (LMBP) to generate source code tokens sequentially. Some others focus on deriving the grammatical structure by generating the program’s abstract syntax tree (AST), i.e., using the grammatical structure-based paradigm (GSBP). Existing studies are trying to generate code through one of the above two models. However, human developers often consider both paradigms: building the grammatical structure of the code and writing source code sentences according to the language model. Therefore, we argue that code generation should consider both GSBP and LMBP. In this paper, we use mutual learning to combine two classes of models to make the two different paradigms train together. To implement the mutual learning framework, we design alignment methods between code and AST. Under this framework, models can be enhanced through shared encoders and knowledge interaction in aligned training steps. We experiment on three Python-based code generation datasets. Experimental results and ablation analysis confirm the effectiveness of our approach. Our results demonstrate that considering both GSBP and LMBP is helpful in improving the performance of code generation.
Yunfei Zhao 0003, Yihong Dong, Ge Li 0001
Internetware2
2023 CODEP: Grammatical Seq2Seq Model for General-Purpose Code Generation
abstract
General-purpose code generation aims to automatically convert the natural language description to code snippets in a general-purpose programming language (GPL) such as Python. In the process of code generation, it is essential to guarantee the generated code satisfies grammar constraints of GPL. However, existing sequence-to-sequence (Seq2Seq) approaches neglect grammar rules when generating GPL code. In this paper, we devise a pushdown automaton (PDA)-based methodology to make the first attempt to consider grammatical Seq2Seq models for general-purpose code generation, exploiting the principle that PL is a subset of PDA recognizable language and code accepted by PDA is grammatical. Specifically, we construct a PDA module and design an algorithm to constrain the generation of Seq2Seq models to ensure grammatical correctness. Guided by this methodology, we further propose CODEP, a code generation framework equipped with a PDA module, to integrate the deduction of PDA into deep learning. This framework leverages the state of PDA deduction (including state representation, state prediction task, and joint prediction with state) to assist models in learning PDA deduction. To comprehensively evaluate CODEP, we construct a PDA for Python and conduct extensive experiments on four public benchmark datasets. CODEP can employ existing sequence-based models as base models, and we show that it achieves 100% grammatical correctness percentage on these benchmark datasets. Consequently, CODEP relatively improves 17% CodeBLEU on CONALA, 8% EM on DJANGO, and 15% CodeBLEU on JUICE-10K compared to base models. Moreover, PDA module also achieves significant improvements on the pre-trained models.
Yihong Dong, Ge Li 0001, Zhi Jin 0001
ISSTA1
2023 An Audio Correlation-Based Graph Neural Network for Depression Recognition
Chenjian Sun, Yihong Dong
PRCV (8)2
2023 Preference-corrected multimodal graph convolutional recommendation network
Xiangen Jia, Yihong Dong, Jiangbo Qian
Appl. Intell.2
2023 Fusing heterogeneous information for multi-modal attributed network embedding
Yang Jieyi, Zhu Feng, Yihong Dong, Jiangbo Qian
Appl. Intell.3
2023 Swin transformer-based supervised hashing
Liangkang Peng, Jiangbo Qian, Chong Wang 0001, Baisong Liu, Yihong Dong
Appl. Intell.5
2023 A gated graph attention network based on dual graph convolution for node embedding
Ruowang Yu, Lanting Wang, Jiangbo Qian, Yihong Dong
Appl. Intell.5
2023 A time sequence coding based node-structure feature model oriented to node classification
Ruowang Yu, Yihong Dong, Jiangbo Qian
Expert Syst. Appl.3
2023 A dual-path U-Net for pulmonary vessel segmentation method based on lightweight 3D attention
Rencheng Wu, Yihong Dong, Jiangbo Qian
Mach. Vis. Appl.3
2023 Multimodal heterogeneous graph attention network
Xiangen Jia, Min Jiang 0005, Yihong Dong, Haocai Lin, Huahui Chen 0001
Neural Comput. Appl.3
2022 Incorporating domain knowledge through task augmentation for front-end JavaScript code generation
abstract
Code generation aims to generate a code snippet automatically from natural language descriptions. Generally, the mainstream code generation methods rely on a large amount of paired training data, including both the natural language description and the code. However, in some domain-specific scenarios, building such a large paired corpus for code generation is difficult because there is no directly available pairing data, and a lot of effort is required to manually write the code descriptions to construct a high-quality training dataset. Due to the limited training data, the generation model cannot be well trained and is likely to be overfitting, making the model's performance unsatisfactory for real-world use. To this end, in this paper, we propose a task augmentation method that incorporates domain knowledge into code generation models through auxiliary tasks and a Subtoken-TranX model by extending the original TranX model to support subtoken-level code generation. To verify our proposed approach, we collect a real-world code generation dataset and conduct experiments on it. Our experimental results demonstrate that the subtoken-level TranX model outperforms the original TranX model and the Transformer model on our dataset, and the exact match accuracy of Subtoken-TranX improves significantly by 12.75% with the help of our task augmentation method. The model performance on several code categories has satisfied the requirements for application in industrial systems. Our proposed approach has been adopted by Alibaba's BizCook platform. To the best of our knowledge, this is the first domain code generation system adopted in industrial development environments.
Sijie Shen, Yihong Dong, Qizhi Guo, Yankun Zhen, Ge Li 0001
ESEC/SIGSOFT FSE3
2022 Mental Disorders Prediction with Heterogeneous Graph Convolutional Network
abstract
In the medical imaging field, Computer-Aided Detection (CADe) has greatly benefited from the recent development of Graph Convolutional Networks (GCNs). GCN-based predictive models require building a population graph to detect the disease states of each subject, based on imaging and non-imaging data. Until now, all existing population-level methods are homogeneous, failing to consider sex differences. To address this issue, we present a heterogeneous population graph convolutional network with hierarchical attention mechanisms, including intra-level and inter-level attention. Specifically, the intra-level attention layer is aimed at learning differences and similarities between the sexes, while the inter-level attention layer is responsible for information integration by assigning weights to different features. The objective is to obtain node embeddings describing individual characteristics completely and provide discriminative inputs to classifiers. Compared to benchmark models, our proposal achieves satisfying prediction results on three datasets, illustrating the framework’s ability to extract predictive attributes from medical multimodal data.
Haocai Lin, Jiacheng Pan, Yihong Dong
SMC3
2022 The deep fusion of topological structure and attribute information for anomaly detection in attributed networks
Jiangjun Su, Yihong Dong, Jiangbo Qian, Jiacheng Pan
Appl. Intell.2
2022 Distracted Driver Detection Based on a CNN With Decreasing Filter Size
abstract
In recent years, the number of traffic accident deaths due to distracted driving has been increasing dramatically. Fortunately, distracted driving can be detected by the rapidly developing deep learning technology. Nevertheless, considering that real-time detection is necessary, three contradictory requirements for an optimized network must be addressed: a small number of parameters, high accuracy, and high speed. We propose a new D-HCNN model based on a decreasing filter size with only 0.76M parameters, a much smaller number of parameters than that used by models in many other studies. D-HCNN uses HOG feature images, L2 weight regularization, dropout and batch normalization to improve the performance. We discuss the advantages and principles of D-HCNN in detail and conduct experimental evaluations on two public datasets, AUC Distracted Driver (AUCD2) and State Farm Distracted Driver Detection (SFD3). The accuracy on AUCD2 and SFD3 is 95.59% and 99.87%, respectively, higher than the accuracy achieved by many other state-of-the-art methods.
Binbin Qin, Jiangbo Qian, Baisong Liu, Yihong Dong
IEEE Trans. Intell. Transp. Syst.5
2021 Semi-Supervised Learning For Signal Recognition With Sparsity And Robust Promotion
abstract
Due to the emergence of deep learning, signal recognition has made great strides in performance improvement. The success of most deep learning methods relies on the accessibility of abundant labelled training data. However, the annotation of signals is quite expensive, making it challenging to train deep learning models substantially. This calls for the development of semi-supervised learning (SSL) method to fully utilize the unlabelled data to assist the training of deep learning models. To achieve this goal, three types of loss function tailored to the task of signal recognition are carefully designed in this paper. Together with the novel design of neural network structure, the proposed SSL method can effectively extract the information from unlabelled training data and thus overcome the difficulty of insufficient training. Extensive numerical results using real-world signal datasets are presented to show the remarkable performance of the proposed SSL method.
Yihong Dong, Lei Cheng 0003, Qingjiang Shi
WCNC1
2021 CapsNet-based supervised hashing
Jiangbo Qian, Xijiong Xie, Yihong Dong
Appl. Intell.5
2020 LTG-LSM: The Optimal Structure in LSM-tree Combined with Reading Hotness
abstract
A growing number of KV storage systems have adopted the Log-Structured-Merge-tree (LSM-tree) due to its excellent write performance. However, the high write amplification in the LSM-tree has always been a difficult problem to solve. The reason is that the design of traditional LSM-tree under-utilizes the data distribution of query, and the design space does not take into account the read and write performance concurrently. As a result, we may sacrifice one to improve another performance. When advancing the writing performance of the LSM-tree, we can only conservatively select the design pattern in the design space to reduce the impact on reading throughputs, resulting in limited improvement. Aiming at the shortcomings of existing methods, a new LSM-tree structure (Leveling-Tiering-Grouped-LSM-tree, LTG-LSM) is proposed by us that combined with reading hotness. The LTG-LSM structure maintains hotness prediction models at each level of the LSM-tree. The structure of the newly generated disk components is determined by the predicted hotness. Finally, a specific compaction algorithm is carried out to handle the compaction between the different structural components and processing workflow hotness changes. Experiments show that the scheme proposed by this paper significantly reduces the write amplification (up to about 71%) of the original LSM-tree with almost no sacrificing reading performance and improves the write throughputs (up to about 24%) in workflows with different configurations.
Jiaping Yu, Huahui Chen 0001, Jiangbo Qian, Yihong Dong
ICPADS4
2012 Enhancing Utility and Privacy-Safety via Semi-homogenous Generalization
Xianmang He, Wei Wang 0009, Huahui Chen 0001, Guang Jin, Yefang Chen, Yihong Dong
DEXA (1)6
2012 Clustering-Based k-Anonymity
Xianmang He, Huahui Chen 0001, Yefang Chen, Yihong Dong, Peng Wang 0027, Zhenhua Huang 0005
PAKDD (1)4