EDBT 2026 Demo / reviewers in the wild / expert
Chen Lyu 0001
dblp:132/7875-1
· DBLP profile ↗
37ranked-venue papers
1as first author
35since 2021 · last 2026
0000-0002-5044-1459ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 16 since 2021Software engineering, systems software and programming languages · 13 · 1 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SkelDPO: A Skeleton-Guided Direct Preference Optimization Framework for Efficient Code Generation
Chen Lyu 0001 |
ICPC | 2 |
| 2026 | Defense against unauthorized distillation in image restoration via feature space perturbation
Zhuoran Zheng, Chen Lyu 0001 |
Neurocomputing | 3 |
| 2026 | Label distribution learning via implicit distribution representation
Zhuoran Zheng, Xin Su 0009, Chen Lyu 0001 |
Neurocomputing | 4 |
| 2025 | SemGuard: Real-Time Semantic Evaluator for Correcting LLM-Generated CodeabstractLarge Language Models (LLMs) can translate natural language requirements into code, yet empirical analyses of representative models reveal that semantic errors—programs that compile but behave incorrectly—constitute the majority of observed faults (e.g., >60% on DeepSeek-Coder-6.7B and QwenCoder-7B). Post-hoc repair pipelines detect such faults only after execution, incurring latency, relying on incomplete test suites, and often mis-localizing the defect. Since semantic drift originates in the autoregressive decoding process, intervening while the code is being generated is a direct way to stop error propagation. Constrained-decoding approaches such as ROCODE attempt this, but still wait until the entire program runs to obtain feedback and use entropy heuristics that do not truly capture semantics. A more effective solution must inject semantic signals—early and precisely—into the decoding process. We present SemGuard, a semantic-evaluator-driven framework that performs real-time, line-level semantic supervision. To train the evaluator, we build SemDiff, the first dataset with fine-grained annotations that mark the exact line where a correct and an incorrect implementation diverge. The evaluator, once embedded in the LLM’s decoder, flags deviations on partial code, rolls back to the faulty line, and guides regeneration—without executing the program or requiring test cases. Across four benchmarks, SemGuard consistently outperforms state-of-the-art baselines. It lowers the semantic error rate by 19.86% on SemDiff relative to ROCODE, and lifts Pass@1 by 48.92% on the realworld LiveCodeBench with CodeLlama-7B. Similar gains hold for StarCoder2-7B on MBPP and for DeepSeekCoder-6.7B on the Java benchmark SemDiff-Java, demonstrating model- and language-agnostic effectiveness. Ruyun Wang, Zhi Jin 0001, Ge Li 0001, Chen Lyu 0001 |
ASE | 7 |
| 2025 | Multi-stage guided code generation for Large Language Models
Yewei Han, Chen Lyu 0001 |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | E-code: Mastering efficient code generation through pretrained models and expert encoder group
Yue Pan 0011, Chen Lyu 0001, Lantian Li, Xiuting Shao |
Inf. Softw. Technol. | 2 |
| 2025 | Measuring code efficiency optimization capabilities with ACEOB
Yue Pan 0011, Xiuting Shao, Chen Lyu 0001 |
J. Syst. Softw. | 3 |
| 2025 | Physics prior-based contrastive learning for low-light image enhancement
Hongxiang Liu, Yunliang Zhuang, Chen Lyu 0001 |
Signal Process. Image Commun. | 3 |
| 2024 | Enhancing Code Generation Performance of Smaller Models by Distilling the Reasoning Ability of LLMsabstractLarge Language Models (LLMs) have recently made significant advances in code generation through the ‘Chain-of-Thought’ prompting technique. This technique empowers the model to autonomously devise “solution plans” to tackle intricate programming challenges, thereby improving its performance in code generation. Nevertheless, smaller models have been struggling to keep up with LLMs in deducing these plans, adversely affecting their code generation capabilities. Given the considerable size and associated deployment costs, along with concerns about data security, many teams opt for deploying smaller models for code generation. Consequently, there arises a compelling need for transferring LLMs’ code generation reasoning abilities to the smaller models. In this paper, we propose the CodePLAN framework, which aims to transfer LLMs’ reasoning capabilities to smaller models through distillation. We adopt a multi-task learning approach, jointly undertaking code generation and solution plan generation tasks, to enhance the code generation capabilities of smaller model. To ensure the superior quality of the solution plans, we advocate for the utilization of backward reasoning and plan sampling strategies. Our experiments show that in comparison to the conventional fine-tuning approach, our approach improves the smaller model’s code generation performance (measured in pass@1 metric) by over 130% on the challenging APPS benchmark. Chen Lyu 0001, Yao Wan 0001, Hongyu Zhang 0002, Ge Li 0001, Zhi Jin 0001 |
LREC/COLING | 2 |
| 2024 | Knowledge-Aware Code Generation with Large Language ModelsabstractLarge Language Models (LLMs) perform well on basic programming problems. However, they encounter challenges when dealing with complex tasks involving the use of diverse algorithmic and data structure skills, particularly programming competition-level problems. Notably, ChatGPT exhibits proficient performance on problems it has encountered during its pre-training phase, but this performance deteriorates when faced with novel problems. Consequently, enhancing the ability of LLMs to address unfamiliar problems has emerged as a pivotal research focus. The problem-solving process of LLMs mirrors human programmers' approach to a certain extent. When confronted with new programming tasks, human programmers engage in task planning and code writing with the previously acquired knowledge about algorithms and data structures. Despite having learned such knowledge, LLMs struggle to effectively apply it when faced with specific new problems. To address this issue, we constructed a novel dataset, CodeF, which contains a portion of programming problems that ChatGPT has not previously encountered. Furthermore, we developed a Knowledge Library tailored for Python programming contest problems and introduced the concept of Knowledge-Aware Code Generation (KareCoder). KareCoder bolsters the models' understanding and problem-solving capabilities by integrating prompt and knowledge from the library into the LLMs' code generation reasoning process, especially on Pass@1 metrics. Upon testing on the CodeF and APPS datasets, KareCoder demonstrated outstanding performance in handling novel problems previously unencountered by LLMs. In contrast with the code directly generated by ChatGPT, KareCoder achieved a relative improvement of 23.3% on the Pass@1 metric on the CodeF post2021-9 dataset. Additionally, it performs well compared to other methods when dealing with problems that LLMs have previously encountered. Our dataset and experiment data are open-sourced and can be accessed at https://github.com/CodeGeneration3/KareCoder. Zhi Jin 0001, Ge Li 0001, Chen Lyu 0001 |
ICPC | 5 |
| 2024 | Sifting through the Chaff: On Utilizing Execution Feedback for Ranking the Generated Code CandidatesabstractLarge Language Models (LLMs), such as GPT-4, StarCoder, and Code Llama, are transforming the way developers approach programming by automatically generating code based on given contexts, such as natural language descriptions or incomplete surrounding code. Despite advancements, generating syntactically and semantically correct code remains challenging, especially for complex programming tasks. Existing approaches typically generate multiple candidate solutions using LLMs to increase the likelihood of producing correct code. However, selecting the correct code from these candidates --- a process known as code ranking --- remains a major challenge. Current research on code ranking can be categorized into execution-based and non-execution-based methods. Execution-based methods, although effective, encounter notable limitations, such as scarcity of quality unit tests and security risks. Non-execution-based methods like CodeRanker, which rely solely on classification labels to train a code ranker, struggle to capture subtle errors and provide detailed error insights. Recognizing the strengths and limitations of both approaches, we propose a new method that integrates the advantages of execution-based and non-execution-based techniques. The key insight of our work is that an effective code ranker is expected to truly comprehend the underlying causes of erroneous code, as relying solely on classification labels is insufficient. Inspired by this, this paper puts forward RankEF, an innovative approach for code ranking that leverages execution feedback. RankEF employs multi-task learning to integrate code classification with execution feedback generation. This approach enables the model to understand the reasons behind incorrect code, distinguishing between correct and incorrect solutions without the need to execute the code during the ranking phase. Experiments on three code generation benchmarks---APPS, MBPP, and HumanEval---demonstrate that RankEF significantly outperforms the state-of-the-art CodeRanker, achieving relative improvements of +30.97%, +31.43%, and +19.51% in Pass@1, Pass@2, and Pass@5 on APPS test, respectively. Yao Wan 0001, Jia Li 0012, Hongyu Zhang 0002, Zhi Jin 0001, Ge Li 0001, Chen Lyu 0001 |
ASE | 7 |
| 2024 | Semantic-aware enhancement: Integrating semantic compensation with 3-Dimensional Lookup Tables for low-light image enhancement
Weizhi Xu 0001, Chen Lyu 0001 |
Eng. Appl. Artif. Intell. | 3 |
| 2024 | Single UHD image dehazing via Interpretable Pyramid Network
Boxue Xiao, Zhuoran Zheng, Yunliang Zhuang, Chen Lyu 0001, Xiuyi Jia |
Signal Process. | 4 |
| 2024 | Dimensional Transformation Mixer for Ultra-High-Definition Industrial Camera DehazingabstractHaze severely affects the reliability of vision-based industrial systems, and most of the current dehazing methods are not applicable to images captured by ultra-high-definition (UHD) industrial cameras. In this article, we propose a novel dimensional transformation mixer (DMixer) model for recovering haze-free images from UHD haze images. In DMixer, the dimensional transformation module encodes the complete image in multiple stages from different perspectives and associates features from different views by permuting the tensor for efficient long-range dependency modeling. In this way, the global perception capabilities of DMixer are complemented, allowing the quality of the reconstructed UHD images to be improved. Furthermore, DMixer employs a dual-stream network framework that combines local and multiscale features, allowing DMixer to better trade off the performance and efficiency for real industrial systems. Our proposed method enables the real-time processing of UHD images ($\sim$59 fps). Extensive results show that the proposed model outperforms current state-of-the-art methods, with PSNR improvements of 1.94 and 0.95 dB on the 4KID and O-HAZE datasets, respectively. Yunliang Zhuang, Zhuoran Zheng, Lei Lyu 0001, Xiuyi Jia, Chen Lyu 0001 |
IEEE Trans. Ind. Informatics | 6 |
| 2023 | A Lightweight Multi-Scale Large Kernel Attention Hierarchical Network for Single Image Deraining
Xin Wang 0156, Chen Lyu 0001 |
ICANN (2) | 2 |
| 2023 | Measuring Efficient Code Generation with GECabstractAlthough efficiency is one of the core metrics in programming, recent large-scale language models often face the issue of “inefficient code” generation, which struggles to meet the real-time requirements of algorithms. However, there is relatively little research on evaluating the selection of efficient algorithms, and it is not easy to rigorously assess a model’s ability to correctly choose efficient algorithm solutions. Furthermore, the selection of efficient algorithm solutions often relies on the appropriate application of problem-solving skills, necessitating more in-depth research on algorithm reasoning. To address this challenge, we introduce the Generation of Efficient Code (GEC) benchmark, which aims to evaluate the ability to select efficient algorithm solutions. Unlike code generation, our benchmark focuses on a model’s ability to generate satisfactory efficient code when given a natural language description and inefficient code. We propose two novel metrics to examine the efficiency of the generated code and assess the model’s ability to generate efficient code. Our benchmark includes 3,712 problems, 31,577 combinations of efficient and inefficient code pairs, and 13,092 alternative efficient codes. We evaluate the performance of mainstream code generation models on the GEC benchmark. As the societal importance of code efficiency increases in the coming years, our benchmark will provide an essential measurement standard for tracking research progress. Our dataset and models are open-source and can be accessed at https://github.com/CodeGeneration2/Efficient-Code-Generation-with-GEC. Yue Pan 0011, Chen Lyu 0001 |
Internetware | 2 |
| 2023 | EnCoSum: enhanced semantic features for multi-scale multi-modal source code summarization
Yuexiu Gao, Hongyu Zhang 0002, Chen Lyu 0001 |
Empir. Softw. Eng. | 3 |
| 2023 | ADGSC: video anomaly detection algorithm based on graph structure change detection in public places
Huaiying Jiang, Chen Lyu 0001, Yuexiu Gao, Yunliang Zhuang, Sanjun Du |
Multim. Tools Appl. | 2 |
| 2022 | MMF3: Neural Code Summarization Based on Multi-Modal Fine-Grained Feature FusionabstractBackground: Code summarization automatically generates the corresponding natural language descriptions according to the input code to characterize the function implemented by source code. Comprehensiveness of code representation is critical to code summarization task. However, most existing approaches typically use coarse-grained fusion methods to integrate multi-modal features. They generally represent different modalities of a piece of code, such as an Abstract Syntax Tree (AST) and a token sequence, as two embeddings and then fuse the two ones at the AST/code levels. Such a coarse integration makes it difficult to learn the correlations between fine-grained code elements across modalities effectively. Aims: This study intends to improve the model’s prediction performance for high-quality code summarization by accurately aligning and fully fusing semantic and syntactic structure information of source code at node/token levels. Method: This paper proposes a Multi-Modal Fine-grained Feature Fusion approach (MMF3) for neural code summarization. The method uses the Transformer architecture. In particular, we introduce a novel fine-grained fusion method, which allows fine-grained fusion of multiple code modalities at the token and node levels. Specifically, we use this method to fuse information from both token and AST modalities and apply the fused features to code summarization. Results: We conduct experiments on one Java and one Python datasets, and evaluate generated summaries using four metrics. The results show that: 1) the performance of our model outperforms the current state-of-the-art models, and 2) the ablation experiments show that our proposed fine-grained fusion method can effectively improve the accuracy of generated summaries. Conclusion: MMF3 can mine the relationships between cross-modal elements and perform accurate fine-grained element-level alignment fusion accordingly. As a result, more clues can be provided to improve the accuracy of the generated code summaries. Yuexiu Gao, Lei Lyu 0001, Chen Lyu 0001 |
ESEM | 4 |
| 2022 | M2TS: multi-scale multi-modal approach based on transformer for source code summarizationabstractSource code summarization aims to generate natural language descriptions of code snippets. Many existing studies learn the syntactic and semantic knowledge of code snippets from their token sequences and Abstract Syntax Trees (ASTs). They use the learned code representations as input to code summarization models, which can accordingly generate summaries describing source code. Traditional models traverse ASTs as sequences or split ASTs into paths as input. However, the former loses the structural properties of ASTs, and the latter destroys the overall structure of ASTs. Therefore, comprehensively capturing the structural features of ASTs in learning code representations for source code summarization remains a challenging problem to be solved. In this paper, we propose M2TS, a Multi-scale Multi-modal approach based on Transformer for source code Summarization. M2TS uses a multi-scale AST feature extraction method, which can extract the structures of ASTs more completely and accurately at multiple local and global levels. To complement missing semantic information in ASTs, we also obtain code token features, and further combine them with the extracted AST features using a cross modality fusion method that not only fuses the syntactic and contextual semantic information of source code, but also highlights the key features of each modality. We conduct experiments on two Java and one Python datasets, and the experimental results demonstrate that M2TS outperforms current state-of-the-art methods. We release our code at https://github.com/TranSMS/M2TS. Yuexiu Gao, Chen Lyu 0001 |
ICPC | 2 |
| 2022 | HELoC: hierarchical contrastive learning of source code representationabstractAbstract syntax trees (ASTs) play a crucial role in source code representation. However, due to the large number of nodes in an AST and the typically deep AST hierarchy, it is challenging to learn the hierarchical structure of an AST effectively. In this paper, we propose HELoC, a hierarchical contrastive learning model for source code representation. To effectively learn the AST hierarchy, we use contrastive learning to allow the network to predict the AST node level and learn the hierarchical relationships between nodes in a self-supervised manner, which makes the representation vectors of nodes with greater differences in AST levels farther apart in the embedding space. By using such vectors, the structural similarities between code snippets can be measured more precisely. In the learning process, a novel GNN (called Residual Self-attention Graph Neural Network, RSGNN) is designed, which enables HELoC to focus on embedding the local structure of an AST while capturing its overall structure. HELoC is self-supervised and can be applied to many source code related downstream tasks such as code classification, code clone detection, and code clustering after pre-training. Our extensive experiments demonstrate that HELoC outperforms the state-of-the-art source code representation models. Hongyu Zhang 0002, Chen Lyu 0001, Zhuoran Zheng, Lei Lyu 0001, Songlin Hu 0001 |
ICPC | 4 |
| 2022 | Context-aware pyramid attention network for crowd counting
Lingyu Gu, Chen Pang 0001, Yanjun Zheng, Chen Lyu 0001, Lei Lyu 0001 |
Appl. Intell. | 4 |
| 2022 | GT-SimNet: Improving code automatic summarization via multi-modal similarity networks
Xuejian Gao, Chen Lyu 0001, Lei Lyu 0001 |
J. Syst. Softw. | 5 |
| 2022 | Adaptive Multiview Graph Difference Analysis for Video SummarizationabstractAdapting detection to different shot types is a significant challenge for video summarization methods based on shot boundary detection. In our recent work, a new graph model was introduced in the feature modelling of frames and analysed for changes in graph structure to improve the detection of shot boundaries. In this paper, we further explore the potential of graph models and propose a more general framework for online, real-time automatic video summarization. The framework develops a novel adaptive multiview graph difference analysis method to improve the algorithm’s robustness in detecting different shot transitions. Previous fusion methods typically used a priori knowledge to assign weights to the various feature differences from videos. In contrast, our framework can weigh and fuse the resulting differences by learning the importance of various video features from the structural changes of the corresponding multiview graphs. Additionally, we propose a new threshold-based adaptive decision method which can dynamically select the most accurate shot boundary decision threshold by analysing a small number of historical frames and learning the tolerance factor in the current shot. The experimental results show that the proposed method outperforms state-of-the-art methods in terms of precision and F-score on the VSUMM and YouTube datasets. Caixia Ma, Lei Lyu 0001, Guoliang Lu, Chen Lyu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2021 | Multi-modal Code Summarization Fusing Local API Dependency Graph and AST
Xuejian Gao, Chen Lyu 0001, Lei Lyu 0001 |
ICONIP (5) | 5 |
| 2021 | MPANet: Multi-level Progressive Aggregation Network for Crowd Counting
Run Han, Chen Pang 0001, Chunmeng Kang, Chen Lyu 0001, Lei Lyu 0001 |
ICONIP (3) | 5 |
| 2021 | Saliency Detection Framework Based on Deep Enhanced Attention Network
Xing Sheng 0001, Zhuoran Zheng, Chunmeng Kang, Yunliang Zhuang, Lei Lyu 0001, Chen Lyu 0001 |
ICONIP (4) | 7 |
| 2021 | Code Representation Based on Hybrid Graph Modelling
Zhuoran Zheng, Xuejian Gao, Chen Lyu 0001, Lei Lyu 0001 |
ICONIP (5) | 5 |
| 2021 | HIANet: Hierarchical Interweaved Aggregation Network for Crowd Counting
Jinyang Xie, Jinfang Zheng, Lingyu Gu, Chen Lyu 0001, Lei Lyu 0001 |
ICONIP (6) | 4 |
| 2021 | SS-CCN: Scale Self-guided Crowd Counting Network
Jinfang Zheng, Jinyang Xie, Chen Lyu 0001, Lei Lyu 0001 |
ICONIP (4) | 3 |
| 2021 | A Dueling-DDPG Architecture for Mobile Robots Path Planning Based on Laser Range Findings
Panpan Zhao, Jinfang Zheng, Qinglin Zhou, Chen Lyu 0001, Lei Lyu 0001 |
PRICAI (1) | 4 |
| 2021 | GCMNet: Gated Cascade Multi-scale Network for Crowd Counting
Jinfang Zheng, Panpan Zhao, Jinyang Xie, Chen Lyu 0001, Lei Lyu 0001 |
PRICAI (2) | 4 |
| 2021 | TreeBERT: A tree-based pre-trained model for programming languageabstractSource code can be parsed into the abstract syntax tree (AST) based on defined syntax rules. However, in pre-training, little work has considered the incorporation of tree structure into the learning process. In this paper, we present TreeBERT, a tree-based pre-trained model for improving programming language-oriented generation tasks. To utilize tree structure, TreeBERT represents the AST corresponding to the code as a set of composition paths and introduces node position embedding. The model is trained by tree masked language modeling (TMLM) and node order prediction (NOP) with a hybrid objective. TMLM uses a novel masking strategy designed according to the tree’s characteristics to help the model understand the AST and infer the missing semantics of the AST. With NOP, TreeBERT extracts the syntactical structure by learning the order constraints of nodes in AST. We pre-trained TreeBERT on datasets covering multiple programming languages. On code summarization and code documentation tasks, TreeBERT outperforms other pre-trained models and state-of-the-art models designed for these tasks. Furthermore, TreeBERT performs well when transferred to the pre-trained unseen programming language. Zhuoran Zheng, Chen Lyu 0001, Lei Lyu 0001 |
UAI | 3 |
| 2021 | Embedding API dependency graph for neural code generation
Chen Lyu 0001, Ruyun Wang, Hongyu Zhang 0002, Hanwen Zhang 0014, Songlin Hu 0001 |
Empir. Softw. Eng. | 1 |
| 2021 | Graph-based structural difference analysis for video summarization
Chunlei Chai, Guoliang Lu, Ruyun Wang, Chen Lyu 0001, Lei Lyu 0001, Peng Zhang 0009, Hong Liu 0013 |
Inf. Sci. | 4 |
| 2019 | Whale optimized mixed kernel function of support vector machine for colorectal cancer diagnosis
Hong Liu 0013, Yuanjie Zheng, Dianjie Lu, Chen Lyu 0001 |
J. Biomed. Informatics | 6 |
| 2018 | Key-frame selection for automatic summarization of surveillance videos: a method of multiple change-point detection
Guoliang Lu, Chen Lyu 0001, Peng Yan 0001 |
Mach. Vis. Appl. | 3 |