VLDB 2026 Research / reviewers in the wild / expert
Peng Lan
dblp:73/8140
· DBLP profile ↗
14ranked-venue papers
2as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 5 · 4 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-step electric load forecasting based on improved dual decomposition and error correction strategy
Peng Lan, Yueyang Gong, Fenggang Sun |
J. Supercomput. | 2 |
| 2025 | FEDDGCN: A Frequency-Enhanced Decoupling Dynamic Graph Convolutional Network for Traffic Flow PredictionabstractAs a core task in Intelligent Transportation Systems (ITS), traffic flow prediction is essential for resource allocation and real-time route planning. Effectively capturing complex temporal correlations and dynamic spatial dependencies in traffic flow data is critical yet challenging for accurate prediction. However, existing approaches are still limited by the insufficient capability for spatial-temporal pattern decoupling and the underutilization of frequency domain information. To address these issues, we propose a novel Frequency-Enhanced Dynamic Decoupling Graph Convolutional Network (FEDDGCN), which introduces a gated decoupling mechanism integrating temporal and spatial embeddings to decouple traffic flow into prominent periodic and perturbative component. It also achieves effective pattern separation by incorporating frequency domain analysis with Fourier filters. Furthermore, a dual-branch spatial-temporal learning module, employing a divide-and-conquer strategy, is designed to achieve separate modeling for the two distinct components. Specially, the dynamic graph convolution modules are utilized to learn spatial dependencies and temporal and frequency attention mechanisms further capture complex temporal correlations for prominent periodic and perturbative components.Extensive experiments on multiple real-world datasets demonstrate that FEDDGCN achieves superior predictive performance compared with state-of-the-art methods. Ruobai Xiang, Zhifang Liao, Peng Lan, Qihao Liang |
CIKM | 4 |
| 2025 | SECMG: Style Enhancement Commit Message GenerationabstractThe commit message (CM) is a summary used in software development to understand code changes (Diff) quickly. A good CM helps improve the efficiency of software development and maintenance. Most existing methods generate CM based on Diff files, but these methods have some limitations. Traditional methods neglect to retrieve the highly similar semantic information that may be embedded in the Commit Message; they didn’t fully consider the similarities and differences between the retrieved CM and the query Diff. To better retrieve relevant commit messages and assist in CM generation, this paper proposes a Style Enhancement Commit Message Generation (SECMG) model. The model divides the CM generation task into two modules, including the Commit Message retrieval module and the Style-Enhancement CM generation module. To maximize the capture of highly relevant retrieval results, we employ a combination of semantic and keyword-based retrieval methods. Additionally, we optimize the retrieval outcomes by training a re-ranking model. In the generation module, we first augment the Diff with structural information at the word level via difflib. Then we use the CodeT5 model as a base to process the Diff and retrieve Commit Messages separately using Bi-Encoder, which are fused through a collaborative attention mechanism. Finally, we introduce a Gram-based style loss function to align the outputs and generate the commit message closest to the target message. We conducted extensive experiments on a real-world dataset, and the experimental results demonstrate the superiority of the present method over other models, improving by 17.50% and 8.73% on BLEU and B-Norm metrics, respectively. Zhifang Liao, Peng Lan |
IJCNN | 3 |
| 2025 | Cmdesum: A Cross-Modal Deliberation Network for Code SummarizationabstractThe task of Code Summarization aims to generate concise natural language descriptions for given source code snippets, thereby assisting developers in reducing the cognitive load of comprehending the code. Existing learning-based models, employing single-pass encoder-decoder frameworks, are unable to harness global information to optimize local content, while those utilizing multi-pass encoder-decoder architectures fail to consider the structural information of the code. To tackle this issue, we propose CMDeSum, a novel deliberation framework that injects cross-modal information in a staged manner, aiming to better balance the code sequence information and Abstract Syntax Tree (AST) structural information. Specifically, we first retrieve comments from code segments similar to the given one as drafts and extract method names and ASTs from the code. Then, in the First-Pass stage, we utilize code sequences and method names to generate initial comments and refine them based on the drafts. In the Second-Pass stage, building upon the results from the First-Pass stage, we utilize additional AST information to modify the comments, producing the final comments. To evaluate our approach, we conducted experiments on existing Java and Python datasets. The experimental results indicate that compared with the state-of-the-art models for code summarization generation, our model has improved by at least 6.3%, 3.0%, and 5.8% in BLEU, ROUGE-L, and METEOR. Zhifang Liao, Peng Lan |
ICPC | 3 |
| 2024 | DupLLM:Duplicate Pull Requests Detection Based on Large Language ModelabstractAs the scale of projects expands, the concurrent development model adopted by the open source community leads to an increasingly prominent problem of repetitive pull requests (PRs). The large number of rejections caused by duplicate pull requests increases the review workload of project maintainers and reduces the efficiency of pull request review. Therefore, it is very necessary to conduct automated duplicate PR detection. In this study, we propose DupLLM, a framework designed to detect duplicate PRs. The framework generates refined sum-maries by feeding the content of individual PRs into a large language model (LLM). Subsequently, the resulting summary is vectorized, converting the textual content into a numerical representation. The similarity between PRs is evaluated by calculating the similarity score between PR summary vectors. Ultimately, the model showed better performance than the best existing model, achieving an effect of 0.929 on P@l. This confirms that LLM can also achieve equivalent results in the field of duplicate PR detection as deep learning is used to train on this task, providing a new direction for the application of LLM in the field of software engineering. Zhifang Liao, Peng Lan |
APSEC | 3 |
| 2024 | Exploring the Potential of Large Language Models in Automatic Pull Request Title Generation: An Empirical StudyabstractPull Requests (PRs) are a collaborative mechanism in GitHub, allowing developers to merge their code changes into another branch of the software repository. The PR title serves as a summary of the PR and needs to accurately and concisely describe the specific changes made, which is useful for reviewers and other developers to review and understand. There are many existing methods for automatically generating PR titles, most of which are based on pre-trained models. Although these methods are effective, pre-trained models often require extensive fine-tuning for specific tasks. Compared to pre-trained models, large language models (LLMs) possess superior semantic understanding capabilities. As a foundational model, they can solve most tasks directly without relying on fine-tuning, providing an alternative solution for PR title generation. However, the capabilities of LLMs in the automatic PR title generation have not been fully explored. To fill this gap, we conducted an empirical study to understand the capabilities of LLMs in PR title generation. Initially, the direct application of LLMs to generate PR titles did not yield satisfactory results. We found that using similar PRs from the dataset as auxiliary information can effectively enhance the title generation capability of LLMs. When the number of most similar PRs used as input increased from 0 to 5, the ROUGE-L F1 score of the titles generated by LLMs increased by an average of 23.48 %, with improvements in other metrics as well. In further experiments, we discovered that setting a lower temperature for the LLMs can bring better performance. We then selected the best parameter configuration and compared it with the existing state-of-the-art methods. Our experimental results show that LLMs outperform the state of the art methods in Precision, Recall, and METEOR metrics on the PRTiger dataset. Additionally, human evaluation results indicate that PR titles generated by LLMs receive higher scores in Correctness, Naturalness, and Comprehensibility. Yitao Zuo, Peng Lan, Zhifang Liao |
APSEC | 2 |
| 2024 | Growing with the Help of Multiple Teachers: Lightweight and Noise-Resistant Student Model for Medical Image Classification
Yucheng Song, Jincan Wang, Yifan Ge, Zhifang Liao, Peng Lan, Lifeng Li |
PRCV (14) | 5 |
| 2023 | Radial-based undersampling approach with adaptive undersampling ratio determination
Peng Lan, Yunsheng Song, Shaomin Mu, Aifeng Li, Haiyan Chen 0001 |
Neurocomputing | 4 |
| 2022 | Walle: An End-to-End, General-Purpose, and Large-Scale Production System for Device-Cloud Collaborative Machine Learning
Chengfei Lv, Chaoyue Niu, Renjie Gu, Xiaotang Jiang, Zhaode Wang, Ziqi Wu, Qiulin Yao, Congyu Huang, Panos Huang, Hui Shu, Jinde Song, Peng Lan, Guohuan Xu, Fei Wu 0001, Shaojie Tang 0001, Fan Wu 0006, Guihai Chen |
OSDI | 15 |
| 2019 | A Recommendation of Crowdsourcing Workers Based on Multi-community Collaboration
Zhifang Liao, Peng Lan, Yan Zhang 0047 |
ICSOC | 3 |
| 2018 | Optimal power allocation for bi-directional full duplex underlay cognitive radio networksabstractCognitive radio and full duplex transmissions are regarded as two effective techniques to improve the spectral efficiency. By combining the two promising techniques, the problem of power allocation for bi‐directional full duplex underlay cognitive spectrum sharing is studied. To protect the primary transmission, consider both the instantaneous and statistical channel state information to adjust the transmit powers of secondary users to avoid causing intolerable interference to the primary users. The corresponding quality‐of‐service (QoS) constraints are determined for the primary system. On this basis, propose optimal power allocation schemes to facilitate the secondary transmission while guaranteeing the primary QoS constraint. Simulation results are provided to verify the efficiency of proposed schemes in improving the system capacity. Peng Lan, Chao Zhai 0001, Lizhen Chen, Bin Gao 0004, Fenggang Sun |
IET Commun. | 1 |
| 2018 | Real-valued DOA estimation with unknown number of sources via reweighted nuclear norm minimization
Fenggang Sun, Qihui Wu 0001, Peng Lan, Guoru Ding, Lizhen Chen |
Signal Process. | 3 |
| 2006 | On the Performance of a New Antenna Selection Algorithm Based on Orthogonal ComponentsabstractAlong with the multiple-input multiple-output (MIMO) system gains, comes a price in hardware complexity and cost due to multiple RF chains. It is possible to alleviate this cost and at the same time maintain many advantages of MIMO systems by a technique known as antenna selection. In this paper, we propose a new receive antenna selection algorithm which is based on the orthogonal components of the rows of channel matrix. It can achieve very nearly outage capacity with the optimal selection method while keep low computational complexity. Besides, the bit error rate (BER) performance under space-time block coding (STBC) transmit scheme demonstrates the significant performance of the proposed selection algorithm Peng Lan, Hongji Xu |
PIMRC | 1 |
| 2006 | Receive Antenna Subsets Selection Based on Orthogonal CompontsabstractMultiple-antenna wireless communication systems have recently attracted significant attention due to their higher capacity and better immunity to fading as compared to systems that employ single-sensor transceiver. Increasing the number of transmit and receive antennas enables to improve system performance but at the price of higher hardware cost and computational burden due to the multiple RF chains. Antenna selection is a low-cost low-complexity alternative to capture many of the advantages of MIMO systems. In this paper, we propose a new antenna selection algorithm with low complexity for wireless multiple-input multiple-output (MIMO) systems. The new algorithm is based on the orthogonal components of the rows of channel matrix. It achieves almost the same outage capacity as the optimal selection method while keeping lower computational complexity. The simulation results for quasi-static Rayleigh flat fading channel demonstrate the significant performance of the proposed selection algorithm Peng Lan, Hongji Xu |
PIMRC | 2 |