EDBT 2026 Demo / reviewers in the wild / expert
Chengru Wu
dblp:390/5507
· DBLP profile ↗
3ranked-venue papers
0as first author
3since 2021 · last 2025
0009-0002-0195-0684ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Software engineering, system software, and programming languages
1 paper |
Program synthesis and code generation · 77% Empirical software engineering · 23% |
Topics — the 2 heaviest of 2, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Program synthesis and code generation
code language model |
0.9 | 1 | 2025 | On the Applicability of Code Language Models to Scientific Computing Programs · IEEE Trans. Software Eng. 2025 |
Empirical software engineering › AI for software engineering
evaluation of language models for code |
0.3 | 1 | 2025 | On the Applicability of Code Language Models to Scientific Computing Programs · IEEE Trans. Software Eng. 2025 |
Methods — techniques the papers use, named apart from their topics
code language model · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CCUP: A Controllable Synthetic Data Generation Pipeline for Pretraining Cloth-Changing Person Re-Identification ModelsabstractDue to the high cost of constructing Cloth-changing person reidentification (CC-ReID) data, the existing data-driven models are hard to train efficiently on limited data, which causes the issue of overfitting. To address this challenge, we propose a low-cost and efficient pipeline specific to CC-ReID tasks for generating controllable and high-quality synthetic data simulating the surveillance scenarios. Particularly, we construct a new self-annotated CC-ReID dataset named Cloth-Changing Unreal Person (CCUP), containing 6,000 IDs, 1,179,976 images, 100 cameras, and 26.5 outfits per individual. Based on this large-scale dataset, we introduce an effective and scalable pretrain-finetune framework for enhancing the generalization of the traditional CC-ReID models. The extensive experimental results demonstrate that our framework could improve the original models such as two typical models TransReID and FIRe2after pretraining on CCUP and finetuning on a benchmark, and outperform other state-of-the-art models. The dataset is available at: https://github.com/yjzhao1019/CCUP. Yujian Zhao, Chengru Wu, Yinong Xu, Xuanzheng Du, Ruiyu Li, Guanglin Niu |
ICME | 2 |
| 2025 | AdaptiveLLM: A Framework for Selecting Optimal Cost-Efficient LLM for Code-Generation Based on CoT LengthabstractWhile Large Language Models (LLMs) have significantly advanced code generation efficiency, they face inherent challenges in balancing performance and inference costs across diverse programming tasks.Dynamically selecting the optimal LLM based on task difficulty and resource constraints offers a promising approach to achieve an optimal balance between efficiency and performance.However, existing model selection methods are resource-intensive and often neglect cost efficiency.Moreover, these approaches rely on human-annotated difficulty labels that are frequently inaccessible in real-world settings and may not align with the LLM's own assessment of task difficulty.In this paper, we introduce Adaptiv-eLLM, a framework that dynamically selects optimal LLMs for a given coding task by automatically assessing task difficulty.Our framework first estimates task difficulty using Chain-of-Thought lengths generated by reasoning model, clusters these into three difficulty levels via k-means, and fine-tunes CodeBERT to embed difficulty-aware features.A trained XGBoost classifier then selects the best model for each problem, optimizing the performance-cost trade-off.Experimental results show that AdaptiveLLM achieves a 7.86% improvement in pass@1 score while reducing resource consumption by 88.9% compared to baseline method ComplexityNet.When compared to a single model, AdaptiveLLM demonstrates an approximately 15% accuracy improvement, while maintaining the same level of cost consumption.Apart from that, the difficulty assessment using CoT provides more reliable selection criteria than human evaluation.Our replication package is available at https://github.com/cjhCoder7/AdaptiveLLM. Junhang Cheng, Fang Liu 0032, Chengru Wu, Li Zhang 0029 |
Internetware | 3 |
| 2025 | On the Applicability of Code Language Models to Scientific Computing ProgramsabstractScientific Computing Programming Languages (SCPLs), like MATLAB and R, are popular and widely used for computational mathematics. In recent years, pre-trained code language models (CLMs) have automated many code-related tasks, covering various general programming languages. SCPLs share many similarities with general programming languages, including similar syntactic structures and the semantics of identifiers. Despite the similarities, there exist many differences between them. For example, lots of numerical operations and dedicated libraries exist in SCPLs. However, there has been little comprehensive work analyzing CLMs’ capabilities in the understanding and generation of pragmatic scientific computing programs. To this end, we investigate the applicability of code language models for the SCPL analysis, especially focus on real-world code in open-source repositories. We first create a benchmark that contains programs and documentation from three widely used scientific computing programming languages, then perform an adequate evaluation of existing advanced code language models on both code understanding and generation tasks using the new benchmark, and study the relations of different training strategies, model types, and model sizes to the performance of different tasks and languages. Evaluation results confirm that, compared to general programming languages, SCPLs are more challenging to understand, and especially to generate, but the use of code language models is nevertheless feasible, and the knowledge obtained from the general languages can be transferred to SCPL analysis. A deeper analysis reveals additional challenges in generating code that incorporates API calls relevant to computational mathematics. We believe that our findings can provide guidance on improving tooling and analyses for the scientific programming language, and also inspire and motivate researchers to improve the robustness of existing code language models. Qianhui Zhao, Fang Liu 0032, Chengru Wu, Li Zhang 0029 |
IEEE Trans. Software Eng. | 4 |