Wenxin Mao

dblp:193/7244 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
3since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Data Dependency-Aware Code Generation from Enhanced UML Sequence Diagrams
abstract
Large language models (LLMs) excel at generating code from natural language (NL) descriptions. However, the plain textual descriptions are inherently ambiguous and often fail to capture complex requirements like intricate system behaviors, conditional logic, and architectural constraints; implicit data dependencies in service-oriented architectures are difficult to infer and handle correctly.To bridge this gap, we propose a novel step-by-step code generation framework named UML2Dep by leveraging unambiguous formal specifications of complex requirements. First, we introduce an enhanced Unified Modeling Language (UML) sequence diagram tailored for service-oriented architectures. This diagram extends traditional visual syntax by integrating decision tables and API specifications, explicitly formalizing structural relationships and business logic flows in service interactions to rigorously eliminate linguistic ambiguity. Second, recognizing the critical role of data flow, we introduce a dedicated data dependency inference (DDI) task. DDI systematically constructs an explicit data dependency graph prior to actual code synthesis. To ensure reliability, we formalize DDI as a constrained mathematical reasoning task through novel prompting strategies, aligning with LLMs’ excellent mathematical strengths. Additional static parsing and dependency pruning further reduce context complexity and cognitive load associated with intricate specifications, thereby enhancing reasoning accuracy and efficiency.Experimental results on our in-house industrial datasets demonstrate the effectiveness of the proposed framework. Specifically, our framework achieves strong performance, with 89.97% recall, 95.06% precision, and 92.33% F1 score on the DDI task. Furthermore, the integration of UML2Dep into the code generation pipeline also improves practical deployment, increasing compilation pass rate by 8.83% and unit test pass rate by 11.66%.
Wenxin Mao, Zhitao Wang, Sirong Chen, Cuiyun Gao 0001, Luyang Cao, Zhi Jin 0001
ASE1
2024 Learning in the Wild: Towards Leveraging Unlabeled Data for Effectively Tuning Pre-trained Code Models
abstract
Pre-trained code models have recently achieved substantial improvements in many code intelligence tasks. These models are first pre-trained on large-scale unlabeled datasets in a task-agnostic manner using self-supervised learning, and then fine-tuned on labeled datasets in downstream tasks. However, the labeled datasets are usually limited in size (i.e., human intensive efforts), which may hinder the performance of pre-trained code models in specific tasks. To mitigate this, one possible solution is to leverage the large-scale unlabeled data in the tuning stage by pseudo-labeling, i.e., generating pseudo labels for unlabeled data and further training the pre-trained code models with the pseudo-labeled data. However, directly employing the pseudo-labeled data can bring a large amount of noise, i.e., incorrect labels, leading to suboptimal performance. How to effectively leverage the noisy pseudo-labeled data is a challenging yet under-explored problem.
Shuzheng Gao, Wenxin Mao, Cuiyun Gao 0001, Li Li 0029, Xing Hu 0008, Xin Xia 0001, Michael R. Lyu
ICSE2
2023 Fuzzy-VGG: A fast deep learning method for predicting the staging of Alzheimer's disease based on brain MRI
Zhaomin Yao, Wenxin Mao, Yizhe Yuan, Zhenning Shi, Gancheng Zhu, Zhiguo Wang 0001, Guoxu Zhang
Inf. Sci.2
2020 Extracting influence relationships in China's industrial ecological transformation using a rough set based machine learning method
abstract
China's industry urgently needs to be transformed from the development patterns driven by traditional production factor to achieving industrial ecological transformation (IET). The IET is influenced by diversified factors including resource input, allocation and flow, environmental regulations and technological innovations in different situation. Revealing the complex influence mechanisms between IET and its influence factors is necessary for effectively analyzing, evaluating and improving the performance of IET. A three stages machine learning method including learning, verification and generalization based on dominance-based rough set approach is presented to extract the influential relationships between the IET and its contextual influence factors. The proposed method excavates and learns the historical panel data of China's 30 provinces, and the cross-validation is conducted to produce a set of highly credible "If-Then" decision rules to generalize the synergistic influential relationships and intensities in IET. The results show that China's investment strength, resource allocation efficiency, command controlled and economic incentive environmental regulations are determinants to enhance the performance of IET, which helps to select the optimal transformation patterns by taking the historical development characteristics as lessons.
Wenxin Mao, Huifang Sun
SMC1
2016 Grey dominance-based rough set approach to decision system with three-parameter interval grey number
abstract
A method of knowledge acquisition for the decision information system whose attribute value of alternatives is three-parameter interval grey number is proposed in this paper. First, in classic rough set, the decision table must be given in advance, but we can only establish information system from the collected data. So, we establish the decision table from information system with the grey relational clustering decision method. Then, we construct the grey dominance relation based on the dominance extent between two three-parameter interval grey numbers, and put forward a method of extracting decision rules and attribute reduction. The last case about the comprehensive evaluation of icebreaking car is given to illustrate the effectiveness of the proposed method.
Dang Luo, Wenxin Mao, Huifang Sun
SMC2