Shukai Ma

dblp:424/7887 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
2 papers
Program synthesis and code generation · 65% Compilers and program optimization · 35%
Artificial intelligence
1 paper
Vision and language · 77% Multi-agent systems · 23%

Topics — the 4 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Compilers and program optimization
code generation
1.012026
SRACG: A Code Generation Framework with Selective Retrieval Augmentation · AAAI 2026
Program synthesis and code generation › code generation with language models
retrieval-augmented code generation
1.012026
SRACG: A Code Generation Framework with Selective Retrieval Augmentation · AAAI 2026
Computer vision › Vision and language › vision-language model
multimodal large language model
0.912025
From Model Diagram to Code: A Benchmark Dataset and Multi-Agent Framework · ACM Multimedia 2025
Knowledge, reasoning and agents › Multi-agent systems › multi-agent collaboration
collaborative agents
0.312025
From Model Diagram to Code: A Benchmark Dataset and Multi-Agent Framework · ACM Multimedia 2025

Methods — techniques the papers use, named apart from their topics

multimodal large language model · 1.7multi-agent framework · 1.7retrieval-augmented generation · 1.0large language model · 1.0
YearPublicationVenuePosition
2026 SRACG: A Code Generation Framework with Selective Retrieval Augmentation
abstract
Large Language Models (LLMs) have demonstrated remarkable performance in code generation, offering new possibilities for translating natural language into executable programs. To further enhance LLMs’ code generation capabilities, Retrieval-Augmented Generation (RAG) has emerged as a promising strategy by retrieving code examples aligned with the generation intent to guide the process. However, existing RAG-based methods often suffer from unnecessary augmentation, preference misalignment, and surface-level mimicry, which undermine the effectiveness of retrieved examples in guiding LLMs toward accurate code generation. To address these challenges, we propose SRACG, a Selective Retrieval-Augmented Code Generation framework. SRACG begins with a necessity-aware selection mechanism to identify generation intents that genuinely require retrieval support, thereby avoiding degradation from indiscriminate augmentation. For intents identified as needing enhancement, it first employs a multi-objective retrieval strategy to select examples that are semantically aligned with the intent. These candidates are then further filtered by assessing their consistency with the LLM’s inherent generation preferences, ensuring alignment in both style and structure. Finally, it extracts execution plans from the filtered examples to uncover their underlying logic, guiding the LLM to better comprehend the examples instead of merely mimicking surface-level content. Experimental results on widely used benchmarks show that SRACG significantly improves the success rate of LLM-generated code and outperforms existing approaches.
Mengzhen Wang, Shukai Ma, Songwen Gong, Jiexin Wang 0002, Ruolin Chen, Liuwen Cao, Yi Cai 0001
AAAI2
2025 From Model Diagram to Code: A Benchmark Dataset and Multi-Agent Framework
abstract
Model Diagram-to-Code Generation aims to translate model diagrams from research papers into implementation code that reconstructs the model's architecture. This task plays a crucial role in accelerating scientific workflows and enhancing the efficiency of industrial model deployment. While recent studies have explored various Image-to-Code Generation tasks using Multimodal Large Language Models (MLLMs), these efforts have primarily focused on reconstructing the visual appearance depicted in input images, leaving this task largely underexplored. The complex structural elements and implicit relationships in model diagrams present greater challenges for MLLMs, particularly in terms of visual reasoning and semantic interpretation. To support this task, we introduce MDCDataset, a dataset designed to evaluate the ability of MLLMs to generate code from model diagrams. It comprises 1,008 instances spanning 16 research domains, each with a model diagram, structured textual content, and the ground-truth code implementation. Furthermore, to address the inherent challenges of this task, we propose MDCAgent, a collaborative multi-agent framework composed of Parsing, Generation, and Check Agents. These agents work in coordination to analyze, extract, and verify complex elements and implicit relationships within model diagrams, thereby enhancing the visual architecture-aware reasoning capabilities of MLLMs. Our extensive experiments confirm the effectiveness of the framework.
Mengzhen Wang, Xunbin Huang, Jiayuan Xie, Shukai Ma, Jiale Men, Dayong Liang, Yi Cai 0001
ACM Multimedia4