Jiajia Ma

dblp:235/0231 · DBLP profile ↗
← Back
9ranked-venue papers
1as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 4 since 2021Software engineering, systems software and programming languages · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 A Human-in-the-Loop Framework for Mongolian Dependency Corpus Construction with Difficulty-Aware Crowdsourcing and Self-training
Jiajia Ma, Wenhong Wu, Guiping Liu, Aodengbala
ICIC (16)1
2026 SQL-Commenter: Aligning Large Language Models for SQL Comment Generation with Direct Preference Optimization
abstract
SQL query comprehension is a significant challenge in database and data analysis environments due to complex syntax, diverse join types, and deep nesting. Despite its critical role in backend development and data science, many queries, particularly within legacy systems, often lack adequate comments, which severely hinders code readability, maintainability, and knowledge transfer. Existing approaches to automated SQL comment generation face two main challenges: limited training datasets that inadequately represent real-world analytical queries involving multi-table joins, window functions, and complex aggregations, and an insufficient understanding of SQL-specific logical semantics and schema-related context by Large Language Models (LLMs), even after standard training. Our empirical analysis shows that even after continual pre-training and supervised fine-tuning, LLMs struggle to precisely understand complex SQL semantics, leading to inaccurate or incomplete comments. To address these challenges, we propose SQL-Commenter, an advanced comment generation method based on LLaMA-3.1-8B. First, we construct a comprehensive dataset containing longer, more complex SQL queries with expert-verified, detailed comments. Second, we perform continual pre-training using a large-scale SQL corpus to enhance the LLM’s understanding of SQL syntax and semantics. Then, we conduct supervised fine-tuning with our high-quality dataset. Finally, we introduce Direct Preference Optimization (DPO), which leverages human feedback to significantly improve comment quality. SQL-Commenter utilizes a preference-based loss function that encourages the LLM to increase the probability of preferred outputs while decreasing the probability of non-preferred outputs, thereby enhancing both fine-grained semantic learning, such as distinguishing between different join types, and context-dependent quality assessment based on business logic. We evaluate SQL-Commenter on the authoritative Spider and Bird benchmarks, where it significantly outperforms state-of-the-art baselines. On average, across these datasets, our method surpasses the strongest baseline (Qwen3-14B) by 9.29, 4.99, and 13.23 percentage points on BLEU-4, METEOR, and ROUGE-L, respectively. Moreover, human evaluation demonstrates the superior quality of comments generated by SQL-Commenter in terms of correctness, completeness, and naturalness.
Li Yang 0015, Changzhi Deng, Jiajia Ma, Fengjun Zhang
ICPC8
2026 Strunkmap: An Abstract Approach to Understand Spatiotemporal Density Distribution
abstract
Visual analysis of spatiotemporal density distributions is crucial for understanding spatiotemporal dynamics. However, existing methods suffer from visual occlusion and information loss when simultaneously displaying multiple density distributions. We present Strunkmap as an abstract approach to address these challenges. We introduce anisotropic kernel density estimation to enhance the accuracy of density generation. We extract the trunks of density distributions to identify the overall spatial patterns. Path scanning and trunk-outline matching strategies are employed to preserve local spatial structure. We design a stacked trunk plot that enables lossless density representation while conserving substantial screen space. Based on the visual design, Strunkmap integrates multiple heatmaps within a single map to effectively display temporal evolution of density distributions without visual occlusion. Ablation studies and comparative experiments validate the superiority of Strunkmap in accuracy and efficiency for hotspot identification and trend exploration. Theoretical analysis demonstrates Strunkmap's scalability, which we further verify through large-scale spatiotemporal data visualization. Color encoding schemes and scaling ratios are discussed to illustrate the flexibility. Our evaluations with user feedback demonstrate that Strunkmap is a viable solution with significant potential to real-world applications.
Zhirong Huang, Jiajia Ma, Shiqi Cheng, Ruize Zhou, Xiaoxiao Ma 0005, Li Yang 0015, Fengjun Zhang
IEEE Trans. Vis. Comput. Graph.5
2025 SAEL: Leveraging Large Language Models with Adaptive Mixture-of-Experts for Smart Contract Vulnerability Detection
abstract
With the increasing security issues in blockchain, smart contract vulnerability detection has become a research focus. Existing vulnerability detection methods have their limitations: 1) Static analysis methods struggle with complex scenarios. 2) Methods based on specialized pre-trained models perform well on specific datasets but have limited generalization capabilities. In contrast, general-purpose Large Language Models (LLMs) demonstrate impressive ability in adapting to new vulnerability patterns. However, they often underperform on specific vulnerability types compared to methods based on specialized pre-trained models. We also observe that explanations generated by generalpurpose LLMs can provide fine-grained code understanding information, contributing to improved detection performance. Inspired by these observations, we propose SAEL, a LLMbased framework for smart contract vulnerability detection. First, we design prompts targeting specific smart contract vulnerabilities to guide general-purpose LLMs in detecting vulnerabilities and providing explanations. The detection results generated by LLMs serve as prediction features. Then, we employ prompt-tuning on CodeT5 and T5 respectively to process contract code and explanations, enhancing model performance on specific tasks. To leverage the strengths of each component, we introduce Adaptive Mixture-of-Experts, a dynamic architecture for smart contract vulnerability detection. This mechanism dynamically adjusts feature weights through a Gating Network, which selects the most relevant features by applying TopK filtering and Softmax normalization, and a Multi-Head Self-Attention mechanism, which enhances cross-feature relationships by processing multiple attention heads in parallel. This design ensures that prediction results for LLMs, explanation features, and contract code features are effectively integrated through gradient optimization. The loss function focuses on the independent prediction performance of each feature and the overall performance of weighted predictions. Experimental results show that SAEL outperforms existing methods in detecting various vulnerabilities.
Shiqi Cheng, Zhirong Huang, Chenjie Shen, Li Yang 0015, Fengjun Zhang, Jiajia Ma
ICSME9
2025 Pre-training and Fine-turning Multi-task Learning Model: An Effective Method for Mongolian Morphological Tagging
abstract
Mongolian1morphological label is the abstract description of Mongolian word morphology, and morphological tagging is an essential pre-processing step for Mongolian natural language processing. The morphological tagging method of Mongolian based on deep sequence annotation has achieved good results. However, the process suffers from two drawbacks: (1)sparse data usually leads to the incomplete training of the deep network. (2)sequential processing of segmentation and tagging causes error diffusion. In recent years, natural language processing has seen a boom in introducing pre-trained knowledge, such as pre-trained word embeddings and large-scale pre-trained models. These approaches have been applied to numerous tasks and have gained much progress. However, as a low-resource language, Mongolian does not have these pre-training resources. This paper proposes a new strategy based on pre-training and fine-tuning multi-task learning model to improve the Mongolian morphological tagging. The proposed strategy constructs a multi-task learning model, which alleviates the problem of error diffusion according to the relationship between Mongolian morphological segmentation and morphological tagging tasks. Then, morphological segmentation knowledge is introduced through pre-training and fine-tuning multi-task learning model to alleviate the problem of data sparsity. Experiments show that the proposed approach outperforms current state-of-the-art methods.
Jiajia Ma, Aodengbala, Guiping Liu
IJCNN2
2025 Dependent syntactic analysis of Mongolian based on semi-supervised self-training
abstract
The analysis of Mongolian dependency syntax has always been an important part of Mongolian semantic analysis, machine translation, semantic role annotation and other tasks. However, the corpus of Mongolian as a low-resource language dependency syntax is very scarce. This paper proposes a new method for analysing Mongolian dependency syntax. The seq2seq model structure is employed for dependency syntax analysis of Mongolian, taking advantage of its morphological features and grammatical structure. However, the scarcity of a Mongolian corpus and the numerous parameters of other deep learning networks can result in overfitting. To address this challenge, a two-stage self-training framework is employed in conjunction with a confidence dynamic threshold setting method to construct a corpus of Mongolian dependent syntax. The experimental results demonstrate that the scores of dependent syntactic analysis in Labeled Attachment Score (LAS) and the Unlabeled Attachment Score (UAS) reach 86.65% and 85.80%, respectively.
Jiajia Ma, Nier Wu, Yatu Ji, Guiping Liu
IJCNN2
2025 EXE-Reviewer: Towards EXplainable and Effective Review Comments Generation
abstract
Modern code review is essential for software quality, but the complexity of codebases and time demands of manual reviews drive interest in automation for greater efficiency and consistency.However, current automated methods often fail to generate meaningful review comments and lack explainability, limiting developers' understanding and trust.This paper presents EXE-Reviewer, aimed at generating more EXplainable and Effective review comments.To enhance effectiveness, we integrate focus information into an existing model to improve its ability to extract key insights, thereby elevating comment quality.To improve explainability, we connect explanatory information (justification behind solutions) to causality, utilizing causality extraction techniques and introducing an explanatory loss.Furthermore, we devise two metrics to assess the quantity and quality of explanatory content, enhancing insight into the model's explanations.We compare EXE-Reviewer to state-ofthe-art methods in terms of effectiveness and explainability of the generated review comments.Experimental results show that EXE-Reviewer achieves a BLEU-4 score of 7.36%, surpassing the state-of-the-art baseline of 18.52%.Meanwhile, both explainability metrics and empirical study demonstrate notable improvements in explainability of the review comments generated by EXE-Reviewer, highlighting the effectiveness of our approach in generating accurate and comprehensible review comments to developers.
Yifei Liu 0002, Li Yang 0015, Xiaoxiao Ma 0005, Jiajia Ma, Fengjun Zhang, Chun Zuo
SEKE7
2023 Who Are the Money Launderers? Money Laundering Detection on Blockchain via Mutual Learning-Based Graph Neural Network
abstract
With the development of blockchain technology, security concerns have become increasingly prominent in recent years. Money laundering through blockchain has been found to generate a significant amount of money and has become a serious threat. Towards money laundering detection in Bitcoin, conventional methods heavily rely on fixed expert rules, leading to low accuracy and poor scalability. Graph convolutional network approaches have improved this issue, but they fail to distinguish the importance of surrounding transactions and the structural information of different transactions. To solve above problems, we propose an approach to detect money laundering on blockchain by mining its transaction records, named AEtransGAT. First, we use a novel approach called transGat as an encoder to determine the significance of surrounding transactions by considering the transaction amount values of transaction flows. The original features and the features after graph embedding are combined to address the issue of feature distortion. Second, we deploy the graph autoencoder as the decoder to learn the overall structural information of different transactions, and the concatenated embedding is used to output the classification results as the detector. Finally, we propose our model based on mutual learning in this task which takes the advantages of both transactions classification loss and structure reconstruction loss. We validate the performance of our model on the Elliptic dataset which is the only large open source dataset in Bitcoin anti-money laundering. The results show that our method outperforms current state-of-the-art methods and is linearly scalable.
Fengjun Zhang, Jiajia Ma, Li Yang 0015, Yuanzhe Yang
IJCNN3
2023 PSCVFinder: A Prompt-Tuning Based Framework for Smart Contract Vulnerability Detection
abstract
With the increasing security issues in the blockchain, smart contract vulnerability detection has gradually become the focus of research. Recently, many approaches have been proposed to detect smart contract vulnerabilities. Despite promising results, these approaches still have three drawbacks: 1) Symbolic execution and static analysis methods are constrained by predefined rules, which limits their adaptability to different vulnerabilities. 2) Most smart contract code contains abundant irrelevant information which is useless for vulnerability detection. 3) Pre-trained models fail to bridge the gap between pre-training and detecting smart contract vulnerabilities.To solve these problems, we propose an approach named PSCVFinder for detecting reentrancy vulnerability and times-tamp dependency vulnerability, which are two severe vulnerabilities in smart contract. To better detect these vulnerabilities, we propose CSCV which is a smart contract slicing method to reduce the irrelevant code. Unlike existing approaches, our model first learns the representation of programming language through the pre-training model, then fully exploits the capacity of large language model with prompt-tuning to precisely detect smart contract vulnerability. We conduct experiments on real-world dataset and the results reflect that PSCVFinder scores 93.83% and 93.49% on two kinds of vulnerabilities in F1-score, surpassing the state-of-the-art baseline by 1.14% and 4.02%, respectively.
Xianglong Liu 0005, Li Yang 0015, Fengjun Zhang, Jiajia Ma
ISSRE6