Guojun Ma

dblp:32/8360 · DBLP profile ↗
← Back
13ranked-venue papers
2as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 6 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 W2S-AlignTree: Weak-to-Strong Inference-Time Alignment for Large Language Models via Monte Carlo Tree Search
abstract
Large Language Models (LLMs) demonstrate impressive capabilities, yet their outputs often suffer from misalignment with human preferences due to the inadequacy of weak supervision and a lack of fine-grained control. Training-time alignment methods like Reinforcement Learning from Human Feedback (RLHF) face prohibitive costs in expert supervision and inherent scalability limitations, offering limited dynamic control during inference. Consequently, there is an urgent need for scalable and adaptable alignment mechanisms. To address this, we propose W2S-AlignTree, a pioneering plug-and-play inference-time alignment framework that synergistically combines Monte Carlo Tree Search (MCTS) with the Weak-to-Strong Generalization paradigm for the first time. W2S-AlignTree formulates LLM alignment as an optimal heuristic search problem within a generative search tree. By leveraging weak model's real-time, step-level signals as alignment proxies and introducing an Entropy-Aware exploration mechanism, W2S-AlignTree enables fine-grained guidance during strong model's generation without modifying its parameters. The approach dynamically balances exploration and exploitation in high-dimensional generation search trees. Experiments across controlled sentiment generation, summarization, and instruction-following show that W2S-AlignTree consistently outperforms strong baselines. Notably, W2S-AlignTree raises the performance of Llama3-8B from 1.89 to 2.19, a relative improvement of 15.9% on the summarization task.
Zhenyu Ding, Tengyue Xiao, Haoying Wang, Guojun Ma, Mingyang Wan, Caigui Jiang, Ning Ding 0006
AAAI5
2025 Route Sparse Autoencoder to Interpret Large Language Models
abstract
Mechanistic interpretability of large language models (LLMs) aims to uncover the internal processes of information propagation and reasoning.Sparse autoencoders (SAEs) have demonstrated promise in this domain by extracting interpretable and monosemantic features.However, prior works primarily focus on feature extraction from a single layer, failing to effectively capture activations that span multiple layers.In this paper, we introduce Route Sparse Autoencoder (RouteSAE), a new framework that integrates a routing mechanism with a shared SAE to efficiently extract features from multiple layers.It dynamically assigns weights to activations from different layers, incurring minimal parameter overhead while achieving high interpretability and flexibility for targeted feature manipulation.We evaluate RouteSAE through extensive experiments on Llama-3.2-1B-Instruct.Specifically, under the same sparsity constraint of 64, RouteSAE extracts 22.5% more features than baseline SAEs while achieving a 22.3% higher interpretability score.These results underscore the potential of RouteSAE as a scalable and effective method for LLM interpretability, with applications in feature discovery and model intervention.Our codes are available at https: //github.com/swei2001/RouteSAEs.
Sihang Li 0002, Mingyang Wan, Guojun Ma, Xiang Wang 0010, Xiangnan He 0001
EMNLP5
2025 KPEE: A Two-Stage Proposal-Based Reformulation of Event Extraction
Hengrui Song, Mingyang Wan, Shannan Yan, Chun Yuan 0003, Guojun Ma
ICIC (23)7
2025 EAV-Mamba: Efficient Audio-Visual Representation Learning for Weakly-Supervised Temporal Action Localization
abstract
Weakly supervised temporal action localization aims to learn to locate actions in videos from video-level or point-level labels, avoiding the need for costly frame-level annotations. Unlike previous work that relies solely on visual modality information, we propose incorporating audio information into the weakly supervised temporal action localization task. While audio-visual localization tasks combine audio and visual information for video localization, temporal action localization often deals with action categories that have weak audio cues. To address this, we propose EAV-Mamba, the first audio-visual perception modeling method based on Mamba. Leveraging Mamba’s powerful audio-visual perception capabilities, we developed modules such as Audio-Perceptive Flow Enhancement, Audio-Perceptive RGB Enhancement, and Audio Self-Perceptive Enhancement. Extensive experiments on two publicly available temporal action localization datasets demonstrate that EAV-Mamba achieves efficient audio-visual perception modeling and state-of-the-art performance in weakly supervised temporal action localization tasks.
Jinwei Fang, Yuxin Qi 0001, Mingyang Wan, Guojun Ma, Ke Zhang 0046, Chun Yuan 0003
ICME5
2025 AnyEdit: Edit Any Knowledge Encoded in Language Models
abstract
Large language models (LLMs) often produce incorrect or outdated information, necessitating efficient and precise knowledge updates. Current model editing methods, however, struggle with long-form knowledge in diverse formats, such as poetry, code snippets, and mathematical derivations. These limitations arise from their reliance on editing a single token’s hidden state, a limitation we term as ``efficacy barrier''. To solve this, we propose \textbf{AnyEdit}, a new autoregressive editing paradigm. It decomposes long-form knowledge into sequential chunks and iteratively edits the key token in each chunk, ensuring consistent and accurate outputs. Theoretically, we ground AnyEdit in the Chain Rule of Mutual Information, showing its ability to update any knowledge within LLMs. Empirically, it outperforms strong baselines by 21.5\% on benchmarks including UnKEBench, AKEW, and our new \textbf{EditEverything} dataset for long-form diverse-formatted knowledge. Additionally, AnyEdit serves as a plug-and-play framework, enabling current editing methods to update knowledge with arbitrary length and format, significantly advancing the scope and practicality of LLM knowledge editing. Our code is available at: \url{https://github.com/jianghoucheng/AnyEdit}.
Houcheng Jiang, Junfeng Fang, Ningyu Zhang 0001, Mingyang Wan, Guojun Ma, Xiang Wang 0010, Xiangnan He 0001, Tat-Seng Chua
ICML5
2025 AGCNet: Improving Inertial Odometry via IMU Accelerometer and Gyroscope Online Compensation
abstract
This paper presents a learning-based online IMU compensation method (AGCNet) that can compensate for run-time errors of the accelerometer and gyroscope to improve inertial odometry. AGCNet employs U-Net architecture with hybrid dilated convolutions to extract multiscale features. It also adopts skip connections and patch-based processing strategy to aggregate local and global information. The network is trained to minimize absolute errors between integration results derived from compensated IMU data and ground truth motion states. The network utilizes IMU measurements from the current time window to correct errors in the subsequent time window, enabling sparser computations. Experiments on two public visual-inertial datasets show that AGCNet can accurately estimate the orientation from IMU measurements, outperforming existing learning-based methods. When applied to Open-VINS, AGCNet improves the accuracy of orientation estimation by an average of 29.8% and position estimation by an average of 37.3%.
Hongyuan Min, Ning Ding 0006, Mingyang Wan, Guojun Ma, Caigui Jiang
IROS4
2025 A Cascaded Pipeline for Self-Directed, Model-Agnostic Unit Test Generation via LLMs
abstract
While existing ML-based unit test generation methods show promising results, they face three key limitations: (1) incomplete test case generation with excessive focus on test oracles, (2) semantic inconsistencies between test components, and (3) dependency on closed-source models compromising data security. In this paper, we propose a novel approach named CasModaTest, a cascaded, model-agnostic, and end-to-end unit test generation framework, to alleviate the above limitations. Specifically, CasModaTest first splits the unit test generation task as two cascaded steps: test prefix generation and test oracle generation. Then, to better stimulate models’ learning ability, we manually build large-scale demo pools to provide CasModaTest with high-quality test prefixes and test oracles examples. Finally, CasModaTest assembles test components and validates their functionality through execution, with error correction during compilation/runtime. Our evaluation on the Defects4J benchmark demonstrates CasModaTest’s superiority over five state-of-the-art approaches, showing significant improvements in both accuracy and focal method coverage. Further validation across $\mathbf{1, 6 2 5}$ methods from six real-world projects reveals that CasModaTest achieves substantially higher code coverage metrics (method/line/branch coverage) compared to the dedicated coverage tool EvoSuite.
Chao Ni 0001, Liushan Chen, Guojun Ma
ISSRE5
2025 Enhancing LLM's Ability to Generate More Repository-Aware Unit Tests Through Precise Context Injection
abstract
Recently, Large Language Models (LLMs) have gained attention for their ability to handle a broad range of tasks, including unit test generation. Despite their success, LLMs may exhibit hallucinations when generating unit tests for focal methods or functions due to their lack of awareness regarding the project’s global context. While many studies have explored the role of context, they often extract fixed patterns of context for different models and focal methods, which may not be suitable for all generation processes (e.g., excessive irrelevant context could lead to redundancy, preventing the model from focusing on essential information).To overcome this limitation, we propose RATester, which integrates language servers to provide dynamic definition lookup to assist the LLM. When RATester encounters an unfamiliar identifier, it first leverages language servers (e.g., Gopls) to fetch relevant definitions and documentation comments, and then uses this global knowledge to guide the LLM. We evaluate the effectiveness and efficiency of RATester by constructing a new Golang dataset from real-world projects. On our Golang dataset, RATester achieves an average line coverage of 26.25%, representing an improvement of 9.10% to 165.69% over the baselines. In mutation testing, RATester shows superior performance by successfully killing 18 to 147 more mutants than the baselines. Additionally, our model-agnostic and generalizability analysis confirms RATester’s effectiveness across different models, programming languages, and model scales, validating its broad applicability.
Chao Ni 0001, Xinrui Li 0004, Liushan Chen, Guojun Ma, Xiaohu Yang 0001
ASE5
2025 Grounding Language with Vision: A Conditional Mutual Information Calibrated Decoding Strategy for Reducing Hallucinations in LVLMs
abstract
Large Vision-Language Models (LVLMs) are susceptible to hallucinations, where generated responses seem semantically plausible yet exhibit little or no relevance to the input image. Previous studies reveal that this issue primarily stems from LVLMs' over-reliance on language priors while disregarding the visual information during decoding. To alleviate this issue, we introduce a novel Conditional Pointwise Mutual Information (C-PMI) calibrated decoding strategy, which adaptively strengthens the mutual dependency between generated texts and input images to mitigate hallucinations. Unlike existing methods solely focusing on text token sampling, we propose to jointly model the contributions of visual and textual tokens to C-PMI, formulating hallucination mitigation as a bi-level optimization problem aimed at maximizing mutual information. To solve it, we design a token purification mechanism that dynamically regulates the decoding process by sampling text tokens remaining maximally relevant to the given image, while simultaneously refining image tokens most pertinent to the generated response. Extensive experiments across various benchmarks reveal that the proposed method significantly reduces hallucinations in LVLMs while preserving decoding efficiency.
Hao Fang 0011, Changle Zhou, Jiawei Kong 0001, Kuofeng Gao, Bin Chen 0011, Guojun Ma, Shutao Xia
NeurIPS7
2024 Smart Issue Detection for Large-Scale Online Service Systems Using Multi-Channel Data
abstract
Abstract Given the scale and complexity of large online service systems and the diversity of environments in which the services are to be invoked, it is inevitable that those service systems contain bugs that affect the users. As a result, it is essential for service providers to discover issues in their systems based on information gathered from users. iFeedback is a state-of-the-art technique for user-feedback-based issue detection. While it has been deployed to help detect issues in real-world service systems, the accuracy of iFeedback’s detection results is relatively low due to limitations in its design. In this paper, we propose theSkyNettechnique and tool that analyzes both user feedback gathered via specific channels and public posts collected from social media platforms to more accurately detect issues in service systems. We have applied the tool to detect issues for three real-world, large-scale online service systems based on their historical data gathered over a ten-month period of time.SkyNetreported in total 2790 issues, among which 93.0% were confirmed by developers as reflecting real problems that deserve their close attention. It also detected 58 out of the 62 severe issues reported during the period, achieving a recall of 93.5% for severe issues. Such results suggestSkyNetis both effective and accurate in issue detection.
Liushan Chen, Yu Pei 0001, Mingyang Wan, Zhihui Fei, Guojun Ma
FASE6
2024 Effective Unit Test Generation for Android Apps
abstract
While the received wisdom says that testing at levels like classes and methods is necessary for detecting bugs in programs, the application of unit testing to Android development in practice is limited so far due to the lack of sufficient technical and tool support. This paper proposes the EvoDroid approach to the automated unit test suite generation for Android code. EvoDroid is inspired by Evoobj, a SOTA test generation technique for object-oriented Java programs based on Evo-SUITE. EvoObj generates unit test suites for Java methods and constructs object construction graphs to guide the synthesis of complex objects as test inputs. In contrast to that, EvoDroid generates test suites for Java classes, and its object synthesis is driven by input structure maps which are comparably effective but much less expensive to construct. EvoDroid also integrates the Robolectric framework to support running Android unit tests on regular Java virtual machines. Experimental evaluation results show that EvoDroid is both effective and efficient in generating unit test suites for Android.
Guojun Ma, Yu Pei 0001, Liushan Chen, Chenqing Gan, Tian Zhang 0001
ICSME1
2023 A probability smoothing Bi-RRT path planning algorithm for indoor robot
Guojun Ma, Yunlong Duan, Zhibin Xie
Future Gener. Comput. Syst.1
2022 Expanding Large Pre-trained Unimodal Models with Multimodal Information Injection for Image-Text Multimodal Classification
abstract
Fine-tuning pre-trained models for downstream tasks is mainstream in deep learning. However, the pre-trained models are limited to be fine-tuned by data from a specific modality. For example, as a visual model, DenseNet cannot directly take the textual data as its input. Hence, although the large pre-trained models such as DenseNet or BERT have a great potential for the downstream recognition tasks, they have weaknesses in leveraging multimodal information, which is a new trend of deep learning. This work focuses on fine-tuning pre-trained unimodal models with multimodal inputs of image-text pairs and expanding them for image-text multimodal recognition. To this end, we propose the Multimodal Information Injection Plug-in (MI2P) which is attached to different layers of the unimodal models (e.g., DenseNet and BERT). The proposed MI2P unit provides the path to integrate the information of other modalities into the unimodal models. Specifically, MI2P performs cross-modal feature transformation by learning the fine-grained correlations between the visual and textual features. Through the proposed MI2P unit, we can inject the language information into the vision backbone by attending the word-wise textual features to different visual channels, as well as inject the visual information into the language backbone by attending the channel-wise visual features to different textual words. Armed with the MI2P attachments, the pre-trained unimodal models can be expanded to process multimodal data without the need to change the network structures.
Guosheng Lin, Mingyang Wan, Tianrui Li 0001, Guojun Ma, Fengmao Lv
CVPR5