Xu Chu 0001

dblp:131/4767-1 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
9since 2021 · last 2026
0009-0002-3330-0037ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 RADO: Reasoning Audit-Driven Optimization for Rigorous Reasoning in High-Stakes Domains
abstract
High-stakes domains such as finance, law, and biomedicine demand both accurate results and rigorous reasoning.Current reinforcement learning paradigms primarily rely on outcomebased rewards, often overlooking latent logical fallacies in intermediate steps.Leveraging the cognitive asymmetry where falsifying local errors is more efficient than generating global correctness, we propose RADO (Reasoning Audit-Driven Optimization).RADO introduces a specialized audit model augmented with external tools to identify local logical ruptures and calibrate reward signals.By integrating Direct Preference Optimization (DPO) with Group Relative Policy Optimization (GRPO), our framework enables explicit supervision over reasoning paths.Experimental results demonstrate that RADO consistently improves final accuracy while significantly enhancing logical rigor in high-stakes domains.
Zhijie Tan, Xu Chu 0001, Guanyu Wang 0002, Weiping Li 0002, Tong Mo
ACL (1)2
2026 MuSe: Multi-Stage Graph Reasoning via Vision-Language Models
abstract
Graph-related tasks are traditionally addressed with Graph Neural Networks (GNNs) or graph transformers, but their task-specific training limits generalization.Large Language Models (LLMs) offer stronger generalization, yet encoding graphs as one-dimensional text struggles to capture multi-hop dependencies and two-dimensional topology.Vision-Language Models (VLMs) provide an alternative by visualizing graphs, but rendering large graphs in a single image causes clutter, occlusion, and distraction, hindering reasoning.We propose MuSe, a novel multi-stage graph reasoning framework based on VLMs.Instead of processing entire graphs at once, MuSe incrementally samples and visualizes task-relevant subgraphs, enabling progressive reasoning.The framework employs a two-stage training paradigm: supervised fine-tuning to acquire local sampling and reasoning skills, followed by reinforcement learning with GRPO to refine the sampling strategy and control dialog length.To support evaluation, we introduce LGVLQA, a new multimodal dataset with larger and more complex graph structures, addressing the scalability limitations of existing benchmarks.Experiments show that MuSe consistently outperforms leading LLM and VLM baselines, demonstrating improved structural understanding and reasoning ability.Our code and data are available at this url.
Guanyu Wang 0002, Xu Chu 0001, Zhijie Tan, Xinrong Chen, Tong Mo, Weiping Li 0002
ACL (1)2
2026 Not All Neighbors are Temporally Relevant: An Adaptive Neighborhood Aggregation Framework for Dynamic Graph Learning
Bingce Wang, Weiping Li 0002, Tong Mo, Xu Chu 0001, Liwen Zhang 0004
DASFAA (2)5
2026 MORE-R1: Guiding LVLM for Multimodal Object-Entity Relation Extraction via Stepwise Reasoning with Reinforcement Learning
Xu Chu 0001, Xinrong Chen, Haochen Li 0001, Zonghong Dai, Hongcheng Fan, Xiaoyue Yuan, Weiping Li 0002, Tong Mo
DASFAA (6)2
2026 Accurate and Efficient Personalized Query Rewriting in Baidu Search
Xu Chu 0001, Wei Li 0336, Zhijie Tan, Dawei Yin 0001, Shuaiqiang Wang, Daiting Shi
WWW1
2025 Adaptive Spatiotemporal Augmentation for Improving Dynamic Graph Learning
abstract
Dynamic graph augmentation is used to improve the performance of dynamic GNNs. Most methods assume temporal locality, meaning that recent edges are more influential than earlier edges. However, for temporal changes in edges caused by random noise, overemphasizing recent edges while neglecting earlier ones may lead to the model capturing noise. To address this issue, we propose STAA (SpatioTemporal Activity-Aware Random Walk Diffusion). STAA identifies nodes likely to have noisy edges in spatiotemporal dimensions. Spatially, it analyzes critical topological positions through graph wavelet coefficients. Temporally, it analyzes edge evolution through graph wavelet coefficient change rates. Then, random walks are used to reduce the weights of noisy edges, deriving a diffusion matrix containing spatiotemporal information as an augmented adjacency matrix for dynamic GNN learning. Experiments on multiple datasets show that STAA outperforms other dynamic graph augmentation methods in node classification and link prediction tasks.
Xu Chu 0001, Hanlin Xue, Bingce Wang, Weiping Li 0002, Tong Mo, Tuoyu Feng, Zhijie Tan
ICASSP1
2025 Few-Shot Object Detection in Satellite Imagery with Feature Fusion Pyramid and Adaptive Region Proposal Networks
abstract
Object detection in satellite imagery presents unique challenges due to the wide variation in object sizes, shapes, and orientations, as well as the limited availability of labeled data for training models. Few-Shot Object Detection (FSOD) aims to address these challenges by enabling models to detect novel objects with only a few labeled examples. However, existing methods struggle to effectively capture multi-scale features and generate flexible region proposals, which are critical for accurate detection in complex aerial scenes. In this paper, we propose FFARPNet, a novel framework specifically designed for FSOD in satellite imagery. Our model introduces two key components: the Feature Fusion Pyramid Network (FFPN), which enhances multi-scale feature representation, and the Adaptive Region Proposal Network (ARPN), which dynamically adjusts region proposals to handle the diverse object scales and shapes found in aerial images. We evaluate FFARPNet on two challenging datasets, DIOR and NWPU VHR-10, and demonstrate significant improvements in detection accuracy across 3-shot, 5-shot, 10-shot, and 20-shot scenarios. Comparative analysis with state-of-the-art methods and ablation studies demonstrate the effectiveness of our proposed model and its core modules. The results highlight the robustness and generalization capability of our model, indicating its potential for remote sensing applications, particularly in scenarios with limited training data.
Tuoyu Feng, Weiping Li 0002, Zhijie Tan, Liwen Zhang 0004, Xu Chu 0001
ICASSP6
2025 Mitigating Hallucinations on Object Attributes using Multiview Images and Negative Instructions
abstract
Current popular Large Vision-Language Models (LVLMs) are suffering from Hallucinations on Object Attributes (HoOA), leading to incorrect determination of fine-grained attributes in the input images. Leveraging significant advancements in 3D generation from a single image, this paper proposes a novel method to mitigate HoOA in LVLMs. This method utilizes multiview images sampled from generated 3D representations as visual prompts for LVLMs, thereby providing more visual information from other viewpoints. Furthermore, we observe the input order of multiple multiview images significantly affects the performance of LVLMs. Consequently, we have devised Multiview Image Augmented VLM (MIAVLM), incorporating a Multiview Attributes Perceiver (MAP) submodule capable of simultaneously eliminating the influence of input image order and aligning visual information from multiview images with Large Language Models (LLMs). Besides, we designed and employed negative instructions to mitigate LVLMs’ bias towards "Yes" responses. Comprehensive experiments demonstrate the effectiveness of our method.
Zhijie Tan, Yuzhi Li, Shengwei Meng, Weiping Li 0002, Tong Mo, Bingce Wang, Xu Chu 0001
ICASSP8
2025 Learn Concepts from Multi-Scale Visual Information for Compositional Zero-Shot Learning
abstract
Compositional Zero-Shot Learning (CZSL) aims at recognizing novel compositions by combining concepts learned from seen compositions. The key to tackle CZSL is disentangling highly coupled attribute-object compositions and learning exclusive concepts. Previous works mainly design networks to learn visual concepts from top-layer representations provided by visual backbones. As visual backbones progressively integrate information layer by layer, some low-level but critical information for concept learning may be lost, and the coupling between attribute and object features deepens. To address these issues, we propose to extract multi-scale visual features and fuse them in an adaptive way by Mixture of Experts (MoE) networks. We also employ feature-level similarity and a maximum entropy regularization term to constrain the model to effectively disentangle and learn concepts from multi-scale visual information. Comprehensive experiments on three CZSL benchmark datasets demonstrate that our method significantly outperforms previous SOTA methods in both closed-world and open-world settings.
Guanyu Wang 0002, Zhijie Tan, Xu Chu 0001, Xinrong Chen, Tong Mo, Weiping Li 0002
MMAsia3