Zhijie Tan

dblp:291/9806 · DBLP profile ↗
← Back
15ranked-venue papers
4as first author
15since 2021 · last 2026
0000-0003-4934-1001ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Towards Order Fairness: Mitigating LLMs Order Sensitivity through Dual Group Advantage Optimization
abstract
Large Language Models (LLMs) suffer from order bias, where their performance is affected by the arrangement order of input elements.This unfairness limits the model's applications in scenarios such as in-context learning and Retrieval-Augmented Generation (RAG).Recent studies attempt to obtain optimal or suboptimal arrangements based on statistical results or using dataset-based search, but these methods increase inference overhead while leaving the model's inherent order bias unresolved.Other studies mitigate order sensitivity through supervised fine-tuning using augmented training sets with multiple order variants, but often at the cost of accuracy, trapping the model in consistent yet incorrect hallucinations.In this paper, we propose Dual Group Advantage Optimization (DGAO), which aims to improve model accuracy and order stability simultaneously.DGAO calculates and balances intragroup relative accuracy advantage and intergroup relative stability advantage, rewarding the policy model for generating order-stable and correct outputs while penalizing ordersensitive or incorrect responses.This marks the first time reinforcement learning has been used to mitigate LLMs' order sensitivity.We also propose two new metrics, Consistency Rate and Overconfidence Rate, to reveal the pseudostability of previous methods and guide more comprehensive evaluation.Extensive experiments demonstrate that DGAO achieves superior order fairness while improving performance on RAG, mathematical reasoning, and classification tasks.Our
Zhijie Tan, Xinrong Chen, Tong Mo
ACL (1)3
2026 RADO: Reasoning Audit-Driven Optimization for Rigorous Reasoning in High-Stakes Domains
abstract
High-stakes domains such as finance, law, and biomedicine demand both accurate results and rigorous reasoning.Current reinforcement learning paradigms primarily rely on outcomebased rewards, often overlooking latent logical fallacies in intermediate steps.Leveraging the cognitive asymmetry where falsifying local errors is more efficient than generating global correctness, we propose RADO (Reasoning Audit-Driven Optimization).RADO introduces a specialized audit model augmented with external tools to identify local logical ruptures and calibrate reward signals.By integrating Direct Preference Optimization (DPO) with Group Relative Policy Optimization (GRPO), our framework enables explicit supervision over reasoning paths.Experimental results demonstrate that RADO consistently improves final accuracy while significantly enhancing logical rigor in high-stakes domains.
Zhijie Tan, Xu Chu 0001, Guanyu Wang 0002, Weiping Li 0002, Tong Mo
ACL (1)1
2026 MuSe: Multi-Stage Graph Reasoning via Vision-Language Models
abstract
Graph-related tasks are traditionally addressed with Graph Neural Networks (GNNs) or graph transformers, but their task-specific training limits generalization.Large Language Models (LLMs) offer stronger generalization, yet encoding graphs as one-dimensional text struggles to capture multi-hop dependencies and two-dimensional topology.Vision-Language Models (VLMs) provide an alternative by visualizing graphs, but rendering large graphs in a single image causes clutter, occlusion, and distraction, hindering reasoning.We propose MuSe, a novel multi-stage graph reasoning framework based on VLMs.Instead of processing entire graphs at once, MuSe incrementally samples and visualizes task-relevant subgraphs, enabling progressive reasoning.The framework employs a two-stage training paradigm: supervised fine-tuning to acquire local sampling and reasoning skills, followed by reinforcement learning with GRPO to refine the sampling strategy and control dialog length.To support evaluation, we introduce LGVLQA, a new multimodal dataset with larger and more complex graph structures, addressing the scalability limitations of existing benchmarks.Experiments show that MuSe consistently outperforms leading LLM and VLM baselines, demonstrating improved structural understanding and reasoning ability.Our code and data are available at this url.
Guanyu Wang 0002, Xu Chu 0001, Zhijie Tan, Xinrong Chen, Tong Mo, Weiping Li 0002
ACL (1)3
2026 Accurate and Efficient Personalized Query Rewriting in Baidu Search
Xu Chu 0001, Wei Li 0336, Zhijie Tan, Dawei Yin 0001, Shuaiqiang Wang, Daiting Shi
WWW5
2026 LLM-SocRec: Enhancing Graph-based Social Recommendation via Collaborative Large Language Models
Zhijie Tan, Weiping Li 0002, Tong Mo
Mach. Learn.2
2026 STLLM-Rec: enhancing explainable recommendation via self-training LLMs
Zhijie Tan, Suhuan Wu, Weiping Li 0002, Tong Mo
World Wide Web (WWW)2
2025 Adaptive Spatiotemporal Augmentation for Improving Dynamic Graph Learning
abstract
Dynamic graph augmentation is used to improve the performance of dynamic GNNs. Most methods assume temporal locality, meaning that recent edges are more influential than earlier edges. However, for temporal changes in edges caused by random noise, overemphasizing recent edges while neglecting earlier ones may lead to the model capturing noise. To address this issue, we propose STAA (SpatioTemporal Activity-Aware Random Walk Diffusion). STAA identifies nodes likely to have noisy edges in spatiotemporal dimensions. Spatially, it analyzes critical topological positions through graph wavelet coefficients. Temporally, it analyzes edge evolution through graph wavelet coefficient change rates. Then, random walks are used to reduce the weights of noisy edges, deriving a diffusion matrix containing spatiotemporal information as an augmented adjacency matrix for dynamic GNN learning. Experiments on multiple datasets show that STAA outperforms other dynamic graph augmentation methods in node classification and link prediction tasks.
Xu Chu 0001, Hanlin Xue, Bingce Wang, Weiping Li 0002, Tong Mo, Tuoyu Feng, Zhijie Tan
ICASSP8
2025 Few-Shot Object Detection in Satellite Imagery with Feature Fusion Pyramid and Adaptive Region Proposal Networks
abstract
Object detection in satellite imagery presents unique challenges due to the wide variation in object sizes, shapes, and orientations, as well as the limited availability of labeled data for training models. Few-Shot Object Detection (FSOD) aims to address these challenges by enabling models to detect novel objects with only a few labeled examples. However, existing methods struggle to effectively capture multi-scale features and generate flexible region proposals, which are critical for accurate detection in complex aerial scenes. In this paper, we propose FFARPNet, a novel framework specifically designed for FSOD in satellite imagery. Our model introduces two key components: the Feature Fusion Pyramid Network (FFPN), which enhances multi-scale feature representation, and the Adaptive Region Proposal Network (ARPN), which dynamically adjusts region proposals to handle the diverse object scales and shapes found in aerial images. We evaluate FFARPNet on two challenging datasets, DIOR and NWPU VHR-10, and demonstrate significant improvements in detection accuracy across 3-shot, 5-shot, 10-shot, and 20-shot scenarios. Comparative analysis with state-of-the-art methods and ablation studies demonstrate the effectiveness of our proposed model and its core modules. The results highlight the robustness and generalization capability of our model, indicating its potential for remote sensing applications, particularly in scenarios with limited training data.
Tuoyu Feng, Weiping Li 0002, Zhijie Tan, Liwen Zhang 0004, Xu Chu 0001
ICASSP3
2025 Mitigating Hallucinations on Object Attributes using Multiview Images and Negative Instructions
abstract
Current popular Large Vision-Language Models (LVLMs) are suffering from Hallucinations on Object Attributes (HoOA), leading to incorrect determination of fine-grained attributes in the input images. Leveraging significant advancements in 3D generation from a single image, this paper proposes a novel method to mitigate HoOA in LVLMs. This method utilizes multiview images sampled from generated 3D representations as visual prompts for LVLMs, thereby providing more visual information from other viewpoints. Furthermore, we observe the input order of multiple multiview images significantly affects the performance of LVLMs. Consequently, we have devised Multiview Image Augmented VLM (MIAVLM), incorporating a Multiview Attributes Perceiver (MAP) submodule capable of simultaneously eliminating the influence of input image order and aligning visual information from multiview images with Large Language Models (LLMs). Besides, we designed and employed negative instructions to mitigate LVLMs’ bias towards "Yes" responses. Comprehensive experiments demonstrate the effectiveness of our method.
Zhijie Tan, Yuzhi Li, Shengwei Meng, Weiping Li 0002, Tong Mo, Bingce Wang, Xu Chu 0001
ICASSP1
2025 Learn Concepts from Multi-Scale Visual Information for Compositional Zero-Shot Learning
abstract
Compositional Zero-Shot Learning (CZSL) aims at recognizing novel compositions by combining concepts learned from seen compositions. The key to tackle CZSL is disentangling highly coupled attribute-object compositions and learning exclusive concepts. Previous works mainly design networks to learn visual concepts from top-layer representations provided by visual backbones. As visual backbones progressively integrate information layer by layer, some low-level but critical information for concept learning may be lost, and the coupling between attribute and object features deepens. To address these issues, we propose to extract multi-scale visual features and fuse them in an adaptive way by Mixture of Experts (MoE) networks. We also employ feature-level similarity and a maximum entropy regularization term to constrain the model to effectively disentangle and learn concepts from multi-scale visual information. Comprehensive experiments on three CZSL benchmark datasets demonstrate that our method significantly outperforms previous SOTA methods in both closed-world and open-world settings.
Guanyu Wang 0002, Zhijie Tan, Xu Chu 0001, Xinrong Chen, Tong Mo, Weiping Li 0002
MMAsia2
2024 LLM-MHR: A LLM-Augmented Multimodal Hashtag Recommendation Algorithm
abstract
The recommendation of suitable hashtags for mi-croposts encompassing multimodal content stands as a pivotal challenge for numerous Social Networking Service (SNS) applications such as Instagram, Weibo, etc. The accuracy of multimodal hashtag recommendation algorithms relies heavily on the comprehension of multimodal information, user historical information, and the reasoning ability based on such information. However, most previous works have not effectively utilized both historical and additional information simultaneously. Large Language Models (LLMs) learn a vast amount of implicit knowledge during the pre-training stage, which can serve as potential knowledge bases while also possessing strong reasoning abilities. Therefore, LLMs can provide additional information to help understand the micropost content and infer suitable hashtags with strong reasoning ability. However, introducing LLMs for multimodal hashtag recommendation faces three main challenges. Firstly, LLMs require an efficient modality alignment module to accept a multimodal input. Secondly, LLMs are highly sensitive to input order, while utilizing user historical information requires accepting multiple historical samples, necessitating the design of a robust historical information processing module to eliminate the influence of input order. Thirdly, fine-tuning LLMs entails substantial computational overheads, necessitating the reduction of additional trainable parameters. To address the first two challenges, this paper designs an efficient modality alignment module capable of processing multiple historical samples, simultaneously addressing the sensitivity of LLMs to input order changes. To tackle the third challenge, a hybrid prompt learning approach utilizing both soft and hard prompts is proposed to achieve parameter-efficient fine-tuning of LLMs. Finally, a LLM-augmented Multimodal Hashtag Recommendation algorithm (LLM-MHR) is implemented. Comprehensive experiments on the representative dataset MACON demonstrate that LLM-MHR has achieved SOTA performances with significant improvements.
Zhijie Tan, Yuzhi Li, Shengwei Meng, Weiping Li 0002, Tong Mo
ICWS1
2024 SEMScene: Semantic-Consistency Enhanced Multi-Level Scene Graph Matching for Image-Text Retrieval
abstract
Image-text retrieval, a fundamental cross-modal task, performs similarity reasoning for images and texts. The primary challenge for image-text retrieval is cross-modal semantic heterogeneity, where the semantic features of visual and textual modalities are rich but distinct. Scene graph is an effective representation for images and texts as it explicitly models objects and their relations. Existing scene graph based methods have not fully taken the features regarding various granularities implicit in scene graph into consideration (e.g., triplets), the inadequate feature matching incurs the absence of non-trivial semantic information (e.g., inner relations among triplets). Therefore, we propose a S emantic-Consistency E nhanced M ulti-Level Scene Graph Matching (SEMScene) network, which exploits the semantic relevance between visual and textual scene graphs from fine-grained to coarse-grained. Firstly, under the scene graph representation, we perform feature matching including low-level node matching, mid-level semantic triplet matching, and high-level holistic scene graph matching. Secondly, to enhance the semantic-consistency for object-fused triplets carrying key correlation information, we propose a dual-step constraint mechanism in mid-level matching. Thirdly, to guide the model to learn the semantic-consistency of matched image-text pairs, we devise effective loss functions for each stage of the dual-step constraint. Comprehensive experiments on Flickr30K and MS-COCO datasets demonstrate that SEMScene achieves state-of-the-art performances with significant improvements.
Haochen Li 0001, Zhijie Tan, Jinsong Huang, Jingjie Xiao, Weiping Li 0002, Tong Mo
ACM Trans. Multim. Comput. Commun. Appl.4
2023 Whisker Analysis Framework for Unrestricted Mice with Neural Networks
Zhijie Tan, Shengwei Meng, Yujia Tan, Tong Mo, Weiping Li 0002
ICANN (5)1
2023 A Semantic-Aware Transmission With Adaptive Control Scheme for Volumetric Video Service
abstract
Volumetric video provides a more immersive holographic virtual experience than conventional video services such as 360-degree and virtual reality (VR) videos. However, due to ultra-high bandwidth requirements, existing compression and transmission technology cannot handle the delivery of real-time volumetric video. Unlike traditional compression methods and the approaches that extend 360-degree video streaming, we propose AITransfer, an AI-powered compression and semantic-aware transmission method for point cloud video data (a popular volumetric data format). AITransfer targets the semantic-level communication beyond transmitting raw point cloud video or compressed video with two outstanding contributions: (1) designing an integrated end-to-end architecture with two fundamental contents of feature extraction and reconstruction to reduce the bandwidth consumption and alleviate the computational pressure; and (2) incorporating the dynamic network condition into end-to-end architecture design and employing a deep reinforcement learning-based adaptive control scheme to provide robust transmission. We conduct extensive experiments on the typical datasets and develop a case study to demonstrate the efficiency and effectiveness. The results show that AITransfer can provide extremely efficient point cloud transmission while maintaining considerable user experience with more than 30.72x compression ratio under the existing network environments.
Yuanwei Zhu, Yakun Huang, Xiuquan Qiao, Zhijie Tan, Boyuan Bai, Huadong Ma, Schahram Dustdar
IEEE Trans. Multim.4
2021 AITransfer: Progressive AI-powered Transmission for Real-Time Point Cloud Video Streaming
abstract
Point cloud video provides a more immersive holographic virtual experience than conventional video services such as 360 degree video and virtual reality (VR) video. However, the existing network bandwidth and transmission technology can not carry real-time point cloud video streaming due to mass data volume, high processing overheads, and extremely bandwidth-consuming. Unlike previous approaches that extend the VR video streaming, we propose AITransfer, an AI-powered bandwidth-aware and adaptive transmission technique driven by extracting and transferring key point cloud features to reduce the bandwidth consumption and alleviate the computational pressure. AITransfer has two outstanding contributions, including (1) incorporating the dynamic network bandwidth into the design of an end-to-end architecture with two fundamental contents of feature extraction and reconstruction, and (2) employing an online adapter to sense the network bandwidth and match the optimal inference model. We conduct extensive experiments on the typical dataset and develop a case study to demonstrate the efficiency and effectiveness. The results show that AITransfer can provide more than 30.72 times compression ratio under the existing network environments.
Yakun Huang, Yuanwei Zhu, Xiuquan Qiao, Zhijie Tan, Boyuan Bai
ACM Multimedia4