VLDB 2026 Research / reviewers in the wild / expert
Yuxuan Fan
dblp:227/4963
· DBLP profile ↗
15ranked-venue papers
0as first author
15since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RoadSceneVQA: Benchmarking Visual Question Answering in Roadside Perception Systems for Intelligent Transportation SystemabstractCurrent roadside perception systems mainly focus on instance-level perception, which fall short in enabling interaction via natural language and reasoning about traffic behaviors in context. To bridge this gap, we introduce RoadSceneVQA, a large-scale and richly annotated visual question answering (VQA) dataset specifically tailored for roadside scenarios. The dataset comprises 34,736 diverse QA pairs collected under varying weather, illumination, and traffic conditions, targeting not only object attributes but also the intent, legality, and interaction patterns of traffic participants. RoadSceneVQA challenges models to perform both explicit recognition and implicit commonsense reasoning, grounded in real-world traffic rules and contextual dependencies. To fully exploit the reasoning potential of Multi-modal Large Language Models (MLLMs), we further propose CogniAnchor Fusion (CAF), a vision-language fusion module inspired by human-like scene anchoring mechanisms. CAF enables precise and efficient cross-modal interaction. Moreover, we propose the Assisted Decoupled Chain-of-Thought (AD-CoT) to enhance the reasoned thinking via CoT prompting and multi-task learning. Experimental results on RoadSceneVQA and CODA-LM benchmark show that the pipeline consistently improves both reasoning accuracy and computational efficiency, allowing the MLLM to achieve state-of-the-art performance in structural traffic perception and reasoning tasks. Runwei Guan, Rongsheng Hu, Shangshu Chen, Ningyuan Xiao, Ziren Tang, Ningwei Ouyang, Shaofeng Liang, Yuxuan Fan, Wanjie Sun, Yutao Yue |
AAAI | 11 |
| 2026 | Deconstructing Pre-training: Knowledge Attribution Analysis in MoE and Dense ModelsabstractMixture-of-Experts (MoE) architectures decouple model capacity from per-token computation, enabling scaling beyond the computational limits imposed by dense scaling laws. Yet how MoE architectures shape knowledge acquisition during pre-training—and how this process differs from dense architectures—remains unknown. To address this issue, we introduce Gated-LPI (Log-Probability Increase), a neuron-level attribution metric that decomposes log-probability increase across neurons. We present a time-resolved comparison of knowledge acquisition dynamics in MoE and dense architectures, tracking checkpoints over 1.2M (~ 5.0T tokens) and 600K (~ 2.5T tokens) training steps, respectively. Our experiments uncover three patterns: (1) Low-entropy backbone. The top approximately 1% of MoE neurons capture over 45% of positive updates, forming a high-utility core, which is absent in the dense baseline. (2) Early consolidation. The MoE model locks into a stable importance profile within < 100K steps, whereas the dense model remains volatile throughout training. (3) Functional robustness. Masking the ten most important MoE attention heads reduces relational HIT@10 by < 10%, compared with > 50% for the dense model, showing that sparsity fosters distributed—rather than brittle—knowledge storage. These patterns collectively demonstrate that sparsity fosters an intrinsically stable and distributed computational backbone from early in training, helping bridge the gap between sparse architectures and training-time interpretability. Junzhuo Li, Yuanlin Chu, Yuxuan Fan, Xuming Hu |
AAAI | 5 |
| 2026 | DIMM: Decoupled Multi-hierarchy Kalman Filter via Reinforcement LearningabstractState estimation is challenging for target tracking with high maneuverability, as the target's state transition function changes rapidly, irregularly, and is unknown to the estimator. Existing work based on interacting multiple model (IMM) achieves more accurate estimation than single-filter approaches through model combination, aligning appropriate models for different motion modes of the target over time. However, two limitations of conventional IMM remain unsolved. First, the solution space of the model combination is constrained as the target's diverse kinematic properties in different directions are ignored. Second, the model combination weights calculated by the observation likelihood are not accurate enough due to the measurement uncertainty. In this paper, we propose a novel framework, DIMM, to effectively combine estimates from different motion models in each direction, thus increasing the target tracking accuracy. First, DIMM extends the model combination solution space of conventional IMM from a hyperplane to a hypercube by designing a 3D-decoupled multi-hierarchy filter bank, which describes the target's motion with various-order linear models. Second, DIMM generates more reliable combination weight matrices through a differentiable adaptive fusion network for importance allocation rather than solely relying on the observation likelihood; it contains an attention-based twin delayed deep deterministic policy gradient (TD3) method with a hierarchical reward. Experiments demonstrate that DIMM significantly improves the tracking accuracy of existing state estimation methods by 31.61%~99.23%. Jirong Zha, Yuxuan Fan, Chen Gao 0001, Xinlei Chen |
AAAI | 2 |
| 2026 | AirCopBench: A Benchmark for Multi-drone Collaborative Embodied Perception and ReasoningabstractMultimodal Large Language Models (MLLMs) have shown promise in single-agent vision tasks, yet benchmarks for evaluating multi-agent collaborative perception remain scarce. This gap is critical, as multi-drone systems provide enhanced coverage, robustness, and collaboration compared to single-sensor setups. Existing multi-image benchmarks mainly target basic perception tasks using high-quality single-agent images, thus failing to evaluate MLLMs in more complex, egocentric collaborative scenarios, especially under real-world degraded perception conditions. To address these challenges, we introduce AirCopBench, the first comprehensive benchmark designed to evaluate MLLMs in embodied aerial collaborative perception under challenging perceptual conditions. AirCopBench includes 14.6k+ questions derived from both simulator and real-world data, spanning four key task dimensions: Scene Understanding, Object Understanding, Perception Assessment, and Collaborative Decision, across 14 task types. We construct the benchmark using data from challenging degraded-perception scenarios with annotated collaborative events, generating large-scale questions through model-, rule-, and human-based methods under rigorous quality control. Evaluations on 40 MLLMs show significant performance gaps in collaborative perception tasks, with the best model trailing humans by 24.38% on average and exhibiting inconsistent results across tasks. Fine-tuning experiments further confirm the feasibility of sim-to-real transfer in aerial collaborative perception. Jirong Zha, Yuxuan Fan, Chen Gao 0001, Xinlei Chen |
AAAI | 2 |
| 2026 | V 2 -Fusion: Virtual voxel enhanced 4D radar-image feature fusion for 3D object detection
Li Wang 0092, Xinyu Zhang 0001, Yuxuan Fan, Tao Xie 0010, Lei Yang 0060, Bin Xu 0003 |
Expert Syst. Appl. | 4 |
| 2025 | How to Enable LLM with 3D Capacity? A Survey of Spatial Reasoning in LLMabstract3D spatial understanding is essential in real-world applications such as robotics, autonomous vehicles, virtual reality, and medical imaging. Recently, Large Language Models (LLMs), having demonstrated remarkable success across various domains, have been leveraged to enhance 3D understanding tasks, showing potential to surpass traditional computer vision methods. In this survey, we present a comprehensive review of methods integrating LLMs with 3D spatial understanding. We propose a taxonomy that categorizes existing methods into three branches: image-based methods deriving 3D understanding from 2D visual data, point cloud-based methods working directly with 3D representations, and hybrid modality-based methods combining multiple data streams. We systematically review representative methods along these categories, covering data representations, architectural modifications, and training strategies that bridge textual and 3D modalities. Finally, we discuss current limitations, including dataset scarcity and computational challenges, while highlighting promising research directions in spatial perception, multi-modal fusion, and real-world applications. Jirong Zha, Yuxuan Fan, Xinlei Chen |
IJCAI | 2 |
| 2025 | EventVAD: Training-Free Event-Aware Video Anomaly DetectionabstractVideo Anomaly Detection (VAD) focuses on identifying anomalies within videos. Supervised methods require an amount of in-domain training data and often struggle to generalize to unseen anomalies. In contrast, training-free methods leverage the intrinsic world knowledge of large language models (LLMs) to detect anomalies but face challenges in localizing fine-grained visual transitions and diverse events. Therefore, we propose EventVAD, an event-aware video anomaly detection framework that combines tailored dynamic graph architectures and multimodal LLMs to perform fine-grained temporal-event reasoning. Specifically, EventVAD first employs dynamic spatiotemporal graph modeling with time-decay constraints to capture event-aware video features. Then, it performs adaptive noise filtering and uses signal ratio thresholding to detect event boundaries via unsupervised statistical features. Finally, it utilizes a hierarchical prompting strategy to guide MLLMs in performing reasoning and making final decisions. We conducted extensive experiments on the UCF-Crime and XD-Violence datasets. The results demonstrate that EventVAD with a 7B MLLM achieves state-of-the-art (SOTA) in training-free settings, outperforming strong baselines that use 7B or larger MLLMs. The code is available at https://github.com/YihuaJerry/EventVAD. Yihua Shao, Haojin He, Siyu Chen 0021, Xinwei Long, Fanhu Zeng, Yuxuan Fan, Muyang Zhang, Ziyang Yan, Ao Ma 0005, Hao Tang 0005, Yan Wang 0105, Shuyan Li |
ACM Multimedia | 7 |
| 2025 | Instantly Learning Preference Alignment via In-context DPOabstractFeifan Song, Yuxuan Fan, Xin Zhang, Peiyi Wang, Houfeng Wang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Feifan Song 0001, Yuxuan Fan, Xin Zhang 0099, Peiyi Wang, Houfeng Wang |
NAACL (Long Papers) | 2 |
| 2025 | Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal LearningabstractMultimodal large language models (MLLMs) have achieved strong performance on vision-language tasks but still struggle with fine-grained visual differences, leading to hallucinations or missed semantic shifts. We attribute this to limitations in both training data and learning objectives. To address these issues, we propose a controlled data generation pipeline that produces minimally edited image pairs with semantically aligned captions. Using this pipeline, we construct the Micro Edit Dataset (MED), containing over 50K image-text pairs spanning 11 fine-grained edit categories, including attribute, count, position, and object presence changes.
Building on MED, we introduce a supervised fine-tuning (SFT) framework with a feature-level consistency loss that promotes stable visual embeddings under small edits. We evaluate our approach on the Micro Edit Detection benchmark, which includes carefully balanced evaluation pairs designed to test sensitivity to subtle visual variations across the same edit categories.
Our method improves difference detection accuracy and reduces hallucinations compared to strong baselines, including GPT-4o. Moreover, it yields consistent gains on standard vision-language tasks such as image captioning and visual question answering. These results demonstrate the effectiveness of combining targeted data and alignment objectives for enhancing fine-grained visual reasoning in MLLMs. Code and datasets are publicly released at https://github.com/Relaxed-System-Lab/hallu_med. Tianyi Bai, Yuxuan Fan, Jiantao Qiu, Fupeng Sun, Junlin Han, Conghui He, Wentao Zhang 0001, Binhang Yuan |
NeurIPS | 2 |
| 2025 | Towards Better Dental AI: A Multimodal Benchmark and Instruction Dataset for Panoramic X-ray AnalysisabstractRecent advances in large vision-language models (LVLMs) have demonstrated strong performance on general-purpose medical tasks. However, their effectiveness in specialized domains such as dentistry remains underexplored. In particular, panoramic X-rays, a widely used imaging modality in oral radiology, pose interpretative challenges due to dense anatomical structures and subtle pathological cues, which are not captured by existing medical benchmarks or instruction datasets. To this end, we introduce MMOral, the first large-scale multimodal instruction dataset and benchmark tailored for panoramic X-ray interpretation. MMOral consists of 20,563 annotated images paired with 1.3 million instruction-following instances across diverse task types, including attribute extraction, report generation, visual question answering, and image-grounded dialogue. In addition, we present MMOral-Bench, a comprehensive evaluation suite covering five key diagnostic dimensions in dentistry. We evaluate 64 LVLMs on MMOral-Bench and find that even the best-performing model, i.e., GPT-4o, only achieves 43.31% accuracy, revealing significant limitations of current models in this domain. To promote the progress of this specific domain, we provide the supervised fine-tuning (SFT) process utilizing our meticulously curated MMOral instruction dataset. Remarkably, a single epoch of SFT yields substantial performance enhancements for LVLMs, e.g., Qwen2.5-VL-7B demonstrates a 24.73% improvement. MMOral holds significant potential as a critical foundation for intelligent dentistry and enables more clinically impactful multimodal AI systems in the dental field. Yuxuan Fan, Yanpeng Sun, Kaixin Guo, Lizhuo Lin, Qi Yong H. Ai, Lun M. Wong, Hao Tang 0005, Kuo Feng Hung |
NeurIPS | 2 |
| 2025 | DTWformer: A DTW-Based Transformer for Multivariate Time Series ForecastingabstractAccurate and robust time series forecasting is critical in numerous domains but remains challenging due to issues such as temporal misalignments, noise, and complex multivariate dependencies. Transformer-based models have demonstrated strong performance in sequential tasks; however, their reliance on dot-product attention renders them sensitive to noise and less effective for misaligned time series. To address these limitations, we propose a novel dynamic time warping (DTW)-based attention mechanism, leveraging a Sakoe-Chiba-constrained softDTW framework to replace the traditional dot-product similarity. This approach enables dynamic sequence alignment, enhancing robustness to temporal misalignments. Building on this innovation, we introduce DTWformer, a multi-scale Transformer model that integrates DTW-attention with adaptive patching to capture dependencies across varying temporal resolutions. DTWformer achieves superior forecasting performance and efficiency, addressing the limitations of existing approaches in handling misaligned and noisy time series. The implementation is publicly available at https://github.com/unihe/DTWformer. Zhongyang Yu, Jidong Yuan, Huiting Pei 0001, Yuxuan Fan, Shijiang Li |
IEEE Trans. Big Data | 5 |
| 2024 | Reinforcement Learning for Efficient Multi-phase Resource AllocationabstractEfficient resource allocation is pivotal for achieving high performance in emerging computer systems, where multiple users and tasks compete for shared resources. This challenge spans various domains, including data centers, multicore processors, cloud computing and edge computing, each requiring nuanced allocation strategies to balance competing demands. Traditional approaches often assume concave utility (performance) functions for users, simplifying optimization but failing to capture the complexities of real-world scenarios where non-concave utility functions prevail. Numerous works in the literature apply the greedy algorithm to nonconcave utility functions, resulting in suboptimal solution due to the short-sighted behaviors. To improve this gap, we propose a novel multiphase resource allocation framework that accurately reflects the non-linear dynamics of these systems. To tackle the NP-complete nature of this problem, we formulate a customized resource allocation Markov Decision Process (MDP) that integrates the characteristics of multi-phase utility functions into a nuanced design of the key MDP components, such as state representations, reward signals, and actions. We explore two reinforcement learning (RL)-based methods, specifically Dueling Deep Q-Network (Dueling DQN) and Proximal Policy Optimization (PPO), to optimize resource allocation over time. Our RL-based strategies outperform the conventional greedy algorithm by approximately 37% in standard environments and up to 73% in specialized environments, highlighting their effectiveness in handling the resource allocation problem with non-concave utility functions and achieving scalable, real-time solutions. Zhenfu Zhang, Haiyan Yin, Liudong Zuo, Xiao Zhang 0006, Jianlin Zhu, Yuxuan Fan, Pan Lai |
HPCC | 6 |
| 2024 | IoT-Based Adaptive Multiplication-Convolution Sparse Denoising for Equipment Edge Condition EvaluationabstractThe advent of equipment condition evaluation at the edge, facilitated by the Internet of Things (IoT), has led to significant reduction in data transmission and improvement in diagnostic efficiency. Nevertheless, the multi-scale components and noise interferences will seriously affect the quality of end-side sensed signals. Additionally, the mismatch between edge-side rigid models and time-varying nature of data features can lead to the failure of edge evaluation. To overcome these issues, this study introduces an IoT-based adaptive multiplication-convolution sparse denoising (AMCSD) method for the accurate equipment edge condition evaluation. Initially, from the perspective of signal processing synergized with deep learning, a lightweight signal denoising network is proposed with an array of multiplication filtering kernels (MFKs). Guiding by fault signal modulation mechanism, the learned MFKs sparse filters can adaptively extracted the fault-related frequency features with irrelevant components suppressed from the time-series differential information distributed in the degeneration process. Subsequently, an end-edge collaborative mechanism framework is designed and deployed on end-edge hardware unit. A compact end-side processing node (EPN) prototype can achieve efficient edge denoising effect with data adaptive compressing. Concurrently, the edge-side device named AlxBoard can implement the model adaptive dynamic updating. This means that adaptive signal denoising employed in the end-side focuses on the improvement of signal quality and sparse filter learning encompassed in the edge-side aims to solve the model mismatch. These advancements are anticipated to provide a monotonic but sensitive evaluation of degradation at the edge, surpassing the capability of conventional approaches. Qihang Wu, Xiaoxi Ding, Wenhao Cheng, Yuxuan Fan |
IEEE Internet Things J. | 4 |
| 2023 | Can We Edit Factual Knowledge by In-Context Learning?abstractPrevious studies have shown that large language models (LLMs) like GPTs store massive factual knowledge in their parameters.However, the stored knowledge could be false or outdated.Traditional knowledge editing methods refine LLMs via fine-tuning on texts containing specific knowledge.However, with the increasing scales of LLMs, these gradient-based approaches bring large computation costs.The trend of model-as-a-service also makes it impossible to modify knowledge in black-box LLMs.Inspired by in-context learning (ICL), a new paradigm based on demonstration contexts without parameter updating, we explore whether ICL can edit factual knowledge.To answer this question, we give a comprehensive empirical study of ICL strategies.Experiments show that in-context knowledge editing (IKE), without any gradient and parameter updating, achieves a competitive success rate compared to gradient-based methods on GPT-J (6B) but with much fewer side effects, including less over-editing on similar but unrelated facts and less knowledge forgetting on previously stored knowledge.We also apply the method to larger LMs with tens or hundreds of parameters like OPT-175B, which shows the scalability of our method.The code is available at https://github.com/pkunlp-icler/IKE. Lei Li 0039, Qingxiu Dong, Yuxuan Fan, Zhiyong Wu 0011, Jingjing Xu 0001, Baobao Chang |
EMNLP | 4 |
| 2023 | TransLink: Transformer-Based Embedding for Tracklets' Global LinkabstractMulti-object tracking (MOT) is essential to many tasks related to the smart transportation. Detecting and tracking humans on the road can give a vital feedback for either the moving vehicle or traffic control to ensure better driving safety and traffic flow. However, most trackers face a common problem of identity (ID) switch, resulting in an incomplete human trajectory prediction. In this paper, we propose a Transformer-based tracklet linking method called TransLink to mitigate the association failures. Specifically, the self-attention mechanism is well exploited to get the feature representation for tracklets, followed by a multilayer perceptron to predict the association likelihood, which can be further used in determining the tracklet association. Experiments on the MOT dataset demonstrate the effectiveness of the proposed module in lifting the tracking performances. Yanting Zhang 0001, Shuanghong Wang, Yuxuan Fan, Gaoang Wang, Cairong Yan |
ICASSP | 3 |