VLDB 2026 Research / reviewers in the wild / expert
Jiaqi Tang 0005
dblp:121/7086-5
· DBLP profile ↗
12ranked-venue papers
5as first author
12since 2021 · last 2026
0009-0003-1251-0825ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 5 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 7 since 2021Computer networks · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Robust-R1: Degradation-Aware Reasoning for Robust Visual UnderstandingabstractMultimodal Large Language Models struggle to maintain reliable performance under extreme real-world visual degradations, which impede their practical robustness. Existing robust MLLMs predominantly rely on implicit training/adaptation that focuses solely on visual encoder generalization, suffering from limited interpretability and isolated optimization. To overcome these limitations, we propose Robust-R1, a novel framework that explicitly models visual degradations through structured reasoning chains. Our approach integrates: (i) supervised fine-tuning for degradation-aware reasoning foundations, (ii) reward-driven alignment for accurately perceiving degradation parameters, and (iii) dynamic reasoning depth scaling adapted to degradation intensity. To facilitate this approach, we introduce a specialized 11K dataset featuring realistic degradations synthesized across four critical real-world visual processing stages, each annotated with structured chains connecting degradation parameters, perceptual influence, pristine semantic reasoning chain, and conclusion. Comprehensive evaluations demonstrate state-of-theart robustness: Robust-R1 outperforms all general and robust baselines on the real-world degradation benchmark R-Bench, while maintaining superior anti-degradation performance under multi-intensity adversarial degradations on MMMB, MMStar, and RealWorldQA. Jiaqi Tang 0005, Jianmin Chen, Wei Wei 0008, Xiaogang Xu 0002, Runtao Liu, Qipeng Xie, Jiafei Wu, Lei Zhang 0001, Qifeng Chen 0001 |
AAAI | 1 |
| 2026 | LongVideoAgent: Multi-Agent Reasoning with Long VideosabstractRecent advances in multimodal LLMs and systems that use tools for long-video QA point to the promise of reasoning over hour-long episodes.However, many methods still compress content into lossy summaries or rely on limited toolsets, weakening temporal grounding and missing fine-grained cues.We propose a multi-agent framework in which a master LLM coordinates a grounding agent to localize question-relevant segments and a vision agent to extract targeted textual observations.The master agent plans with a step limit, and is trained with reinforcement learning to encourage concise, correct, and efficient multi-agent cooperation.This design helps the master agent focus on relevant clips via grounding, complements subtitles with visual detail, and yields interpretable trajectories.On our proposed LongTVQA and LongTVQA+ which are episode-level datasets aggregated from TVQA/TVQA+, our multiagent system significantly outperforms strong non-agent baselines.Experiments also show reinforcement learning further strengthens reasoning and planning for the trained agent. Runtao Liu, Jiaqi Tang 0005, Yue Ma 0016, Renjie Pi, Qifeng Chen 0001 |
ACL (1) | 3 |
| 2026 | Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video UnderstandingabstractKe Ma, Jiaqi Tang, Bin Guo, Xueting Han, Ruonan Xu, Qingfeng He, Ziheng Wang, Xu Wang, Qifeng Chen, Zhiwen Yu, Yunhao Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Jiaqi Tang 0005, Bin Guo 0001, Xueting Han, Ruonan Xu, Qingfeng He, Xu Wang 0018, Qifeng Chen 0001, Zhiwen Yu 0001, Yunhao Liu 0001 |
ACL (1) | 2 |
| 2026 | Responsive Test-Time Model Adaptation for Mobile Applications via Runtime-Efficient Sparse UpdatesabstractAdapting the on-device deep learning (DL) models to continual and unpredictable domain shifts is crucial for mobile applications like autonomous driving and augmented reality, which require seamless user experiences in ever-changing environments. Test-time adaptation (TTA) offers a promising solution that fine-tunes models withunlabeledreal-time data just before making predictions. However, TTA introduces a significant challenge, its forward-backward-reforward pipeline increases latency, compromising responsiveness in time-sensitive mobile scenarios. In this paper, we present AdaShadow, a responsive test-time adaptation framework designed for non-stationary mobile data distributions and resource dynamics. AdaShadow focuses on selectively updating only adaptation-critical layers, minimizing the latency impact. While this idea has been explored in general on-device training, the unsupervised and real-time nature of TTA presents unique challenges in assessing layer importance, predicting latency, and planning layer updates efficiently. AdaShadow addresses these challenges through abackpropagation-free importance assessorfor rapid identification of critical layers, a unit-basedruntime latency predictorthat accounts for resource constraints, and anonline layer update schedulerto ensure prompt retraining. Furthermore, a memory I/O-aware computation reuse scheme optimizes the reforward pass to further reduce delays. Our evaluations show that AdaShadow reduces adaptation latency by up to 3.7× (achieving ms-level response time) and improves accuracy by up to 25.4%, all while maintaining low memory and energy costs across CNN, and ViT models, outperforming state-of-the-art methods. Bin Guo 0001, Sicong Liu 0005, Zimu Zhou, Jiaqi Tang 0005, Shiyan Luo, Geyang Song, Zhiwen Yu 0001 |
IEEE Trans. Mob. Comput. | 5 |
| 2025 | SURGEON: Memory-Adaptive Fully Test-Time Adaptation via Dynamic Activation SparsityabstractDespite the growing integration of deep models into mobile terminals, the accuracy of these models declines significantly due to various deployment interferences. Test-time adaptation (TTA) has emerged to improve the performance of deep models by adapting them to unlabeled target data online. Yet, the significant memory cost, particularly in resource-constrained terminals, impedes the effective deployment of most backward-propagation-based TTA methods. To tackle memory constraints, we introduce Surgeon, a method that substantially reduces memory cost while preserving comparable accuracy improvements during fully test-time adaptation (FTTA) without relying on specific network architectures or modifications to the original training procedure. Specifically, we propose a novel dynamic activation sparsity strategy that directly prunes activations at layer-specific dynamic ratios during adaptation, allowing for flexible control of learning ability and memory cost in a data-sensitive manner. Among this, two metrics, Gradient Importance and Layer Activation Memory, are considered to determine the layer-wise pruning ratios, reflecting accuracy contribution and memory efficiency, respectively. Experimentally, our method surpasses the baselines by not only reducing memory usage but also achieving superior accuracy, delivering SOTA performance across diverse datasets, architectures, and tasks. Jiaqi Tang 0005, Bin Guo 0001, Fan Dang 0001, Sicong Liu 0005, Zhui Zhu, Ying-Cong Chen, Zhiwen Yu 0001, Yunhao Liu 0001 |
CVPR | 2 |
| 2025 | Co-Painter: Fine-Grained Controllable Image Stylization via Implicit Decoupling and Adaptive Injection
Wei Wei 0008, Jiaqi Tang 0005, Jiangtao Nie, Yanyu Ye, Xiaogang Xu 0002, Ying-Cong Chen, Lei Zhang 0001 |
ICCV | 3 |
| 2025 | Rhythmguassian: Repurposing Generalizable Gaussian Model for Remote Physiological Measurement
Hao Lu 0009, Yuting Zhang 0008, Jiaqi Tang 0005, Wenhang Ge, Wei Wei 0008, Kaishun Wu, Ying-Cong Chen |
ICCV | 3 |
| 2024 | Learning to Remove Wrinkled Transparent Film with Polarized PriorabstractIn this paper, we study a new problem, Film Removal (FR), which attempts to remove the interference of wrinkled transparent films and reconstruct the original information under films for industrial recognition systems. We first physically model the imaging of industrial materials covered by the film. Considering the specular highlight from the film can be effectively recorded by the polarized camera, we build a practical dataset with polarization information containing paired data with and without transparent film. We aim to remove interference from the film (specular highlights and other degradations) with an end-to-end framework. To locate the specular highlight, we use an angle estimation network to optimize the polarization angle with the minimized specular highlight. The image with minimized specular highlight is set as a prior for supporting the reconstruction network. Based on the prior and the polarized images, the reconstruction network can decouple all degradations from the film. Extensive experiments show that our framework achieves SOTA performance in both image reconstruction and industrial downstream tasks. Our code will be released at https://github.com/jqtangust/FilmRemoval. Jiaqi Tang 0005, Ruizheng Wu, Xiaogang Xu 0002, Sixing Hu, Ying-Cong Chen |
CVPR | 1 |
| 2024 | An Incremental Unified Framework for Small Defect Inspection
Jiaqi Tang 0005, Hao Lu 0009, Xiaogang Xu 0002, Ruizheng Wu, Sixing Hu, Tong Zhang 0001, Tsz Wa Cheng, Ming Ge, Ying-Cong Chen, Fugee Tsung |
ECCV (31) | 1 |
| 2024 | HAWK: Learning to Understand Open-World Video AnomaliesabstractVideo Anomaly Detection (VAD) systems can autonomously monitor and identify disturbances, reducing the need for manual labor and associated costs. However, current VAD systems are often limited by their superficial semantic understanding of scenes and minimal user interaction. Additionally, the prevalent data scarcity in existing datasets restricts their applicability in open-world scenarios.
In this paper, we introduce HAWK, a novel framework that leverages interactive large Visual Language Models (VLM) to interpret video anomalies precisely. Recognizing the difference in motion information between abnormal and normal videos, HAWK explicitly integrates motion modality to enhance anomaly identification. To reinforce motion attention, we construct an auxiliary consistency loss within the motion and video space, guiding the video branch to focus on the motion modality. Moreover, to improve the interpretation of motion-to-language, we establish a clear supervisory relationship between motion and its linguistic representation. Furthermore, we have annotated over 8,000 anomaly videos with language descriptions, enabling effective training across diverse open-world scenarios, and also created 8,000 question-answering pairs for users' open-world questions. The final results demonstrate that HAWK achieves SOTA performance, surpassing existing baselines in both video description generation and question-answering. Our codes/dataset/demo will be released at https://github.com/jqtangust/hawk. Jiaqi Tang 0005, Hao Lu 0009, Ruizheng Wu, Xiaogang Xu 0002, Bin Guo 0001, Jiangbo Lu, Qifeng Chen 0001, Ying-Cong Chen |
NeurIPS | 1 |
| 2024 | AdaShadow: Responsive Test-time Model Adaptation in Non-stationary Mobile EnvironmentsabstractOn-device adapting to continual, unpredictable domain shifts is essential for mobile applications like autonomous driving and augmented reality to deliver seamless user experiences in evolving environments. Test-time adaptation (TTA) emerges as a promising solution by tuning model parameters with unlabeled live data immediately before prediction. However, TTA's unique forward-backward-reforward pipeline notably increases the latency over standard inference, undermining the responsiveness in time-sensitive mobile applications. This paper presents AdaShadow, a responsive test-time adaptation framework for non-stationary mobile data distribution and resource dynamics via selective updates of adaptation-critical layers. Although the tactic is recognized in generic on-device training, TTA's unsupervised and online context presents unique challenges in estimating layer importance and latency, as well as scheduling the optimal layer update plan. AdaShadow addresses these challenges with a backpropagation-free assessor to rapidly identify critical layers, a unit-based runtime predictor to account for resource dynamics in latency estimation, and an online scheduler for prompt layer update planning. Also, AdaShadow incorporates a memory I/O-aware computation reuse scheme to further reduce latency in the reforwardpass. Results show that AdaShadow achieves the best accuracy-latency balance under continual shifts. At low memory and energy costs, Adashadow provides a 2x to 3.5x speedup (ms-level) over state-of-the-art TTA methods with comparable accuracy and a 14.8% to 25.4% accuracy boost over efficient supervised methods with similar latency. Sicong Liu 0005, Zimu Zhou, Bin Guo 0001, Jiaqi Tang 0005, Zhiwen Yu 0001 |
SenSys | 5 |
| 2023 | High Dynamic Range Image Reconstruction via Deep Explicit Polynomial Curve EstimationabstractDue to limited camera capacities, digital images usually have a narrower dynamic illumination range than real-world scene radiance. To resolve this problem, High Dynamic Range (HDR) reconstruction is proposed to recover the dynamic range to better represent real-world scenes. However, due to different physical imaging parameters, the tone-mapping functions between images and real radiance are highly diverse, which makes HDR reconstruction extremely challenging. Existing solutions can not explicitly clarify a corresponding relationship between the tone-mapping function and the generated HDR image, but this relationship is vital when guiding the reconstruction of HDR images. To address this problem, we propose a method to explicitly estimate the tone mapping function and its corresponding HDR image in one network. Firstly, based on the characteristics of the tone mapping function, we construct a model by a polynomial to describe the trend of the tone curve. To fit this curve, we use a learnable network to estimate the coefficients of the polynomial. This curve will be automatically adjusted according to the tone space of the Low Dynamic Range (LDR) image, and reconstruct the real HDR image. Besides, since all current datasets do not provide the corresponding relationship between the tone mapping function and the LDR image, we construct a new dataset with both synthetic and real images. Extensive experiments show that our method generalizes well under different tone-mapping functions and achieves SOTA performance. The code/dataset is available at https://github.com/jqtangust/EPCE-HDR.git. Jiaqi Tang 0005, Xiaogang Xu 0002, Sixing Hu, Ying-Cong Chen |
ECAI | 1 |