Tianwen Qian

dblp:298/5209 · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
12since 2021 · last 2026
0000-0002-3881-4857ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 EgoCross: Benchmarking Multimodal Large Language Models for Cross-Domain Egocentric Video Question Answering
abstract
Recent advances in Multimodal Large Language Models (MLLMs) have significantly pushed the frontier of egocentric video question answering (EgocentricQA). However, existing benchmarks and studies are mainly limited to common daily activities such as cooking and cleaning. In contrast, real-world deployment inevitably encounters domain shifts, where target domains differ substantially in both visual style and semantic content. To bridge this gap, we introduce EgoCross, a comprehensive benchmark designed to evaluate the cross-domain generalization of MLLMs in EgocentricQA. EgoCross covers four diverse and challenging domains, including surgery, industry, extreme sports, and animal perspective, representing realistic and high-impact application scenarios. It comprises approximately 1,000 QA pairs across 798 video clips, spanning four key QA tasks: prediction, recognition, localization, and counting. Each QA pair provides both OpenQA and CloseQA formats to support fine-grained evaluation. Extensive experiments show that most existing MLLMs, whether general-purpose or egocentric-specialized, struggle to generalize to domains beyond daily life, highlighting the limitations of current models. Furthermore, we conduct several pilot studies, e.g., fine-tuning and reinforcement learning, to explore potential improvements. We hope EgoCross and our accompanying analysis will serve as a foundation for advancing domain-adaptive, robust egocentric video understanding.
Yuqian Fu, Tianwen Qian, Qi'ao Xu, Silong Dai, Danda Pani Paudel, Luc Van Gool, Xiaoling Wang 0004
AAAI3
2026 Guiding LLMs to decode text via aligning semantics in EEG signals and language
Huanran Zheng, Yuanbin Wu, Tianwen Qian, Wenjing Yue, Xiaoling Wang 0004
Expert Syst. Appl.3
2026 SeGDP: Source-free Cross-domain Few-shot Learning via Semantic Guided Diversity Prompting
abstract
Source-free cross-domain few-shot learning (SF-CDFSL) aims to transfer pre-trained models to target domains with minimal samples, eliminating the need for source domain data. However, limited samples constrain visual diversity and cross-domain images lack inherent semantic context or prior knowledge, impairing the feature discriminability and generalization of large-scale pre-trained models, affecting transfer performance. To tackle these problems, this article introduces Semantic Guided Diversity Prompting (SeGDP), a method that utilizes semantic guided visual prompts to enhance input diversity. Specifically, SeGDP obtains additional diversity features by concatenating different visual prompts to each support sample, guided by randomly combined and sampled text descriptions during training. Additionally, deep prompt tuning and adapter are introduced to learn static knowledge and further enhance the model’s cross-domain adaptation capability. Extensive experimental results across multiple benchmarks demonstrate that the proposed SeGDP achieves state-of-the-art (SOTA) performance under SF-CDFSL task, and rivals the performance of leading source-utilized models. Our code is available on https://github.com/qwzlh/TOMM_submission .
Linhai Zhuo, Zheng Wang 0059, Tianwen Qian, Yuqian Fu
ACM Trans. Multim. Comput. Commun. Appl.3
2025 NeighborRetr: Balancing Hub Centrality in Cross-Modal Retrieval
abstract
Cross-modal retrieval aims to bridge the semantic gap between different modalities, such as visual and textual data, enabling accurate retrieval across them. Despite significant advancements with models like CLIP that align cross-modal representations, a persistent challenge remains: the hubness problem, where a small subset of samples (hubs) dominate as nearest neighbors, leading to biased representations and degraded retrieval accuracy. Existing methods often mitigate hubness through post-hoc normalization techniques, relying on prior data distributions that may not be practical in real-world scenarios. In this paper, we directly mitigate hubness during training and introduce NeighborRetr, a novel method that effectively balances the learning of hubs and adaptively adjusts the relations of various kinds of neighbors. Our approach not only mitigates the hubness problem but also enhances retrieval performance, achieving state-of-the-art results on multiple cross-modal retrieval benchmarks. Furthermore, Neighbor-Retr demonstrates robust generalization to new domains with substantial distribution shifts, highlighting its effectiveness in real-world applications. We make our code publicly available at: https://github.com/NeighborRetr.
Zengrong Lin, Zheng Wang 0059, Tianwen Qian, Pan Mu, Sixian Chan 0001, Cong Bai
CVPR3
2025 HSACNet: Hierarchical Scale-Aware Consistency Regularized Semi-Supervised Change Detection
abstract
Semi-Supervised change detection (SSCD) aims to detect changes between bi-temporal remote sensing images by utilizing limited labeled data and abundant unlabeled data. Existing methods struggle in complex scenarios, exhibiting poor performance when confronted with noisy data. They typically neglect intra-layer multi-scale features while emphasizing inter-layer fusion, harming the integrity of change objects with different scales. In this paper, we propose HSACNet, a Hierarchical Scale-Aware Consistency regularized Network for SSCD. Specifically, we integrate Segment Anything Model 2 (SAM2), using its Hiera backbone as the encoder to extract inter-layer multi-scale features and applying adapters for parameter-efficient fine-tuning. Moreover, we design a Scale-Aware Differential Attention Module (SADAM) that can precisely capture intra-layer multi-scale change features and suppress noise. Additionally, a dual-augmentation consistency regularization strategy is adopted to effectively utilize the unlabeled data. Extensive experiments across four CD benchmarks demonstrate that our HSACNet achieves state-of-the-art performance, with reduced parameters and computational cost.
Qi'ao Xu, Pengfei Wang 0009, Tianwen Qian, Xiaoling Wang 0004
ICME4
2025 Look Before You Decide: Prompting Active Deduction of MLLMs for Assumptive Reasoning
Yian Li, Wentao Tian, Tianwen Qian, Na Zhao 0004, Bin Zhu 0006, Jingjing Chen 0001, Yu-Gang Jiang 0001
ACM Multimedia4
2025 Domain-RAG: Retrieval-Guided Compositional Image Generation for Cross-Domain Few-Shot Object Detection
abstract
Cross-Domain Few-Shot Object Detection (CD-FSOD) aims to detect novel objects with only a handful of labeled samples from previously unseen domains. While data augmentation and generative methods have shown promise in few-shot learning, their effectiveness for CD-FSOD remains unclear due to the need for both visual realism and domain alignment. Existing strategies, such as copy-paste augmentation and text-to-image generation, often fail to preserve the correct object category or produce backgrounds coherent with the target domain, making them non-trivial to apply directly to CD-FSOD. To address these challenges, we propose Domain-RAG, a training-free, retrieval-guided compositional image generation framework tailored for CD-FSOD. Domain-RAG consists of three stages: domain-aware background retrieval, domain-guided background generation, and foreground-background composition. Specifically, the input image is first decomposed into foreground and background regions. We then retrieve semantically and stylistically similar images to guide a generative model in synthesizing a new background, conditioned on both the original and retrieved contexts. Finally, the preserved foreground is composed with the newly generated domain-aligned background to form the generated image. Without requiring any additional supervision or training, Domain-RAG produces high-quality, domain-consistent samples across diverse tasks, including CD-FSOD, remote sensing FSOD, and camouflaged FSOD. Extensive experiments show consistent improvements over strong baselines and establish new state-of-the-art results. Codes will be released upon acceptance.The source code and instructions are available at https://github.com/LiYu0524/Domain-RAG.
Yu Li 0007, Xingyu Qiu, Yuqian Fu, Tianwen Qian, Xu Zheng 0002, Danda Pani Paudel, Yanwei Fu 0001, Xuanjing Huang 0001, Luc Van Gool, Yu-Gang Jiang 0001
NeurIPS5
2024 NuScenes-QA: A Multi-Modal Visual Question Answering Benchmark for Autonomous Driving Scenario
abstract
We introduce a novel visual question answering (VQA) task in the context of autonomous driving, aiming to answer natural language questions based on street-view clues. Compared to traditional VQA tasks, VQA in autonomous driving scenario presents more challenges. Firstly, the raw visual data are multi-modal, including images and point clouds captured by camera and LiDAR, respectively. Secondly, the data are multi-frame due to the continuous, real-time acquisition. Thirdly, the outdoor scenes exhibit both moving foreground and static background. Existing VQA benchmarks fail to adequately address these complexities. To bridge this gap, we propose NuScenes-QA, the first benchmark for VQA in the autonomous driving scenario, encompassing 34K visual scenes and 460K question-answer pairs. Specifically, we leverage existing 3D detection annotations to generate scene graphs and design question templates manually. Subsequently, the question-answer pairs are generated programmatically based on these templates. Comprehensive statistics prove that our NuScenes-QA is a balanced large-scale benchmark with diverse question formats. Built upon it, we develop a series of baselines that employ advanced 3D detection and VQA techniques. Our extensive experiments highlight the challenges posed by this new task. Codes and dataset are available at https://github.com/qiantianwen/NuScenes-QA.
Tianwen Qian, Jingjing Chen 0001, Linhai Zhuo, Yu-Gang Jiang 0001
AAAI1
2024 An Experiment on 5G Open-RAN Platform with Sub-THz Backhauling
abstract
The deployment of fifth-generation (5G) open radio access network (RAN) introduces new opportunities for wireless communication, particularly in the context of x-Haul wireless links. This paper presents a comprehensive demonstration of how photonic-generated sub- THz at 120 GHz was utilized for wireless backhaul within the 5G open RAN platform. The experimental setup highlighted the feasibility and effectiveness of establishing a sub- THz link as the backhaul, marking a significant advancement in this emerging field. Through extensive performance analysis, we evaluated 5G network's throughput, achieving approximately 957 Mbps for downlink with UDP and 269 Mbps with TCP. In the case of uplink, both UDP and TCP exhibited comparable speeds, both exceeding 80 Mbps. This work marks a significant achievement in advancing the development of photonic-generated sub- THz wireless links, holding the potential to enhance wireless transport capabilities, particularly in the context of advanced 5G networks and beyond..
Norshahida Saba, Efstathios Andrianopoulos, Nikolaos K. Lyras, Saimanoj Katta, Mohammedadem Abdulkadir, Evangelos Pikasis, Garrit Schwanke, Tianwen Qian, Milan Deumer, Simon Nellen, Elias D. Tsirbas, Eleftherios K. Loghis, Georgios Megas, Christos Tsokos, Panos Groumas, David de Felipe, Robert B. Kohlhaas, Maria Massaouti, Ch. Kouloumentas, Dimitrios Kritharidis, Norbert Keil, Hercules Avramopoulos, José Costa-Requena, Riku Jäntti
WCNC8
2024 Locate Before Answering: Answer Guided Question Localization for Video Question Answering
abstract
Video question answering (VideoQA) is an essential task in vision-language understanding, which has attracted numerous research attention recently. Nevertheless, existing works mostly achieve promising performances on short videos of duration within 15 seconds. For VideoQA on minute-level long-term videos, those methods are likely to fail because of lacking the ability to deal with noise and redundancy caused by scene changes and multiple actions in the video. Considering the fact that the question often remains concentrated in a short temporal range, we propose to first locate the question to a segment in the video and then infer the answer using the located segment only. Under this scheme, we propose “Locate before Answering” (LocAns), a novel approach that integrates a question localization module and an answer prediction module into an end-to-end model. During the training phase, the available answer label not only serves as the supervision signal of the answer prediction module, but also is used to generate pseudo temporal labels for the question localization module. Moreover, we design a decoupled alternative training strategy to update the two modules separately. In the experiments, LocAns achieves state-of-the-art performance on three modern long-term VideoQA datasets, NExT-QA, ActivityNet-QA, and AGQA. Its qualitative examples show the reliable performance of the question localization.
Tianwen Qian, Ran Cui, Jingjing Chen 0001, Yu-Gang Jiang 0001
IEEE Trans. Multim.1
2023 Scene Graph Refinement Network for Visual Question Answering
abstract
Visual Question Answering aims to answer the free-form natural language question based on the visual clues in a given image. It is a difficult problem as it requires understanding the fine-grained structured information of both language and image for compositional reasoning. To establish the compositional reasoning, recent works attempt to introduce the scene graph in VQA. However, as the generated scene graphs are usually quite noisy, it greatly limits the performance of question answering. Therefore, this paper proposes to refine the scene graphs for improving the effectiveness. Specifically, we present a novelSceneGraphRefinement network (SGR), which introduces a transformer-based refinement network to enhance the object and relation features for better classification. Moreover, as the question provides valuable clues for distinguishing whether the$\left\langle \mathit{subject, predicate, object} \right\rangle$triplets are helpful or not, the SGR network exploits the semantic information presented in the questions to select the most relevant relations for question answering. Extensive experiments are conducted on the GQA benchmark demonstrate the effectiveness of our method.
Tianwen Qian, Jingjing Chen 0001, Shaoxiang Chen 0001, Bo Wu 0018, Yu-Gang Jiang 0001
IEEE Trans. Multim.1
2022 Video Moment Retrieval from Text Queries via Single Frame Annotation
abstract
Video moment retrieval aims at finding the start and end timestamps of a moment (part of a video) described by a given natural language query. Fully supervised methods need complete temporal boundary annotations to achieve promising results, which is costly since the annotator needs to watch the whole moment. Weakly supervised methods only rely on the paired video and query, but the performance is relatively poor. In this paper, we look closer into the annotation process and propose a new paradigm called "glance annotation". This paradigm requires the timestamp of only one single random frame, which we refer to as a "glance", within the temporal boundary of the fully supervised counterpart. We argue this is beneficial because comparing to weak supervision, trivial cost is added yet more potential in performance is provided. Under the glance annotation setting, we propose a method named as Video moment retrieval via Glance Annotation (ViGA) based on contrastive learning. ViGA cuts the input video into clips and contrasts between clips and queries, in which glance guided Gaussian distributed weights are assigned to all clips. Our extensive experiments indicate that ViGA achieves better results than the state-of-the-art weakly supervised methods by a large margin, even comparable to fully supervised methods in some cases.
Ran Cui, Tianwen Qian, Jingjing Chen 0001, Huyang Sun, Yu-Gang Jiang 0001
SIGIR2