Zongheng Tang

dblp:278/2959 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
9since 2021 · last 2026
0000-0002-9903-802XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021
YearPublicationVenuePosition
2026 AerialVLA: A Vision-Language-Action Model for Aerial Navigation with Online Dialogue
abstract
Visual Dialogue Navigation (VDN) aims to enable agents to reach target locations through dialogue with humans. The integration of VDN into Unmanned Aerial Vehicle (UAV) systems enhances human-machine interaction by enabling intuitive, hands-free operation, thereby unlocking vast applications. However, existing VDN models for UAVs can only perform navigation based on dialogue history, lacking proactive interaction capabilities to correct trajectories. Moreover, their sequential observation history recording mechanism struggles to accurately localize landmarks observed in the historical context, leading to ineffective utilization of referential information in new user instructions.To address these, we present AerialVLA, an end-to-end UAV navigation framework integrating dialogue comprehension, action decision-making, and navigational question generation. AerialVLA comprises three core components: i) we propose the Progress-Driven Navigation-Query Alternation mechanism to determine optimal questioning timing through navigation progress estimation autonomously. ii) To effectively model long-horizon history observation sequences, we develop the History Spatial-Temporal Fusion module that extracts discriminative spatial-temporal representations from historical observations. iii) Furthermore, to overcome data scarcity in training, we devise the Online Task-Driven Augmentation strategy that enhances learning through action-conditioned data augmentation. Experimental results demonstrate that AerialVLA achieves state-of-the-art navigation performance while exhibiting effective dialogue capabilities.Moreover, to better evaluate the agent's proactive dialogue and navigation abilities, our evaluation benchmark, named UAV Navigation with Online Dialogue (UNOD), incorporates an online dialogue interaction module. The UNOD assesses UAV agents' real-time questioning capabilities by leveraging an Air Commander Large Language Model to simulate human-UAV interactions during testing.
Zongheng Tang, Xiaoduo Li, Si Liu 0001
AAAI3
2025 Unleashing the Temporal-Spatial Reasoning Capacity of GPT for Training-Free Audio and Language Referenced Video Object Segmentation
abstract
In this paper, we propose an Audio-Language-Referenced SAM 2 (AL-Ref-SAM 2) pipeline to explore the training-free paradigm for audio and language-referenced video object segmentation, namely AVS and RVOS tasks. The intuitive solution leverages GroundingDINO to identify the target object from a single frame and SAM 2 to segment the identified object throughout the video, which is less robust to spatiotemporal variations due to a lack of video context exploration. Thus, in our AL-Ref-SAM 2 pipeline, we propose a novel GPT-assisted Pivot Selection (GPT-PS) module to instruct GPT-4 to perform two-step temporal-spatial reasoning for sequentially selecting pivot frames and pivot boxes, thereby providing SAM 2 with a high-quality initial object prompt. Within GPT-PS, two task-specific Chain-of-Thought prompts are designed to unleash GPT’s temporal-spatial reasoning capacity by guiding GPT to make selections based on a comprehensive understanding of video and reference information. Furthermore, we propose a Language-Binded Reference Unification (LBRU) module to convert audio signals into language-formatted references, thereby unifying the formats of AVS and RVOS tasks in the same pipeline. Extensive experiments show that our training-free AL-Ref-SAM 2 pipeline achieves performances comparable to or even better than fully-supervised fine-tuning methods.
Shaofei Huang 0001, Rui Ling, Tianrui Hui, Zongheng Tang, Xiaoming Wei, Jizhong Han, Si Liu 0001
AAAI5
2025 CoST: Efficient Collaborative Perception from Unified Spatiotemporal Perspective
abstract
Collaborative perception shares information among different agents and helps solving problems that individual agents may face, e.g., occlusions and small sensing range. Prior methods usually separate the multi-agent fusion and multi-time fusion into two consecutive steps. In contrast, this paper proposes an efficient collaborative perception that aggregates the observations from different agents (space) and different times into a unified spatio-temporal space simultanesouly. The unified spatio-temporal space brings two benefits, i.e., efficient feature transmission and superior feature fusion. 1) Efficient feature transmission: each static object yields a single observation in the spatial temporal space, and thus only requires transmission only once (whereas prior methods re-transmit all the object features multiple times). 2) superior feature fusion: merging the multi-agent and multi-time fusion into a unified spatial-temporal aggregation enables a more holistic perspective, thereby enhancing perception performance in challenging scenarios. Consequently, our Collaborative perception with Spatio-temporal Transformer (CoST) gains improvement in both efficiency and accuracy. Notably, CoST is not tied to any specific method and is compatible with a majority of previous methods, enhancing their accuracy while reducing the transmission bandwidth.
Zongheng Tang, Yi Liu 0070, Yifan Sun 0003, Yulu Gao, Runsheng Xu, Si Liu 0001
ICCV1
2024 MI3C: Mining intra- and inter-image context for person search
Zongheng Tang, Yulu Gao, Tianrui Hui, Fengguang Peng, Si Liu 0001
Pattern Recognit.1
2024 Linker: Learning Long Short-term Associations for Robust Visual Tracking
abstract
iamese and Transformer trackers have demon strated exceptional performance in visual object tracking. These methods utilize initial and potentially online templates to locate the target in subsequent frames. Despite their success, these trackers are vulnerable to changes in the target's appearance due to slow template updates and interference from similar objects, resulting from the absence of scene information. To address these issues, we introduce a reference region within our tracker. The reference region is updated rapidly, providing short-term scene information. By associating the initial template, reference region, and current search region, we enhance the tracker's ability to adapt to changes in target appearance and discriminate between the target and other objects. Additionally, we propose a novel Reference-Enhance (RE) module, which aggregates contextually relevant information from the reference region to enhance the template feature. Extensive experiments show our method achieves state-of-the-art performance on six popular visual object tracking benchmarks while running at over 40 FPS.
Zizheng Xun, Shangzhe Di, Yulu Gao, Zongheng Tang, Gang Wang 0031, Si Liu 0001, Bo Li 0006
IEEE Trans. Multim.4
2023 DETR with Additional Global Aggregation for Cross-domain Weakly Supervised Object Detection
abstract
This paper presents a DETR-based method for cross-domain weakly supervised object detection (CDWSOD), aiming at adapting the detector from source to target domain through weak supervision. We think DETR has strong potential for CDWSOD due to an insight: the encoder and the decoder in DETR are both based on the attention mechanism and are thus capable of aggregating semantics across the entire image. The aggregation results, i.e., image-level predictions, can naturally exploit the weak supervision for domain alignment. Such motivated, we propose DETR with additional Global Aggregation (DETR-GA), a CDWSOD detector that simultaneously makes “instance-level + image-level” predictions and utilizes “strong + weak” supervisions. The key point of DETR-GA is very simple: for the encoder / decoder, we respectively add multiple class queries / a foreground query to aggregate the semantics into image-level predictions. Our query-based aggregation has two advantages. First, in the encoder, the weakly-supervised class queries are capable of roughly locating the corresponding positions and excluding the distraction from non-relevant regions. Second, through our design, the object queries and the foreground query in the decoder share consensus on the class semantics, therefore making the strong and weak supervision mutually benefit each other for domain alignment. Extensive experiments on four popular cross-domain benchmarks show that DETR-GA significantly improves cross-domain detection accuracy (e.g., 29.0% → 79.4% mAP on PASCAL VOC → Clipartalldataset) and advances the states of the art.
Zongheng Tang, Yifan Sun 0003, Si Liu 0001, Yi Yang 0001
CVPR1
2023 Unified Transformer With Isomorphic Branches for Natural Language Tracking
abstract
Natural language tracking aims to localize the target object referred to by a language description using a sequence of bounding boxes in video frames. Compared with traditional visual single object tracking initialized with only a bounding box (BBox), this task introduces high-level semantic information to reduce the ambiguity of BBox and enhance the ability to retrieve the target in a global manner. Thus, it can yield more accurate and robust tracking results. Previous methods usually adopt off-the-shelf grounding and tracking branches to tackle this task, where feature representations are learned in isolation without benefiting each other. The two branches can associate with each other to discover crucial clues since the language description and the template image provide information from different sources. Therefore, we propose a unified transformer method for natural language tracking named TransNLT, which utilizes isomorphic Transformer structures for grounding and tracking branches where collaborative learning is enabled to construct comprehensive features for the target. In addition, we propose a Selective Feature Gathering (SFG) module which can integrate cross-modal global information of the tracking target from the visual template and language description. Through effective interaction of visual and language information, we can achieve better results than tracking with only a single modality. Extensive experiments on three popular natural language tracking benchmarks show our proposed TransNLT outperforms previous state-of-the-art methods.
Zongheng Tang, Qianli Zhou, Tianrui Hui, Quange Tan, Si Liu 0001
IEEE Trans. Circuits Syst. Video Technol.2
2022 PIC'22: 4th Person in Context Workshop
abstract
Understanding human and the surrounding context is crucial for the perception of the image and video. It benefits many related applications, such as person search, virtual tryon/makeup, abnormal action detection. In the proposed 4th Person in Context (PIC) workshop, to further promote the progress in the above-mentioned areas, we hold three human-centric perception and cognition challenges including Make-up Temporal Video Grounding (MTVG), Make-up Dense Video Caption (MDVC) and Human-centric Spatio-Temporal Video Grounding (HC-STVG). All the human-centric challenges focus on understanding the human behavior, interactions and relationships in video sequences, which requires understanding both visual and linguistic information, as well as complicated multimodal reasoning. The three sub-problems are complementary and collaboratively contribute to a unified human-centric perception and cognition solution.
Si Liu 0001, Qin Jin, Luoqi Liu, Zongheng Tang, Linli Lin
ACM Multimedia4
2022 Human-Centric Spatio-Temporal Video Grounding With Visual Transformers
abstract
In this work, we introduce a novel task – Human-centric Spatio-Temporal Video Grounding (HC-STVG). Unlike the existing referring expression tasks in images or videos, by focusing on humans, HC-STVG aims to localize a spatio-temporal tube of the target person from an untrimmed video based on a given textural description. This task is useful, especially for healthcare and security related applications, where the surveillance videos can be extremely long but only a specific person during a specific period is concerned. HC-STVG is a video grounding task that requires both spatial (where) and temporal (when) localization. Unfortunately, the existing grounding methods cannot handle this task well. We tackle this task by proposing an effective baseline method named Spatio-Temporal Grounding with Visual Transformers (STGVT), which utilizes Visual Transformers to extract cross-modal representations for video-sentence matching and temporal localization. To facilitate this task, we also contribute an HC-STVG datasetThe new dataset is available athttps://github.com/tzhhhh123/HC-STVG. consisting of 5,660 video-sentence pairs on complex multi-person scenes. Specifically, each video lasts for 20 seconds, pairing with a natural query sentence with an average of 17.25 words. Extensive experiments are conducted on this dataset, demonstrating that the newly-proposed method outperforms the existing baseline methods.
Zongheng Tang, Yue Liao, Si Liu 0001, Guanbin Li, Xiaojie Jin 0004, Qian Yu 0002, Dong Xu 0001
IEEE Trans. Circuits Syst. Video Technol.1