EDBT 2026 Demo / reviewers in the wild / expert
Yuankai Qi
dblp:136/5491
· DBLP profile ↗
83ranked-venue papers
11as first author
64since 2021 · last 2026
0000-0003-4312-5682ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 58 · 9 first-author · 43 since 2021Artificial intelligence and machine learning · 52 · 7 first-author · 41 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Tracking the Unstable: Appearance-Guided Motion Modeling for Robust Multi-Object Tracking in UAV-Captured VideosabstractMulti-object tracking (MOT) aims to track multiple objects while maintaining consistent identities across frames of a given video. In unmanned aerial vehicle (UAV) recorded videos, frequent viewpoint changes and complex UAV-ground relative motion dynamics pose significant challenges, which often lead to unstable affinity measurement and ambiguous association. Existing methods typically model motion and appearance cues separately, overlooking their spatio-temporal interplay and resulting in suboptimal tracking performance. In this work, we propose AMOT, which jointly exploits appearance and motion cues through two key components: an Appearance-Motion Consistency (AMC) matrix and a Motion-aware Track Continuation (MTC) module. Specifically, the AMC matrix computes bi-directional spatial consistency under the guidance of appearance features, enabling more reliable and context-aware identity association. The MTC module complements AMC by reactivating unmatched tracks through appearance-guided predictions that align with Kalman-based predictions, thereby reducing broken trajectories caused by missed detections. Extensive experiments on three UAV benchmarks, including VisDrone2019, UAVDT, and VT-MOT-UAV, demonstrate that our AMOT outperforms current state-of-the-art methods and generalizes well in a plug-and-play and training-free manner. Jianbo Ma 0001, Hui Luo 0002, Qi Chen 0014, Yuankai Qi, Yumei Sun, Amin Beheshti, Jianlin Zhang 0001, Ming-Hsuan Yang 0001 |
AAAI | 4 |
| 2026 | InstructDubber: Instruction-based Alignment for Zero-shot Movie DubbingabstractMovie dubbing seeks to synthesize speech from a given script using a specific voice, while ensuring accurate lip synchronization and emotion-prosody alignment with the character’s visual performance. However, existing alignment approaches based on visual features face two key limitations: (1) they rely on complex, handcrafted visual preprocessing pipelines, including facial landmark detection and feature extraction; and (2) they generalize poorly to unseen visual domains, often resulting in degraded alignment and dubbing quality. To address these issues, we propose InstructDubber, a novel instruction-based alignment dubbing method for both robust in-domain and zero-shot movie dubbing. Specifically, we first feed the video, script, and corresponding prompts into a multimodal large language model to generate natural language dubbing instructions regarding the speaking rate and emotion state depicted in the video, which is robust to visual domain variations. Second, we design an instructed duration distilling module to mine discriminative duration cues from speaking rate instructions to predict lip-aligned phoneme-level pronunciation duration. Third, for emotion-prosody alignment, we devise an instructed emotion calibrating module, which fine-tunes an LLM-based instruction analyzer using ground truth dubbing emotion as supervision and predicts prosody based on the calibrated emotion analysis. Finally, the predicted duration and prosody, together with the script, are fed into the audio decoder to generate video-aligned dubbing. Extensive experiments on three major benchmarks demonstrate that InstructDubber outperforms state‑of‑the‑art approaches across both in‑domain and zero‑shot scenarios. Zhedong Zhang, Liang Li 0003, Gaoxiang Cong 0001, Chunshan Liu, Xiaowan Wang, Tao Gu 0001, Yuankai Qi |
AAAI | 8 |
| 2026 | Teacher Agent: A Knowledge Distillation-Free Framework for Rehearsal-Based Video Incremental Learning
Shengqin Jiang, Yaoyu Fang, Haokui Zhang, Qingshan Liu 0001, Yuankai Qi, Yang Yang 0002, Peng Wang 0023 |
Int. J. Comput. Vis. | 5 |
| 2026 | Collaborative Temporal Consistency Learning for Point-supervised Natural Language Video Localization
Zhuo Tao, Liang Li 0003, Qi Chen 0014, Yunbin Tu, Zhengjun Zha, Amin Beheshti, Qingming Huang, Yuankai Qi, Ming-Hsuan Yang 0001 |
Int. J. Comput. Vis. | 8 |
| 2026 | SeqCount: A sequence modeling framework for class-agnostic countingabstractClass-agnostic counting aims to count the number of objects in any category with only a few exemplars. It is crucial for solving the challenge of counting any visual class without re-finetuning, which in turn lowers deployment costs across diverse scenarios. Existing methods count the number of exemplar objects by integrating them over the density map smoothed by Gaussian kernels. However, designing generic kernels to generate density maps is challenging due to the different sizes and shapes of objects. To solve this problem, we propose SeqCount, which eliminates the need for density maps and treats object counting as a sequence generation problem. Specifically, we consider an input image as $N\times N$ patches and propose a serialization scheme. Then, we use an encoding-decoding structure to exploit the correlation among the patches. Experimental results on five challenging datasets demonstrate that our method performs favorably against the state-of-the-art models. Guorong Li, Xinyan Liu 0008, Zhenjun Han, Yuankai Qi |
J. Vis. Commun. Image Represent. | 6 |
| 2026 | Consistency-Aware Anchor Pyramid Network for Crowd LocalizationabstractCrowd localization aims to predict the positions of humans in images of crowded scenes. While existing methods have made significant progress, two primary challenges remain: (i) a fixed number of evenly distributed anchors can cause excessive or insufficient predictions across regions in an image with varying crowd densities, and (ii) ranking inconsistency of predictions between the testing and training phases leads to the model being sub-optimal in inference. To address these issues, we propose a Consistency-Aware Anchor Pyramid Network (CAAPN) comprising two key components: an Adaptive Anchor Generator (AAG) and a Localizer with Augmented Matching (LAM). The AAG module adaptively generates anchors based on estimated crowd density in local regions to alleviate the anchor deficiency or excess problem. It also considers the spatial distribution prior to heads for better performance. The LAM module is designed to augment the predictions which are used to optimize the neural network during training by introducing an extra set of target candidates and correctly matching them to the ground truth. The proposed method achieves favorable performance against state-of-the-art approaches on five challenging datasets: ShanghaiTech A and B, UCF-QNRF, JHU-CROWD++, and NWPU-Crowd. Xinyan Liu 0008, Guorong Li, Yuankai Qi, Zhenjun Han, Anton van den Hengel, Nicu Sebe, Ming-Hsuan Yang 0001, Qingming Huang |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2026 | Dynamic example network for class-agnostic object counting
Xinyan Liu 0008, Guorong Li, Yuankai Qi, Ziheng Yan, Weigang Zhang, Laiyun Qing, Qingming Huang |
Pattern Recognit. | 3 |
| 2026 | RETTA: Retrieval-enhanced test-time adaptation for zero-shot video captioningabstractDespite the significant progress of fully-supervised video captioning, zero-shot methods remain much less explored. In this paper, we propose a novel zero-shot video captioning framework named R etrieval- E nhanced T est- T ime A daptation (RETTA), which takes advantage of existing pre-trained large-scale vision and language models to directly generate captions with test-time adaptation. Specifically, we bridge video and text using four key models: a general video-text retrieval model XCLIP, a general image-text matching model CLIP, a text alignment model AnglE, and a text generation model GPT-2, due to their source-code availability. The main challenge is how to enable the text generation model to be sufficiently aware of the content in a given video so as to generate corresponding captions. To address this problem, we propose using learnable tokens as a communication medium among these four frozen models GPT-2, XCLIP, CLIP, and AnglE. Different from the conventional way that trains these tokens with training data, we propose to learn these tokens with soft targets of the inference data under several carefully crafted loss functions, which enable the tokens to absorb video information catered for GPT-2. This adaptation requires only a few iterations ( e.g. , 16) and does not require ground truth data. Extensive experimental on MSR-VTT, MSVD, and VATEX, show absolute 5.1 % ∼ 32.4 % improvements in CIDEr scores compared to several state-of-the-art zero-shot video captioning methods. Yunchuan Ma, Laiyun Qing, Guorong Li, Yuankai Qi, Amin Beheshti, Quan Z. Sheng, Qingming Huang |
Pattern Recognit. | 4 |
| 2026 | Parameter-efficient action planning with large language models for vision-and-language navigationabstract• In REVERIE, the agent needs fine-grained instructions to act more efficiently. • Large Language Models (LLMs) show great potential for the navigation task. • Generating goal plans and single-step instructions improves the agent’s performance. • Parameter-efficient fine-tuning the LLMs results in more accurate navigation plans. • Designing precise prompts alongside fine-tuning mitigates hallucination generation. The remote embodied referring expression (REVERIE) task requires an agent to navigate through complex indoor environments and localize a remote object specified by high-level instructions, such as “ bring me a spoon ”, without pre-exploration. Hence, an efficient navigation plan is essential for the final success. This paper proposes a novel parameter-efficient action planner using large language models (PEAP-LLM) to generate a single-step instruction at each location. The proposed model consists of two modules, LLM goal planner (LGP) and LoRA action planner (LAP). Initially, LGP extracts the goal-oriented plan from REVERIE instructions, including the target object and room. Then, LAP generates a single-step instruction with the goal-oriented plan, high-level instruction, and current visual observation as input. PEAP-LLM enables the embodied agent to interact with LAP as the path planner on the fly. A simple direct application of LLMs hardly achieves good performance. Moreover, existing hard-prompt-based methods are prone to errors in complex scenarios and require human intervention. To address these issues and prevent the LLM from generating hallucinations and biased information, we propose a novel two-stage method for fine-tuning the LLM, consisting of supervised fine-tuning (SFT) and direct preference optimization (DPO). SFT improves the quality of the generated instructions, while DPO incorporates environmental feedback into fine-tuning. Experimental results demonstrate a noticeable enhancement over the baseline model on the REVERIE benchmark. Bahram Mohammadi, Ehsan Abbasnejad, Yuankai Qi, Qi Wu 0001, Anton van den Hengel, Qinfeng Shi |
Pattern Recognit. | 3 |
| 2026 | Link prediction on multi-relational graphs from an influence propagation perspective
Zidu Yin, Yuankai Qi, Dong Gong, Ehsan Abbasnejad, Kun Yue, Qinfeng Shi |
Pattern Recognit. | 2 |
| 2026 | Ghost-Free HDR Imaging via Latent Low-Frequency Priors and Deformable Attention AlignmentabstractRecovering ghost-free High Dynamic Range (HDR) images from multiple Low Dynamic Range (LDR) images becomes challenging when the LDR images exhibit saturation and significant motion. Recent Diffusion Models (DMs) have been introduced in HDR imaging field, showing promising performance, particularly in achieving visually perceptible better results compared to previous DNN-based methods. However, DMs require extensive iterations with large models to estimate entire images, resulting in inefficiency that hinders their practical application. To address this challenge, we propose the Low-Frequency aware Diffusion (LF-Diff) model for ghost-free HDR imaging. The key idea of LF-Diff is implementing the DMs in a highly compacted latent space and integrating it into a regression-based model to enhance the details of reconstructed images. Specifically, as low-frequency information is closely related to human visual perception we propose to utilize DMs to create compact low-frequency priors for the reconstruction process. These priors are integrated into a carefully designed Dynamic HDR Reconstruction Network (DHRNet), which employs a regression-based approach to produce high-quality HDR images. Furthermore, we introduce the Attention-guided Deformable Alignment Module (ADAM) that utilizes correlation-driven feature matching to learn deformable receptive fields for self-attention, enabling efficient pre-alignment of LDR images by focusing on salient regions. Extensive experiments on synthetic and real-world benchmark datasets demonstrate that our LF-Diff performs favorably against several state-of-the-art methods and is $10\times $ faster than previous DM-based methods. Tao Hu 0013, Qingsen Yan, Wei Dong 0010, Peng Wu 0015, Yuankai Qi, Weisi Lin, Yanning Zhang 0001 |
IEEE Trans. Image Process. | 6 |
| 2025 | Incomplete Multi-View Multi-Label Classification via Diffusion-Guided Redundancy RemovalabstractIncomplete multi-view multi-label classification aims to accurately predict labels for each sample in the face of some missing views. Due to its widespread presence in real-world scenarios, it has become an extensively researched topic. In addition to the challenges brought by missing views, it also encounters issues caused by redundant views, whose inclusion fails to make a positive contribution to performance. In this paper, we make the first attempt to take advantage of diffusion models to address the missing view problem and design a strategy to identify and remove redundant views. Specifically, we train a diffusion model conditioned on the pseudo-labels to recover information of missing views. The learned diffusion model can carry data distribution knowledge in training split to the data. Regarding redundant identification strategy, it is designed by considering both the additional information of views and the classification difficulty level of samples, thereby adaptively identifying and removing redundant views. We conduct extensive experiments on five datasets, and the proposed method achieves favorable performance against several state-of-the-art methods on the multi-view multi-label classification task. Shilong Ou, Zhe Xue, Lixiong Qin, Yawen Li 0001, Meiyu Liang, Junjiang Wu, Xuyun Zhang, Amin Beheshti, Yuankai Qi |
AAAI | 9 |
| 2025 | Generating Synthetic Data for Unsupervised Federated Learning of Cross-Modal RetrievalabstractUnsupervised federated learning for cross-modal retrieval has received increasing attention in recent years as it can free the requirement for annotations and avoid uploading original clients’ data to servers. Most existing methods focus on how to learn better local models and their aggregation to overcome data distribution drift across clients. Unlike prior works, we propose to address the data distribution problem by generating synthetic data, which can benefit existing federated learning methods. Specifically, we train a WGAN generator with three newly designed loss constraints on each client to improve the quality of the generated data. We first compute cluster prototypes to address the problem of lack of labels. Then, a direct contrastive loss between generated image and text features, an indirect contrastive loss with reference to cluster prototypes, and a Jensen-Shannon Divergence (JSD) loss also with reference to cluster prototypes work together to constrain the WGAN. The locally trained generators and local prototypes are sent to the server to generate and filter synthetic data with consideration of data distribution across all clients. The filtered data are used to train the aggregated global retrieval model, which is later sent to clients. The final global model becomes robust to all clients after several rounds of client-server iteration. Extensive experiments using four baselines across three datasets demonstrate that our method performs favourably against state-of-the-art methods. Tianlong Zhang, Zhe Xue, Mahmood Adnan, Junping Du 0001, Yuchen Dong, Shilong Ou, Lang Feng 0006, Ming-Hsuan Yang 0001, Yuankai Qi |
AAAI | 9 |
| 2025 | STGS: Spatio-temporal Graph Sparsification Using Reinforcement LearningabstractSpatio-temporal graphs encode dynamic interactions across space and time, but their size and complexity pose challenges for analysis and computation. Graph sparsification provides an effective solution to these issues by reducing the number of edges while preserving the essential structural and dynamic properties of the network. This reduction is crucial for enhancing the interpretability of complex graphs, revealing hidden patterns, and enabling more efficient computational analysis. However, real-world graphs often exhibit continuous spatial and temporal evolution, which most existing sparsification algorithms, primarily designed for static graphs, fail to address. We introduce STGS (Spatio-Temporal Graph Sparsification), a reinforcement learning-based framework for sparsifying spatio-temporal graphs. By learning to prune edges while preserving key spatio-temporal patterns, STGS enables efficient analysis of evolving systems. Experiments on real-world datasets demonstrate that STGS outperforms existing methods in both structural preservation and downstream forecasting tasks. Nasrin Shabani, Amin Beheshti, Yuankai Qi, Venus Haghighi, Jin Foo, Jia Wu 0001 |
CIKM | 3 |
| 2025 | EmoDubber: Towards High Quality and Emotion Controllable Movie DubbingabstractGiven a piece of text, a video clip, and a reference audio, the movie dubbing task aims to generate speech that aligns with the video while cloning the desired voice. The existing methods have two primary deficiencies: (1) They struggle to simultaneously hold audio-visual sync and achieve clear pronunciation; (2) They lack the capacity to express user-defined emotions. To address these problems, we propose EmoDubber, an emotion-controllable dubbing architecture that allows users to specify emotion type and emotional intensity while satisfying high-quality lip sync and pronunciation. Specifically, we first design Lip-related Prosody Aligning (LPA), which focuses on learning the inherent consistency between lip motion and prosody variation by duration level contrastive learning to incorporate reasonable alignment. Then, we design Pronunciation Enhancing (PE) strategy to fuse the video-level phoneme sequences by efficient conformer to improve speech intelligibility. Next, the speaker identity adapting module decodes acoustics prior and inject the speaker style embedding. After that, the proposed Flow-based User Emotion Controlling (FUEC) is used to synthesize waveform by flow matching prediction network conditioned on acoustics prior. In this process, the FUEC determines the gradient direction and guidance scale based on the user’s emotion instructions by the positive and negative guidance mechanism, which focuses on amplifying the desired emotion while suppressing others. Extensive experimental results demonstrate favorable performance compared to several state-of-the-art methods. The code and trained models will be made available at https://github.com/GalaxyCong/DubFlow. Gaoxiang Cong 0001, Jiadong Pan, Liang Li 0003, Yuankai Qi, Yuxin Peng 0001, Anton van den Hengel, Jian Yang 0001, Qingming Huang |
CVPR | 4 |
| 2025 | Medusa: A Multi-Scale High-order Contrastive Dual-Diffusion Approach for Multi-View ClusteringabstractDeep multi-view clustering methods utilize information from multiple views to achieve enhanced clustering results and have gained increasing popularity in recent years. Most existing methods typically focus on either inter-view or intra-view relationships, aiming to align information across views or analyze structural patterns within individual views. However, they often incorporate inter-view complementary information in a simplistic manner, while overlooking the complex, high-order relationships within multi-view data and the interactions among samples, resulting in an incomplete utilization of the rich information available. Instead, we propose a multi-scale approach that exploits all of the available information. We first introduce a dual graph diffusion module guided by a consensus graph. This module leverages inter-view information to enhance the representation of both nodes and edges within each view. Secondly, we propose a novel contrastive loss function based on hypergraphs to more effectively model and leverage complex intra-view data relationships. Finally, we propose to adaptively learn fusion weights at the sample level, which enables a more flexible and dynamic aggregation of multi-view information. Extensive experiments on eight datasets show favorable performance of the proposed method compared to state-of-the-art approaches, demonstrating its effectiveness across diverse scenarios. Liang Chen 0030, Zhe Xue, Yawen Li 0001, Meiyu Liang, Yan Wang 0002, Anton van den Hengel, Yuankai Qi |
CVPR | 7 |
| 2025 | Visual and Semantic Prompt Collaboration for Generalized Zero-Shot LearningabstractGeneralized zero-shot learning aims to recognize both seen and unseen classes with the help of semantic information that is shared among different classes. It inevitably requires consistent visual-semantic alignment. Existing approaches fine-tune the visual backbone by seen-class data to obtain semantic-related visual features, which may cause overfitting on seen classes with a limited number of training images. This paper proposes a novel visual and semantic prompt collaboration framework, which utilizes prompt tuning techniques for efficient feature adaptation. Specifically, we design a visual prompt to integrate the visual information for discriminative feature learning and a semantic prompt to integrate the semantic formation for visual-semantic alignment. To achieve effective prompt information integration, we further design a weak prompt fusion mechanism for the shallow layers and a strong prompt fusion mechanism for the deep layers in the network. Through the collaboration of visual and semantic prompts, we can obtain discriminative semantic-related features for generalized zero-shot image recognition. Extensive experiments demonstrate that our framework consistently achieves favorable performance in both conventional zero-shot learning and generalized zero-shot learning benchmarks compared to other state-of-the-art methods. Huajie Jiang, Zhengxian Li, Xiaohan Yu 0001, Yongli Hu, Jian Yang 0001, Yuankai Qi |
CVPR | 7 |
| 2025 | Separation of Powers: On Segregating Knowledge from Observation in LLM-enabled Knowledge-based Visual Question AnsweringabstractKnowledge-Based visual question answering (KBVQA) separates image interpretation and knowledge retrieval into separate processes, motivated in part by the fact that they are very different tasks. In this paper, we transform the KB-VQA into linguistic question-answering tasks so that we can leverage the rich world knowledge and strong reasoning abilities of Large Language Models (LLMs). The caption-then-question approach to KBVQA has been effective, but relies on the captioning method to describe the detail required to answer every possible question. We propose instead a Question-Aware Captioner (QACap), which uses the question as guidance to extract correlated visual information from the image and generate a question-related caption. To train such a model, we utilize GPT-4 to build a corresponding high-quality question-aware caption dataset on top of existing KBVQA datasets. Extensive experiments demonstrate that our QACap model and dataset significantly improve KBVQA performance. Our method, QA-Cap, achieves 68.2% accuracy on the OKVQA validation set, 73.4% on the direct-answer part of the A-OKVQA validation set, and 74.8% on the multiple-choice part, all setting new SOTA benchmarks. Zhuo Tao, Qi Chen 0014, Liang Li 0003, Yuankai Qi, Anton van den Hengel, Qingming Huang |
CVPR | 5 |
| 2025 | Prosody-Enhanced Acoustic Pre-training and Acoustic-Disentangled Prosody Adapting for Movie DubbingabstractMovie dubbing describes the process of transforming a script into speech that aligns temporally and emotionally with a given movie clip while exemplifying the speaker’s voice demonstrated in a short reference audio clip. This task demands the model bridge character performances and complicated prosody structures to build a high-quality video-synchronized dubbing track. The limited scale of movie dubbing datasets, along with the background noise inherent in audio data, hinder the acoustic modeling performance of trained models. To address these issues, we propose an acoustic-prosody disentangled two-stage method to achieve high-quality dubbing generation with precise prosody alignment. First, we propose a prosody-enhanced acoustic pre-training to develop robust acoustic modeling capabilities. Then, we freeze the pre-trained acoustic system and design an acoustic-disentangled framework to model prosodic text features and dubbing style while maintaining acoustic quality. Additionally, we incorporate an in-domain emotion analysis module to reduce the impact of visual domain shifts across different movies, thereby enhancing emotion-prosody alignment. Extensive experiments show that our method performs favorably against the state-of-the-art models on two primary benchmarks. The project is available at https://zzdoog.github.io/ProDubber/. Zhedong Zhang, Liang Li 0003, Chenggang Yan 0001, Chunshan Liu, Anton van den Hengel, Yuankai Qi |
CVPR | 6 |
| 2025 | Seeing the Trees for the Forest: Rethinking Weakly-Supervised Medical Visual GroundingabstractVisual grounding (VG) is the capability to identify the specific regions in an image associated with a particular text description. In medical imaging, VG enhances interpretability by highlighting relevant pathological features corresponding to textual descriptions, improving model transparency and trustworthiness for wider adoption of deep learning models in clinical practice. Current models struggle to associate textual descriptions with disease regions due to inefficient attention mechanisms and a lack of fine-grained token representations. In this paper, we empirically demonstrate two key observations. First, current VLMs assign high norms to background tokens, diverting the model's attention from regions of disease. Second, the global tokens used for cross-modal learning are not representative of local disease tokens. This hampers identifying correlations between the text and disease tokens. To address this, we introduce simple, yet effective Disease-Aware Prompting (DAP) process, which uses the explainability map of a VLM to identify the appropriate image features. This simple strategy amplifies disease-relevant regions while suppressing background interference. Without any additional pixel-level annotations, DAP improves visual grounding accuracy by 20.74% compared to state-of-the-art methods across three major chest X-ray datasets. Ta Duc Huy, Duy Anh Huynh, Yutong Xie 0001, Yuankai Qi, Qi Chen 0014, Phi-Le Nguyen, Sen Kim Tran, Son Lam Phung, Anton van den Hengel, Zhibin Liao, Minh-Son To, Johan Verjans, Vu Minh Hieu Phan |
ICCV | 4 |
| 2025 | Hierarchical Prompt-Guided Alignment for Multi-view Clustering
Shifeng Bao, Zhe Xue, Shilong Ou, Amin Beheshti, Yuankai Qi |
ICIC (12) | 5 |
| 2025 | Language-Conditioned Waypoint Predictor for Continuous Vision-and-Language NavigationabstractWaypoint prediction is a popular technique for Vision-and-Language Navigation in Continuous Environments (VLN-CE), which abstracts navigable locations as waypoints to ease the subsequent action prediction. Nevertheless, we found current waypoint predictors are not always accurate, limiting navigation’s overall performance. One possible reason may be the lack of language context, leading to the failure to generate corresponding waypoints for critical locations mentioned in the instructions. To that end, we propose a novel framework to enable the training of the language-conditioned waypoint predictor. First, as the VLN-CE agents ground instructions with the environment when navigating, we employ a pre-trained agent to encode language for the waypoint predictor. Second, the language-conditioned waypoint predictor is trained with the data collected using the same agent. Third, we train the new VLN-CE navigation agent with the proposed waypoint predictor. Fourth, the disparity between the language encoder agent and the navigation agent drives us to devise a cycle training scheme to alternately train the agent and the waypoint predictor, further enhancing the performance of both the waypoint predictor and navigation agent. Experimental results show that our waypoint predictor’s performance surpasses all existing ones. With better waypoints, the gap between waypoint-based methods and their upper bound narrows by about 60%. Yuankai Qi, Xu Yang 0004, Zhaoxiang Zhang 0001 |
ICME | 2 |
| 2025 | FlowDubber: Movie Dubbing with LLM-based Semantic-aware Learning and Flow Matching based Voice EnhancingabstractMovie Dubbing aims to convert scripts into speeches that align with the given movie clip in both temporal and emotional aspects while preserving the vocal timbre of a given brief reference audio. Existing methods focus primarily on reducing the word error rate while ignoring the importance of lip-sync and acoustic quality. To address these issues, we propose a novel dubbing architecture based on Large Language Model (LLM) and Conditional Flow Matching (CFM), named FlowDubber, which achieves high-quality audio-visual sync and pronunciation by incorporating a large speech language model with dual contrastive alignment while improving acoustic quality via Flow-based Voice Enhancing (FVE). First, we introduce Qwen2.5 as the backbone of large speech language model to learn the in-context sequence from movie scripts and reference audio. Second, the proposed semantic-aware learning focuses on capturing LLM semantic knowledge at the phoneme level, which facilitates mutual alignment with lip movement from silent video via Dual Contrastive Alignment (DCA). Third, the FVE introduces an LLM-based acoustics flow matching guidance to strengthen clarity by decoupling Classifier-Free Guidance (CFG) enhancement. Extensive experiments demonstrate that our method outperforms several state-of-the-art methods on two primary benchmarks. The demos are available at https://galaxycong.github.io/LLM-Flow-Dubber/. Gaoxiang Cong 0001, Liang Li 0003, Jiadong Pan, Zhedong Zhang, Amin Beheshti, Anton van den Hengel, Yuankai Qi, Qingming Huang |
ACM Multimedia | 7 |
| 2025 | CausalMVC: Causal Content-Style Representation Learning for Deep Multi-View Clustering
Shifeng Bao, Zhe Xue, Qi Chen 0014, Shilong Ou, Amin Beheshti, Quan Z. Sheng, Anton van den Hengel, Yuankai Qi |
ACM Multimedia | 8 |
| 2025 | SDVPT: Semantic-Driven Visual Prompt Tuning for Open-world Object CountingabstractOpen-world object counting leverages the robust text-image alignment of pre-trained vision-language models (VLMs) to enable counting of arbitrary categories in images specified by textual queries. However, widely adopted naive fine-tuning strategies concentrate exclusively on text-image consistency for categories contained in, which leads to limited generalizability for unseen categories. In this work, we propose a plug-and-play Semantic-Driven Visual Prompt Tuning framework (SDVPT) that transfers knowledge from the training set to unseen categories with minimal overhead in parameters and inference time. First, we introduce a two-stage visual prompt learning strategy composed of Category-Specific Prompt Initialization (CSPI) and Topology-Guided Prompt Refinement (TGPR). The CSPI generates category-specific visual prompts, and then TGPR distills latent structural patterns from the VLM's text encoder to refine these prompts. During inference, we dynamically synthesize the visual prompts for unseen categories based on the semantic correlation between unseen and training categories, facilitating robust text-image alignment for unseen categories. Extensive experiments integrating SDVPT with all available open-world object counting models demonstrate its effectiveness and adaptability across three widely used datasets: FSC-147, CARPK, and PUCPR+. Code is available https://github.com/Eamon-0v0/SDVPT Guorong Li, Laiyun Qing, Amin Beheshti, Jian Yang 0001, Quan Z. Sheng, Yuankai Qi, Qingming Huang |
ACM Multimedia | 7 |
| 2025 | Dubbing Movies via Hierarchical Phoneme Modeling and Acoustic Diffusion DenoisingabstractGiven a piece of text, a video clip, and reference audio, the movie dubbing (also known as Visual Voice Cloning, V2C) task aims to generate speeches that clone reference voice and align well with the video in both emotion and lip movement, which is more challenging than conventional text-to-speech synthesis tasks. To align the generated speech with the inherent lip motion of the given silent video, most existing works utilize each video frame to query textual phonemes. However, such an attention operation usually leads to mumble speech because different phonemes are fused for video frames corresponding to one phoneme (video frames are finer-grained than phonemes). To address this issue, we propose a diffusion-based movie dubbing architecture, which improves pronunciation by Hierarchical Phoneme Modeling (HPM) and generates better mel-spectrogram through Acoustic Diffusion Denoising (ADD). We term our model as HD-Dubber. Specifically, our HPM bridges the visual information and corresponding speech prosody from three aspects: (1) aligning lip movement with the speech duration based on each phoneme unit by contrastive learning; (2) conveying facial expression to phoneme-level energy and pitch; and (3) injecting global emotions captured from video scenes into prosody. On the other hand, ADD exploits a denoising diffusion framework to transform the noise signal into a mel-spectrogram via a parameterized Markov chain conditioned on textual phonemes and reference audio. ADD has two novel denoisers, the Style-adaptive Residual Denoiser (SRD) and the Phoneme-enhanced U-net Denoiser (PUD), to enhance speaker similarity and improve pronunciation quality. Extensive experimental results on the three benchmark datasets demonstrate the state-of-the-art performance of the proposed method. The source code and trained models will be made available to the public. Liang Li 0003, Gaoxiang Cong 0001, Yuankai Qi, Zhengjun Zha, Qi Wu 0001, Quan Z. Sheng, Qingming Huang, Ming-Hsuan Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | Boosting UAV Detection via Memory-Enhanced Attention and Contrastive LearningabstractWith unmanned aerial vehicles (UAVs) having emerged in diverse application domains, visual detection of UAVs has become a critical research focus in recent years. However, most existing methods are limited in capturing small UAVs and may not perform well in complex backgrounds. To address these challenges, we propose a novel detection framework that integrates newly designed memory mechanism and contrastive loss to improve UAV detection. Specifically, we first utilize a clustering algorithm to gather representative UAV prototypes, which are then utilized to construct a reliable memory bank. Then, we design a UAV Memory-Enhanced Attention (UMEA) module to propagate high-confidence prototypes from the memory bank, thereby enhancing the appearance features of UAVs in the input frame. Furthermore, we introduce a Memory-Driven Contrastive Learning (MDCL) loss function to pull UAVs closer in the feature space while pushing them further away from the background. Extensive experiments conducted on three challenging datasets, NPS-Drones, ARD-MAV and Drone-vs-Bird demonstrate that the proposed method outperforms several state-of-the-art models in terms of the main metric AP with a large absolute margin, 2.1%, 3.6%, and 4.4%, respectively. Yunchuan Ma, Yuankai Qi, Laiyun Qing, Guorong Li |
IEEE Signal Process. Lett. | 3 |
| 2025 | Dual Prototype Contrastive Network for Generalized Zero-Shot LearningabstractGeneralized zero-shot learning (GZSL) requires that models are able to recognize classes they were trained on, and new classes they haven't seen before. Feature-generation approaches are popular due to their effectiveness in mitigating overfitting to the training classes. Existing generative approaches usually adopt simple discriminators for distribution or classification supervision, however, thus limiting their ability to generate visual features that are discriminative of and transferable to novel categories. To overcome this limitation and improve the quality of generated features, we propose a dual prototype contrastive augmented discriminator for the generative adversarial network. Specifically, we design a Dual Prototype Contrastive Network (DPCN), which leverages complementary information between visual space and semantic space through multi-task prototype contrastive learning. Contrastive learning of the visual prototypes enhances the ability of the generated features to distinguish between classes, while the contrastive learning of the semantic prototypes improves their transferability. Furthermore, we introduce margins into the contrastive learning process to ensure both intra-class compactness and inter-class separation. To demonstrate the effectiveness of the proposed approach, we conduct experiments on three widely-used zero-shot learning benchmark datasets, where DPCN achieves state-of-the-art performance for GZSL. Huajie Jiang, Zhengxian Li, Yongli Hu, Jian Yang 0001, Anton van den Hengel, Ming-Hsuan Yang 0001, Yuankai Qi |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2025 | Input-Regulated Remote Sensing Counting With Region UnderstandingabstractRemote sensing counting aims to automatically estimate the number of objects of interest from high-resolution aerial or satellite imagery, providing critical decision-making support in areas such as urban planning, traffic monitoring, and disaster response. While most existing methods leverage pre-trained models to enhance feature generalization, their performance is often hindered by the severe scarcity of annotated remote sensing data. This limits their generalizability in complex scenarios. To address these challenges, we propose a novel remote sensing counting network that effectively captures informative signals from relatively limited annotated data. Specifically, we first introduce a graph-driven input regulator that constructs a graph structure by modeling relationships among input features, effectively capturing intrinsic contextual dependencies. This structure allows the regulator to assign adaptive pixel-level weights to network inputs, prioritizing relevant signals while mitigating the risk of overfitting to a fixed data distribution. Second, we design a dynamic region-aware module that leverages fuzzy logic to adaptively identify and enhance highly discriminative local regions. In this way, it improves the robustness of the feature representations. Extensive experiments demonstrate the effectiveness of the proposed method compared with several state-of-the-art methods. Shengqin Jiang, Haojian Long, Fengna Cheng, Yuankai Qi, Xiaobo Lu, Qingshan Liu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Spatial-Temporal Interleaved Network for Efficient Action RecognitionabstractThe decomposition of 3D convolution will considerably reduce the computing complexity of 3D convolutional neural networks, yet simple stacking restricts the performance of neural networks. To this end, we propose a spatial-temporal interleaved network for efficient action recognition. By deeply analyzing this task, it revisits the structure of 3D neural networks in action recognition from the following perspectives. To enhance the learning of robust spatial-temporal features, we initially propose an interleaved feature interaction module to comprehensively explore cross-layer features and capture the most discriminative information among them. With regards to being lightweight, a boosted parallel pseudo-3D module is introduced with the goal of circumventing a substantial number of computations from the lower to middle levels while enhancing temporal and spatial features in parallel at high levels. Furthermore, we exploit a spatial-temporal differential attention mechanism to suppress redundant features in different dimensions while reaping the benefits of nearly negligible parameters. Lastly, extensive experiments on four action recognition benchmarks are given to show the advantages and efficiency of our proposed method. Specifically, our method attains a 15.2% improvement in Top-1 accuracy compared to our baseline, a stack of full 3D convolutional layers, on the Something-Something V1 dataset while utilizing only 18.2% of the parameters. Shengqin Jiang, Haokui Zhang, Yuankai Qi, Qingshan Liu 0001 |
IEEE Trans. Ind. Informatics | 3 |
| 2025 | Self-Reflection Neural Network for Class-Incremental Object CountingabstractIn crowded scenarios, achieving the counting task of dynamically evolving categories is extremely challenging. In addition to grappling with challenges such as scale variations, severe occlusion and complex backgrounds, it is imperative to mitigate the issue of catastrophic forgetting. Previous approaches have heavily relied on leveraging historical data for knowledge distillation to tackle these difficulties. However, this strategy encounters two prominent obstacles: 1) Employing the teacher network from the previous stage for distillation incurs additional computational overhead during the training stage. 2) Although knowledge distillation can facilitate effective knowledge transfer, some inaccurate predictions from the teacher network may affect the knowledge acquisition in the current stage. To overcome these issues, we introduce a novel solution: a self-reflection neural network for class-incremental object counting. First, we construct a global-aware incremental regression branch that uses stacked transformer layers as backends to capture global information, while the final regression layers dynamically expand as categories increase. Furthermore, we introduce an uncertain estimation branch that selectively isolates certain feature maps to avoid some neurons updated with excessive gradient information, thereby enhancing the network plasticity while preserving stability. The output of this branch functions as a regularization signal, steering the learning process of the incremental regression branch. To foster a more robust retention of past knowledge, we propose a self-reflection loss. It employs the rectified outputs of global-aware incremental regression branch to encourage the network to reflect upon and refine its grasp of historical knowledge, effectively averting the pitfalls of inaccurate information. Our extensive experiments validate the effectiveness of our proposed method, achieving state-of-the-art results. Shengqin Jiang, Linfei Li, Fengna Cheng, Yuankai Qi, Qingshan Liu 0001 |
IEEE Trans. Multim. | 4 |
| 2025 | Dynamic Erasing Network With Adaptive Temporal Modeling for Weakly Supervised Video Anomaly DetectionabstractThe weakly supervised video anomaly detection aims to learn a detection model using only video-level labeled data. Prior studies ignore the complexity or duration of anomalies present in abnormal videos during temporal modeling. Moreover, existing works usually detect the most abnormal segments, potentially overlooking the completeness of anomalies. We propose a dynamic erasing network (DE-Net) for weakly supervised video anomaly detection, which learns video-specific temporal features via adaptive temporal modeling (ATM) to address these limitations. Specifically, to handle duration variations of abnormal events, we propose an ATM module capable of adaptively selecting and aggregating the most appropriate K temporal scale features for each video. Then, we design a dynamic erasing (DE) strategy that dynamically assesses the completeness of the detected anomalies and erases prominent abnormal segments to encourage the model to discover gentle abnormal segments. The proposed method achieves favorable performance compared to several state-of-the-art approaches on the widely used XD-Violence, TAD, and UCF-Crime datasets. Chen Zhang 0013, Guorong Li, Yuankai Qi, Hanhua Ye, Laiyun Qing, Ming-Hsuan Yang 0001, Qingming Huang |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Augmented Commonsense Knowledge for Remote Object GroundingabstractThe vision-and-language navigation (VLN) task necessitates an agent to perceive the surroundings, follow natural language instructions, and act in photo-realistic unseen environments. Most of the existing methods employ the entire image or object features to represent navigable viewpoints. However, these representations are insufficient for proper action prediction, especially for the REVERIE task, which uses concise high-level instructions, such as “Bring me the blue cushion in the master bedroom”. To address enhancing representation, we propose an augmented commonsense knowledge model (ACK) to leverage commonsense information as a spatio-temporal knowledge graph for improving agent navigation. Specifically, the proposed approach involves constructing a knowledge base by retrieving commonsense information from ConceptNet, followed by a refinement module to remove noisy and irrelevant knowledge. We further present ACK which consists of knowledge graph-aware cross-modal and concept aggregation modules to enhance visual representation and visual-textual data alignment by integrating visible objects, commonsense knowledge, and concept history, which includes object and knowledge temporal information. Moreover, we add a new pipeline for the commonsense-based decision-making process which leads to more accurate local action prediction. Experimental results demonstrate our proposed model noticeably outperforms the baseline and archives the state-of-the-art on the REVERIE benchmark. The source code is available at https://github.com/Bahram-Mohammadi/ACK. Bahram Mohammadi, Yicong Hong, Yuankai Qi, Qi Wu 0001, Shirui Pan, Qinfeng Shi |
AAAI | 3 |
| 2024 | Generating Content for HDR Deghosting from Frequency ViewabstractRecovering ghost-free High Dynamic Range (HDR) images from multiple Low Dynamic Range (LDR) images becomes challenging when the LDR images exhibit saturation and significant motion. Recent Diffusion Models (DMs) have been introduced in HDR imaging field, demonstrating promising performance, particularly in achieving visually perceptible results compared to previous DNN-based methods. However, DMs require extensive iterations with large models to estimate entire images, resulting in inefficiency that hinders their practical application. To address this challenge, we propose the Low-Frequency aware Diffusion (LF-Diff) model for ghost-free HDR imaging. The key idea of LF-Diff is implementing the DMs in a highly compacted latent space and integrating it into a regression-based model to enhance the details of reconstructed images. Specifically, as low-frequency information is closely related to human visual perception we propose to utilize DMs to create compact low-frequency priors for the reconstruction process. In addition, to take full advantage of the above low-frequency priors, the Dynamic HDR Reconstruction Network (DHRNet) is carried out in a regression-based manner to obtain final HDR images. Extensive experiments conducted on synthetic and real-world benchmark datasets demonstrate that our LF-Diff performs favorably against several state-of-the-art methods and is 10x faster than previous DM-based methods. Tao Hu 0013, Qingsen Yan, Yuankai Qi, Yanning Zhang 0001 |
CVPR | 3 |
| 2024 | Weakly Supervised Video Individual CountingabstractVideo Individual Counting (VIC) aims to predict the number of unique individuals in a single video. Existing methods learn representations based on trajectory labels for individuals, which are annotation-expensive. To provide a more realistic reflection of the underlying practical challenge, we introduce a weakly supervised VIC task, wherein trajectory labels are not provided. Instead, two types of labels are provided to indicate traffic entering the field of view (inflow) and leaving the field view (outflow). We also propose the first solution as a baseline that formulates the task as a weakly supervised contrastive learning problem under group-level matching. In doing so, we devise an end-to-end trainable soft contrastive loss to drive the network to distin-guish inflow, outflow, and the remaining. To facilitate future study in this direction, we generate annotations from the existing VIC datasets Sense Crowd and CroHD and also build a new dataset, UAVVIC. Extensive results show that our baseline weakly supervised method outperforms supervised methods, and thus, little information is lost in the transition to the more practically relevant weakly supervised task. The code and trained model can be found at CGNet. Xinyan Liu 0008, Guorong Li, Yuankai Qi, Ziheng Yan, Zhenjun Han, Anton van den Hengel, Ming-Hsuan Yang 0001, Qingming Huang |
CVPR | 3 |
| 2024 | Decomposing Disease Descriptions for Enhanced Pathology Detection: A Multi-Aspect Vision-Language Pre-Training FrameworkabstractMedical vision language pre-training (VLP) has emerged as a frontier of research, enabling zero-shot pathological recognition by comparing the query image with the textual descriptions for each disease. Due to the complex semantics of biomedical texts, current methods struggle to align medical images with key pathological findings in un-structured reports. This leads to the misalignment with the target disease's textual representation. In this paper, we introduce a novel VLP framework designed to dissect disease descriptions into their fundamental aspects, leveraging prior knowledge about the visual manifestations of pathologies. This is achieved by consulting a large language model and medical experts. Integrating a Transformer module, our approach aligns an input image with the diverse elements of a disease, generating aspect-centric image representations. By consolidating the matches from each aspect, we improve the compatibility between an image and its associated disease. Additionally, capitalizing on the aspect-oriented representations, we present a dual-head Transformer tailored to process known and unknown diseases, optimizing the comprehensive detection efficacy. Conducting experiments on seven downstream datasets, ours improves the accuracy of recent methods by up to 8.56% and 17.26% for seen and unseen categories, respectively. Our code is released at https://github.com/HieuPhan33/MAVL. Vu Minh Hieu Phan, Yutong Xie 0001, Yuankai Qi, Lingqiao Liu, Liyang Liu, Bowen Zhang 0009, Zhibin Liao, Qi Wu 0001, Minh-Son To, Johan Verjans |
CVPR | 3 |
| 2024 | Generating High-Quality Symbolic Music Using Fine-Grained Discriminators
Zhedong Zhang, Liang Li 0003, Hongkui Wang, Chenggang Yan 0001, Jian Yang 0001, Yuankai Qi |
ICPR (20) | 8 |
| 2024 | Structural Attention: Rethinking Transformer for Unpaired Medical Image Synthesis
Vu Minh Hieu Phan, Yutong Xie 0001, Bowen Zhang 0009, Yuankai Qi, Zhibin Liao, Antonios Perperidis, Son Lam Phung, Johan Verjans, Minh-Son To |
MICCAI (7) | 4 |
| 2024 | From Speaker to Dubber: Movie Dubbing with Prosody and Duration Consistency Learning
Zhedong Zhang, Liang Li 0003, Gaoxiang Cong 0001, Haibing Yin, Chenggang Yan 0001, Anton van den Hengel, Yuankai Qi |
ACM Multimedia | 8 |
| 2024 | Style-aware two-stage learning framework for video captioningabstractSignificant progress has been made in video captioning in recent years. However, most existing methods directly learn from all given captions without distinguishing the styles of captions. The large diversity in these captions might bring ambiguity to the model learning. To address this issue, we propose a style-aware two-stage learning framework. In the first stage, the model is trained with captions of separate styles, including length style (short, medium, long), action style (single action or multiple actions), and object style (one object or more). For efficiency, a shared model with multiple individual style vectors is learned. In the second stage, a video style encoder is devised to capture style information from the input video, and it outputs a guidance signal of how to utilize the style vectors for the final caption generation. Without whistles and bells, our method achieves state-of-the-art performance on three widely-used public datasets, MSVD, MSR-VTT and VATEX. The source code and trained models will be made available to the public. Yunchuan Ma, Yuankai Qi, Amin Beheshti, Laiyun Qing, Guorong Li |
Knowl. Based Syst. | 3 |
| 2024 | Learning Hierarchical Modular Networks for Video CaptioningabstractVideo captioning aims to generate natural language descriptions for a given video clip. Existing methods mainly focus on end-to-end representation learning via word-by-word comparison between predicted captions and ground-truth texts. Although significant progress has been made, such supervised approaches neglect semantic alignment between visual and linguistic entities, which may negatively affect the generated captions. In this work, we propose a hierarchical modular network to bridge video representations and linguistic semantics at four granularities before generating captions: entity, verb, predicate, and sentence. Each level is implemented by one module to embed corresponding semantics into video representations. Additionally, we present a reinforcement learning module based on the scene graph of captions to better measure sentence similarity. Extensive experimental results show that the proposed method performs favorably against the state-of-the-art models on three widely-used benchmark datasets, including microsoft research video description corpus (MSVD), MSR-video to text (MSR-VTT), and video-and-TEXt (VATEX). Guorong Li, Hanhua Ye, Yuankai Qi, Shuhui Wang, Laiyun Qing, Qingming Huang, Ming-Hsuan Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Rethink video retrieval representation for video captioning
Mingkai Tian, Guorong Li, Yuankai Qi, Shuhui Wang, Quan Z. Sheng, Qingming Huang |
Pattern Recognit. | 3 |
| 2024 | A Unified Object Counting Network With Object Occupation PriorabstractThe counting task, which plays a fundamental role in numerous applications (e.g., crowd counting, traffic statistics), aims to predict the number of objects with various densities. Existing object counting tasks are designed for a single object class. However, it is inevitable to encounter newly coming data with new classes in our real world. We name this scenario as evolving object counting. In this paper, we build the first evolving object counting dataset and propose a unified object counting network as the first attempt to address this task. The proposed network consists of two key components: a class-agnostic mask module and a class-incremental module. The class-agnostic mask module learns generic object occupation prior by predicting a class-agnostic binary mask (e.g., 1 denotes there exists an object at the considering position in an image and 0 otherwise). The class-incremental module is used to handle new classes and provides discriminative class guidance for density map prediction. The combined outputs of the class-agnostic mask module and image feature extractor are used to predict the final density map. When new classes arrive, we first add new neural nodes to the last regression and classification layers of the class-incremental module. Then, instead of retraining the model from scratch, we utilize knowledge distillation to help the model retain and consolidate what it has previously learned. We also employ a support sample bank to store a small number of typical training samples for each class, which are used to prevent the model from forgetting key information from old data. With this design, our model can efficiently and effectively adapt to new classes while maintaining good performance on already-seen data without large-scale retraining. Extensive experiments on the collected dataset demonstrate favorable performance. The dataset and code will be available at:https://github.com/Tanyjiang/EOCO. Shengqin Jiang, Fengna Cheng, Yuankai Qi, Qingshan Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Progressive Multi-Resolution Loss for Crowd CountingabstractCrowd counting is usually handled in a density map regression fashion, which is supervised via an L2 loss between the predicted density map and ground truth. To effectively regulate models, various improved L2 loss functions have been developed to find a better correspondence between predicted density and annotation positions. In this paper, we propose to predict the density map at one resolution but measure its quality via a derived log-formed loss at multiple resolutions. Unlike existing methods that assume density maps at different resolutions are independent, our loss is obtained by modeling the likelihood function inspired by the relationship of density maps across multi-resolutions. We find that the traditional single-resolution L2 loss is a particular case of our derived log-likelihood. We mathematically prove it is superior to a single-resolution L2 loss. Without bells and whistles, the proposed loss substantially improves several baselines and performs favorably compared to state-of-the-art methods on five crowd counting datasets: NWPU-Crowd, ShanghaiTech A & B, UCF-QNRF, and JHU-Crowd++. The source code and trained models are released athttps://github.com/streamer-AP/PML_Loss.git. Ziheng Yan, Yuankai Qi, Guorong Li, Xinyan Liu 0008, Weigang Zhang, Ming-Hsuan Yang 0001, Qingming Huang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Rethinking Attentive Object Detection via Neural Attention LearningabstractVisual attention advances object detection by attending neural networks to object representations. While existing methods incorporate empirical modules to empower network attention, we rethink attentive object detection from the network learning perspective in this work. We propose a NEural Attention Learning approach (NEAL) which consists of two parts. During the back-propagation of each training iteration, we first calculate the partial derivatives (a.k.a. the accumulated gradients) of the classification output with respect to the input features. We refine these partial derivatives to obtain attention response maps whose elements reflect the contributions to the final network predictions. Then, we formulate the attention response maps as extra objective functions, which are combined together with the original detection loss to train detectors in an end-to-end manner. In this way, we succeed in learning an attentive CNN model without introducing additional network structures. We apply NEAL to the two-stage object detection frameworks, which are usually composed of a CNN feature backbone, a region proposal network (RPN), and a classifier. We show that the proposed NEAL not only helps the RPN attend to objects but also enables the classifier to pay more attention to the premier positive samples. To this end, the localization (proposal generation) and classification mutually benefit from each other in our proposed method. Extensive experiments on large-scale benchmark datasets, including MS COCO 2017 and Pascal VOC 2012, demonstrate that the proposed NEAL algorithm advances the two-stage object detector over state-of-the-art approaches. Chongjian Ge, Yibing Song, Chao Ma 0004, Yuankai Qi, Ping Luo 0002 |
IEEE Trans. Image Process. | 4 |
| 2023 | Learning to Dub Movies via Hierarchical Prosody ModelsabstractGiven a piece of text, a video clip and a reference audio, the movie dubbing (also known as visual voice clone, V2C) task aims to generate speeches that match the speaker's emotion presented in the video using the desired speaker voice as reference. V2C is more challenging than conventional text-to-speech tasks as it additionally requires the generated speech to exactly match the varying emotions and speaking speed presented in the video. Unlike previous works, we propose a novel movie dubbing architecture to tackle these problems via hierarchical prosody modeling, which bridges the visual information to corresponding speech prosody from three aspects: lip, face, and scene. Specifically, we align lip movement to the speech duration, and convey facial expression to speech energy and pitch via attention mechanism based on valence and arousal representations inspired by the psychology findings. Moreover, we design an emotion booster to capture the atmosphere from global video scenes. All these embeddings are used together to generate mel-spectrogram, which is then converted into speech waves by an existing vocoder. Extensive experimental results on the V2C and Chem benchmark datasets demonstrate the favourable performance of the proposed method. The code and trained models will be made available at https://github.com/GalaxyCong/HPMDubbing Gaoxiang Cong 0001, Liang Li 0003, Yuankai Qi, Zhengjun Zha, Qi Wu 0001, Bin Jiang 0011, Ming-Hsuan Yang 0001, Qingming Huang |
CVPR | 3 |
| 2023 | Exploiting Completeness and Uncertainty of Pseudo Labels for Weakly Supervised Video Anomaly DetectionabstractWeakly supervised video anomaly detection aims to identify abnormal events in videos using only video-level labels. Recently, two-stage self-training methods have achieved significant improvements by self-generating pseudo labels and self-refining anomaly scores with these labels. As the pseudo labels play a crucial role, we propose an enhancement framework by exploiting completeness and uncertainty properties for effective self-training. Specifically, we first design a multi-head classification module (each head serves as a classifier) with a diversity loss to maximize the distribution differences of predicted pseudo labels across heads. This encourages the generated pseudo labels to cover as many abnormal events as possible. We then devise an iterative uncertainty pseudo label refinement strategy, which improves not only the initial pseudo labels but also the updated ones obtained by the desired classifier in the second stage. Extensive experimental results demonstrate the proposed method performs favorably against state-of-the-art approaches on the UCF-Crime, TAD, and XD-Violence benchmark datasets. Chen Zhang 0013, Guorong Li, Yuankai Qi, Shuhui Wang, Laiyun Qing, Qingming Huang, Ming-Hsuan Yang 0001 |
CVPR | 3 |
| 2023 | AerialVLN: Vision-and-Language Navigation for UAVsabstractRecently emerged Vision-and-Language Navigation (VLN) tasks have drawn significant attention in both computer vision and natural language processing communities. Existing VLN tasks are built for agents that navigate on the ground, either indoors or outdoors. However, many tasks require intelligent agents to carry out in the sky, such as UAV-based goods delivery, traffic/security patrol, and scenery tour, to name a few. Navigating in the sky is more complicated than on the ground because agents need to consider the flying height and more complex spatial relationship reasoning. To fill this gap and facilitate research in this field, we propose a new task named AerialVLN, which is UAV-based and towards outdoor environments. We develop a 3D simulator rendered by near-realistic pictures of 25 city-level scenarios. Our simulator supports continuous navigation, environment extension and configuration. We also proposed an extended baseline model based on the widely-used cross-modal-alignment (CMA) navigation methods. We find that there is still a significant gap between the baseline model and human performance, which suggests AerialVLN is a new challenging task. Dataset and code is available at https://github.com/AirVLN/AirVLN. Yuankai Qi, Peng Wang 0015, Yanning Zhang 0001, Qi Wu 0001 |
ICCV | 3 |
| 2023 | March in Chat: Interactive Prompting for Remote Embodied Referring ExpressionabstractMany Vision-and-Language Navigation (VLN) tasks have been proposed in recent years, from room-based to object-based and indoor to outdoor. The REVERIE (Remote Embodied Referring Expression) is interesting since it only provides high-level instructions to the agent, which are closer to human commands in practice. Nevertheless, this poses more challenges than other VLN tasks since it requires agents to infer a navigation plan only based on a short instruction. Large Language Models (LLMs) show great potential in robot action planning by providing proper prompts. Still, this strategy has not been explored under the REVERIE settings. There are several new challenges. For example, the LLM should be environment-aware so that the navigation plan can be adjusted based on the current visual observation. Moreover, the LLM planned actions should be adaptable to the much larger and more complex REVERIE environment. This paper proposes a March-in-Chat (MiC) model that can talk to the LLM on the fly and plan dynamically based on a newly proposed Room-and-Object Aware Scene Perceiver (ROASP). Our MiC model outperforms the previous state-of-the-art by large margins by SPL and RGSPL metrics on the REVERIE benchmark. The source code is available at https://github.com/YanyuanQiao/MiC Yanyuan Qiao, Yuankai Qi, Zheng Yu 0004, Jing Liu 0001, Qi Wu 0001 |
ICCV | 2 |
| 2023 | Mind the Gap: Improving Success Rate of Vision-and-Language Navigation by Revisiting Oracle Success RoutesabstractVision-and-Language Navigation (VLN) aims to navigate to the target location by following a given instruction. Unlike existing methods focused on predicting a more accurate action at each step in navigation, in this paper, we make the first attempt to tackle a long-ignored problem in VLN: narrowing the gap between Success Rate (SR) and Oracle Success Rate (OSR). We observe a consistently large gap (up to 9%) on four state-of-the-art VLN methods across two benchmark datasets: R2R and REVERIE. The high OSR indicates the robot agent passes the target location, while the low SR suggests the agent actually fails to stop at the target location at last. Instead of predicting actions directly, we propose to mine the target location from a trajectory given by off-the-shelf VLN models. Specially, we design a multi-module transformer-based model for learning compact discriminative trajectory viewpoint representation, which is used to predict the confidence of being a target location as described in the instruction. The proposed method is evaluated on three widely-adopted datasets: R2R, REVERIE and NDH, and shows promising results, demonstrating the potential for more future research. Chongyang Zhao 0003, Yuankai Qi, Qi Wu 0001 |
ACM Multimedia | 2 |
| 2023 | CALM: An Enhanced Encoding and Confidence Evaluating Framework for Trustworthy Multi-view LearningabstractMulti-view learning aims to leverage data acquired from multiple sources to achieve better performance compared to using a single view. However, the performance of multi-view learning can be negatively impacted by noisy or corrupted views in certain real-world situations. As a result, it is crucial to assess the confidence of predictions and obtain reliable learning outcomes. In this paper, we introduce CALM, an enhanced encoding and confidence evaluation framework for trustworthy multi-view classification. Our method comprises enhanced multi-view encoding, multi-view confidence-aware fusion, and multi-view classification regularization, enabling the simultaneous evaluation of prediction confidence and the yielding trustworthy classifications. Enhanced multi-view encoding takes advantage of cross-view consistency and class diversity to improve the efficacy of the learned latent representation, facilitating more reliable classification results. Multi-view confidence-aware fusion utilizes a confidence-aware estimator to evaluate the confidence scores of classification outcomes. The final multi-view classification results are then derived through confidence-aware fusion. To achieve reliable and accurate confidence scores, multivariate Gaussian distributions are employed to model the prediction distribution. The advantage of CALM lies in its ability to evaluate the quality of each view, reducing the influence of low-quality views on the multi-view fusion process and ultimately leading to improved classification performance and confidence evaluation. Comprehensive experimental results demonstrate that our method outperforms other trusted multi-view learning methods in terms of effectiveness, reliability, and robustness. Zhe Xue, Boang Li, Junping Du 0001, Meiyu Liang, Yuankai Qi |
ACM Multimedia | 7 |
| 2023 | HOP+: History-Enhanced and Order-Aware Pre-Training for Vision-and-Language NavigationabstractRecent works attempt to employ pre-training in Vision-and-Language Navigation (VLN). However, these methods neglect the importance of historical contexts or ignore predicting future actions during pre-training, limiting the learning of visual-textual correspondence and the capability of decision-making. To address these problems, we present a history-enhanced and order-aware pre-training with the complementing fine-tuning paradigm (HOP+) for VLN. Specifically, besides the common Masked Language Modeling (MLM) and Trajectory-Instruction Matching (TIM) tasks, we design three novel VLN-specific proxy tasks: Action Prediction with History (APH) task, Trajectory Order Modeling (TOM) task and Group Order Modeling (GOM) task. APH task takes into account the visual perception trajectory to enhance the learning of historical knowledge as well as action prediction. The two temporal visual-textual alignment tasks, TOM and GOM further improve the agent's ability to order reasoning. Moreover, we design a memory network to address the representation inconsistency of history context between the pre-training and the fine-tuning stages. The memory network effectively selects and summarizes historical information for action prediction during fine-tuning, without costing huge extra computation consumption for downstream VLN tasks. HOP+ achieves new state-of-the-art performance on four downstream VLN tasks (R2R, REVERIE, RxR, and NDH), which demonstrates the effectiveness of our proposed method. Yanyuan Qiao, Yuankai Qi, Yicong Hong, Zheng Yu 0004, Peng Wang 0015, Qi Wu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | V2C: Visual Voice CloningabstractExisting Voice Cloning (VC) tasks aim to convert a para-graph text to a speech with desired voice specified by a ref-erence audio. This has significantly boosted the development of artificial speech applications. However, there also exist many scenarios that cannot be well reflected by these VC tasks, such as movie dubbing, which requires the speech to be with emotions consistent with the movie plots. To fill this gap, in this work we propose a new task named Vi-sual Voice Cloning (V2C), which seeks to convert a para-graph of text to a speech with both desired voice speci-fied by a reference audio and desired emotion specified by a reference video. To facilitate research in this field, we construct a dataset, V2C-Animation, and propose a strong baseline based on existing state-of-the-art (SoTA) VC techniques. Our dataset contains 10,217 animated movie clips covering a large variety of genres (e.g., Comedy, Fantasy) and emotions (e.g., happy, sad). We further design a set of evaluation metrics, named MCD-DTW-SL, which help eval-uate the similarity between ground-truth speeches and the synthesised ones. Extensive experimental results show that even SoTA VC methods cannot generate satisfying speeches for our V2C task. We hope the proposed new task together with the constructed dataset and evaluation metric will fa-cilitate the research in the field of voice cloning and broader vision-and-language community. Source code and dataset will be released in https://github.com/chenqi008/V2C. Qi Chen 0014, Mingkui Tan, Yuankai Qi, Jiaqiu Zhou, Yuanqing Li 0001, Qi Wu 0001 |
CVPR | 3 |
| 2022 | HOP: History-and-Order Aware Pretraining for Vision-and-Language NavigationabstractPretraining has been adopted in a few of recent works for Vision-and-Language Navigation (VLN). However, pre-vious pre-training methods for VLN either lack the ability to predict future actions or ignore the trajectory contexts, which are essential for a greedy navigation process. In this work, to promote the learning of spatio-temporal visual-textual correspondence as well as the agent's capability of decision making, we propose a novel history-and-order aware pre-training paradigm (HOP) with VLN-specific objectives that exploit the past observations and support future action prediction. Specifically, in addition to the commonly used Masked Language Modeling (MLM) and Trajectory-Instruction Matching (TIM), we design two proxy tasks to model temporal order information: Trajectory Order Modeling (TOM) and Group Order Modeling (GOM). Moreover, our navigation action prediction is also enhanced by intro-ducing the task of Action Prediction with History (APH), which takes into account the history visual perceptions. Extensive experimental results on four downstream VLN tasks (R2R, REVERIE, NDH, RxR) demonstrate the effectiveness of our proposed method compared against several state-of-the-art agents. Yanyuan Qiao, Yuankai Qi, Yicong Hong, Zheng Yu 0004, Peng Wang 0015, Qi Wu 0001 |
CVPR | 2 |
| 2022 | Hierarchical Modular Network for Video CaptioningabstractVideo captioning aims to generate natural language descriptions according to the content, where representation learning plays a crucial role. Existing methods are mainly developed within the supervised learning framework via word-by-word comparison of the generated caption against the ground-truth text without fully exploiting linguistic semantics. In this work, we propose a hierarchical modular network to bridge video representations and linguistic semantics from three levels before generating captions. In particular, the hierarchy is composed of: (I) Entity level, which highlights objects that are most likely to be mentioned in captions. (II) Predicate level, which learns the actions conditioned on highlighted objects and is supervised by the predicate in captions. (III) Sentence level, which learns the global semantic representation and is supervised by the whole caption. Each level is implemented by one module. Extensive experimental results show that the proposed method performs favorably against the state-of-the-art models on the two widely-used benchmarks: MSVD 104.0% and MSR-VTT 51.5% in CIDEr score. Code will be made available at https://github.com/MarcusNerva/HMN. Hanhua Ye, Guorong Li, Yuankai Qi, Shuhui Wang, Qingming Huang, Ming-Hsuan Yang 0001 |
CVPR | 3 |
| 2022 | Multi-Attention Network for Compressed Video Referring Object SegmentationabstractReferring video object segmentation aims to segment the object referred by a given language expression. Existing works typically require compressed video bitstream to be decoded to RGB frames before being segmented, which increases computation and storage requirements and ultimately slows the inference down. This may hamper its application in real-world computing resource limited scenarios, such as autonomous cars and drones. To alleviate this problem, in this paper, we explore the referring object segmenta- tion task on compressed videos, namely on the original video data flow. Besides the inherent difficulty of the video referring object segmentation task itself, obtaining discriminative representation from compressed video is also rather challenging. To address this problem, we propose a multi-attention network which consists of dual-path dual-attention module and a query-based cross-modal Transformer module. Specifically, the dual-path dual-attention module is designed to extract effective representation from compressed data in three modalities, i.e., I-frame, Motion Vector and Residual. The query-based cross-modal Transformer firstly models the corre- lation between linguistic and visual modalities, and then the fused multi-modality features are used to guide object queries to generate a content-aware dynamic kernel and to predict final segmentation masks. Different from previous works, we propose to learn just one kernel, which thus removes the complicated post mask-matching procedure of existing methods. Extensive promising experimental results on three challenging datasets show the effectiveness of our method compared against several state-of-the-art methods which are proposed for processing RGB data. Source code is available at: https://github.com/DexiangHong/MANet. Weidong Chen 0013, Dexiang Hong, Yuankai Qi, Zhenjun Han, Shuhui Wang, Laiyun Qing, Qingming Huang, Guorong Li |
ACM Multimedia | 3 |
| 2022 | Diagnosing Vision-and-Language Navigation: What Really MattersabstractWanrong Zhu, Yuankai Qi, Pradyumna Narayana, Kazoo Sone, Sugato Basu, Xin Wang, Qi Wu, Miguel Eckstein, William Yang Wang. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Wanrong Zhu, Yuankai Qi, Pradyumna Narayana, Kazoo Sone, Sugato Basu, Xin Wang 0061, Qi Wu 0001, Miguel P. Eckstein, William Yang Wang |
NAACL-HLT | 2 |
| 2021 | VLN BERT: A Recurrent Vision-and-Language BERT for NavigationabstractAccuracy of many visiolinguistic tasks has benefited significantly from the application of vision-and-language (V&L) BERT. However, its application for the task of vision-and-language navigation (VLN) remains limited. One reason for this is the difficulty adapting the BERT architecture to the partially observable Markov decision process present in VLN, requiring history-dependent attention and decision making. In this paper we propose a recurrent BERT model that is time-aware for use in VLN. Specifically, we equip the BERT model with a recurrent function that maintains cross-modal state information for the agent. Through extensive experiments on R2R and REVERIE we demonstrate that our model can replace more complex encoder-decoder models to achieve state-of-the-art results. Moreover, our approach can be generalised to other transformer-based architectures, supports pre-training, and is capable of solving navigation and referring expression tasks simultaneously. Yicong Hong, Qi Wu 0001, Yuankai Qi, Cristian Rodriguez Opazo, Stephen Gould |
CVPR | 3 |
| 2021 | The Road to Know-Where: An Object-and-Room Informed Sequential BERT for Indoor Vision-Language NavigationabstractVision-and-Language Navigation (VLN) requires an agent to find a path to a remote location on the basis of natural-language instructions and a set of photo-realistic panoramas. Most existing methods take the words in the instructions and the discrete views of each panorama as the minimal unit of encoding. However, this requires a model to match different nouns (e.g., TV, table) against the same input view feature. In this work, we propose an object-informed sequential BERT to encode visual perceptions and linguistic instructions at the same fine-grained level, namely objects and words. Our sequential BERT also enables the visual-textual clues to be interpreted in light of the temporal context, which is crucial to multi-round VLN tasks. Additionally, we enable the model to identify the relative direction (e.g., left/right/front/back) of each navigable location and the room type (e.g., bedroom, kitchen) of its current and final navigation goal, as such information is widely mentioned in instructions implying the desired next and final locations. We thus enable the model to know-where the objects lie in the images, and to know-where they stand in the scene. Extensive experiments demonstrate the effectiveness compared against several state-of-the-art methods on three indoor VLN tasks: REVERIE, NDH, and R2R. Project repository: https://github.com/YuankaiQi/ORIST Yuankai Qi, Zizheng Pan, Yicong Hong, Ming-Hsuan Yang 0001, Anton van den Hengel, Qi Wu 0001 |
ICCV | 1 |
| 2021 | Neighbor-view Enhanced Model for Vision and Language NavigationabstractVision and Language Navigation (VLN) requires an agent to navigate to a target location by following natural language instructions. Most of existing works represent a navigation candidate by the feature of the corresponding single view where the candidate lies in. However, an instruction may mention landmarks out of the single view as references, which might lead to failures of textual-visual matching of existing methods. In this work, we propose a multi-module Neighbor-View Enhanced Model (NvEM) to adaptively incorporate visual contexts from neighbor views for better textual-visual matching. Specifically, our NvEM utilizes a subject module and a reference module to collect contexts from neighbor views. The subject module fuses neighbor views at a global level, and the reference module fuses neighbor objects at a local level. Subjects and references are adaptively determined via attention mechanisms. Our model also includes an action module to utilize the strong orientation guidance (e.g., "turn left'') in instructions. Each module predicts navigation action separately and their weighted sum is used for predicting the final action. Extensive experimental results demonstrate the effectiveness of the proposed method on the R2R and R4R benchmarks against several state-of-the-art navigators, and NvEM even beats some pre-training ones. Our code is available at https://github.com/MarSaKi/NvEM. Dong An 0002, Yuankai Qi, Yan Huang 0008, Qi Wu 0001, Liang Wang 0001, Tieniu Tan |
ACM Multimedia | 2 |
| 2021 | R-GAN: Exploring Human-like Way for Reasonable Text-to-Image Synthesis via Generative Adversarial NetworksabstractDespite recent significant progress on generative models, context-rich text-to-image synthesis depicting multiple complex objects is still non-trivial. The main challenges lie in the ambiguous semantic of a complex description and the intricate scene of an image with various objects, different positional relationship and diverse appearances. To address these challenges, we propose R-GAN, which can generate reasonable images according to the given text in a human-like way. Specifically, just like humans will first find and settle the essential elements to create a simple sketch, we first capture a monolithic-structural text representation by building a scene graph to find the essential semantic elements. Then, based on this representation, we design a bounding box generator to estimate the layout with position and size of target objects, and a following shape generator, which draws a fine-detailed shape for each object. Different from previous work only generating coarse shapes blindly, we introduce a coarse-to-fine shape generator based on a shape knowledge base. At last, to finish the final image synthesis, we propose a multi-modal geometry-aware spatially-adaptive generator conditioned on the monolithic-structural text representation and the geometry-aware map of the shapes. Extensive experiments on the real-world dataset MSCOCO show the superiority of our method in terms of both quantitative and qualitative metrics. Yanyuan Qiao, Qi Chen 0014, Chaorui Deng, Yuankai Qi, Mingkui Tan, Xincheng Ren, Qi Wu 0001 |
ACM Multimedia | 5 |
| 2021 | Image editing with varying intensities of processing
Yasi Wang, Yuankai Qi, Hongxun Yao, Dong Gong, Qi Wu 0001 |
Comput. Vis. Image Underst. | 2 |
| 2021 | Light fixed-time control for cluster synchronization of complex networks
Shengqin Jiang, Yuankai Qi, Shuiming Cai, Xiaobo Lu |
Neurocomputing | 2 |
| 2021 | D3D: Dual 3-D Convolutional Network for Real-Time Action RecognitionabstractThree-dimensional convolutional neural networks (3D CNNs) have been explored to learn spatio-temporal information for video-based human action recognition. Expensive computational cost and memory demand resulted from standard 3D CNNs, however, hinder their application in practical scenarios. In this article, we address the aforementioned limitations by proposing a novel dual 3-D convolutional network (D3DNet) with two complementary lightweight branches. A coarse branch maintains large temporal receptive field by a fast temporal downsampling strategy and simulates the expensive 3-D convolutions using a combination of more efficient spatial convolutions and temporal convolutions. Meanwhile, a fine branch progressively downsamples the video in the temporal domain and adopts 3-D convolutional units with reduced channel capacities to capture multiresolution spatio-temporal information. Instead of learning these two branches independently, a shallow spatiotemporal downsampling module is shared for these two branches for efficient low-level feature learning. Besides, lateral connections are learned to effectively fuse the information from the two branches at multiple stages. The proposed network makes good balance between inference speed and action recognition performance. Based on RGB information only, it achieves competing performance on five popular video-based action recognition datasets, with inference speed of 3200 FPS on a single NVIDIA GTX 2080Ti card. Shengqin Jiang, Yuankai Qi, Haokui Zhang, Zongwen Bai, Xiaobo Lu, Peng Wang 0023 |
IEEE Trans. Ind. Informatics | 2 |
| 2020 | Release the Power of Online-Training for Robust Visual TrackingabstractConvolutional neural networks (CNNs) have been widely adopted in the visual tracking community, significantly improving the state-of-the-art. However, most of them ignore the important cues lying in the distribution of training data and high-level features that are tightly coupled with the target/background classification. In this paper, we propose to improve the tracking accuracy via online training. On the one hand, we squeeze redundant training data by analyzing the dataset distribution in low-level feature space. On the other hand, we design statistic-based losses to increase the inter-class distance while decreasing the intra-class variance of high-level semantic features. We demonstrate the effectiveness on top of two high-performance tracking methods: MDNet and DAT. Experimental results on the challenging large-scale OTB2015 and UAVDT demonstrate the outstanding performance of our tracking method. Guorong Li, Yuankai Qi, Qingming Huang |
AAAI | 3 |
| 2020 | Overwater Image Dehazing via Cycle-Consistent Generative Adversarial Network
Shunyuan Zheng, Jiamin Sun, Qinglin Liu, Yuankai Qi, Shengping Zhang |
ACCV (2) | 4 |
| 2020 | REVERIE: Remote Embodied Visual Referring Expression in Real Indoor EnvironmentsabstractOne of the long-term challenges of robotics is to enable robots to interact with humans in the visual world via natural language, as humans are visual animals that communicate through language. Overcoming this challenge requires the ability to perform a wide variety of complex tasks in response to multifarious instructions from humans. In the hope that it might drive progress towards more flexible and powerful human interactions with robots, we propose a dataset of varied and complex robot tasks, described in natural language, in terms of objects visible in a large set of real images. Given an instruction, success requires navigating through a previously-unseen environment to identify an object. This represents a practical challenge, but one that closely reflects one of the core visual problems in robotics. Several state-of-the-art vision-and-language navigation, and referring-expression models are tested to verify the difficulty of this new task, but none of them show promising results because there are many fundamental differences between our task and previous ones. A novel Interactive Navigator-Pointer model is also proposed that provides a strong baseline on the task. The proposed model especially achieves the best performance on the unseen test split, but still leaves substantial room for improvement compared to the human performance. Repository: https://github.com/YuankaiQi/REVERIE. Yuankai Qi, Qi Wu 0001, Peter Anderson 0001, Xin Wang 0061, William Yang Wang, Chunhua Shen, Anton van den Hengel |
CVPR | 1 |
| 2020 | Object-and-Action Aware Model for Visual Language Navigation
Yuankai Qi, Zizheng Pan, Shengping Zhang, Anton van den Hengel, Qi Wu 0001 |
ECCV (10) | 1 |
| 2020 | Language and Visual Entity Relationship Graph for Agent NavigationabstractVision-and-Language Navigation (VLN) requires an agent to navigate in a real-world environment following natural language instructions. From both the textual and visual perspectives, we find that the relationships among the scene, its objects, and directional cues are essential for the agent to interpret complex instructions and correctly perceive the environment. To capture and utilize the relationships, we propose a novel Language and Visual Entity Relationship Graph for modelling the inter-modal relationships between text and vision, and the intra-modal relationships among visual entities. We propose a message passing algorithm for propagating information between language elements and visual entities in the graph, which we then combine to determine the next action to take. Experiments show that by taking advantage of the relationships we are able to improve over state-of-the-art. On the Room-to-Room (R2R) benchmark, our method achieves the new best performance on the test unseen split with success rate weighted by path length of 52%. On the Room-for-Room (R4R) dataset, our method significantly improves the previous best from 13% to 34% on the success weighted by normalized dynamic time warping. Yicong Hong, Cristian Rodriguez Opazo, Yuankai Qi, Qi Wu 0001, Stephen Gould |
NeurIPS | 3 |
| 2020 | Siamese Local and Global Networks for Robust Face TrackingabstractConvolutional neural networks (CNNs) have achieved great success in several face-related tasks, such as face detection, alignment and recognition. As a fundamental problem in computer vision, face tracking plays a crucial role in various applications, such as video surveillance, human emotion detection and human-computer interaction. However, few CNN-based approaches are proposed for face (bounding box) tracking. In this paper, we propose a face tracking method based on Siamese CNNs, which takes advantages of powerful representations of hierarchical CNN features learned from massive face images. The proposed method captures discriminative face information at both local and global levels. At the local level, representations for attribute patches (i.e:, eyes, nose and mouth) are learned to distinguish a face from another one, which are robust to pose changes and occlusions. At the global level, representations for each whole face are learned, which take into account the spatial relationships among local patches and facial characters, such as skin color and nevus. In addition, we build a new largescale challenging face tracking dataset to evaluate face tracking methods and to facilitate the research forward in this field. Extensive experiments on the collected dataset demonstrate the effectiveness of our method in comparison to several state-of-theart visual tracking methods. Yuankai Qi, Shengping Zhang, Feng Jiang 0001, Huiyu Zhou 0001, Dacheng Tao, Xuelong Li 0001 |
IEEE Trans. Image Process. | 1 |
| 2019 | Learning Attribute-Specific Representations for Visual TrackingabstractIn recent years, convolutional neural networks (CNNs) have achieved great success in visual tracking. Most of existing methods train or fine-tune a binary classifier to distinguish the target from its background. However, they may suffer from the performance degradation due to insufficient training data. In this paper, we show that attribute information (e.g., illumination changes, occlusion and motion) in the context facilitates training an effective classifier for visual tracking. In particular, we design an attribute-based CNN with multiple branches, where each branch is responsible for classifying the target under a specific attribute. Such a design reduces the appearance diversity of the target under each attribute and thus requires less data to train the model. We combine all attributespecific features via ensemble layers to obtain more discriminative representations for the final target/background classification. The proposed method achieves favorable performance on the OTB100 dataset compared to state-of-the-art tracking methods. After being trained on the VOT datasets, the proposed network also shows a good generalization ability on the UAV-Traffic dataset, which has significantly different attributes and target appearances with the VOT datasets. Yuankai Qi, Shengping Zhang, Weigang Zhang, Li Su 0003, Qingming Huang, Ming-Hsuan Yang 0001 |
AAAI | 1 |
| 2019 | High Performance Gesture Recognition via Effective and Efficient Temporal ModelingabstractState-of-the-art hand gesture recognition methods have investigated the spatiotemporal features based on 3D convolutional neural networks (3DCNNs) or convolutional long short-term memory (ConvLSTM). However, they often suffer from the inefficiency due to the high computational complexity of their network structures. In this paper, we focus instead on the 1D convolutional neural networks and propose a simple and efficient architectural unit, Multi-Kernel Temporal Block (MKTB), that models the multi-scale temporal responses by explicitly applying different temporal kernels. Then, we present a Global Refinement Block (GRB), which is an attention module for shaping the global temporal features based on the cross-channel similarity. By incorporating the MKTB and GRB, our architecture can effectively explore the spatiotemporal features within tolerable computational cost. Extensive experiments conducted on public datasets demonstrate that our proposed model achieves the state-of-the-art with higher efficiency. Moreover, the proposed MKTB and GRB are plug-and-play modules and the experiments on other tasks, like video understanding and video-based person re-identification, also display their good performance in efficiency and capability of generalization. Feng Ni, Yuexin Ma, Xinge Zhu, Yuankai Qi, Riming Qiu, Yongtao Wang |
IJCAI | 5 |
| 2019 | Robust visual tracking via scale-and-state-awareness
Yuankai Qi, Shengping Zhang, Qingming Huang, Hongxun Yao |
Neurocomputing | 1 |
| 2019 | Hedging Deep Features for Visual TrackingabstractConvolutional Neural Networks (CNNs) have been applied to visual tracking with demonstrated success in recent years. Most CNN-based trackers utilize hierarchical features extracted from a certain layer to represent the target. However, features from a certain layer are not always effective for distinguishing the target object from the backgrounds especially in the presence of complicated interfering factors (e.g., heavy occlusion, background clutter, illumination variation, and shape deformation). In this work, we propose a CNN-based tracking algorithm which hedges deep features from different CNN layers to better distinguish target objects and background clutters. Correlation filters are applied to feature maps of each CNN layer to construct a weak tracker, and all weak trackers are hedged into a strong one. For robust visual tracking, we propose a hedge method to adaptively determine weights of weak classifiers by considering both the difference between the historical as well as instantaneous performance, and the difference among all weak trackers over time. In addition, we design a Siamese network to define the loss of each weak tracker for the proposed hedge method. Extensive experiments on large benchmark datasets demonstrate the effectiveness of the proposed algorithm against the state-of-the-art tracking methods. Yuankai Qi, Shengping Zhang, Qingming Huang, Hongxun Yao, Jongwoo Lim, Ming-Hsuan Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2018 | The Unmanned Aerial Vehicle Benchmark: Object Detection and Tracking
Dawei Du, Yuankai Qi, Hongyang Yu 0001, Kaiwen Duan, Guorong Li, Weigang Zhang, Qingming Huang, Qi Tian 0001 |
ECCV (10) | 2 |
| 2018 | Plant identification based on very deep convolutional neural networks
Heyan Zhu, Qinglin Liu, Yuankai Qi, Feng Jiang 0001, Shengping Zhang |
Multim. Tools Appl. | 3 |
| 2018 | BoMW: Bag of Manifold Words for One-Shot Learning Gesture Recognition From KinectabstractIn this paper, we study one-shot learning gesture recognition on RGB-D data recorded from Microsoft's Kinect. To this end, we propose a novel bag of manifold words (BoMW)-based feature representation on symmetric positive definite (SPD) manifolds. In particular, we use covariance matrices to extract local features from RGB-D data due to its compact representation ability as well as the convenience of fusing both RGB and depth information. Since covariance matrices are SPD matrices and the space spanned by them is the SPD manifold, traditional learning methods in the Euclidean space, such as sparse coding, cannot be directly applied to them. To overcome this problem, we propose a unified framework to transfer the sparse coding on SPD manifolds to the one on the Euclidean space, which enables any existing learning method to be used. After building BoMW representation on a video from each gesture class, a nearest neighbor classifier is adopted to perform the one-shot learning gesture recognition. Experimental results on the ChaLearn gesture data set demonstrate the outstanding performance of the proposed one-shot learning gesture recognition method compared against the state-of-the-art methods. The effectiveness of the proposed feature extraction method is also validated on a new RGB-D action recognition data set. Lei Zhang 0036, Shengping Zhang, Feng Jiang 0001, Yuankai Qi, Jun Zhang 0017, Yuliang Guo, Huiyu Zhou 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2018 | Structure-Aware Local Sparse Coding for Visual TrackingabstractSparse coding has been applied to visual tracking and related vision problems with demonstrated success in recent years. Existing tracking methods based on local sparse coding sample patches from a target candidate and sparsely encode these using a dictionary consisting of patches sampled from target template images. The discriminative strength of existing methods based on local sparse coding is limited as spatial structure constraints among the template patches are not exploited. To address this problem, we propose a structure-aware local sparse coding algorithm, which encodes a target candidate using templates with both global and local sparsity constraints. For robust tracking, we show the local regions of a candidate region should be encoded only with the corresponding local regions of the target templates that are the most similar from the global view. Thus, a more precise and discriminative sparse representation is obtained to account for appearance changes. To alleviate the issues with tracking drifts, we design an effective template update scheme. Extensive experiments on challenging image sequences demonstrate the effectiveness of the proposed algorithm against numerous state-of-the-art methods. Yuankai Qi, Jian Zhang 0018, Shengping Zhang, Qingming Huang, Ming-Hsuan Yang 0001 |
IEEE Trans. Image Process. | 1 |
| 2018 | Point-to-Set Distance Metric Learning on Deep Representations for Visual TrackingabstractFor autonomous driving application, a car shall be able to track objects in the scene in order to estimate where and how they will move such that the tracker embedded in the car can efficiently alert the car for effective collision-avoidance. Traditional discriminative object tracking methods usually train a binary classifier via a support vector machine (SVM) scheme to distinguish the target from its background. Despite demonstrated success, the performance of the SVM-based trackers is limited because the classification is carried out only depending on support vectors (SVs) but the target's dynamic appearance may look similar to the training samples that have not been selected as SVs, especially when the training samples are not linearly classifiable. In such cases, the tracker may drift to the background and fail to track the target eventually. To address this problem, in this paper, we propose to integrate the point-to-set/image-to-imageSet distance metric learning (DML) into visual tracking tasks and take full advantage of all the training samples when determining the best target candidate. The point-to-set DML is conducted on convolutional neural network features of the training data extracted from the starting frames. When a new frame comes, target candidates are first projected to the common subspace using the learned mapping functions, and then the candidate having the minimal distance to the target template sets is selected as the tracking result. Extensive experimental results show that even without model update the proposed method is able to achieve favorable performance on challenging image sequences compared with several state-of-the-art trackers. Shengping Zhang, Yuankai Qi, Feng Jiang 0001, Xiangyuan Lan, Pong C. Yuen, Huiyu Zhou 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2017 | Robust Visual Tracking via Basis MatchingabstractMost existing tracking approaches are based on either the tracking by detection framework or the tracking by matching framework. The former needs to learn a discriminative classifier using positive and negative samples, which will cause tracking drift due to unreliable samples. The latter usually performs tracking by matching local interest points between a target candidate and the tracked target, which is not robust to target appearance changes over time. In this paper, we propose a novel tracking by matching framework for robust tracking based on basis matching rather than point matching. In particular, we learn the target model from target images using a set of Gabor basis functions, which have large responses on the corresponding spatial positions after a max pooling. During tracking, a target candidate is evaluated by computing the responses of the Gabor basis functions on their corresponding spatial positions. The experimental results on a set of challenging sequences validate that the performance of the proposed tracking method outperforms those of several state-of-the-art methods. Shengping Zhang, Xiangyuan Lan, Yuankai Qi, Pong C. Yuen |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2016 | Hedged Deep TrackingabstractIn recent years, several methods have been developed to utilize hierarchical features learned from a deep convolutional neural network (CNN) for visual tracking. However, as features from a certain CNN layer characterize an object of interest from only one aspect or one level, the performance of such trackers trained with features from one layer (usually the second to last layer) can be further improved. In this paper, we propose a novel CNN based tracking framework, which takes full advantage of features from different CNN layers and uses an adaptive Hedge method to hedge several CNN based trackers into a single stronger one. Extensive experiments on a benchmark dataset of 100 challenging image sequences demonstrate the effectiveness of the proposed algorithm compared to several state-of-theart trackers. Yuankai Qi, Shengping Zhang, Hongxun Yao, Qingming Huang, Jongwoo Lim, Ming-Hsuan Yang 0001 |
CVPR | 1 |
| 2014 | Structure-aware multi-object discovery for weakly supervised trackingabstractRecent progress on tracking has focused on designing robust statistical model or proposing effective appearance features to improve precision. This paper addresses another problem, namely the discovery and tracking of generic multi-object which have the similar appearance and motion pattern based on limited human annotations. We present a model-free tracking method that can automatically discover and track multi-object sharing the same spatial and motion structure, and update the structure during the tracking without prior acknowledge. The candidate objects are first selected by a SVM classifier trained on histogram-of-gradient (HOG) features. Then a segment algorithm is exploited to decide the suitable sizes of tracking boxes. The structure constrains are updated in a real-time manner according to the motion measure among the specified object and corresponding candidates. Experimental results reveal significant convenience and remarkable performance of our approach for the task of structure preserving multi-object discovery and tracking. Yuankai Qi, Hongxun Yao, Xiaoshuai Sun, Xin Sun 0003, Yanhao Zhang 0001, Qingming Huang |
ICIP | 1 |
| 2013 | 3D Segmentation of the Lung Based on the Neighbor Information and CurvatureabstractA novel method for the automatic segmentation of the lung in X-ray computed tomography (CT) images is presented. In this paper, a maximum a posteriori (MAP) estimation framework, combining neighbor prior information and image gray level information, is used to extract the boundary of lung. The relationship of the left lung and the right lung is represented as a joint density function. We use the principal component analysis (PCA) to build the neighbor prior model in a set of training images. A double dimension reduction algorithm is developed to improve the efficiency. The model is formulated in terms of level set functions, and the surfaces evolve according to the associated Euler-Lagrange equations. Then we propose a new algorithm to refine the rough boundary generated by the MAP framework. This algorithm consists of two stages: 1. automatically detecting and rough fitting the region of lung hilum, 2. refining the fitting curve based on the curvature information. Yuankai Qi, Kaikun Dong |
ICIG | 1 |