Yuwei Wu 0001

dblp:63/5298-1 · DBLP profile ↗
← Back
98ranked-venue papers
12as first author
49since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 64 · 6 first-author · 41 since 2021Graphics, computer vision, multimedia, augmented reality and games · 64 · 7 first-author · 32 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-authorSystems, architecture and hardware · 1Computer networks · 1 · 1 since 2021Security and privacy · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 LongSplat: Online Generalizable 3D Gaussian Splatting from Long Sequence Images
abstract
3D Gaussian Splatting (3DGS) achieves high-fidelity novel view synthesis, but its application in online long-sequence scenarios is still restricted. Existing methods either rely on slow per-scene optimization or lack efficient frame-wise 3DGS updates, making them unsuitable for online long-sequence videos. In this paper, we propose LongSplat, an online real-time 3D Gaussian reconstruction framework designed for long-sequence image input. The core idea of LongSplat is to maintain a global 3DGS set and design a streaming 3DGS update mechanism that selectively compressing redundant historical Gaussians and introducing new Gaussians by comparing the current observations with the historical Gaussian. To achieve this goal, we design a Gaussian-Image Representation (GIR), which encodes 3D Gaussian parameters into a structured, image-like 2D format. GIR simultaneously enables identity-aware redundancy compression as well as the fusion of current view and historical Gaussians, which are used for online reconstruction and adapt the model to long sequences without overwhelming memory or computational costs. Extensive experiments demonstrate that LongSplat achieves state-of-the-art efficiency-quality trade-offs in real-time novel view synthesis, delivering real-time reconstruction while reducing Gaussian counts by 44% compared to per-pixel prediction paradigms.
Guichen Huang, Ruoyu Wang 0014, Xiangjun Gao, Che Sun, Yuwei Wu 0001, Shenghua Gao, Yunde Jia
AAAI5
2026 Composition-Incremental Learning for Compositional Generalization
abstract
Compositional generalization has achieved substantial progress in computer vision on pre-collected training data. Nonetheless, real-world data continually emerges, with possible compositions being nearly infinite, long-tailed, and not entirely visible. Thus, an ideal model is supposed to gradually improve the capability of compositional generalization in an incremental manner. In this paper, we explore Composition-Incremental Learning for Compositional Generalization (CompIL) in the context of the compositional zero-shot learning (CZSL) task, where models need to continually learn new compositions, intending to improve their compositional generalization capability progressively. To quantitatively evaluate CompIL, we develop a benchmark construction pipeline leveraging existing datasets, yielding MIT-States-CompIL and C-GQA-CompIL. Furthermore, we propose a pseudo-replay framework utilizing a visual synthesizer to synthesize visual representations of learned compositions and a linguistic primitive distillation mechanism to maintain aligned primitive representations across the learning process. Extensive experiments demonstrate the effectiveness of the proposed framework.
Zhen Li 0026, Yuwei Wu 0001, Chenchen Jing, Che Sun, Chuanhao Li 0001, Yunde Jia
AAAI2
2026 3D scene reconstruction from a limited number of viewpoints
Zhidan Liu 0005, Che Sun, Yuwei Wu 0001
Comput. Vis. Image Underst.4
2026 Robust video anomaly detection via causal feature-guided data augmentation
Chenrui Shi, Che Sun, Yuwei Wu 0001
Comput. Vis. Image Underst.4
2026 Riemannian Implicit Differentiation via a Fixed-Point Equation for Riemannian Bilevel Optimization
abstract
Various Riemannian optimization tasks, such as Riemannian metaoptimization (RMO) and Riemannian metalearning, can be formulated as Riemannian bilevel optimization problems (i.e., the inner-level and outer-level optimization). Implicit differentiation has shown effectiveness in solving RMO, which decouples the computation of outer gradients from the inner-level process, avoiding huge computational burdens. However, extending implicit differentiation to other Riemannian bilevel optimization tasks is nontrivial because it requires much expert involvement for case-by-case derivations. In this article, we propose a Riemannian implicit differentiation method that provides a unified expression for outer gradients, leading to flexible application to other tasks with less expert involvement. Specifically, we formulate the inner-level optimization as a root-finding process of a fixed-point equation, through which the inner-level optimization among different tasks is formulated in a unified way. By differentiating the fixed-point equation, we derive a unified expression for outer gradients, circumventing the case-by-case derivations for different tasks. Then, we present convergence analysis and approximation error analysis, which guarantee the effectiveness of our method in various Riemannian optimization tasks. We further conduct experiments on multiple Riemannian optimization tasks, and the experimental results confirm the effectiveness.
Xiaomeng Fan, Yuwei Wu 0001, Zhi Gao 0002, Zhipeng Lu 0003, Mehrtash Harandi, Yunde Jia
IEEE Trans. Neural Networks Learn. Syst.2
2025 Consistency of Compositional Generalization Across Multiple Levels
abstract
Compositional generalization is the capability of a model to understand novel compositions composed of seen concepts. There are multiple levels of novel compositions including phrase-phrase level, phrase-word level, and word-word level. Existing methods achieve promising compositional generalization, but the consistency of compositional generalization across multiple levels of novel compositions remains unexplored. The consistency refers to that a model should generalize to a phrase-phrase level novel composition, and phrase-word/word-word level novel compositions that can be derived from it simultaneously. In this paper, we propose a meta-learning based framework, for achieving consistent compositional generalization across multiple levels. The basic idea is to progressively learn compositions from simple to complex for consistency. Specifically, we divide the original training set into multiple validation sets based on compositional complexity, and introduce multiple meta-weight-nets to generate sample weights for samples in different validation sets. To fit the validation sets in order of increasing compositional complexity, we optimize the parameters of each meta-weight-net independently and sequentially in a multilevel optimization manner. We build a GQA-CCG dataset to quantitatively evaluate the consistency. Experimental results on visual question answering and temporal video grounding, demonstrate the effectiveness of the proposed framework.
Chuanhao Li 0001, Zhen Li 0026, Chenchen Jing, Xiaomeng Fan, Wenbo Ye, Yuwei Wu 0001, Yunde Jia
AAAI6
2025 World Knowledge-Enhanced Reasoning Using Instruction-Guided Interactor in Autonomous Driving
abstract
The Multi-modal Large Language Models (MLLMs) with extensive world knowledge have revitalized autonomous driving, particularly in reasoning tasks within perceivable regions. However, when faced with perception-limited areas (dynamic or static occlusion regions), MLLMs struggle to effectively integrate perception ability with world knowledge for reasoning. These perception-limited regions can conceal crucial safety information, especially for vulnerable road users. In this paper, we propose a framework, which aims to improve autonomous driving performance under perception-limited conditions by enhancing the integration of perception capabilities and world knowledge. Specifically, we propose a plug-and-play instruction-guided interaction module that bridges modality gaps and significantly reduces the input sequence length, allowing it to adapt effectively to multi-view video inputs. Furthermore, to better integrate world knowledge with driving-related tasks, we have collected and refined a large-scale multi-modal dataset that includes 2 million natural language QA pairs, 1.7 million grounding task data. To evaluate the model’s utilization of world knowledge, we introduce an object-level risk assessment dataset comprising 200K QA pairs, where the questions necessitate multi-step reasoning leveraging world knowledge for resolution. Extensive experiments validate the effectiveness of our proposed method.
Mingliang Zhai, Zengyuan Guo, Ningrui Yang, Xiameng Qin, Sanyuan Zhao, Junyu Han, Ji Tao, Yuwei Wu 0001, Yunde Jia
AAAI9
2025 Diving into the Fusion of Monocular Priors for Generalized Stereo Matching
abstract
The matching formulation makes it naturally hard for the stereo matching to handle ill-posed regions like occlusions and non-Lambertian surfaces. Fusing monocular priors has been proven helpful for ill-posed matching, but the biased monocular prior learned from small stereo datasets constrains the generalization. Recently, stereo matching has progressed by leveraging the unbiased monocular prior from the vision foundation model (VFM) to improve the generalization in ill-posed regions. We dive into the fusion process and observe three main problems limiting the fusion of the VFM monocular prior. The first problem is the misalignment between affine-invariant relative monocular depth and absolute depth of disparity. Besides, when we use the monocular feature in an iterative update structure, the over-confidence in the disparity update leads to local optima results. A direct fusion of a monocular depth map could alleviate the local optima problem, but noisy disparity results computed at the first several iterations will misguide the fusion. In this paper, we propose a binary local ordering map to guide the fusion, which converts the depth map into a binary relative format, unifying the relative and absolute depth representation. The computed local ordering map is also used to re-weight the initial disparity update, resolving the local optima and noisy problem. In addition, we formulate the final direct fusion of monocular depth to the disparity as a registration problem, where a pixel-wise linear regression module can globally and adaptively align them. Our method fully exploits the monocular prior to support stereo matching results effectively and efficiently. We significantly improve the performance from the experiments when generalizing from SceneFlow to Middlebury and Booster datasets while barely reducing the efficiency.
Chengtang Yao, Lidong Yu, Zhidan Liu 0005, Jiaxi Zeng, Yuwei Wu 0001, Yunde Jia
ICCV5
2025 Multi-modal Agent Tuning: Building a VLM-Driven Agent for Efficient Tool Usage
abstract
The advancement of large language models (LLMs) prompts the development of multi-modal agents, which are used as a controller to call external tools, providing a feasible way to solve practical tasks. In this paper, we propose a multi-modal agent tuning method that automatically generates multi-modal tool-usage data and tunes a vision-language model (VLM) as the controller for powerful tool-usage reasoning. To preserve the data quality, we prompt the GPT-4o mini model to generate queries, files, and trajectories, followed by query-file and trajectory verifiers. Based on the data synthesis pipeline, we collect the MM-Traj dataset that contains 20K tasks with trajectories of tool usage. Then, we develop the T3-Agent via Trajectory Tuning on VLMs for Tool usage using MM-Traj. Evaluations on the GTA and GAIA benchmarks show that the T3-Agent consistently achieves improvements on two popular VLMs: MiniCPM-V-8.5B and Qwen2-VL-7B, which outperforms untrained VLMs by 20%, showing the effectiveness of the proposed data synthesis pipeline, leading to high-quality data for tool-usage capabilities.
Zhi Gao 0002, Bofei Zhang, Pengxiang Li 0002, Xiaojian Ma 0001, Yuwei Wu 0001, Yunde Jia, Song-Chun Zhu, Qing Li 0003
ICLR7
2025 Multi-Sourced Compositional Generalization in Visual Question Answering
abstract
Compositional generalization is the ability of generalizing novel compositions from seen primitives, and has received much attention in vision-and-language (V&L) recently. Due to the multi-modal nature of V&L tasks, the primitives composing compositions source from different modalities, resulting in multi-sourced novel compositions. However, the generalization ability over multi-sourced novel compositions, i.e., multi-sourced compositional generalization (MSCG) remains unexplored. In this paper, we explore MSCG in the context of visual question answering (VQA), and propose a retrieval-augmented training framework to enhance the MSCG ability of VQA models by learning unified representations for primitives from different modalities. Specifically, semantically equivalent primitives are retrieved for each primitive in the training samples, and the retrieved features are aggregated with the original primitive to refine the model. This process helps the model learn consistent representations for the same semantic primitives across different modalities. To evaluate the MSCG ability of VQA models, we construct a new GQA-MSCG dataset based on the GQA dataset, in which samples include three types of novel compositions composed of primitives from different modalities. The GQA-MSCG dataset is available at https://github.com/NeverMoreLCH/MSCG.
Chuanhao Li 0001, Wenbo Ye, Zhen Li 0026, Yuwei Wu 0001, Yunde Jia
IJCAI4
2025 Beyond the Seen: Bounded Distribution Estimation for Open-Vocabulary Learning
abstract
Open-vocabulary learning requires modeling the data distribution in open environments, which consists of both seen-class and unseen-class data. Existing methods estimate the distribution in open environments using seen-class data, where the absence of unseen classes makes the estimation error inherently unidentifiable. Intuitively, learning beyond the seen classes is crucial for distribution estimation to bound the estimation error. We theoretically demonstrate that the distribution can be effectively estimated by generating unseen-class data, through which the estimation error is upper-bounded. Building on this theoretical insight, we propose a novel open-vocabulary learning method, which generates unseen-class data for estimating the distribution in open environments. The method consists of a class-domain-wise data generation pipeline and a distribution alignment algorithm. The data generation pipeline generates unseen-class data under the guidance of a hierarchical semantic tree and domain information inferred from the seen-class data, facilitating accurate distribution estimation. With the generated data, the distribution alignment algorithm estimates and maximizes the posterior probability to enhance generalization in open-vocabulary learning. Extensive experiments on 11 datasets demonstrate that our method outperforms baseline approaches by up to 14%, highlighting its effectiveness and superiority.
Xiaomeng Fan, Yuchuan Mao, Zhi Gao 0002, Yuwei Wu 0001, Jin Chen 0009, Yunde Jia
NeurIPS4
2025 Iterative Tool Usage Exploration for Multimodal Agents via Step-wise Preference Tuning
abstract
Multimodal agents, which integrate a controller (e.g., a vision language model) with external tools, have demonstrated remarkable capabilities in tackling complex multimodal tasks. Existing approaches for training these agents, both supervised fine-tuning and reinforcement learning, depend on extensive human-annotated task-answer pairs and tool trajectories. However, for complex multimodal tasks, such annotations are prohibitively expensive or impractical to obtain. In this paper, we propose an iterative tool usage exploration method for multimodal agents without any pre-collected data, namely SPORT, via step-wise preference optimization to refine the trajectories of tool usage. Our method enables multimodal agents to autonomously discover effective tool usage strategies through self-exploration and optimization, eliminating the bottleneck of human annotation. SPORT has four iterative components: task synthesis, step sampling, step verification, and preference tuning. We first synthesize multimodal tasks using language models. Then, we introduce a novel trajectory exploration scheme, where step sampling and step verification are executed alternately to solve synthesized tasks. In step sampling, the agent tries different tools and obtains corresponding results. In step verification, we employ a verifier to provide AI feedback to construct step-wise preference data. The data is subsequently used to update the controller for tool usage through preference tuning, producing a SPORT agent. By interacting with real environments, the SPORT agent gradually evolves into a more refined and capable system. Evaluation in the GTA and GAIA benchmarks shows that the SPORT agent achieves 6.41% and 3.64% improvements, underscoring the generalization and effectiveness introduced by our method.
Pengxiang Li 0002, Zhi Gao 0002, Bofei Zhang, Yapeng Mi, Xiaojian Ma 0001, Chenrui Shi, Yuwei Wu 0001, Yunde Jia, Song-Chun Zhu, Qing Li 0003
NeurIPS8
2025 Sekai: A Video Dataset towards World Exploration
abstract
Video generation techniques have made remarkable progress, promising to be the foundation of interactive world exploration.However, existing video generation datasets are not well-suited for world exploration training as they suffer from some limitations: limited locations, short duration, static scenes, and a lack of annotations about exploration and the world.In this paper, we introduce Sekai (meaning "world" in Japanese), a high-quality first-person view worldwide video dataset with rich annotations for world exploration. It consists of over 5,000 hours of walking or drone view (FPV and UVA) videos from over 100 countries and regions across 750 cities. We develop an efficient and effective toolbox to collect, pre-process and annotate videos with location, scene, weather, crowd density, captions, and camera trajectories.Comprehensive analyses and experiments demonstrate the dataset’s scale, diversity, annotation quality, and effectiveness for training video generation models.We believe Sekai will benefit the area of video generation and world exploration, and motivate valuable applications.
Zhen Li 0026, Chuanhao Li 0001, Xiaofeng Mao, Shaoheng Lin, Ming Li 0010, Shitian Zhao, Zhaopan Xu, Xinyue Li 0001, Yukang Feng, Zizhen Li, Fanrui Zhang, Jiaxin Ai, Yuwei Wu 0001, Tong He 0001, Yunde Jia, Kaipeng Zhang
NeurIPS15
2025 3D Visual Illusion Depth Estimation
abstract
3D visual illusion is a perceptual phenomenon where a two-dimensional plane is manipulated to simulate three-dimensional spatial relationships, making a flat artwork or object look three-dimensional in the human visual system. In this paper, we reveal that the machine visual system is also seriously fooled by 3D visual illusions, including monocular and binocular depth estimation. In order to explore and analyze the impact of 3D visual illusion on depth estimation, we collect a large dataset containing almost 3k scenes and 200k images to train and evaluate SOTA monocular and binocular depth estimation methods. We also propose a 3D visual illusion depth estimation framework that uses common sense from the vision language model to adaptively fuse depth from binocular disparity and monocular depth. Experiments show that SOTA monocular, binocular, and multi-view depth estimation approaches are all fooled by various 3D visual illusions, while our method achieves SOTA performance.
Chengtang Yao, Zhidan Liu 0005, Jiaxi Zeng, Lidong Yu, Yuwei Wu 0001, Yunde Jia
NeurIPS5
2025 Large-scale Riemannian meta-optimization via subspace adaptation
Peilin Yu 0001, Yuwei Wu 0001, Zhi Gao 0002, Xiaomeng Fan, Yunde Jia
Comput. Vis. Image Underst.2
2025 Curvature Learning for Generalization of Hyperbolic Neural Networks
Xiaomeng Fan, Yuwei Wu 0001, Zhi Gao 0002, Mehrtash Harandi, Yunde Jia
Int. J. Comput. Vis.2
2025 Inter-Scale Similarity Guided Cost Aggregation for Stereo Matching
abstract
Stereo matching aims to estimate 3D geometry by computing disparity from a rectified image pair. Most deep learning based stereo matching methods aggregate multi-scale cost volumes computed by downsampling and achieve good performance. However, their effectiveness in fine-grained areas is limited by significant detail loss during downsampling and the use of fixed weights in upsampling. In this paper, we propose an inter-scale similarity-guided cost aggregation method that dynamically upsamples the cost volumes according to the content of images for stereo matching. The method consists of two modules: inter-scale similarity measurement and stereo-content-aware cost aggregation. Specifically, we use inter-scale similarity measurement to generate similarity guidance from feature maps in adjacent scales. The guidance, generated from both reference and target images, is then used to aggregate the cost volumes from low-resolution to high-resolution via stereo-content-aware cost aggregation. We further split the 3D aggregation into 1D disparity and 2D spatial aggregation to reduce the computational cost. Experimental results on various benchmarks (e.g., SceneFlow, KITTI, Middlebury and ETH3D-two-view) show that our method achieves consistent performance gain on multiple models (e.g., PSM-Net, HSM-Net, CF-Net, FastAcv, and FactAcvPlus). The code can be found athttps://github.com/Pengxiang-Li/issga-stereo.
Pengxiang Li 0002, Chengtang Yao, Yunde Jia, Yuwei Wu 0001
IEEE Trans. Circuits Syst. Video Technol.4
2024 Residual Hyperbolic Graph Convolution Networks
abstract
Hyperbolic graph convolutional networks (HGCNs) have demonstrated representational capabilities of modeling hierarchical-structured graphs. However, as in general GCNs, over-smoothing may occur as the number of model layers increases, limiting the representation capabilities of most current HGCN models. In this paper, we propose residual hyperbolic graph convolutional networks (R-HGCNs) to address the over-smoothing problem. We introduce a hyperbolic residual connection function to overcome the over-smoothing problem, and also theoretically prove the effectiveness of the hyperbolic residual function. Moreover, we use product manifolds and HyperDrop to facilitate the R-HGCNs. The distinctive features of the R-HGCNs are as follows: (1) The hyperbolic residual connection preserves the initial node information in each layer and adds a hyperbolic identity mapping to prevent node features from being indistinguishable. (2) Product manifolds in R-HGCNs have been set up with different origin points in different components to facilitate the extraction of feature information from a wider range of perspectives, which enhances the representing capability of R-HGCNs. (3) HyperDrop adds multiplicative Gaussian noise into hyperbolic representations, such that perturbations can be added to alleviate the over-fitting problem without deconstructing the hyperbolic geometry. Experiment results demonstrate the effectiveness of R-HGCNs under various graph convolution layers and different structures of product manifolds.
Yangkai Xue, Jindou Dai, Zhipeng Lu 0003, Yuwei Wu 0001, Yunde Jia
AAAI4
2024 Compositional Substitutivity of Visual Reasoning for Visual Question Answering
Chuanhao Li 0001, Zhen Li 0026, Chenchen Jing, Yuwei Wu 0001, Mingliang Zhai, Yunde Jia
ECCV (48)4
2024 Temporally Consistent Stereo Matching
Jiaxi Zeng, Chengtang Yao, Yuwei Wu 0001, Yunde Jia
ECCV (31)3
2024 In-Context Compositional Generalization for Large Vision-Language Models
abstract
Recent work has revealed that in-context learning for large language models exhibits compositional generalization capacity, which can be enhanced by selecting in-context demonstrations similar to test cases to provide contextual information.However, how to exhibit in-context compositional generalization (ICCG) of large vision-language models (LVLMs) is non-trival.Due to the inherent asymmetry between visual and linguistic modalities, ICCG in LVLMs faces an inevitable challenge-redundant information on the visual modality.The redundant information affects in-context learning from two aspects: (1) Similarity calculation may be dominated by redundant information, resulting in sub-optimal demonstration selection.(2) Redundant information in in-context demonstrations brings misleading contextual information to in-context learning.To alleviate these problems, we propose a demonstration selection method to achieve ICCG for LVLMs, by considering two key factors of demonstrations: content and structure, from a multimodal perspective.Specifically, we design a diversity-coverage-based matching score to select demonstrations with maximum coverage, and avoid selecting demonstrations with redundant information via their content redundancy and structural complexity.We build a GQA-ICCG dataset to simulate the ICCG setting, and conduct experiments on GQA-ICCG and the VQA v2 dataset.Experimental results demonstrate the effectiveness of our method.
Chuanhao Li 0001, Chenchen Jing, Zhen Li 0026, Mingliang Zhai, Yuwei Wu 0001, Yunde Jia
EMNLP5
2024 FIRE: A Dataset for Feedback Integration and Refinement Evaluation of Multimodal Models
abstract
Vision language models (VLMs) have achieved impressive progress in diverse applications, becoming a prevalent research direction. In this paper, we build FIRE, a feedback-refinement dataset, consisting of 1.1M multi-turn conversations that are derived from 27 source datasets, empowering VLMs to spontaneously refine their responses based on user feedback across diverse tasks. To scale up the data collection, FIRE is collected in two components: FIRE-100K and FIRE-1M, where FIRE-100K is generated by GPT-4V, and FIRE-1M is freely generated via models trained on FIRE-100K. Then, we build FIRE-Bench, a benchmark to comprehensively evaluate the feedback-refining capability of VLMs, which contains 11K feedback-refinement conversations as the test data, two evaluation settings, and a model to provide feedback for VLMs. We develop the FIRE-LLaVA model by fine-tuning LLaVA on FIRE-100K and FIRE-1M, which shows remarkable feedback-refining capability on FIRE-Bench and outperforms untrained VLMs by 50%, making more efficient user-agent interactions and underscoring the significance of the FIRE dataset.
Pengxiang Li 0002, Zhi Gao 0002, Bofei Zhang, Yuwei Wu 0001, Mehrtash Harandi, Yunde Jia, Song-Chun Zhu, Qing Li 0003
NeurIPS5
2024 SearchLVLMs: A Plug-and-Play Framework for Augmenting Large Vision-Language Models by Searching Up-to-Date Internet Knowledge
abstract
Large vision-language models (LVLMs) are ignorant of the up-to-date knowledge, such as LLaVA series, because they cannot be updated frequently due to the large amount of resources required, and therefore fail in many cases. For example, if a LVLM was released on January 2024, and it wouldn't know the singer of the theme song for the new Detective Conan movie, which wasn't released until April 2024. To solve the problem, a promising solution motivated by retrieval-augmented generation (RAG) is to provide LVLMs with up-to-date knowledge via internet search during inference, i.e., internet-augmented generation (IAG), which is already integrated in some closed-source commercial LVLMs such as GPT-4V. However, the specific mechanics underpinning them remain a mystery. In this paper, we propose a plug-and-play framework, for augmenting existing LVLMs in handling visual question answering (VQA) about up-to-date knowledge, dubbed SearchLVLMs. A hierarchical filtering model is trained to effectively and efficiently find the most helpful content from the websites returned by a search engine to prompt LVLMs with up-to-date knowledge. To train the model and evaluate our framework's performance, we propose a pipeline to automatically generate news-related VQA samples to construct a dataset, dubbed UDK-VQA. A multi-model voting mechanism is introduced to label the usefulness of website/content for VQA samples to construct the training set. Experimental results demonstrate the effectiveness of our framework, outperforming GPT-4o by $\sim$30\% in accuracy.
Chuanhao Li 0001, Zhen Li 0026, Chenchen Jing, Wenqi Shao, Yuwei Wu 0001, Ping Luo 0002, Yu Qiao 0001, Kaipeng Zhang
NeurIPS6
2024 Visual-Guided Reasoning Path Generation for Visual Question Answering
Chenchen Jing, Mingliang Zhai, Yuwei Wu 0001, Yunde Jia
PRCV (1)4
2024 Adversarial Sample Synthesis for Visual Question Answering
abstract
Language prior is a major block to improving the generalization of visual question answering (VQA) models. Recent work has revealed that synthesizing extra training samples to balance training sets is a promising way to alleviate language priors. However, most existing methods synthesize extra samples in a manner independent of training processes, which neglect the fact that the language priors memorized by VQA models are changing during training, resulting in insufficient synthesized samples. In this article, we propose an adversarial sample synthesis method, which synthesizes different adversarial samples by adversarial masking at different training epochs to cope with the changing memorized language priors. The basic idea behind our method is to use adversarial masking to synthesize adversarial samples that will cause the model to make wrong answers. To this end, we design a generative module to carry out adversarial masking by attacking the VQA model and introduce a bias-oriented objective to supervise the training of the generative module. We couple the sample synthesis with the training process of the VQA model, which ensures that the synthesized samples at different training epochs are beneficial to the VQA model. We incorporated the proposed method into three VQA models including UpDn, LMH, and LXMERT and conducted experiments on three datasets including VQA-CP v1, VQA-CP v2, and VQA v2. Experimental results demonstrate that a large improvement of our method, such as 16.22% gains on LXMERT in the overall accuracy of VQA-CP v2.
Chuanhao Li 0001, Chenchen Jing, Zhen Li 0026, Yuwei Wu 0001, Yunde Jia
ACM Trans. Multim. Comput. Commun. Appl.4
2023 Learning Event-Relevant Factors for Video Anomaly Detection
abstract
Most video anomaly detection methods discriminate events that deviate from normal patterns as anomalies. However, these methods are prone to interferences from event-irrelevant factors, such as background textures and object scale variations, incurring an increased false detection rate. In this paper, we propose to explicitly learn event-relevant factors to eliminate the interferences from event-irrelevant factors on anomaly predictions. To this end, we introduce a causal generative model to separate the event-relevant factors and event-irrelevant ones in videos, and learn the prototypes of event-relevant factors in a memory augmentation module. We design a causal objective function to optimize the causal generative model and develop a counterfactual learning strategy to guide anomaly predictions, which increases the influence of the event-relevant factors. The extensive experiments show the effectiveness of our method for video anomaly detection.
Che Sun, Chenrui Shi, Yunde Jia, Yuwei Wu 0001
AAAI4
2023 Exploring Data Geometry for Continual Learning
abstract
Continual learning aims to efficiently learn from a non-stationary stream of data while avoiding forgetting the knowledge of old data. In many practical applications, data complies with non-Euclidean geometry. As such, the commonly used Euclidean space cannot gracefully capture non-Euclidean geometric structures of data, leading to in-ferior results. In this paper, we study continual learning from a novel perspective by exploring data geometry for the non-stationary stream of data. Our method dynamically expands the geometry of the underlying space to match growing geometric structures induced by new data, and pre-vents forgetting by keeping geometric structures of old data into account. In doing so, making use of the mixed cur-vature space, we propose an incremental search scheme, through which the growing geometric structures are en-coded. Then, we introduce an angular-regularization loss and a neighbor-robustness loss to train the model, capa-ble of penalizing the change of global geometric structures and local geometric structures. Experiments show that our method achieves better performance than baseline methods designed in Euclidean space.
Zhi Gao 0002, Yunde Jia, Mehrtash Harandi, Yuwei Wu 0001
CVPR6
2023 Exploring the Effect of Primitives for Compositional Generalization in Vision-and-Language
abstract
Compositionality is one of the fundamental properties of human cognition (Fodor & Pylyshyn, 1988). Compositional generalization is critical to simulate the compositional capability of humans, and has received much attention in the vision-and-language (V&L) community. It is essential to understand the effect of the primitives, including words, image regions, and video frames, to improve the compositional generalization capability. In this paper, we explore the effect of primitives for compositional generalization in V&L. Specifically, we present a self-supervised learning based framework that equips existing V&L methods with two characteristics: semantic equivariance and semantic invariance. With the two characteristics, the methods understand primitives by perceiving the effect of primitive changes on sample semantics and ground-truth. Experimental results on two tasks: temporal video grounding and visual question answering, demonstrate the effectiveness of our framework.
Chuanhao Li 0001, Zhen Li 0026, Chenchen Jing, Yunde Jia, Yuwei Wu 0001
CVPR5
2023 Video Anomaly Detection via Sequentially Learning Multiple Pretext Tasks
abstract
Learning multiple pretext tasks is a popular approach to tackle the nonalignment problem in unsupervised video anomaly detection. However, the conventional learning method of simultaneously learning multiple pretext tasks, is prone to sub-optimal solutions, incurring sharp performance drops. In this paper, we propose to sequentially learn multiple pretext tasks according to their difficulties in an ascending manner to improve the performance of anomaly detection. The core idea is to relax the learning objective by starting with easy pretext tasks in the early stage and gradually refine it by involving more challenging pretext tasks later on. In this way, our method is able to reduce the difficulties of learning and avoid converging to sub-optimal solutions. Specifically, we design a tailored sequential learning order for three widely-used pretext tasks. It starts with frame prediction task, then moves on to frame reconstruction task and last ends with frame-order classification task. We further introduce a new contrastive loss which makes the learned representations of normality more discriminative by pushing normal and pseudo-abnormal samples apart. Extensive experiments on three datasets demonstrate the effectiveness of our method.
Chenrui Shi, Che Sun, Yuwei Wu 0001, Yunde Jia
ICCV3
2023 Sparse Point Guided 3D Lane Detection
abstract
3D lane detection usually builds a dense correspondence between the front-view space and the BEV space to estimate lane points in the 3D space. 3D lanes only occupy a small ratio of the dense correspondence, while most correspondence belongs to the redundant background. This sparsity phenomenon bottlenecks valuable computation and raises the computation cost of building a high-resolution correspondence for accurate results. In this paper, we propose a sparse point-guided 3D lane detection, focusing on points related to 3D lanes. Our method runs in a coarse-to-fine manner, including coarse-level lane detection and iterative fine-level sparse point refinements. In coarse-level lane detection, we build a dense but efficient correspondence between the front view and BEV space at a very low resolution to compute coarse lanes. Then in fine-level sparse point refinement, we sample sparse points around coarse lanes to extract local features from the high-resolution front-view feature map. The high-resolution local information brought by sparse points refines 3D lanes in the BEV space hierarchically from low resolution to high resolution. The sparse point guides a more effective information flow and greatly promotes the SOTA result by 3 points on the overall F1-score and 6 points on several hard situations while reducing almost half memory cost and speeding up 2 times.
Chengtang Yao, Lidong Yu, Yuwei Wu 0001, Yunde Jia
ICCV3
2023 Parameterized Cost Volume for Stereo Matching
abstract
Stereo matching becomes computationally challenging when dealing with a large disparity range. Prior methods mainly alleviate the computation through dynamic cost volume by focusing on a local disparity space, but it requires many iterations to get close to the ground truth due to the lack of a global view. We find that the dynamic cost volume approximately encodes the disparity space as a single Gaussian distribution with a fixed and small variance at each iteration, which results in an inadequate global view over disparity space and a small update step at every iteration. In this paper, we propose a parameterized cost volume to encode the entire disparity space using multi-Gaussian distribution. The disparity distribution of each pixel is parameterized by weights, means, and variances. The means and variances are used to sample disparity candidates for cost computation, while the weights and means are used to calculate the disparity output. The above parameters are computed through a JS-divergence-based optimization, which is realized as a gradient descent update in a feed-forward differential module. Experiments show that our method speeds up the runtime of RAFT-Stereo by 4 ~ 15 times, achieving real-time performance and comparable accuracy. The code is available at https://github.com/jiaxiZeng/Parameterized-Cost-Volume-for-Stereo-Matching.
Jiaxi Zeng, Chengtang Yao, Lidong Yu, Yuwei Wu 0001, Yunde Jia
ICCV4
2023 Fast-StrucTexT: An Efficient Hourglass Transformer with Modality-guided Dynamic Token Merge for Document Understanding
abstract
Transformers achieve promising performance in document understanding because of their high effectiveness and still suffer from quadratic computational complexity dependency on the sequence length. General efficient transformers are challenging to be directly adapted to model document. They are unable to handle the layout representation in documents, e.g. word, line and paragraph, on different granularity levels and seem hard to achieve a good trade-off between efficiency and performance. To tackle the concerns, we propose Fast-StrucTexT, an efficient multi-modal framework based on the StrucTexT algorithm with an hourglass transformer architecture, for visual document understanding. Specifically, we design a modality-guided dynamic token merging block to make the model learn multi-granularity representation and prunes redundant tokens. Additionally, we present a multi-modal interaction module called Symmetry Cross-Attention (SCA) to consider multi-modal fusion and efficiently guide the token mergence. The SCA allows one modality input as query to calculate cross attention with another modality in a dual phase. Extensive experiments on FUNSD, SROIE, and CORD datasets demonstrate that our model achieves the state-of-the-art performance and almost 1.9x faster inference time than the state-of-the-art methods.
Mingliang Zhai, Yulin Li 0004, Xiameng Qin, Qunyi Xie, Chengquan Zhang, Yuwei Wu 0001, Yunde Jia
IJCAI8
2023 Learning to Optimize on Riemannian Manifolds
abstract
Many learning tasks are modeled as optimization problems with nonlinear constraints, such as principal component analysis and fitting a Gaussian mixture model. A popular way to solve such problems is resorting to Riemannian optimization algorithms, which yet heavily rely on both human involvement and expert knowledge about Riemannian manifolds. In this paper, we propose a Riemannian meta-optimization method to automatically learn a Riemannian optimizer. We parameterize the Riemannian optimizer by a novel recurrent network and utilize Riemannian operations to ensure that our method is faithful to the geometry of manifolds. The proposed method explores the distribution of the underlying data by minimizing the objective of updated parameters, and hence is capable of learning task-specific optimizations. We introduce a Riemannian implicit differentiation training scheme to achieve efficient training in terms of numerical stability and computational cost. Unlike conventional meta-optimization training schemes that need to differentiate through the whole optimization trajectory, our training scheme is only related to the final two optimization steps. In this way, our training scheme avoids the exploding gradient problem, and significantly reduces the computational load and memory footprint. We discuss experimental results across various constrained problems, including principal component analysis on Grassmann manifolds, face recognition, person re-identification, and texture image classification on Stiefel manifolds, clustering and similarity learning on symmetric positive definite manifolds, and few-shot learning on hyperbolic manifolds.
Zhi Gao 0002, Yuwei Wu 0001, Xiaomeng Fan, Mehrtash Harandi, Yunde Jia
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Curvature-Adaptive Meta-Learning for Fast Adaptation to Manifold Data
abstract
Meta-learning methods are shown to be effective in quickly adapting a model to novel tasks. Most existing meta-learning methods represent data and carry out fast adaptation in euclidean space. In fact, data of real-world applications usually resides in complex and various Riemannian manifolds. In this paper, we propose a curvature-adaptive meta-learning method that achieves fast adaptation to manifold data by producing suitable curvature. Specifically, we represent data in the product manifold of multiple constant curvature spaces and build a product manifold neural network as the base-learner. In this way, our method is capable of encoding complex manifold data into discriminative and generic representations. Then, we introduce curvature generation and curvature updating schemes, through which suitable product manifolds for various forms of data manifolds are constructed via few optimization steps. The curvature generation scheme identifies task-specific curvature initialization, leading to a shorter optimization trajectory. The curvature updating scheme automatically produces appropriate learning rate and search direction for curvature, making a faster and more adaptive optimization paradigm compared to hand-designed optimization schemes. We evaluate our method on a broad set of problems including few-shot classification, few-shot regression, and reinforcement learning tasks. Experimental results show that our method achieves substantial improvements as compared to meta-learning methods ignoring the geometry of the underlying space.
Zhi Gao 0002, Yuwei Wu 0001, Mehrtash Harandi, Yunde Jia
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 Efficient Riemannian Meta-Optimization by Implicit Differentiation
abstract
To solve optimization problems with nonlinear constrains, the recently developed Riemannian meta-optimization methods show promise, which train neural networks as an optimizer to perform optimization on Riemannian manifolds. A key challenge is the heavy computational and memory burdens, because computing the meta-gradient with respect to the optimizer involves a series of time-consuming derivatives, and stores large computation graphs in memory. In this paper, we propose an efficient Riemannian meta-optimization method that decouples the complex computation scheme from the meta-gradient. We derive Riemannian implicit differentiation to compute the meta-gradient by establishing a link between Riemannian optimization and the implicit function theorem. As a result, the updating our optimizer is only related to the final two iterations, which in turn speeds up our method and reduces the memory footprint significantly. We theoretically study the computational load and memory footprint of our method for long optimization trajectories, and conduct an empirical study to demonstrate the benefits of the proposed method. Evaluations of three optimization problems on different Riemannian manifolds show that our method achieves state-of-the-art performance in terms of the convergence speed and the quality of optima.
Xiaomeng Fan, Yuwei Wu 0001, Zhi Gao 0002, Yunde Jia, Mehrtash Harandi
AAAI2
2022 Learning the Dynamics of Visual Relational Reasoning via Reinforced Path Routing
abstract
Reasoning is a dynamic process. In cognitive theories, the dynamics of reasoning refers to reasoning states over time after successive state transitions. Modeling the cognitive dynamics is of utmost importance to simulate human reasoning capability. In this paper, we propose to learn the reasoning dynamics of visual relational reasoning by casting it as a path routing task. We present a reinforced path routing method that represents an input image via a structured visual graph and introduces a reinforcement learning based model to explore paths (sequences of nodes) over the graph based on an input sentence to infer reasoning results. By exploring such paths, the proposed method represents reasoning states clearly and characterizes state transitions explicitly to fully model the reasoning dynamics for accurate and transparent visual relational reasoning. Extensive experiments on referring expression comprehension and visual question answering demonstrate the effectiveness of our method.
Chenchen Jing, Yunde Jia, Yuwei Wu 0001, Chuanhao Li 0001, Qi Wu 0001
AAAI3
2022 Maintaining Reasoning Consistency in Compositional Visual Question Answering
abstract
A compositional question refers to a question that contains multiple visual concepts (e.g., objects, attributes, and relationships) and requires compositional reasoning to answer. Existing VQA models can answer a compositional question well, but cannot work well in terms of reasoning consistency in answering the compositional question and its sub-questions. For example, a compositional question for an image is: “Are there any elephants to the right of the white bird?” and one of its sub-questions is “Is any bird visible in the scene?”. The models may answer “yes” to the compositional question, but “no” to the sub-question. This paper presents a dialog-like reasoning method for maintaining reasoning consistency in answering a compositional question and its sub-questions. Our method integrates the reasoning processes for the sub-questions into the reasoning process for the compositional question like a dialog task, and uses a consistency constraint to penalize inconsistent answer predictions. In order to enable quantitative evaluation of reasoning consistency, we construct a GQA-Sub dataset based on the well-organized GQA dataset. Experimental results on the GQA dataset and the GQA-Sub dataset demonstrate the effectiveness of our method.
Chenchen Jing, Yunde Jia, Yuwei Wu 0001, Qi Wu 0001
CVPR3
2022 Global-Aware Registration of Less-Overlap RGB-D Scans
abstract
We propose a novel method of registering less-overlap RGB-D scans. Our method learns global information of a scene to construct a panorama, and aligns RGB-D scans to the panorama to perform registration. Different from existing methods that use local feature points to register less-overlap RGB-D scans and mismatch too much, we use global information to guide the registration, thereby allevi-ating the mismatching problem by preserving global consis-tency of alignments. To this end, we build a scene inference network to construct the panorama representing global in-formation. We introduce a reinforcement learning strategy to iteratively align RGB-D scans with the panorama and re-fine the panorama representation, which reduces the noise of global information and preserves global consistency of both geometric and photometric alignments. Experimental results on benchmark datasets including SUNCG, Matterport, and ScanNet show the superiority of our method.
Che Sun, Yunde Jia, Yuwei Wu 0001
CVPR4
2022 Evidential Reasoning for Video Anomaly Detection
abstract
Video anomaly detection aims to discriminate events that deviate from normal patterns in a video. Modeling the decision boundaries of anomalies is challenging, due to the uncertainty in the probability of deviating from normal patterns. In this paper, we propose a deep evidential reasoning method that explicitly learns the uncertainty to model the boundaries. Our method encodes various visual cues as evidences representing potential deviations, assigns beliefs to the predicted probability of deviating from normal patterns based on the evidences, and estimates the uncertainty from the remained beliefs to model the boundaries. To do this, we build a deep evidential reasoning network to encode evidence vectors and estimate uncertainty by learning evidence distributions and deriving beliefs from the distributions. We introduce an unsupervised strategy to train our network by minimizing an energy function of the deep Gaussian mixed model (GMM). Experimental results show that our uncertainty score is beneficial for modeling the boundaries of video anomalies on three benchmark datasets.
Che Sun, Yunde Jia, Yuwei Wu 0001
ACM Multimedia3
2022 Hyperbolic Feature Augmentation via Distribution Estimation and Infinite Sampling on Manifolds
abstract
Learning in hyperbolic spaces has attracted growing attention recently, owing to their capabilities in capturing hierarchical structures of data. However, existing learning algorithms in the hyperbolic space tend to overfit when limited data is given. In this paper, we propose a hyperbolic feature augmentation method that generates diverse and discriminative features in the hyperbolic space to combat overfitting. We employ a wrapped hyperbolic normal distribution to model augmented features, and use a neural ordinary differential equation module that benefits from meta-learning to estimate the distribution. This is to reduce the bias of estimation caused by the scarcity of data. We also derive an upper bound of the augmentation loss, which enables us to train a hyperbolic model by using an infinite number of augmentations. Experiments on few-shot learning and continual learning tasks show that our method significantly improves the performance of hyperbolic algorithms in scarce data regimes.
Zhi Gao 0002, Yuwei Wu 0001, Yunde Jia, Mehrtash Harandi
NeurIPS2
2022 Synthesizing Counterfactual Samples for Overcoming Moment Biases in Temporal Video Grounding
Mingliang Zhai, Chuanhao Li 0001, Chenchen Jing, Yuwei Wu 0001
PRCV (1)4
2022 Stitching images from a conventional camera and a fisheye camera based on nonrigid warping
Yanmei Dong, Mingtao Pei, Yuwei Wu 0001, Yunde Jia
Multim. Tools Appl.3
2022 Infinite-dimensional feature aggregation via a factorized bilinear model
Jindou Dai, Yuwei Wu 0001, Zhi Gao 0002, Yunde Jia
Pattern Recognit.2
2021 Learning a Gradient-free Riemannian Optimizer on Tangent Spaces
Xiaomeng Fan, Zhi Gao 0002, Yuwei Wu 0001, Yunde Jia, Mehrtash Harandi
AAAI3
2021 A Hyperbolic-to-Hyperbolic Graph Convolutional Network
abstract
Hyperbolic graph convolutional networks (GCNs) demonstrate powerful representation ability to model graphs with hierarchical structure. Existing hyperbolic GCNs resort to tangent spaces to realize graph convolution on hyperbolic manifolds, which is inferior because tangent space is only a local approximation of a manifold. In this paper, we propose a hyperbolic-to-hyperbolic graph convolutional network (H2H-GCN) that directly works on hyperbolic manifolds. Specifically, we developed a manifold-preserving graph convolution that consists of a hyperbolic feature transformation and a hyperbolic neighborhood aggregation. The hyperbolic feature transformation works as linear transformation on hyperbolic manifolds. It ensures the transformed node representations still lie on the hyperbolic manifold by imposing the orthogonal constraint on the transformation sub-matrix. The hyperbolic neighborhood aggregation updates each node representation via the Einstein midpoint. The H2H-GCN avoids the distortion caused by tangent space approximations and keeps the global hyperbolic structure. Extensive experiments show that the H2H-GCN achieves substantial improvements on the link prediction, node classification, and graph classification tasks.
Jindou Dai, Yuwei Wu 0001, Zhi Gao 0002, Yunde Jia
CVPR2
2021 A Decomposition Model for Stereo Matching
abstract
In this paper, we present a decomposition model for stereo matching to solve the problem of excessive growth in computational cost (time and memory cost) as the resolution increases. In order to reduce the huge cost of stereo matching at the original resolution, our model only runs dense matching at a very low resolution and uses sparse matching at different higher resolutions to recover the disparity of lost details scale-by-scale. After the decomposition of stereo matching, our model iteratively fuses the sparse and dense disparity maps from adjacent scales with an occlusion-aware mask. A refinement network is also applied to improving the fusion result. Compared with high-performance methods like PSMNet and GANet, our method achieves 10−100× speed increase while obtaining comparable disparity estimation results.
Chengtang Yao, Yunde Jia, Huijun Di, Pengxiang Li 0002, Yuwei Wu 0001
CVPR5
2021 Curvature Generation in Curved Spaces for Few-Shot Learning
abstract
Few-shot learning describes the challenging problem of recognizing samples from unseen classes given very few labeled examples. In many cases, few-shot learning is cast as learning an embedding space that assigns test samples to their corresponding class prototypes. Previous methods assume that data of all few-shot learning tasks comply with a fixed geometrical structure, mostly a Euclidean structure. Questioning this assumption that is clearly difficult to hold in real-world scenarios and incurs distortions to data, we propose to learn a task-aware curved embedding space by making use of the hyperbolic geometry. As a result, task-specific embedding spaces where suitable curvatures are generated to match the characteristics of data are constructed, leading to more generic embedding spaces. We then leverage on intra-class and inter-class context information in the embedding space to generate class prototypes for discriminative classification. We conduct a comprehensive set of experiments on inductive and transductive few-shot learning, demonstrating the benefits of our proposed method over existing embedding methods.
Zhi Gao 0002, Yuwei Wu 0001, Yunde Jia, Mehrtash Harandi
ICCV2
2021 Multi-homography Estimation and Inference Driven by Contour Alignment
Yunde Jia, Huijun Di, Yuwei Wu 0001
ICIG (1)4
2021 Adversarial 3D Convolutional Auto-Encoder for Abnormal Event Detection in Videos
abstract
Abnormal event detection aims to identify the events that deviate from expected normal patterns. Existing methods usually extract normal spatio-temporal patterns of appearance and motion in a separate manner, which ignores low-level correlations between appearance and motion patterns and may fall short of capturing fine-grained spatio-temporal patterns. In this paper, we propose to simultaneously learn appearance and motion to obtain fine-grained spatio-temporal patterns. To this end, we present an adversarial 3D convolutional auto-encoder to learn the normal spatio-temporal patterns and then identify abnormal events by diverging them from the learned normal patterns in videos. The encoder captures the low-level correlations between spatial and temporal dimensions of videos, and generates distinctive features representing visual spatio-temporal information. The decoder reconstrucccts the original video from the encoded features representing by 3D de-convolutions and learns the normal spatio-temporal patterns in an unsupervised manner. We introduce the denoising reconstruction error and adversarial learning strategy to train the 3D convolutional auto-encoder to implicitly learn accurate data distributions that are considered normal patterns, which benefits enhancing the reconstruction ability of the auto-encoder to discriminate abnormal events. Both the theoretical analysis and the extensive experiments on four publicly available datasets demonstrate the effectiveness of our method.
Che Sun, Yunde Jia, Hao Song 0002, Yuwei Wu 0001
IEEE Trans. Multim.4
2020 Revisiting Bilinear Pooling: A Coding Perspective
abstract
Bilinear pooling has achieved state-of-the-art performance on fusing features in various machine learning tasks, owning to its ability to capture complex associations between features. Despite the success, bilinear pooling suffers from redundancy and burstiness issues, mainly due to the rank-one property of the resulting representation. In this paper, we prove that bilinear pooling is indeed a similarity-based coding-pooling formulation. This establishment then enables us to devise a new feature fusion algorithm, the factorized bilinear coding (FBC) method, to overcome the drawbacks of the bilinear pooling. We show that FBC can generate compact and discriminative representations with substantially fewer parameters. Experiments on two challenging tasks, namely image classification and visual question answering, demonstrate that our method surpasses the bilinear pooling technique by a large margin.
Zhi Gao 0002, Yuwei Wu 0001, Xiaoxun Zhang, Jindou Dai, Yunde Jia, Mehrtash Harandi
AAAI2
2020 Overcoming Language Priors in VQA via Decomposed Linguistic Representations
abstract
Most existing Visual Question Answering (VQA) models overly rely on language priors between questions and answers. In this paper, we present a novel method of language attention-based VQA that learns decomposed linguistic representations of questions and utilizes the representations to infer answers for overcoming language priors. We introduce a modular language attention mechanism to parse a question into three phrase representations: type representation, object representation, and concept representation. We use the type representation to identify the question type and the possible answer set (yes/no or specific concepts such as colors or numbers), and the object representation to focus on the relevant region of an image. The concept representation is verified with the attended region to infer the final answer. The proposed method decouples the language-based concept discovery and vision-based concept verification in the process of answer inference to prevent language priors from dominating the answering process. Experiments on the VQA-CP dataset demonstrate the effectiveness of our method.
Chenchen Jing, Yuwei Wu 0001, Xiaoxun Zhang, Yunde Jia, Qi Wu 0001
AAAI2
2020 Learning to Optimize on SPD Manifolds
abstract
Many tasks in computer vision and machine learning are modeled as optimization problems with constraints in the form of Symmetric Positive Definite (SPD) matrices. Solving such optimization problems is challenging due to the non-linearity of the SPD manifold, making optimization with SPD constraints heavily relying on expert knowledge and human involvement. In this paper, we propose a meta-learning method to automatically learn an iterative optimizer on SPD manifolds. Specifically, we introduce a novel recurrent model that takes into account the structure of input gradients and identifies the updating scheme of optimization. We parameterize the optimizer by the recurrent model and utilize Riemannian operations to ensure that our method is faithful to the geometry of SPD manifolds. Compared with existing SPD optimizers, our optimizer effectively exploits the underlying data distribution and learns a better optimization trajectory in a data-driven manner. Extensive experiments on various computer vision tasks including metric nearness, clustering, and similarity learning demonstrate that our optimizer outperforms existing state-of-the-art methods consistently.
Zhi Gao 0002, Yuwei Wu 0001, Yunde Jia, Mehrtash Harandi
CVPR2
2020 Visual-Semantic Graph Matching for Visual Grounding
abstract
Visual Grounding is the task of associating entities in a natural language sentence with objects in an image. In this paper, we formulate visual grounding as a graph matching problem to find node correspondences between a visual scene graph and a language scene graph. These two graphs are heterogeneous, representing structure layouts of the sentence and image, respectively. We learn unified contextual node representations of the two graphs by using a cross-modal graph convolutional network to reduce their discrepancy. The graph matching is thus relaxed as a linear assignment problem because the learned node representations characterize both node information and structure information. A permutation loss and a semantic cycle-consistency loss are further introduced to solve the linear assignment problem with or without ground-truth correspondences. Experimental results on two visual grounding tasks, i.e., referring expression comprehension and phrase localization, demonstrate the effectiveness of our method.
Chenchen Jing, Yuwei Wu 0001, Mingtao Pei, Yao Hu 0002, Yunde Jia, Qi Wu 0001
ACM Multimedia2
2020 Scene-Aware Context Reasoning for Unsupervised Abnormal Event Detection in Videos
abstract
In this paper, we propose a scene-aware context reasoning method that exploits context information from visual features for unsupervised abnormal event detection in videos, which bridges the semantic gap between visual context and the meaning of abnormal events. In particular, we build na spatio-temporal context graph to model visual context information including appearances of objects, spatio-temporal relationships among objects and scene types. The context information is encoded into the nodes and edges of the graph, and their states are iteratively updated by using multiple RNNs with message passing for context reasoning. To infer the spatio-temporal context graph in various scenes, we develop a graph-based deep Gaussian mixture model for scene clustering in an unsupervised manner. We then compute frame-level anomaly scores based on the context information to discriminate abnormal events in various scenes. Evaluations on three challenging datasets, including the UCF-Crime, Avenue, and ShanghaiTech datasets, demonstrate the effectiveness of our method.
Che Sun, Yunde Jia, Yao Hu 0002, Yuwei Wu 0001
ACM Multimedia4
2020 Can You Easily Perceive the Local Environment? A User Interface with One Stitched Live Video for Mobile Robotic Telepresence Systems
abstract
Many existing mobile robotic telepresence systems have equipped with two cameras, one is a forward-facing camera for video communication, and the other is a downward-facing camera for robot navigation. However, the two live videos from these two cameras would cause some confusion which makes it difficult for a remote operator to perceive the local environment. In this paper, we propose to use a user interface with one stitched live video instead of two live videos for mobile robotic telepresence systems. We used a video stitching algorithm to stitch the two live videos into one live video through which a remote operator can well perceive the local environment. We conducted a user study to investigate the difference between one stitched live video and two separate live videos in the user interface. The results show that the user interface with one stitched live video improves task efficiency, the number of errors, and remote operators’ feelings of presence, and enables remote operators to concentrate on the work they are doing.
Yanmei Dong, Yunde Jia, Weichao Shen, Yuwei Wu 0001
Int. J. Hum. Comput. Interact.4
2020 Face Spoofing Detection Using Relativity Representation on Riemannian Manifold
abstract
Face recognition and verification systems are susceptible to spoofing attacks using photographs, videos or masks. Most existing methods focus on spoofing detection in Euclidean space, and ignore the features' manifold structure and interrelationships, thus limiting their capabilities of discrimination and generalization. In this paper, we propose a relativity representation on Riemannian manifold for face spoofing detection. The relativity representation improves generalization capability while ensuring discriminability, at both levels of feature description and classification score. The feature-level relativity representation generalizes information by modeling interrelationships among basic features, and would not depend too much on characteristics of a particular dataset. The score-level relativity representation makes decisions relatively, not absolutely, according to interrelationships (via Riemannian metric) and competitions (via example reweighting) among data samples on Riemannian manifold. The discriminability is ensured by the high-order nature of the feature-level relativity representation as well as Riemannian reweighted discriminative learning of the score-level relativity representation. Moreover, we integrate an attack-sensitive SVM classifier in Euclidean space to improve spoofing detection. Experiments demonstrate the effectiveness of our method on both intra-dataset and cross-dataset testing.
Chengtang Yao, Yunde Jia, Huijun Di, Yuwei Wu 0001
IEEE Trans. Inf. Forensics Secur.4
2020 A Robust Distance Measure for Similarity-Based Classification on the SPD Manifold
abstract
The symmetric positive definite (SPD) matrices, forming a Riemannian manifold, are commonly used as visual representations. The non-Euclidean geometry of the manifold often makes developing learning algorithms (e.g., classifiers) difficult and complicated. The concept of similarity-based learning has been shown to be effective to address various problems on SPD manifolds. This is mainly because the similarity-based algorithms are agnostic to the geometry and purely work based on the notion of similarities/distances. However, existing similarity-based models on SPD manifolds opt for holistic representations, ignoring characteristics of information captured by SPD matrices. To circumvent this limitation, we propose a novel SPD distance measure for the similarity-based algorithm. Specifically, we introduce the concept of point-to-set transformation, which enables us to learn multiple lower dimensional and discriminative SPD manifolds from a higher dimensional one. For lower dimensional SPD manifolds obtained by the point-to-set transformation, we propose a tailored set-to-set distance measure by making use of the family of alpha-beta divergences. We further propose to learn the point-to-set transformation and the set-to-set distance measure jointly, yielding a powerful similarity-based algorithm on SPD manifolds. Our thorough evaluations on several visual recognition tasks (e.g., action classification and face recognition) suggest that our algorithm comfortably outperforms various state-of-the-art algorithms.
Zhi Gao 0002, Yuwei Wu 0001, Mehrtash Harandi, Yunde Jia
IEEE Trans. Neural Networks Learn. Syst.2
2019 3D Shape Reconstruction From Images in the Frequency Domain
abstract
Reconstructing the high-resolution volumetric 3D shape from images is challenging due to the cubic growth of computational cost. In this paper, we propose a Fourier-based method that reconstructs a 3D shape from images in a 2D space by predicting slices in the frequency domain. According to the Fourier slice projection theorem, we introduce a thickness map to bridge the domain gap between images in the spatial domain and slices in the frequency domain. The thickness map is the 2D spatial projection of the 3D shape, which is easily predicted from the input image by a general convolutional neural network. Each slice in the frequency domain is the Fourier transform of the corresponding thickness map. All slices constitute a 3D descriptor and the 3D shape is the inverse Fourier transform of the descriptor. Using slices in the frequency domain, our method can transfer the 3D shape reconstruction from the 3D space into the 2D space, which significantly reduces the computational cost. The experiment results on the ShapeNet dataset demonstrate that our method achieves competitive reconstruction accuracy and computational efficiency compared with the state-of-the-art reconstruction methods.
Weichao Shen, Yunde Jia, Yuwei Wu 0001
CVPR3
2019 Tracker-Level Decision by Deep Reinforcement Learning for Robust Visual Tracking
Wenju Huang, Yuwei Wu 0001, Yunde Jia
ICIG (1)2
2019 Person-following for Telepresence Robots Using Web Cameras
abstract
Many existing mobile robotic telepresence systems have equipped with two web cameras, one is a forward-facing camera (FF camera) for video communication, and the other is a downward-facing camera (DF camera) for robot navigation. In this paper, we present a new framework of autonomous person-following for telepresence robots using the two web cameras. Based on correlation filters tracking methods, we use the FF camera to track the upper body of a person and the DF camera to localize and track the person's feet. We improve the robustness of feet trackers, consisting of a left foot tracker and a right foot tracker, by making full use of the spatial constraints of the human body parts. We conducted experiments on tracking in different environmental situations and real person-following scenario to evaluate the effectiveness of our method.
Xianda Cheng, Yunde Jia, Jingyu Su, Yuwei Wu 0001
IROS4
2019 Temporal Invariant Factor Disentangled Model for Representation Learning
Weichao Shen, Yuwei Wu 0001, Yunde Jia
PRCV (2)2
2019 Unsupervised deep quantization for object instance search
Yuwei Wu 0001, Chenchen Jing, Yunde Jia
Neurocomputing2
2019 Deep convolutional network with locality and sparsity constraints for texture classification
Xingyuan Bu, Yuwei Wu 0001, Zhi Gao 0002, Yunde Jia
Pattern Recognit.2
2019 Learning a robust representation via a deep network on symmetric positive definite manifolds
Zhi Gao 0002, Yuwei Wu 0001, Xingyuan Bu, Junsong Yuan 0001, Yunde Jia
Pattern Recognit.2
2019 Robust Distracter-Resistive Tracker via Learning a Multi-Component Discriminative Dictionary
abstract
Discriminative dictionary learning (DDL) provides an appealing paradigm for appearance modeling in visual tracking. However, most existing DDL-based trackers cannot handle drastic appearance changes, especially for scenarios with background cluster and/or similar object interference. One reason is that they often suffer from the loss of subtle visual information, which is critical to distinguish an object from distracters. In this paper, we explore the use of activations from the convolutional layer of a convolutional neural network to improve the object representation and then propose a robust distracter-resistive tracker via learning a multi-component discriminative dictionary. The proposed method exploits both the intra-class and inter-class visual information to learn shared atoms and the class-specific atoms. By imposing several constraints into the objective function, the learned dictionary is reconstructive, compressive, and discriminative, and thus can better distinguish an object from the background. In addition, our convolutional features have structural information for object localization and balance the discriminative power and semantic information of the object. Tracking is carried out within a Bayesian inference framework where a joint decision measure is used to construct the observation model. To alleviate the drift problem, the reliable tracking results obtained online are accumulated to update the dictionary. Both the qualitative and quantitative results on the CVPR2013 benchmark, the VOT2015 data set, and the SPOT data set demonstrate that our tracker achieves substantially better overall performance against the state-of-the-art approaches.
Weichao Shen, Yuwei Wu 0001, Junsong Yuan 0001, Ling-Yu Duan, Jian Zhang 0002, Yunde Jia
IEEE Trans. Circuits Syst. Video Technol.2
2019 Temporal Action Localization in Untrimmed Videos Using Action Pattern Trees
abstract
In this paper, we present a novel framework of automatically localizing action instances based on action pattern trees (AP-Trees) in a long untrimmed video. For localizing action instances in videos with varied temporal lengths, we first split videos into sequential segments and then use the AP-Trees to produce precise temporal boundaries of action instances. The AP-Trees can exploit the temporal information between segments of videos based on the label vectors of segments, by learning the occurrence frequency and order of segments. In AP-Trees, nodes stand for action class labels of segments and edges represent the temporal relationships between two consecutive segments. Thus, we can discover the occurrence frequencies of segments by searching paths of AP-Trees. In order to obtain accurate labels of video segments, we introduce deep neural networks to annotate the segments by simultaneously leveraging the spatio-temporal information and the high-level semantic feature of segments. In the networks, informative action maps are generated by a global average pooling layer to retain the spatio-temporal information of segments. An overlap loss function is employed to further improve the precision of label vectors of segments by considering the temporal overlap between segments and the ground truth. The experiments on THUMOS2014, MSR ActionII, and MPII Cooking datasets demonstrate the effectiveness of the method.
Hao Song 0002, Xinxiao Wu, Yuwei Wu 0001, Yunde Jia
IEEE Trans. Multim.4
2019 Codebook-Free Compact Descriptor for Scalable Visual Search
abstract
The MPEG compact descriptors for visual search (CDVS) is a standard toward image matching and retrieval. To achieve high retrieval accuracy over a large scale image/video dataset, recent research efforts have demonstrated that employing extremely high-dimensional descriptors such as the Fisher vector (FV) and the vector of locally aggregated descriptors (VLAD) can yield good performance. Since the FV (or VLAD) possesses high discriminability but small visual vocabulary, it has been adopted by CDVS to construct a global compact descriptor. In this paper, we study the development of global compact descriptors in the completed CDVS standard and the emerging compact descriptors for video analysis (CDVA) standard, in which we formulate the FV (or VLAD) compression as a resource-constrained optimization problem. Accordingly, we propose a codebook-free aggregation method via dual selection to generate a global compact visual descriptor, which supports fast and accurate feature matching free of large visual codebooks, fulfilling the low memory requirement of mobile visual search at significantly reduced latency. Specifically, we investigate both sample-specific Gaussian component redundancy and bit dependency within a binary aggregated descriptor to produce compact binary codes. Our technique contributes to the scalable compressed Fisher vector (SCFV) adopted by the CDVS standard. Moreover, the SCFV descriptor is currently serving as the frame-level hand-crafted video feature, which inspires the inheritance of CDVS descriptors for the emerging CDVA standard. Furthermore, we investigate the positive complementary effect of our standard compliant compact descriptor and deep learning based features extracted from convolutional neural networks with significant mean average precision gains. Extensive evaluation over benchmark databases shows the significant merits of the codebook-free binary codes for scalable visual search.
Yuwei Wu 0001, Feng Gao 0014, Jie Lin 0001, Vijay Chandrasekhar 0001, Junsong Yuan 0001, Ling-Yu Duan
IEEE Trans. Multim.1
2018 Deep Stereo Matching With Explicit Cost Aggregation Sub-Architecture
abstract
Deep neural networks have shown excellent performance for stereo matching. Many efforts focus on the feature extraction and similarity measurement of the matching cost computation step while less attention is paid on cost aggregation which is crucial for stereo matching. In this paper, we present a learning-based cost aggregation method for stereo matching by a novel sub-architecture in the end-to-end trainable pipeline. We reformulate the cost aggregation as a learning process of the generation and selection of cost aggregation proposals which indicate the possible cost aggregation results. The cost aggregation sub-architecture is realized by a two-stream network: one for the generation of cost aggregation proposals, the other for the selection of the proposals. The criterion for the selection is determined by the low-level structure information obtained from a light convolutional network. The two-stream network offers a global view guidance for the cost aggregation to rectify the mismatching value stemming from the limited view of the matching cost computation. The comprehensive experiments on challenge datasets such as KITTI and Scene Flow show that our method outperforms the state-of-the-art methods.
Lidong Yu, Yuwei Wu 0001, Yunde Jia
AAAI3
2018 Multi-layer CNN Features Aggregation for Real-time Visual Tracking
abstract
In this paper, we propose a novel convolutional neural network (CNN) based tracking framework, which aggregates multiple CNN features from different layers into a robust representation and realizes real-time tracking. We found that some feature maps have interference for effectively representing objects. Instead of using original features, we build an end-to-end feature aggregation network (FAN) which suppresses the noisy feature maps of CNN layers. The feature significantly benefits to represent objects with both coarse semantic information and fine details. The FAN, as a light-weight network, can run at real-time. The highlighted region of feature maps obtained from the FAN is the tracking result. Our method performs at a real-time speed of 24 fps while maintaining a promising accuracy compared with state-of-the-art methods on existing tracking benchmarks.
Lijia Zhang, Yanmei Dong, Yuwei Wu 0001
ICPR3
2018 Set-to-Set Distance Metric Learning on SPD Manifolds
Zhi Gao 0002, Yuwei Wu 0001, Yunde Jia
PRCV (3)2
2018 Minimizing Reconstruction Bias Hashing via Joint Projection Learning and Quantization
abstract
Hashing, a widely-studied solution to the approximate nearest neighbor (ANN) search, aims to map data points in the high-dimensional Euclidean space to the low-dimensional Hamming space while preserving the similarity between original points. As directly learning binary codes can be NP-hard due to discrete constraints, a two-stage scheme, namely "projection and quantization", has already become a standard paradigm for learning similarity-preserving hash codes. However, most existing hashing methods typically separate these two stages and thus fail to investigate complementary effects of both stages. In this paper, we systematically study the relationship between "projection and quantization", and propose a novel minimal reconstruction bias hashing (MRH) method to learn compact binary codes, in which the projection learning and quantization optimizing are jointly performed. By introducing a lower bound analysis, we design an effective ternary search algorithm to solve the corresponding optimization problem. Furthermore, we conduct some insightful discussions on the proposed MRH approach, including the theoretical proof, and computational complexity. Distinct from previous works, MRH can adaptively adjust the projection dimensionality to balance the information loss between projection and quantization. The proposed framework not only provides a unique perspective to view traditional hashing methods but also evokes some other researches, e.g., guiding the design of the loss functions in deep networks. Extensive experiment results have shown that the proposed MRH significantly outperforms a variety of state-of-the-art methods over eight widely used benchmarks.
Ling-Yu Duan, Yuwei Wu 0001, Zhe Wang 0019, Junsong Yuan 0001, Wen Gao 0001
IEEE Trans. Image Process.2
2018 Group-Sensitive Triplet Embedding for Vehicle Reidentification
abstract
The widespread use of surveillance cameras toward smart and safe cities poses the critical but challenging problem of vehicle reidentification (Re-ID). The state-of-the-art research work performed vehicle Re-ID relying on deep metric learning with a triplet network. However, most existing methods basically ignore the impact of intraclass variance-incorporated embedding on the performance of vehicle reidentification, in which robust fine-grained features for large-scale vehicle Re-ID have not been fully studied. In this paper, we propose a deep metric learning method, group-sensitive-triplet embedding (GS-TRE), to recognize and retrieve vehicles, in which intraclass variance is elegantly modeled by incorporating an intermediate representation “group” between samples and each individual vehicle in the triplet network learning. To capture the intraclass variance attributes of each individual vehicle, we utilize an online grouping method to partition samples within each vehicle ID into a few groups, and build up the triplet samples at multiple granularities across different vehicle IDs as well as different groups within the same vehicle ID to learn fine-grained features. In particular, we construct a large-scale vehicle database “PKU-Vehicle,” consisting of 10 million vehicle images captured by different surveillance cameras in several cities, to evaluate the vehicle Re-ID performance in real-world video surveillance applications. Extensive experiments over benchmark datasets VehicleID, VeRI, and CompCar have shown that the proposed GS-TRE significantly outperforms the state-of-the-art approaches for vehicle Re-ID.
Yihang Lou, Feng Gao 0014, Shiqi Wang 0001, Yuwei Wu 0001, Ling-Yu Duan
IEEE Trans. Multim.5
2017 Deep Manifold Learning of Symmetric Positive Definite Matrices with Application to Face Recognition
abstract
In this paper, we aim to construct a deep neural network which embeds high dimensional symmetric positive definite (SPD) matrices into a more discriminative low dimensional SPD manifold. To this end, we develop two types of basic layers: a 2D fully connected layer which reduces the dimensionality of the SPD matrices, and a symmetrically clean layer which achieves non-linear mapping. Specifically, we extend the classical fully connected layer such that it is suitable for SPD matrices, and we further show that SPD matrices with symmetric pair elements setting zero operations are still symmetric positive definite. Finally, we complete the construction of the deep neural network for SPD manifold learning by stacking the two layers. Experiments on several face datasets demonstrate the effectiveness of the proposed method.
Zhen Dong 0002, Su Jia, Chi Zhang 0063, Mingtao Pei, Yuwei Wu 0001
AAAI5
2017 Efficient Object Instance Search Using Fuzzy Objects Matching
abstract
Recently, global features aggregated from local convolutional features of the convolutional neural network have shown to be much more effective in comparison with hand-crafted features for image retrieval. However, the global feature might not effectively capture the relevance between the query object and reference images in the object instance search task, especially when the query object is relatively small and there exist multiple types of objects in reference images. Moreover, the object instance search requires to localize the object in the reference image, which may not be achieved through global representations. In this paper, we propose a Fuzzy Objects Matching (FOM) framework to effectively and efficiently capture the relevance between the query object and reference images in the dataset. In the proposed FOM scheme, object proposals are utilized to detect the potential regions of the query object in reference images. To achieve high search efficiency, we factorize the feature matrix of all the object proposals from one reference image into the product of a set of fuzzy objects and sparse codes. In addition, we refine the feature of the generated fuzzy objects according to its neighborhood in the feature space to generate more robust representation. The experimental results demonstrate that the proposed FOM framework significantly outperforms the state-of-the-art methods in precision with less memory and computational cost on three public datasets.
Yuwei Wu 0001, Sreyasee Das Bhattacharjee, Junsong Yuan 0001
AAAI2
2017 HOPE: Hierarchical Object Prototype Encoding for Efficient Object Instance Search in Videos
abstract
This paper tackles the problem of efficient and effective object instance search in videos. To effectively capture the relevance between a query and video frames and precisely localize the particular object, we leverage the object proposals to improve the quality of object instance search in videos. However, hundreds of object proposals obtained from each frame could result in unaffordable memory and computational cost. To this end, we present a simple yet effective hierarchical object prototype encoding (HOPE) model to accelerate the object instance search without sacrificing accuracy, which exploits both the spatial and temporal self-similarity property existing in object proposals generated from video frames. We design two types of sphere k-means methods, i.e., spatially-constrained sphere k-means and temporally-constrained sphere k-means to learn frame-level object prototypes and dataset-level object prototypes, respectively. In this way, the object instance search problem is cast to the sparse matrix-vector multiplication problem. Thanks to the sparsity of the codes, both the memory and computational cost are significantly reduced. Experimental results on two video datasets demonstrate that our approach significantly improves the performance of video object instance search over other state-of-the-art fast search schemes.
Yuwei Wu 0001, Junsong Yuan 0001
CVPR2
2017 Compact discriminative object representation via weakly supervised learning for real-time visual tracking
abstract
Object representations are of great importance for robust visual tracking. Although the high‐dimensional representation can effectively encode the input data with more information, exploiting it in a real‐time tracking system would be intractable and infeasible due to the high computational cost and memory requirements. In this study, the authors propose a compact discriminative object representation to achieve both good tracking accuracy and efficiency. An ensemble of weak training sets is generated based on the self‐representative ability of tracking samples, which is applied to learn discriminative functions. Each candidate is represented by the concatenation of project values on all the weak training sets. Tracking is then carried out within a Bayesian inference framework where the classification score of the support vector machine is used to construct the observation model. The evaluations on TB50 benchmark dataset demonstrate that the proposed algorithm is much more computationally efficient than the state‐of‐the‐art methods with comparable accuracy.
Weichao Shen, Yuwei Wu 0001, Yunde Jia
IET Comput. Vis.2
2017 A Hybrid Data Association Framework for Robust Online Multi-Object Tracking
abstract
Global optimization algorithms have shown impressive performance in data-association-based multi-object tracking, but handling online data remains a difficult hurdle to overcome. In this paper, we present a hybrid data association framework with a min-cost multi-commodity network flow for robust online multi-object tracking. We build local target-specific models interleaved with global optimization of the optimal data association over multiple video frames. More specifically, in the min-cost multi-commodity network flow, the target-specific similarities are online learned to enforce the local consistency for reducing the complexity of the global data association. Meanwhile, the global data association taking multiple video frames into account alleviates irrecoverable errors caused by the local data association between adjacent frames. To ensure the efficiency of online tracking, we give an efficient near-optimal solution to the proposed min-cost multi-commodity flow problem, and provide the empirical proof of its sub-optimality. The comprehensive experiments on real data demonstrate the superior tracking performance of our approach in various challenging situations.
Min Yang 0003, Yuwei Wu 0001, Yunde Jia
IEEE Trans. Image Process.2
2016 Learning a Multi-class Discriminative Dictionary with Nonredundancy Constraints for Visual Classification
abstract
Recent studies have demonstrated advantages of sparse representation in providing an appealing paradigm for visual classification tasks. However, how to effectively learn a compact dictionary of superior reconstruction and discrimination power is still a challenging problem. In this paper, we concurrently exploit both the intra-class and the inter-class visual correlations to learn a multi-class discriminative dictionary. The intra-nonredundancy constraint prevents zero entities from appearing in the class-specific bases, thereby making the learned dictionary more stable. The inter-nonredundancy constraint effectively separates the common visual patterns from all the class-specific bases, yielding a more compact dictionary. Combining nonredundancy constraints with the reconstruction error and the classification error to form a unified objective function, our method can learn a superior dictionary and an optimal linear classifier simultaneously. Extensive experimental results demonstrate that the proposed algorithm achieves notable improvement over the state-of-the-art methods in image classification and visual tracking tasks.
Yuwei Wu 0001, Junsong Yuan 0001, Yap-Peng Tan
ACM Multimedia2
2016 A Compact Binary Aggregated Descriptor via Dual Selection for Visual Search
abstract
To achieve high retrieval accuracy over a large scale image/video dataset, recent research efforts have demonstrated that employing extremely high-dimensional descriptors such as the Fisher Vector (FV) and the Vector of Locally Aggregated Descriptors (VLAD) can yield good performance. To enable fast search, the FV (or VLAD) is usually compressed by product quantization (PQ) or hashing. However, compressing high-dimensional descriptors via PQ or hashing may become intractable and infeasible due to both the storage and computation requirements for the linear/nonlinear projection of PQ or hashing methods. We develop a novel compact aggregated descriptor via dual selection for visual search. We utilize both sample-specific Gaussian component redundancy and bit dependency within a binary aggregated descriptor to produce its compact binary codes. The proposed method can effectively reduce the codesize of the raw aggregated descriptors, without degrading the search accuracy or introducing additional memory footprint. We demonstrate the significant advantages of the proposed binary codes in solving the approximate nearest neighbor (ANN) visual search problem. Experimental results on extensive datasets show that our method outperforms the state-of-the-art methods.
Yuwei Wu 0001, Zhe Wang 0019, Junsong Yuan 0001, Ling-Yu Duan
ACM Multimedia1
2016 Nonnegative correlation coding for image classification
Zhen Dong 0002, Wei Liang 0008, Yuwei Wu 0001, Mingtao Pei, Yunde Jia
Sci. China Inf. Sci.3
2016 Discriminative Action States Discovery for Online Action Recognition
abstract
In this paper, we provide an approach for online human action recognition, where the videos are represented by frame-level descriptors. To address the large intraclass variations of frame-level descriptors, we propose an action states discovery method to discover the different distributions of frame-level descriptors while training a classifier. A positive sample set is treated as multiple clusters called action states. The action states model can be effectively learned by clustering the positive samples and optimizing the decision boundary of each state simultaneously. Experimental results show that our method not only outperforms the state-of-the-art methods, but also can predict the video by an on-going process with a real-time speed.
Junsong Yuan 0001, Yuwei Wu 0001
IEEE Signal Process. Lett.3
2016 Online Discriminative Tracking With Active Example Selection
abstract
Most existing discriminative tracking algorithms use a sampling-and-labeling strategy to collect examples and treat the training example collection as a task that is independent of classifier learning. However, the examples collected directly by sampling are neither necessarily informative nor intended to be useful for classifier learning. Updating the classifier with these examples might introduce ambiguity to the tracker. In this paper, we present a novel online discriminative tracking framework that explicitly couples the objectives of example collection and classifier learning. Our method uses Laplacian regularized least squares (LapRLS) to learn a robust classifier that can sufficiently exploit unlabeled data and preserve the local geometrical structure of the feature space. To ensure the high classification confidence of the classifier, we propose an active example selection approach to automatically select the most informative examples for LapRLS. Part of the selected examples that satisfy strict constraints are labeled to enhance the adaptivity of our tracker, which actually provides robust supervisory information to guide semisupervised learning. With active example selection, we are able to avoid the ambiguity introduced by an independent example collection strategy and to alleviate the drift problem caused by misaligned examples. Comparison with the state-of-the-art trackers on the comprehensive benchmark demonstrates that our tracking algorithm is more effective and accurate.
Min Yang 0003, Yuwei Wu 0001, Mingtao Pei, Bo Ma 0001, Yunde Jia
IEEE Trans. Circuits Syst. Video Technol.2
2015 Discriminative Neighborhood Preserving Dictionary Learning for Image Classification
Shiye Zhang, Zhen Dong 0002, Yuwei Wu 0001, Mingtao Pei
ICIG (2)3
2015 Learning online structural appearance model for robust object tracking
Min Yang 0003, Mingtao Pei, Yuwei Wu 0001, Yunde Jia
Sci. China Inf. Sci.3
2015 Online visual tracking by integrating spatio-temporal cues
abstract
The performance of online visual trackers has improved significantly, but designing an effective appearance‐adaptive model is still a challenging task because of the accumulation of errors during the model updating with newly obtained results, which will cause tracker drift. In this study, the authors propose a novel online tracking algorithm by integrating spatio‐temporal cues to alleviate the drift problem. The authors' goal is to develop a more robust way of updating an adaptive appearance model. The model consists of multiple modules called temporal cues, and these modules are updated in an alternate way which can keep both the historical and current information of the tracked object to handle drastic appearance change. Each module is represented by several fragments called spatial cues. In order to incorporate all the spatial and temporal cues, the authors develop an efficient cue quality evaluation criterion that combines appearance and motion information. Then the tracking results are obtained by a two‐stage dynamic integration mechanism. Both qualitative and quantitative evaluations on challenging video sequences demonstrate that the proposed algorithm performs more favourably against the state‐of‐the‐art methods.
Yang He 0004, Mingtao Pei, Min Yang 0003, Yuwei Wu 0001, Yunde Jia
IET Comput. Vis.4
2015 Differential tracking with a kernel-based region covariance descriptor
Yuwei Wu 0001, Bo Ma 0001, Yunde Jia
Pattern Anal. Appl.1
2015 Manifold Kernel Sparse Representation of Symmetric Positive-Definite Matrices and Its Applications
abstract
The symmetric positive-definite (SPD) matrix, as a connected Riemannian manifold, has become increasingly popular for encoding image information. Most existing sparse models are still primarily developed in the Euclidean space. They do not consider the non-linear geometrical structure of the data space, and thus are not directly applicable to the Riemannian manifold. In this paper, we propose a novel sparse representation method of SPD matrices in the data-dependent manifold kernel space. The graph Laplacian is incorporated into the kernel space to better reflect the underlying geometry of SPD matrices. Under the proposed framework, we design two different positive definite kernel functions that can be readily transformed to the corresponding manifold kernels. The sparse representation obtained has more discriminating power. Extensive experimental results demonstrate good performance of manifold kernel sparse codes in image classification, face recognition, and visual tracking.
Yuwei Wu 0001, Yunde Jia, Peihua Li, Jian Zhang 0002, Junsong Yuan 0001
IEEE Trans. Image Process.1
2015 Robust Discriminative Tracking via Landmark-Based Label Propagation
abstract
The appearance of an object could be continuously changing during tracking, thereby being not independent identically distributed. A good discriminative tracker often needs a large number of training samples to fit the underlying data distribution, which is impractical for visual tracking. In this paper, we present a new discriminative tracker via landmark-based label propagation (LLP) that is nonparametric and makes no specific assumption about the sample distribution. With an undirected graph representation of samples, the LLP locally approximates the soft label of each sample by a linear combination of labels on its nearby landmarks. It is able to effectively propagate a limited amount of initial labels to a large amount of unlabeled samples. To this end, we introduce a local landmarks approximation method to compute the cross-similarity matrix between the whole data and landmarks. Moreover, a soft label prediction function incorporating the graph Laplacian regularizer is used to diffuse the known labels to all the unlabeled vertices in the graph, which explicitly considers the local geometrical structure of all samples. Tracking is then carried out within a Bayesian inference framework, where the soft label prediction value is used to construct the observation model. Both qualitative and quantitative evaluations on the benchmark data set containing 51 challenging image sequences demonstrate that the proposed algorithm outperforms the state-of-the-art methods.
Yuwei Wu 0001, Mingtao Pei, Min Yang 0003, Junsong Yuan 0001, Yunde Jia
IEEE Trans. Image Process.1
2015 Vehicle Type Classification Using a Semisupervised Convolutional Neural Network
abstract
In this paper, we propose a vehicle type classification method using a semisupervised convolutional neural network from vehicle frontal-view images. In order to capture rich and discriminative information of vehicles, we introduce sparse Laplacian filter learning to obtain the filters of the network with large amounts of unlabeled data. Serving as the output layer of the network, the softmax classifier is trained by multitask learning with small amounts of labeled data. For a given vehicle image, the network can provide the probability of each type to which the vehicle belongs. Unlike traditional methods by using handcrafted visual features, our method is able to automatically learn good features for the classification task. The learned features are discriminative enough to work well in complex scenes. We build the challenging BIT-Vehicle dataset, including 9850 high-resolution vehicle frontal-view images. Experimental results on our own dataset and a public dataset demonstrate the effectiveness of the proposed method.
Zhen Dong 0002, Yuwei Wu 0001, Mingtao Pei, Yunde Jia
IEEE Trans. Intell. Transp. Syst.2
2014 Landmark-Based Inductive Model for Robust Discriminative Tracking
Yuwei Wu 0001, Mingtao Pei, Min Yang 0003, Yang He 0004, Yunde Jia
ACCV (5)1
2014 Coupling Semi-supervised Learning and Example Selection for Online Object Tracking
Min Yang 0003, Yuwei Wu 0001, Mingtao Pei, Bo Ma 0001, Yunde Jia
ACCV (4)2
2014 Learning distance metric for object contour tracking
Yuwei Wu 0001, Bo Ma 0001
Pattern Anal. Appl.1
2014 Metric Learning Based Structural Appearance Model for Robust Visual Tracking
abstract
Appearance modeling is a key issue for the success of a visual tracker. Sparse representation based appearance modeling has received an increasing amount of interest in recent years. However, most of existing work utilizes reconstruction errors to compute the observation likelihood under the generative framework, which may give poor performance, especially for significant appearance variations. In this paper, we advocate an approach to visual tracking that seeks an appropriate metric in the feature space of sparse codes and propose a metric learning based structural appearance model for more accurate matching of different appearances. This structural representation is acquired by performing multiscale max pooling on the weighted local sparse codes of image patches. An online multiple instance metric learning algorithm is proposed that learns a discriminative and adaptive metric, thereby better distinguishing the visual object of interest from the background. The multiple instance setting is able to alleviate the drift problem potentially caused by misaligned training examples. Tracking is then carried out within a Bayesian inference framework, in which the learned metric and the structure object representation are used to construct the observation model. Comprehensive experiments on challenging image sequences demonstrate qualitatively and quantitatively that the proposed algorithm outperforms the state-of-the-art methods.
Yuwei Wu 0001, Bo Ma 0001, Min Yang 0003, Jian Zhang 0002, Yunde Jia
IEEE Trans. Circuits Syst. Video Technol.1
2013 Segmentation of the left ventricle in cardiac cine MRI using a shape-constrained snake model
Yuwei Wu 0001, Yuanquan Wang 0001, Yunde Jia
Comput. Vis. Image Underst.1
2013 Adaptive diffusion flow active contours for image segmentation
Yuwei Wu 0001, Yuanquan Wang 0001, Yunde Jia
Comput. Vis. Image Underst.1
2012 A variational method for contour tracking via covariance matching
Yuwei Wu 0001, Bo Ma 0001
Sci. China Inf. Sci.1
2011 Covariance Matching for PDE-based Contour Tracking
abstract
This paper presents a novel formulation for object tracking. We model the second-order statistics of image regions and perform covariance matching under the variational level set framework. Specifically, covariance matrix is adopted as a visual object representation for partial differential equation (PDE) based contour tracking. Log-Euclidean calculus is used as a covariance distance metric instead of Euclidean distance which is unsuitable for measuring the similarities between covariance matrices, because the matrices typically lie on a non-Euclidean manifold. A novel image energy functional is formulated by minimizing the distance metrics between the candidate object region and a given template, and maximizing the ones between the background region and the template. The corresponding gradient flow is then derived according to a variational approach, enabling PDE-based visual tracking. Experiments on synthetic and real video sequences prove the validity of the proposed method.
Bo Ma 0001, Yuwei Wu 0001
ICIG2
2010 Adaptive Diffusion Flow for Parametric Active Contours
abstract
This paper proposes a novel external force for active contours, called adaptive diffusion flow (ADF). We reconsider the generative mechanism of gradient vector flow (GVF) diffusion process from the perspective of image restoration, and exploit a harmonic hyper surface minimal function to substitute smoothness energy term of GVF for alleviating the possible leakage problem. Meanwhile, a ∞- laplacian functional is incorporated in the ADF framework to ensure that the vector flow diffuses mainly along normal direction in homogenous regions of an image. Experiments on synthetic and real images demonstrate the good properties of the ADF snake, including noise robustness, weak edge preserving, and concavity convergence.
Yuwei Wu 0001, Yunde Jia, Yuanquan Wang 0001
ICPR1