VLDB 2026 Research / reviewers in the wild / expert
Yunde Jia
dblp:71/2334
· DBLP profile ↗
230ranked-venue papers
4as first author
52since 2021 · last 2026
0000-0003-1900-8945ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 147 · 2 first-author · 39 since 2021Graphics, computer vision, multimedia, augmented reality and games · 139 · 2 first-author · 36 since 2021Applied, interdisciplinary, general and emerging computing · 10Human-computer interaction and ubiquitous computing · 9 · 1 since 2021Systems, architecture and hardware · 3Computer networks · 3 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LongSplat: Online Generalizable 3D Gaussian Splatting from Long Sequence Imagesabstract3D Gaussian Splatting (3DGS) achieves high-fidelity novel view synthesis, but its application in online long-sequence scenarios is still restricted. Existing methods either rely on slow per-scene optimization or lack efficient frame-wise 3DGS updates, making them unsuitable for online long-sequence videos. In this paper, we propose LongSplat, an online real-time 3D Gaussian reconstruction framework designed for long-sequence image input. The core idea of LongSplat is to maintain a global 3DGS set and design a streaming 3DGS update mechanism that selectively compressing redundant historical Gaussians and introducing new Gaussians by comparing the current observations with the historical Gaussian. To achieve this goal, we design a Gaussian-Image Representation (GIR), which encodes 3D Gaussian parameters into a structured, image-like 2D format. GIR simultaneously enables identity-aware redundancy compression as well as the fusion of current view and historical Gaussians, which are used for online reconstruction and adapt the model to long sequences without overwhelming memory or computational costs. Extensive experiments demonstrate that LongSplat achieves state-of-the-art efficiency-quality trade-offs in real-time novel view synthesis, delivering real-time reconstruction while reducing Gaussian counts by 44% compared to per-pixel prediction paradigms. Guichen Huang, Ruoyu Wang 0014, Xiangjun Gao, Che Sun, Yuwei Wu 0001, Shenghua Gao, Yunde Jia |
AAAI | 7 |
| 2026 | Composition-Incremental Learning for Compositional GeneralizationabstractCompositional generalization has achieved substantial progress in computer vision on pre-collected training data. Nonetheless, real-world data continually emerges, with possible compositions being nearly infinite, long-tailed, and not entirely visible. Thus, an ideal model is supposed to gradually improve the capability of compositional generalization in an incremental manner. In this paper, we explore Composition-Incremental Learning for Compositional Generalization (CompIL) in the context of the compositional zero-shot learning (CZSL) task, where models need to continually learn new compositions, intending to improve their compositional generalization capability progressively. To quantitatively evaluate CompIL, we develop a benchmark construction pipeline leveraging existing datasets, yielding MIT-States-CompIL and C-GQA-CompIL. Furthermore, we propose a pseudo-replay framework utilizing a visual synthesizer to synthesize visual representations of learned compositions and a linguistic primitive distillation mechanism to maintain aligned primitive representations across the learning process. Extensive experiments demonstrate the effectiveness of the proposed framework. Zhen Li 0026, Yuwei Wu 0001, Chenchen Jing, Che Sun, Chuanhao Li 0001, Yunde Jia |
AAAI | 6 |
| 2026 | Riemannian Implicit Differentiation via a Fixed-Point Equation for Riemannian Bilevel OptimizationabstractVarious Riemannian optimization tasks, such as Riemannian metaoptimization (RMO) and Riemannian metalearning, can be formulated as Riemannian bilevel optimization problems (i.e., the inner-level and outer-level optimization). Implicit differentiation has shown effectiveness in solving RMO, which decouples the computation of outer gradients from the inner-level process, avoiding huge computational burdens. However, extending implicit differentiation to other Riemannian bilevel optimization tasks is nontrivial because it requires much expert involvement for case-by-case derivations. In this article, we propose a Riemannian implicit differentiation method that provides a unified expression for outer gradients, leading to flexible application to other tasks with less expert involvement. Specifically, we formulate the inner-level optimization as a root-finding process of a fixed-point equation, through which the inner-level optimization among different tasks is formulated in a unified way. By differentiating the fixed-point equation, we derive a unified expression for outer gradients, circumventing the case-by-case derivations for different tasks. Then, we present convergence analysis and approximation error analysis, which guarantee the effectiveness of our method in various Riemannian optimization tasks. We further conduct experiments on multiple Riemannian optimization tasks, and the experimental results confirm the effectiveness. Xiaomeng Fan, Yuwei Wu 0001, Zhi Gao 0002, Zhipeng Lu 0003, Mehrtash Harandi, Yunde Jia |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2025 | Consistency of Compositional Generalization Across Multiple LevelsabstractCompositional generalization is the capability of a model to understand novel compositions composed of seen concepts. There are multiple levels of novel compositions including phrase-phrase level, phrase-word level, and word-word level. Existing methods achieve promising compositional generalization, but the consistency of compositional generalization across multiple levels of novel compositions remains unexplored. The consistency refers to that a model should generalize to a phrase-phrase level novel composition, and phrase-word/word-word level novel compositions that can be derived from it simultaneously. In this paper, we propose a meta-learning based framework, for achieving consistent compositional generalization across multiple levels. The basic idea is to progressively learn compositions from simple to complex for consistency. Specifically, we divide the original training set into multiple validation sets based on compositional complexity, and introduce multiple meta-weight-nets to generate sample weights for samples in different validation sets. To fit the validation sets in order of increasing compositional complexity, we optimize the parameters of each meta-weight-net independently and sequentially in a multilevel optimization manner. We build a GQA-CCG dataset to quantitatively evaluate the consistency. Experimental results on visual question answering and temporal video grounding, demonstrate the effectiveness of the proposed framework. Chuanhao Li 0001, Zhen Li 0026, Chenchen Jing, Xiaomeng Fan, Wenbo Ye, Yuwei Wu 0001, Yunde Jia |
AAAI | 7 |
| 2025 | World Knowledge-Enhanced Reasoning Using Instruction-Guided Interactor in Autonomous DrivingabstractThe Multi-modal Large Language Models (MLLMs) with extensive world knowledge have revitalized autonomous driving, particularly in reasoning tasks within perceivable regions. However, when faced with perception-limited areas (dynamic or static occlusion regions), MLLMs struggle to effectively integrate perception ability with world knowledge for reasoning. These perception-limited regions can conceal crucial safety information, especially for vulnerable road users. In this paper, we propose a framework, which aims to improve autonomous driving performance under perception-limited conditions by enhancing the integration of perception capabilities and world knowledge. Specifically, we propose a plug-and-play instruction-guided interaction module that bridges modality gaps and significantly reduces the input sequence length, allowing it to adapt effectively to multi-view video inputs. Furthermore, to better integrate world knowledge with driving-related tasks, we have collected and refined a large-scale multi-modal dataset that includes 2 million natural language QA pairs, 1.7 million grounding task data. To evaluate the model’s utilization of world knowledge, we introduce an object-level risk assessment dataset comprising 200K QA pairs, where the questions necessitate multi-step reasoning leveraging world knowledge for resolution. Extensive experiments validate the effectiveness of our proposed method. Mingliang Zhai, Zengyuan Guo, Ningrui Yang, Xiameng Qin, Sanyuan Zhao, Junyu Han, Ji Tao, Yuwei Wu 0001, Yunde Jia |
AAAI | 10 |
| 2025 | Diving into the Fusion of Monocular Priors for Generalized Stereo MatchingabstractThe matching formulation makes it naturally hard for the stereo matching to handle ill-posed regions like occlusions and non-Lambertian surfaces. Fusing monocular priors has been proven helpful for ill-posed matching, but the biased monocular prior learned from small stereo datasets constrains the generalization. Recently, stereo matching has progressed by leveraging the unbiased monocular prior from the vision foundation model (VFM) to improve the generalization in ill-posed regions. We dive into the fusion process and observe three main problems limiting the fusion of the VFM monocular prior. The first problem is the misalignment between affine-invariant relative monocular depth and absolute depth of disparity. Besides, when we use the monocular feature in an iterative update structure, the over-confidence in the disparity update leads to local optima results. A direct fusion of a monocular depth map could alleviate the local optima problem, but noisy disparity results computed at the first several iterations will misguide the fusion. In this paper, we propose a binary local ordering map to guide the fusion, which converts the depth map into a binary relative format, unifying the relative and absolute depth representation. The computed local ordering map is also used to re-weight the initial disparity update, resolving the local optima and noisy problem. In addition, we formulate the final direct fusion of monocular depth to the disparity as a registration problem, where a pixel-wise linear regression module can globally and adaptively align them. Our method fully exploits the monocular prior to support stereo matching results effectively and efficiently. We significantly improve the performance from the experiments when generalizing from SceneFlow to Middlebury and Booster datasets while barely reducing the efficiency. Chengtang Yao, Lidong Yu, Zhidan Liu 0005, Jiaxi Zeng, Yuwei Wu 0001, Yunde Jia |
ICCV | 6 |
| 2025 | Multi-modal Agent Tuning: Building a VLM-Driven Agent for Efficient Tool UsageabstractThe advancement of large language models (LLMs) prompts the development of multi-modal agents, which are used as a controller to call external tools, providing a feasible way to solve practical tasks. In this paper, we propose a multi-modal agent tuning method that automatically generates multi-modal tool-usage data and tunes a vision-language model (VLM) as the controller for powerful tool-usage reasoning. To preserve the data quality, we prompt the GPT-4o mini model to generate queries, files, and trajectories, followed by query-file and trajectory verifiers. Based on the data synthesis pipeline, we collect the MM-Traj dataset that contains 20K tasks with trajectories of tool usage. Then, we develop the T3-Agent via Trajectory Tuning on VLMs for Tool usage using MM-Traj. Evaluations on the GTA and GAIA benchmarks show that the T3-Agent consistently achieves improvements on two popular VLMs: MiniCPM-V-8.5B and Qwen2-VL-7B, which outperforms untrained VLMs by 20%, showing the effectiveness of the proposed data synthesis pipeline, leading to high-quality data for tool-usage capabilities. Zhi Gao 0002, Bofei Zhang, Pengxiang Li 0002, Xiaojian Ma 0001, Yuwei Wu 0001, Yunde Jia, Song-Chun Zhu, Qing Li 0003 |
ICLR | 8 |
| 2025 | Multi-Sourced Compositional Generalization in Visual Question AnsweringabstractCompositional generalization is the ability of generalizing novel compositions from seen primitives, and has received much attention in vision-and-language (V&L) recently. Due to the multi-modal nature of V&L tasks, the primitives composing compositions source from different modalities, resulting in multi-sourced novel compositions. However, the generalization ability over multi-sourced novel compositions, i.e., multi-sourced compositional generalization (MSCG) remains unexplored. In this paper, we explore MSCG in the context of visual question answering (VQA), and propose a retrieval-augmented training framework to enhance the MSCG ability of VQA models by learning unified representations for primitives from different modalities. Specifically, semantically equivalent primitives are retrieved for each primitive in the training samples, and the retrieved features are aggregated with the original primitive to refine the model. This process helps the model learn consistent representations for the same semantic primitives across different modalities. To evaluate the MSCG ability of VQA models, we construct a new GQA-MSCG dataset based on the GQA dataset, in which samples include three types of novel compositions composed of primitives from different modalities. The GQA-MSCG dataset is available at https://github.com/NeverMoreLCH/MSCG. Chuanhao Li 0001, Wenbo Ye, Zhen Li 0026, Yuwei Wu 0001, Yunde Jia |
IJCAI | 5 |
| 2025 | From Notation to Gesture: Virtual Conductor Gesture Generation in VR Via Structured Score SemanticsabstractConductor avatar plays a dual role in immersive Virtual Reality (VR) interactive systems by interpreting musical scores and guiding orchestral performance. Rule-based score-driven methods ensure precise synchronization with predefined conducting templates or videos, but are constrained by pre-authored data. Audio-driven frameworks offer greater adaptability through real-time gesture generation but often fail to capture the symbolic semantics of musical scores. To overcome these limitations, we propose a novel score-driven gesture generation framework that translates symbolic musical representations into plausible conducting gestures. Our approach adopts a two-stage architecture, combining a comparative learning stage for pre-training a score encoder with a generative learning stage for gesture synthesis. The score encoder explicitly models musical features such as tempo, chord, intensity, and cycle semantics, directly informing gesture generation. To support this research, we introduce Multimodal Symphonic Conducting Dataset (MSCD), the first synchronized dataset comprising conducting gestures, performance audio, and editable symbolic scores, effectively bridging the gap between musical semantics and gesture synthesis. Qualitative and quantitative analyses are provided to demonstrate the effectiveness of our approach, while a user study is designed to identify the strengths and limitations of the current work. Haozhe Ma, Yuxin Shen, Yunde Jia |
ISMAR | 4 |
| 2025 | Beyond the Seen: Bounded Distribution Estimation for Open-Vocabulary LearningabstractOpen-vocabulary learning requires modeling the data distribution in open environments, which consists of both seen-class and unseen-class data.
Existing methods estimate the distribution in open environments using seen-class data, where the absence of unseen classes makes the estimation error inherently unidentifiable.
Intuitively, learning beyond the seen classes is crucial for distribution estimation to bound the estimation error.
We theoretically demonstrate that the distribution can be effectively estimated by generating unseen-class data, through which the estimation error is upper-bounded.
Building on this theoretical insight, we propose a novel open-vocabulary learning method, which generates unseen-class data for estimating the distribution in open environments.
The method consists of a class-domain-wise data generation pipeline and a distribution alignment algorithm.
The data generation pipeline generates unseen-class data under the guidance of a hierarchical semantic tree and domain information inferred from the seen-class data, facilitating accurate distribution estimation.
With the generated data, the distribution alignment algorithm estimates and maximizes the posterior probability to enhance generalization in open-vocabulary learning.
Extensive experiments on 11 datasets demonstrate that our method outperforms baseline approaches by up to 14%, highlighting its effectiveness and superiority. Xiaomeng Fan, Yuchuan Mao, Zhi Gao 0002, Yuwei Wu 0001, Jin Chen 0009, Yunde Jia |
NeurIPS | 6 |
| 2025 | Iterative Tool Usage Exploration for Multimodal Agents via Step-wise Preference TuningabstractMultimodal agents, which integrate a controller (e.g., a vision language model) with external tools, have demonstrated remarkable capabilities in tackling complex multimodal tasks.
Existing approaches for training these agents, both supervised fine-tuning and reinforcement learning, depend on extensive human-annotated task-answer pairs and tool trajectories.
However, for complex multimodal tasks, such annotations are prohibitively expensive or impractical to obtain.
In this paper, we propose an iterative tool usage exploration method for multimodal agents without any pre-collected data, namely SPORT, via step-wise preference optimization to refine the trajectories of tool usage. Our method enables multimodal agents to autonomously discover effective tool usage strategies through self-exploration and optimization, eliminating the bottleneck of human annotation.
SPORT has four iterative components: task synthesis, step sampling, step verification, and preference tuning.
We first synthesize multimodal tasks using language models.
Then, we introduce a novel trajectory exploration scheme, where step sampling and step verification are executed alternately to solve synthesized tasks.
In step sampling, the agent tries different tools and obtains corresponding results.
In step verification, we employ a verifier to provide AI feedback to construct step-wise preference data.
The data is subsequently used to update the controller for tool usage through preference tuning, producing a SPORT agent.
By interacting with real environments, the SPORT agent gradually evolves into a more refined and capable system.
Evaluation in the GTA and GAIA benchmarks shows that the SPORT agent achieves 6.41% and 3.64% improvements, underscoring the generalization and effectiveness introduced by our method. Pengxiang Li 0002, Zhi Gao 0002, Bofei Zhang, Yapeng Mi, Xiaojian Ma 0001, Chenrui Shi, Yuwei Wu 0001, Yunde Jia, Song-Chun Zhu, Qing Li 0003 |
NeurIPS | 9 |
| 2025 | Sekai: A Video Dataset towards World ExplorationabstractVideo generation techniques have made remarkable progress, promising to be the foundation of interactive world exploration.However, existing video generation datasets are not well-suited for world exploration training as they suffer from some limitations: limited locations, short duration, static scenes, and a lack of annotations about exploration and the world.In this paper, we introduce Sekai (meaning "world" in Japanese), a high-quality first-person view worldwide video dataset with rich annotations for world exploration. It consists of over 5,000 hours of walking or drone view (FPV and UVA) videos from over 100 countries and regions across 750 cities. We develop an efficient and effective toolbox to collect, pre-process and annotate videos with location, scene, weather, crowd density, captions, and camera trajectories.Comprehensive analyses and experiments demonstrate the dataset’s scale, diversity, annotation quality, and effectiveness for training video generation models.We believe Sekai will benefit the area of video generation and world exploration, and motivate valuable applications. Zhen Li 0026, Chuanhao Li 0001, Xiaofeng Mao, Shaoheng Lin, Ming Li 0010, Shitian Zhao, Zhaopan Xu, Xinyue Li 0001, Yukang Feng, Zizhen Li, Fanrui Zhang, Jiaxin Ai, Yuwei Wu 0001, Tong He 0001, Yunde Jia, Kaipeng Zhang |
NeurIPS | 17 |
| 2025 | 3D Visual Illusion Depth Estimationabstract3D visual illusion is a perceptual phenomenon where a two-dimensional plane is manipulated to simulate three-dimensional spatial relationships, making a flat artwork or object look three-dimensional in the human visual system. In this paper, we reveal that the machine visual system is also seriously fooled by 3D visual illusions, including monocular and binocular depth estimation. In order to explore and analyze the impact of 3D visual illusion on depth estimation, we collect a large dataset containing almost 3k scenes and 200k images to train and evaluate SOTA monocular and binocular depth estimation methods. We also propose a 3D visual illusion depth estimation framework that uses common sense from the vision language model to adaptively fuse depth from binocular disparity and monocular depth. Experiments show that SOTA monocular, binocular, and multi-view depth estimation approaches are all fooled by various 3D visual illusions, while our method achieves SOTA performance. Chengtang Yao, Zhidan Liu 0005, Jiaxi Zeng, Lidong Yu, Yuwei Wu 0001, Yunde Jia |
NeurIPS | 6 |
| 2025 | Large-scale Riemannian meta-optimization via subspace adaptation
Peilin Yu 0001, Yuwei Wu 0001, Zhi Gao 0002, Xiaomeng Fan, Yunde Jia |
Comput. Vis. Image Underst. | 5 |
| 2025 | Curvature Learning for Generalization of Hyperbolic Neural Networks
Xiaomeng Fan, Yuwei Wu 0001, Zhi Gao 0002, Mehrtash Harandi, Yunde Jia |
Int. J. Comput. Vis. | 5 |
| 2025 | A method of embedding a high-resolution image into a large field-of-view image
Yanmei Dong, Mingtao Pei, Yunde Jia |
Multim. Tools Appl. | 3 |
| 2025 | Inter-Scale Similarity Guided Cost Aggregation for Stereo MatchingabstractStereo matching aims to estimate 3D geometry by computing disparity from a rectified image pair. Most deep learning based stereo matching methods aggregate multi-scale cost volumes computed by downsampling and achieve good performance. However, their effectiveness in fine-grained areas is limited by significant detail loss during downsampling and the use of fixed weights in upsampling. In this paper, we propose an inter-scale similarity-guided cost aggregation method that dynamically upsamples the cost volumes according to the content of images for stereo matching. The method consists of two modules: inter-scale similarity measurement and stereo-content-aware cost aggregation. Specifically, we use inter-scale similarity measurement to generate similarity guidance from feature maps in adjacent scales. The guidance, generated from both reference and target images, is then used to aggregate the cost volumes from low-resolution to high-resolution via stereo-content-aware cost aggregation. We further split the 3D aggregation into 1D disparity and 2D spatial aggregation to reduce the computational cost. Experimental results on various benchmarks (e.g., SceneFlow, KITTI, Middlebury and ETH3D-two-view) show that our method achieves consistent performance gain on multiple models (e.g., PSM-Net, HSM-Net, CF-Net, FastAcv, and FactAcvPlus). The code can be found athttps://github.com/Pengxiang-Li/issga-stereo. Pengxiang Li 0002, Chengtang Yao, Yunde Jia, Yuwei Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Residual Hyperbolic Graph Convolution NetworksabstractHyperbolic graph convolutional networks (HGCNs) have demonstrated representational capabilities of modeling hierarchical-structured graphs. However, as in general GCNs, over-smoothing may occur as the number of model layers increases, limiting the representation capabilities of most current HGCN models. In this paper, we propose residual hyperbolic graph convolutional networks (R-HGCNs) to address the over-smoothing problem. We introduce a hyperbolic residual connection function to overcome the over-smoothing problem, and also theoretically prove the effectiveness of the hyperbolic residual function. Moreover, we use product manifolds and HyperDrop to facilitate the R-HGCNs. The distinctive features of the R-HGCNs are as follows: (1) The hyperbolic residual connection preserves the initial node information in each layer and adds a hyperbolic identity mapping to prevent node features from being indistinguishable. (2) Product manifolds in R-HGCNs have been set up with different origin points in different components to facilitate the extraction of feature information from a wider range of perspectives, which enhances the representing capability of R-HGCNs. (3) HyperDrop adds multiplicative Gaussian noise into hyperbolic representations, such that perturbations can be added to alleviate the over-fitting problem without deconstructing the hyperbolic geometry. Experiment results demonstrate the effectiveness of R-HGCNs under various graph convolution layers and different structures of product manifolds. Yangkai Xue, Jindou Dai, Zhipeng Lu 0003, Yuwei Wu 0001, Yunde Jia |
AAAI | 5 |
| 2024 | Compositional Substitutivity of Visual Reasoning for Visual Question Answering
Chuanhao Li 0001, Zhen Li 0026, Chenchen Jing, Yuwei Wu 0001, Mingliang Zhai, Yunde Jia |
ECCV (48) | 6 |
| 2024 | Temporally Consistent Stereo Matching
Jiaxi Zeng, Chengtang Yao, Yuwei Wu 0001, Yunde Jia |
ECCV (31) | 4 |
| 2024 | In-Context Compositional Generalization for Large Vision-Language ModelsabstractRecent work has revealed that in-context learning for large language models exhibits compositional generalization capacity, which can be enhanced by selecting in-context demonstrations similar to test cases to provide contextual information.However, how to exhibit in-context compositional generalization (ICCG) of large vision-language models (LVLMs) is non-trival.Due to the inherent asymmetry between visual and linguistic modalities, ICCG in LVLMs faces an inevitable challenge-redundant information on the visual modality.The redundant information affects in-context learning from two aspects: (1) Similarity calculation may be dominated by redundant information, resulting in sub-optimal demonstration selection.(2) Redundant information in in-context demonstrations brings misleading contextual information to in-context learning.To alleviate these problems, we propose a demonstration selection method to achieve ICCG for LVLMs, by considering two key factors of demonstrations: content and structure, from a multimodal perspective.Specifically, we design a diversity-coverage-based matching score to select demonstrations with maximum coverage, and avoid selecting demonstrations with redundant information via their content redundancy and structural complexity.We build a GQA-ICCG dataset to simulate the ICCG setting, and conduct experiments on GQA-ICCG and the VQA v2 dataset.Experimental results demonstrate the effectiveness of our method. Chuanhao Li 0001, Chenchen Jing, Zhen Li 0026, Mingliang Zhai, Yuwei Wu 0001, Yunde Jia |
EMNLP | 6 |
| 2024 | FIRE: A Dataset for Feedback Integration and Refinement Evaluation of Multimodal ModelsabstractVision language models (VLMs) have achieved impressive progress in diverse applications, becoming a prevalent research direction. In this paper, we build FIRE, a feedback-refinement dataset, consisting of 1.1M multi-turn conversations that are derived from 27 source datasets, empowering VLMs to spontaneously refine their responses based on user feedback across diverse tasks. To scale up the data collection, FIRE is collected in two components: FIRE-100K and FIRE-1M, where FIRE-100K is generated by GPT-4V, and FIRE-1M is freely generated via models trained on FIRE-100K. Then, we build FIRE-Bench, a benchmark to comprehensively evaluate the feedback-refining capability of VLMs, which contains 11K feedback-refinement conversations as the test data, two evaluation settings, and a model to provide feedback for VLMs. We develop the FIRE-LLaVA model by fine-tuning LLaVA on FIRE-100K and FIRE-1M, which shows remarkable feedback-refining capability on FIRE-Bench and outperforms untrained VLMs by 50%, making more efficient user-agent interactions and underscoring the significance of the FIRE dataset. Pengxiang Li 0002, Zhi Gao 0002, Bofei Zhang, Yuwei Wu 0001, Mehrtash Harandi, Yunde Jia, Song-Chun Zhu, Qing Li 0003 |
NeurIPS | 7 |
| 2024 | Visual-Guided Reasoning Path Generation for Visual Question Answering
Chenchen Jing, Mingliang Zhai, Yuwei Wu 0001, Yunde Jia |
PRCV (1) | 5 |
| 2024 | Source-Free Image-Text Matching via Uncertainty-Aware LearningabstractWhen applying a trained image-text matching model to a new scenario, the performance may largely degrade due to domain shift, which makes it impractical in real-world applications. In this paper, we make the first attempt on adapting the image-text matching model well-trained on a labeled source domain to an unlabeled target domain in the absence of source data, namely, source-free image-text matching. This task is challenging since it has no direct access to the source data when learning to reduce the doma in shift. To address this challenge, we propose a simple yet effective method that introduces uncertainty-aware learning to generate high-quality pseudo-pairs of image and text for target adaptation. Specifically, starting with using the pre-trained source model to retrieve several top-ranked image-text pairs from the target domain as pseudo-pairs, we then model uncertainty of each pseudo-pair by calculating the variance of retrieved texts (resp. images) given the paired image (resp. text) as query, and finally incorporate the uncertainty into an objective function to down-weight noisy pseudo-pairs for better training, thereby enhancing adaptation. This uncertainty-aware training approach can be generally applied on all existing models. Extensive experiments on the COCO and Flickr30K datasets demonstrate the effectiveness of the proposed method. Mengxiao Tian, Shuo Yang 0002, Xinxiao Wu, Yunde Jia |
IEEE Signal Process. Lett. | 4 |
| 2024 | Adversarial Sample Synthesis for Visual Question AnsweringabstractLanguage prior is a major block to improving the generalization of visual question answering (VQA) models. Recent work has revealed that synthesizing extra training samples to balance training sets is a promising way to alleviate language priors. However, most existing methods synthesize extra samples in a manner independent of training processes, which neglect the fact that the language priors memorized by VQA models are changing during training, resulting in insufficient synthesized samples. In this article, we propose an adversarial sample synthesis method, which synthesizes different adversarial samples by adversarial masking at different training epochs to cope with the changing memorized language priors. The basic idea behind our method is to use adversarial masking to synthesize adversarial samples that will cause the model to make wrong answers. To this end, we design a generative module to carry out adversarial masking by attacking the VQA model and introduce a bias-oriented objective to supervise the training of the generative module. We couple the sample synthesis with the training process of the VQA model, which ensures that the synthesized samples at different training epochs are beneficial to the VQA model. We incorporated the proposed method into three VQA models including UpDn, LMH, and LXMERT and conducted experiments on three datasets including VQA-CP v1, VQA-CP v2, and VQA v2. Experimental results demonstrate that a large improvement of our method, such as 16.22% gains on LXMERT in the overall accuracy of VQA-CP v2. Chuanhao Li 0001, Chenchen Jing, Zhen Li 0026, Yuwei Wu 0001, Yunde Jia |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2023 | Learning Event-Relevant Factors for Video Anomaly DetectionabstractMost video anomaly detection methods discriminate events that deviate from normal patterns as anomalies. However, these methods are prone to interferences from event-irrelevant factors, such as background textures and object scale variations, incurring an increased false detection rate. In this paper, we propose to explicitly learn event-relevant factors to eliminate the interferences from event-irrelevant factors on anomaly predictions. To this end, we introduce a causal generative model to separate the event-relevant factors and event-irrelevant ones in videos, and learn the prototypes of event-relevant factors in a memory augmentation module. We design a causal objective function to optimize the causal generative model and develop a counterfactual learning strategy to guide anomaly predictions, which increases the influence of the event-relevant factors. The extensive experiments show the effectiveness of our method for video anomaly detection. Che Sun, Chenrui Shi, Yunde Jia, Yuwei Wu 0001 |
AAAI | 3 |
| 2023 | Exploring Data Geometry for Continual LearningabstractContinual learning aims to efficiently learn from a non-stationary stream of data while avoiding forgetting the knowledge of old data. In many practical applications, data complies with non-Euclidean geometry. As such, the commonly used Euclidean space cannot gracefully capture non-Euclidean geometric structures of data, leading to in-ferior results. In this paper, we study continual learning from a novel perspective by exploring data geometry for the non-stationary stream of data. Our method dynamically expands the geometry of the underlying space to match growing geometric structures induced by new data, and pre-vents forgetting by keeping geometric structures of old data into account. In doing so, making use of the mixed cur-vature space, we propose an incremental search scheme, through which the growing geometric structures are en-coded. Then, we introduce an angular-regularization loss and a neighbor-robustness loss to train the model, capa-ble of penalizing the change of global geometric structures and local geometric structures. Experiments show that our method achieves better performance than baseline methods designed in Euclidean space. Zhi Gao 0002, Yunde Jia, Mehrtash Harandi, Yuwei Wu 0001 |
CVPR | 4 |
| 2023 | Exploring the Effect of Primitives for Compositional Generalization in Vision-and-LanguageabstractCompositionality is one of the fundamental properties of human cognition (Fodor & Pylyshyn, 1988). Compositional generalization is critical to simulate the compositional capability of humans, and has received much attention in the vision-and-language (V&L) community. It is essential to understand the effect of the primitives, including words, image regions, and video frames, to improve the compositional generalization capability. In this paper, we explore the effect of primitives for compositional generalization in V&L. Specifically, we present a self-supervised learning based framework that equips existing V&L methods with two characteristics: semantic equivariance and semantic invariance. With the two characteristics, the methods understand primitives by perceiving the effect of primitive changes on sample semantics and ground-truth. Experimental results on two tasks: temporal video grounding and visual question answering, demonstrate the effectiveness of our framework. Chuanhao Li 0001, Zhen Li 0026, Chenchen Jing, Yunde Jia, Yuwei Wu 0001 |
CVPR | 4 |
| 2023 | Video Anomaly Detection via Sequentially Learning Multiple Pretext TasksabstractLearning multiple pretext tasks is a popular approach to tackle the nonalignment problem in unsupervised video anomaly detection. However, the conventional learning method of simultaneously learning multiple pretext tasks, is prone to sub-optimal solutions, incurring sharp performance drops. In this paper, we propose to sequentially learn multiple pretext tasks according to their difficulties in an ascending manner to improve the performance of anomaly detection. The core idea is to relax the learning objective by starting with easy pretext tasks in the early stage and gradually refine it by involving more challenging pretext tasks later on. In this way, our method is able to reduce the difficulties of learning and avoid converging to sub-optimal solutions. Specifically, we design a tailored sequential learning order for three widely-used pretext tasks. It starts with frame prediction task, then moves on to frame reconstruction task and last ends with frame-order classification task. We further introduce a new contrastive loss which makes the learned representations of normality more discriminative by pushing normal and pseudo-abnormal samples apart. Extensive experiments on three datasets demonstrate the effectiveness of our method. Chenrui Shi, Che Sun, Yuwei Wu 0001, Yunde Jia |
ICCV | 4 |
| 2023 | Sparse Point Guided 3D Lane Detectionabstract3D lane detection usually builds a dense correspondence between the front-view space and the BEV space to estimate lane points in the 3D space. 3D lanes only occupy a small ratio of the dense correspondence, while most correspondence belongs to the redundant background. This sparsity phenomenon bottlenecks valuable computation and raises the computation cost of building a high-resolution correspondence for accurate results. In this paper, we propose a sparse point-guided 3D lane detection, focusing on points related to 3D lanes. Our method runs in a coarse-to-fine manner, including coarse-level lane detection and iterative fine-level sparse point refinements. In coarse-level lane detection, we build a dense but efficient correspondence between the front view and BEV space at a very low resolution to compute coarse lanes. Then in fine-level sparse point refinement, we sample sparse points around coarse lanes to extract local features from the high-resolution front-view feature map. The high-resolution local information brought by sparse points refines 3D lanes in the BEV space hierarchically from low resolution to high resolution. The sparse point guides a more effective information flow and greatly promotes the SOTA result by 3 points on the overall F1-score and 6 points on several hard situations while reducing almost half memory cost and speeding up 2 times. Chengtang Yao, Lidong Yu, Yuwei Wu 0001, Yunde Jia |
ICCV | 4 |
| 2023 | Parameterized Cost Volume for Stereo MatchingabstractStereo matching becomes computationally challenging when dealing with a large disparity range. Prior methods mainly alleviate the computation through dynamic cost volume by focusing on a local disparity space, but it requires many iterations to get close to the ground truth due to the lack of a global view. We find that the dynamic cost volume approximately encodes the disparity space as a single Gaussian distribution with a fixed and small variance at each iteration, which results in an inadequate global view over disparity space and a small update step at every iteration. In this paper, we propose a parameterized cost volume to encode the entire disparity space using multi-Gaussian distribution. The disparity distribution of each pixel is parameterized by weights, means, and variances. The means and variances are used to sample disparity candidates for cost computation, while the weights and means are used to calculate the disparity output. The above parameters are computed through a JS-divergence-based optimization, which is realized as a gradient descent update in a feed-forward differential module. Experiments show that our method speeds up the runtime of RAFT-Stereo by 4 ~ 15 times, achieving real-time performance and comparable accuracy. The code is available at https://github.com/jiaxiZeng/Parameterized-Cost-Volume-for-Stereo-Matching. Jiaxi Zeng, Chengtang Yao, Lidong Yu, Yuwei Wu 0001, Yunde Jia |
ICCV | 5 |
| 2023 | Fast-StrucTexT: An Efficient Hourglass Transformer with Modality-guided Dynamic Token Merge for Document UnderstandingabstractTransformers achieve promising performance in document understanding because of their high effectiveness and still suffer from quadratic computational complexity dependency on the sequence length. General efficient transformers are challenging to be directly adapted to model document. They are unable to handle the layout representation in documents, e.g. word, line and paragraph, on different granularity levels and seem hard to achieve a good trade-off between efficiency and performance. To tackle the concerns, we propose Fast-StrucTexT, an efficient multi-modal framework based on the StrucTexT algorithm with an hourglass transformer architecture, for visual document understanding. Specifically, we design a modality-guided dynamic token merging block to make the model learn multi-granularity representation and prunes redundant tokens. Additionally, we present a multi-modal interaction module called Symmetry Cross-Attention (SCA) to consider multi-modal fusion and efficiently guide the token mergence. The SCA allows one modality input as query to calculate cross attention with another modality in a dual phase. Extensive experiments on FUNSD, SROIE, and CORD datasets demonstrate that our model achieves the state-of-the-art performance and almost 1.9x faster inference time than the state-of-the-art methods. Mingliang Zhai, Yulin Li 0004, Xiameng Qin, Qunyi Xie, Chengquan Zhang, Yuwei Wu 0001, Yunde Jia |
IJCAI | 9 |
| 2023 | Learning to Optimize on Riemannian ManifoldsabstractMany learning tasks are modeled as optimization problems with nonlinear constraints, such as principal component analysis and fitting a Gaussian mixture model. A popular way to solve such problems is resorting to Riemannian optimization algorithms, which yet heavily rely on both human involvement and expert knowledge about Riemannian manifolds. In this paper, we propose a Riemannian meta-optimization method to automatically learn a Riemannian optimizer. We parameterize the Riemannian optimizer by a novel recurrent network and utilize Riemannian operations to ensure that our method is faithful to the geometry of manifolds. The proposed method explores the distribution of the underlying data by minimizing the objective of updated parameters, and hence is capable of learning task-specific optimizations. We introduce a Riemannian implicit differentiation training scheme to achieve efficient training in terms of numerical stability and computational cost. Unlike conventional meta-optimization training schemes that need to differentiate through the whole optimization trajectory, our training scheme is only related to the final two optimization steps. In this way, our training scheme avoids the exploding gradient problem, and significantly reduces the computational load and memory footprint. We discuss experimental results across various constrained problems, including principal component analysis on Grassmann manifolds, face recognition, person re-identification, and texture image classification on Stiefel manifolds, clustering and similarity learning on symmetric positive definite manifolds, and few-shot learning on hyperbolic manifolds. Zhi Gao 0002, Yuwei Wu 0001, Xiaomeng Fan, Mehrtash Harandi, Yunde Jia |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Curvature-Adaptive Meta-Learning for Fast Adaptation to Manifold DataabstractMeta-learning methods are shown to be effective in quickly adapting a model to novel tasks. Most existing meta-learning methods represent data and carry out fast adaptation in euclidean space. In fact, data of real-world applications usually resides in complex and various Riemannian manifolds. In this paper, we propose a curvature-adaptive meta-learning method that achieves fast adaptation to manifold data by producing suitable curvature. Specifically, we represent data in the product manifold of multiple constant curvature spaces and build a product manifold neural network as the base-learner. In this way, our method is capable of encoding complex manifold data into discriminative and generic representations. Then, we introduce curvature generation and curvature updating schemes, through which suitable product manifolds for various forms of data manifolds are constructed via few optimization steps. The curvature generation scheme identifies task-specific curvature initialization, leading to a shorter optimization trajectory. The curvature updating scheme automatically produces appropriate learning rate and search direction for curvature, making a faster and more adaptive optimization paradigm compared to hand-designed optimization schemes. We evaluate our method on a broad set of problems including few-shot classification, few-shot regression, and reinforcement learning tasks. Experimental results show that our method achieves substantial improvements as compared to meta-learning methods ignoring the geometry of the underlying space. Zhi Gao 0002, Yuwei Wu 0001, Mehrtash Harandi, Yunde Jia |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Adaptive Latent Graph Representation Learning for Image-Text MatchingabstractImage-text matching is a challenging task due to the modality gap. Many recent methods focus on modeling entity relationships to learn a common embedding space of image and text. However, these methods suffer from distractions of entity relationships such as irrelevant visual regions in an image and noisy textual words in a text. In this paper, we propose an adaptive latent graph representation learning method to reduce the distractions of entity relationships for image-text matching. Specifically, we use an improved graph variational autoencoder to separate the distracting factors and latent factor of relationships and jointly learn latent textual graph representations, latent visual graph representations, and a visual-textual graph embedding space. We also introduce an adaptive cross-attention mechanism to perform feature attending on the latent graph representations across images and texts, thus further narrowing the modality gap to boost the matching performance. Extensive experiments on two public datasets, Flickr30K and COCO, show the effectiveness of our method. Mengxiao Tian, Xinxiao Wu, Yunde Jia |
IEEE Trans. Image Process. | 3 |
| 2023 | TIUI: Touching Live Video for Telepresence OperationabstractThis paper presents a framework for telepresence operation by touching live video on a touchscreen. Our goal is to enable users to use a smartpad to teleoperate everyday objects by touching the objects’ live video they are watching. To this end, we coined the term “teleinteractive device” to describe such an object with an identity, an actuator, and a communication network. We developed a touchable live video image-based user interface (TIUI) that empowers users to teleoperate any teleinteractive device by touching its live video with touchscreen gestures. The TIUI contains four modules — touch, control, recognition, and knowledge — to perform live video understanding, communication, and control for telepresence operation. We implemented a telepresence operation system that consists of a telepresence robot and teleinteractive devices at a local site, a smartpad with the TIUI at a remote site, and communication networks connecting the two sites. We demonstrated potential applications of the system in remotely controlling telepresence robots, opening doors with access control panels, and pushing power wheelchairs. We conducted user studies to show the effectiveness of the proposed framework. Yunde Jia, Yanmei Dong, Bin Xu 0019, Che Sun |
IEEE Trans. Mob. Comput. | 1 |
| 2022 | Efficient Riemannian Meta-Optimization by Implicit DifferentiationabstractTo solve optimization problems with nonlinear constrains, the recently developed Riemannian meta-optimization methods show promise, which train neural networks as an optimizer to perform optimization on Riemannian manifolds. A key challenge is the heavy computational and memory burdens, because computing the meta-gradient with respect to the optimizer involves a series of time-consuming derivatives, and stores large computation graphs in memory. In this paper, we propose an efficient Riemannian meta-optimization method that decouples the complex computation scheme from the meta-gradient. We derive Riemannian implicit differentiation to compute the meta-gradient by establishing a link between Riemannian optimization and the implicit function theorem. As a result, the updating our optimizer is only related to the final two iterations, which in turn speeds up our method and reduces the memory footprint significantly. We theoretically study the computational load and memory footprint of our method for long optimization trajectories, and conduct an empirical study to demonstrate the benefits of the proposed method. Evaluations of three optimization problems on different Riemannian manifolds show that our method achieves state-of-the-art performance in terms of the convergence speed and the quality of optima. Xiaomeng Fan, Yuwei Wu 0001, Zhi Gao 0002, Yunde Jia, Mehrtash Harandi |
AAAI | 4 |
| 2022 | Learning the Dynamics of Visual Relational Reasoning via Reinforced Path RoutingabstractReasoning is a dynamic process. In cognitive theories, the dynamics of reasoning refers to reasoning states over time after successive state transitions. Modeling the cognitive dynamics is of utmost importance to simulate human reasoning capability. In this paper, we propose to learn the reasoning dynamics of visual relational reasoning by casting it as a path routing task. We present a reinforced path routing method that represents an input image via a structured visual graph and introduces a reinforcement learning based model to explore paths (sequences of nodes) over the graph based on an input sentence to infer reasoning results. By exploring such paths, the proposed method represents reasoning states clearly and characterizes state transitions explicitly to fully model the reasoning dynamics for accurate and transparent visual relational reasoning. Extensive experiments on referring expression comprehension and visual question answering demonstrate the effectiveness of our method. Chenchen Jing, Yunde Jia, Yuwei Wu 0001, Chuanhao Li 0001, Qi Wu 0001 |
AAAI | 2 |
| 2022 | Maintaining Reasoning Consistency in Compositional Visual Question AnsweringabstractA compositional question refers to a question that contains multiple visual concepts (e.g., objects, attributes, and relationships) and requires compositional reasoning to answer. Existing VQA models can answer a compositional question well, but cannot work well in terms of reasoning consistency in answering the compositional question and its sub-questions. For example, a compositional question for an image is: “Are there any elephants to the right of the white bird?” and one of its sub-questions is “Is any bird visible in the scene?”. The models may answer “yes” to the compositional question, but “no” to the sub-question. This paper presents a dialog-like reasoning method for maintaining reasoning consistency in answering a compositional question and its sub-questions. Our method integrates the reasoning processes for the sub-questions into the reasoning process for the compositional question like a dialog task, and uses a consistency constraint to penalize inconsistent answer predictions. In order to enable quantitative evaluation of reasoning consistency, we construct a GQA-Sub dataset based on the well-organized GQA dataset. Experimental results on the GQA dataset and the GQA-Sub dataset demonstrate the effectiveness of our method. Chenchen Jing, Yunde Jia, Yuwei Wu 0001, Qi Wu 0001 |
CVPR | 2 |
| 2022 | Global-Aware Registration of Less-Overlap RGB-D ScansabstractWe propose a novel method of registering less-overlap RGB-D scans. Our method learns global information of a scene to construct a panorama, and aligns RGB-D scans to the panorama to perform registration. Different from existing methods that use local feature points to register less-overlap RGB-D scans and mismatch too much, we use global information to guide the registration, thereby allevi-ating the mismatching problem by preserving global consis-tency of alignments. To this end, we build a scene inference network to construct the panorama representing global in-formation. We introduce a reinforcement learning strategy to iteratively align RGB-D scans with the panorama and re-fine the panorama representation, which reduces the noise of global information and preserves global consistency of both geometric and photometric alignments. Experimental results on benchmark datasets including SUNCG, Matterport, and ScanNet show the superiority of our method. Che Sun, Yunde Jia, Yuwei Wu 0001 |
CVPR | 2 |
| 2022 | Evidential Reasoning for Video Anomaly DetectionabstractVideo anomaly detection aims to discriminate events that deviate from normal patterns in a video. Modeling the decision boundaries of anomalies is challenging, due to the uncertainty in the probability of deviating from normal patterns. In this paper, we propose a deep evidential reasoning method that explicitly learns the uncertainty to model the boundaries. Our method encodes various visual cues as evidences representing potential deviations, assigns beliefs to the predicted probability of deviating from normal patterns based on the evidences, and estimates the uncertainty from the remained beliefs to model the boundaries. To do this, we build a deep evidential reasoning network to encode evidence vectors and estimate uncertainty by learning evidence distributions and deriving beliefs from the distributions. We introduce an unsupervised strategy to train our network by minimizing an energy function of the deep Gaussian mixed model (GMM). Experimental results show that our uncertainty score is beneficial for modeling the boundaries of video anomalies on three benchmark datasets. Che Sun, Yunde Jia, Yuwei Wu 0001 |
ACM Multimedia | 2 |
| 2022 | Hyperbolic Feature Augmentation via Distribution Estimation and Infinite Sampling on ManifoldsabstractLearning in hyperbolic spaces has attracted growing attention recently, owing to their capabilities in capturing hierarchical structures of data. However, existing learning algorithms in the hyperbolic space tend to overfit when limited data is given. In this paper, we propose a hyperbolic feature augmentation method that generates diverse and discriminative features in the hyperbolic space to combat overfitting. We employ a wrapped hyperbolic normal distribution to model augmented features, and use a neural ordinary differential equation module that benefits from meta-learning to estimate the distribution. This is to reduce the bias of estimation caused by the scarcity of data. We also derive an upper bound of the augmentation loss, which enables us to train a hyperbolic model by using an infinite number of augmentations. Experiments on few-shot learning and continual learning tasks show that our method significantly improves the performance of hyperbolic algorithms in scarce data regimes. Zhi Gao 0002, Yuwei Wu 0001, Yunde Jia, Mehrtash Harandi |
NeurIPS | 3 |
| 2022 | Stitching images from a conventional camera and a fisheye camera based on nonrigid warping
Yanmei Dong, Mingtao Pei, Yuwei Wu 0001, Yunde Jia |
Multim. Tools Appl. | 4 |
| 2022 | Towards a Weakly Supervised Framework for 3D Point Cloud Object Detection and AnnotationabstractIt is quite laborious and costly to manually label LiDAR point cloud data for training high-quality 3D object detectors. This work proposes a weakly supervised framework which allows learning 3D detection from a few weakly annotated examples. This is achieved by a two-stage architecture design. Stage-1 learns to generate cylindrical object proposals under inaccurate and inexact supervision, obtained by our proposed BEV center-click annotation strategy, where only the horizontal object centers are click-annotated in bird's view scenes. Stage-2 learns to predict cuboids and confidence scores in a coarse-to-fine, cascade manner, under incomplete supervision, i.e., only a small portion of object cuboids are precisely annotated. With KITTI dataset, using only 500 weakly annotated scenes and 534 precisely labeled vehicle instances, our method achieves 86-97 percent the performance of current top-leading, fully supervised detectors (which require 3,712 exhaustively annotated scenes with 15,654 instances). More importantly, with our elaborately designed network architecture, our trained model can be applied as a 3D object annotator, supporting both automatic and active (human-in-the-loop) working modes. The annotations generated by our model can be used to train 3D object detectors, achieving over 95 percent of their original performance (with manually labeled training data). Our experiments also show our model's potential in boosting performance when given more training data. The above designs make our approach highly practical and open-up opportunities for learning 3D detection at reduced annotation cost. Qinghao Meng, Wenguan Wang, Tianfei Zhou, Jianbing Shen, Yunde Jia, Luc Van Gool |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Infinite-dimensional feature aggregation via a factorized bilinear model
Jindou Dai, Yuwei Wu 0001, Zhi Gao 0002, Yunde Jia |
Pattern Recognit. | 4 |
| 2022 | Exploiting Informative Video Segments for Temporal Action LocalizationabstractWe propose a novel method of exploiting informative video segments by learning segment weights for temporal action localization in untrimmed videos. Informative video segments represent the intrinsic motion and appearance of an action, and thus contribute crucially to action localization. The learned segment weights represent the informativeness of video segments to recognize actions and help infer the boundaries required to temporally localize actions. We build a supervised temporal attention network (STAN) that includes a supervised segment-level attention module to dynamically learn the weights of video segments, and a feature-level attention module to effectively fuse multiple features of segments. Through the cascade of the attention modules, STAN exploits informative video segments and generates descriptive and discriminative video representations. We use a proposal generator and a classifier to estimate the boundaries of actions and classify the classes of actions. Extensive experiments are conducted on two public benchmarks, i.e., THUMOS2014 and ActivityNet1.3. The results demonstrate that our proposed method achieves competitive performance compared with existing state-of-the-art methods. Moreover, compared with the baseline method that treats video segments equally, STAN achieves significant improvements with an increase of the mean average precision from 30.4% to 39.8% on the THUMOS2014 dataset, and from 31.4% to 35.9% on the ActivityNet1.3 dataset, demonstrating the effectiveness of learning informative video segments for temporal action localization. Che Sun, Hao Song 0002, Xinxiao Wu, Yunde Jia, Jiebo Luo 0001 |
IEEE Trans. Multim. | 4 |
| 2021 | Learning a Gradient-free Riemannian Optimizer on Tangent Spaces
Xiaomeng Fan, Zhi Gao 0002, Yuwei Wu 0001, Yunde Jia, Mehrtash Harandi |
AAAI | 4 |
| 2021 | A Hyperbolic-to-Hyperbolic Graph Convolutional NetworkabstractHyperbolic graph convolutional networks (GCNs) demonstrate powerful representation ability to model graphs with hierarchical structure. Existing hyperbolic GCNs resort to tangent spaces to realize graph convolution on hyperbolic manifolds, which is inferior because tangent space is only a local approximation of a manifold. In this paper, we propose a hyperbolic-to-hyperbolic graph convolutional network (H2H-GCN) that directly works on hyperbolic manifolds. Specifically, we developed a manifold-preserving graph convolution that consists of a hyperbolic feature transformation and a hyperbolic neighborhood aggregation. The hyperbolic feature transformation works as linear transformation on hyperbolic manifolds. It ensures the transformed node representations still lie on the hyperbolic manifold by imposing the orthogonal constraint on the transformation sub-matrix. The hyperbolic neighborhood aggregation updates each node representation via the Einstein midpoint. The H2H-GCN avoids the distortion caused by tangent space approximations and keeps the global hyperbolic structure. Extensive experiments show that the H2H-GCN achieves substantial improvements on the link prediction, node classification, and graph classification tasks. Jindou Dai, Yuwei Wu 0001, Zhi Gao 0002, Yunde Jia |
CVPR | 4 |
| 2021 | A Decomposition Model for Stereo MatchingabstractIn this paper, we present a decomposition model for stereo matching to solve the problem of excessive growth in computational cost (time and memory cost) as the resolution increases. In order to reduce the huge cost of stereo matching at the original resolution, our model only runs dense matching at a very low resolution and uses sparse matching at different higher resolutions to recover the disparity of lost details scale-by-scale. After the decomposition of stereo matching, our model iteratively fuses the sparse and dense disparity maps from adjacent scales with an occlusion-aware mask. A refinement network is also applied to improving the fusion result. Compared with high-performance methods like PSMNet and GANet, our method achieves 10−100× speed increase while obtaining comparable disparity estimation results. Chengtang Yao, Yunde Jia, Huijun Di, Pengxiang Li 0002, Yuwei Wu 0001 |
CVPR | 2 |
| 2021 | Curvature Generation in Curved Spaces for Few-Shot LearningabstractFew-shot learning describes the challenging problem of recognizing samples from unseen classes given very few labeled examples. In many cases, few-shot learning is cast as learning an embedding space that assigns test samples to their corresponding class prototypes. Previous methods assume that data of all few-shot learning tasks comply with a fixed geometrical structure, mostly a Euclidean structure. Questioning this assumption that is clearly difficult to hold in real-world scenarios and incurs distortions to data, we propose to learn a task-aware curved embedding space by making use of the hyperbolic geometry. As a result, task-specific embedding spaces where suitable curvatures are generated to match the characteristics of data are constructed, leading to more generic embedding spaces. We then leverage on intra-class and inter-class context information in the embedding space to generate class prototypes for discriminative classification. We conduct a comprehensive set of experiments on inductive and transductive few-shot learning, demonstrating the benefits of our proposed method over existing embedding methods. Zhi Gao 0002, Yuwei Wu 0001, Yunde Jia, Mehrtash Harandi |
ICCV | 3 |
| 2021 | Multi-homography Estimation and Inference Driven by Contour Alignment
Yunde Jia, Huijun Di, Yuwei Wu 0001 |
ICIG (1) | 2 |
| 2021 | Adversarial 3D Convolutional Auto-Encoder for Abnormal Event Detection in VideosabstractAbnormal event detection aims to identify the events that deviate from expected normal patterns. Existing methods usually extract normal spatio-temporal patterns of appearance and motion in a separate manner, which ignores low-level correlations between appearance and motion patterns and may fall short of capturing fine-grained spatio-temporal patterns. In this paper, we propose to simultaneously learn appearance and motion to obtain fine-grained spatio-temporal patterns. To this end, we present an adversarial 3D convolutional auto-encoder to learn the normal spatio-temporal patterns and then identify abnormal events by diverging them from the learned normal patterns in videos. The encoder captures the low-level correlations between spatial and temporal dimensions of videos, and generates distinctive features representing visual spatio-temporal information. The decoder reconstrucccts the original video from the encoded features representing by 3D de-convolutions and learns the normal spatio-temporal patterns in an unsupervised manner. We introduce the denoising reconstruction error and adversarial learning strategy to train the 3D convolutional auto-encoder to implicitly learn accurate data distributions that are considered normal patterns, which benefits enhancing the reconstruction ability of the auto-encoder to discriminate abnormal events. Both the theoretical analysis and the extensive experiments on four publicly available datasets demonstrate the effectiveness of our method. Che Sun, Yunde Jia, Hao Song 0002, Yuwei Wu 0001 |
IEEE Trans. Multim. | 2 |
| 2020 | Revisiting Bilinear Pooling: A Coding PerspectiveabstractBilinear pooling has achieved state-of-the-art performance on fusing features in various machine learning tasks, owning to its ability to capture complex associations between features. Despite the success, bilinear pooling suffers from redundancy and burstiness issues, mainly due to the rank-one property of the resulting representation. In this paper, we prove that bilinear pooling is indeed a similarity-based coding-pooling formulation. This establishment then enables us to devise a new feature fusion algorithm, the factorized bilinear coding (FBC) method, to overcome the drawbacks of the bilinear pooling. We show that FBC can generate compact and discriminative representations with substantially fewer parameters. Experiments on two challenging tasks, namely image classification and visual question answering, demonstrate that our method surpasses the bilinear pooling technique by a large margin. Zhi Gao 0002, Yuwei Wu 0001, Xiaoxun Zhang, Jindou Dai, Yunde Jia, Mehrtash Harandi |
AAAI | 5 |
| 2020 | Joint Commonsense and Relation Reasoning for Image and Video CaptioningabstractExploiting relationships between objects for image and video captioning has received increasing attention. Most existing methods depend heavily on pre-trained detectors of objects and their relationships, and thus may not work well when facing detection challenges such as heavy occlusion, tiny-size objects, and long-tail classes. In this paper, we propose a joint commonsense and relation reasoning method that exploits prior knowledge for image and video captioning without relying on any detectors. The prior knowledge provides semantic correlations and constraints between objects, serving as guidance to build semantic graphs that summarize object relationships, some of which cannot be directly perceived from images or videos. Particularly, our method is implemented by an iterative learning algorithm that alternates between 1) commonsense reasoning for embedding visual regions into the semantic space to build a semantic graph and 2) relation reasoning for encoding semantic graphs to generate sentences. Experiments on several benchmark datasets validate the effectiveness of our prior knowledge-based approach. Jingyi Hou, Xinxiao Wu, Xiaoxun Zhang, Yayun Qi, Yunde Jia, Jiebo Luo 0001 |
AAAI | 5 |
| 2020 | Overcoming Language Priors in VQA via Decomposed Linguistic RepresentationsabstractMost existing Visual Question Answering (VQA) models overly rely on language priors between questions and answers. In this paper, we present a novel method of language attention-based VQA that learns decomposed linguistic representations of questions and utilizes the representations to infer answers for overcoming language priors. We introduce a modular language attention mechanism to parse a question into three phrase representations: type representation, object representation, and concept representation. We use the type representation to identify the question type and the possible answer set (yes/no or specific concepts such as colors or numbers), and the object representation to focus on the relevant region of an image. The concept representation is verified with the attended region to infer the final answer. The proposed method decouples the language-based concept discovery and vision-based concept verification in the process of answer inference to prevent language priors from dominating the answering process. Experiments on the VQA-CP dataset demonstrate the effectiveness of our method. Chenchen Jing, Yuwei Wu 0001, Xiaoxun Zhang, Yunde Jia, Qi Wu 0001 |
AAAI | 4 |
| 2020 | Learning to Optimize on SPD ManifoldsabstractMany tasks in computer vision and machine learning are modeled as optimization problems with constraints in the form of Symmetric Positive Definite (SPD) matrices. Solving such optimization problems is challenging due to the non-linearity of the SPD manifold, making optimization with SPD constraints heavily relying on expert knowledge and human involvement. In this paper, we propose a meta-learning method to automatically learn an iterative optimizer on SPD manifolds. Specifically, we introduce a novel recurrent model that takes into account the structure of input gradients and identifies the updating scheme of optimization. We parameterize the optimizer by the recurrent model and utilize Riemannian operations to ensure that our method is faithful to the geometry of SPD manifolds. Compared with existing SPD optimizers, our optimizer effectively exploits the underlying data distribution and learns a better optimization trajectory in a data-driven manner. Extensive experiments on various computer vision tasks including metric nearness, clustering, and similarity learning demonstrate that our optimizer outperforms existing state-of-the-art methods consistently. Zhi Gao 0002, Yuwei Wu 0001, Yunde Jia, Mehrtash Harandi |
CVPR | 3 |
| 2020 | Deep 3D Portrait From a Single ImageabstractIn this paper, we present a learning-based approach for recovering the 3D geometry of human head from a single portrait image. Our method is learned in an unsupervised manner without any ground-truth 3D data. We represent the head geometry with a parametric 3D face model together with a depth map for other head regions including hair and ear. A two-step geometry learning scheme is proposed to learn 3D head reconstruction from in-the-wild face images, where we first learn face shape on single images using self-reconstruction and then learn hair and ear geometry using pairs of images in a stereo-matching fashion. The second step is based on the output of the first to not only improve the accuracy but also ensure the consistency of overall head geometry. We evaluate the accuracy of our method both in 3D and with pose manipulation tasks on 2D images. We alter pose based on the recovered geometry and apply a refinement network trained with adversarial learning to ameliorate the reprojected images and translate them to the real image domain. Extensive evaluations and comparison with previous methods show that our new method can produce high-fidelity 3D head geometry and head pose manipulation results. Sicheng Xu, Jiaolong Yang, Dong Chen 0003, Fang Wen 0001, Yu Deng 0006, Yunde Jia, Xin Tong 0001 |
CVPR | 6 |
| 2020 | Visual-Semantic Graph Matching for Visual GroundingabstractVisual Grounding is the task of associating entities in a natural language sentence with objects in an image. In this paper, we formulate visual grounding as a graph matching problem to find node correspondences between a visual scene graph and a language scene graph. These two graphs are heterogeneous, representing structure layouts of the sentence and image, respectively. We learn unified contextual node representations of the two graphs by using a cross-modal graph convolutional network to reduce their discrepancy. The graph matching is thus relaxed as a linear assignment problem because the learned node representations characterize both node information and structure information. A permutation loss and a semantic cycle-consistency loss are further introduced to solve the linear assignment problem with or without ground-truth correspondences. Experimental results on two visual grounding tasks, i.e., referring expression comprehension and phrase localization, demonstrate the effectiveness of our method. Chenchen Jing, Yuwei Wu 0001, Mingtao Pei, Yao Hu 0002, Yunde Jia, Qi Wu 0001 |
ACM Multimedia | 5 |
| 2020 | Scene-Aware Context Reasoning for Unsupervised Abnormal Event Detection in VideosabstractIn this paper, we propose a scene-aware context reasoning method that exploits context information from visual features for unsupervised abnormal event detection in videos, which bridges the semantic gap between visual context and the meaning of abnormal events. In particular, we build na spatio-temporal context graph to model visual context information including appearances of objects, spatio-temporal relationships among objects and scene types. The context information is encoded into the nodes and edges of the graph, and their states are iteratively updated by using multiple RNNs with message passing for context reasoning. To infer the spatio-temporal context graph in various scenes, we develop a graph-based deep Gaussian mixture model for scene clustering in an unsupervised manner. We then compute frame-level anomaly scores based on the context information to discriminate abnormal events in various scenes. Evaluations on three challenging datasets, including the UCF-Crime, Avenue, and ShanghaiTech datasets, demonstrate the effectiveness of our method. Che Sun, Yunde Jia, Yao Hu 0002, Yuwei Wu 0001 |
ACM Multimedia | 2 |
| 2020 | Incremental transfer learning for video annotation via grouped heterogeneous sourcesabstractHere, the authors focus on incrementally acquiring heterogeneous knowledge from both internet and publicly available datasets to reduce the tedious and expensive labelling efforts required in video annotation. An incremental transfer learning framework is presented to integrate heterogeneous source knowledge and update the annotation model incrementally during the transfer learning process. Under this framework, web images and existing action videos form the source domain to provide labelled static and motion information of the target domain videos, respectively. Moreover, according to the semantic of the source domain data, all the source domain data are partitioned into several groups. Different from traditional methods, which compare the entire target domain videos with each source group from the source domain, the authors treat the group weights as sample‐specific variables and optimise them along with new adding data. Two regularisers are used to prevent the incremental learning process from negative transfer. Experimental results on the two large‐scale consumer video datasets (i.e. multimedia event detection (MED) and Columbia consumer video (CCV)) show the effectiveness of the proposed method. Hao Song 0002, Xinxiao Wu, Yunde Jia |
IET Comput. Vis. | 4 |
| 2020 | Can You Easily Perceive the Local Environment? A User Interface with One Stitched Live Video for Mobile Robotic Telepresence SystemsabstractMany existing mobile robotic telepresence systems have equipped with two cameras, one is a forward-facing camera for video communication, and the other is a downward-facing camera for robot navigation. However, the two live videos from these two cameras would cause some confusion which makes it difficult for a remote operator to perceive the local environment. In this paper, we propose to use a user interface with one stitched live video instead of two live videos for mobile robotic telepresence systems. We used a video stitching algorithm to stitch the two live videos into one live video through which a remote operator can well perceive the local environment. We conducted a user study to investigate the difference between one stitched live video and two separate live videos in the user interface. The results show that the user interface with one stitched live video improves task efficiency, the number of errors, and remote operators’ feelings of presence, and enables remote operators to concentrate on the work they are doing. Yanmei Dong, Yunde Jia, Weichao Shen, Yuwei Wu 0001 |
Int. J. Hum. Comput. Interact. | 2 |
| 2020 | Online maximum a posteriori tracking of multiple objects using sequential trajectory prior
Min Yang 0003, Mingtao Pei, Yunde Jia |
Image Vis. Comput. | 3 |
| 2020 | Face Spoofing Detection Using Relativity Representation on Riemannian ManifoldabstractFace recognition and verification systems are susceptible to spoofing attacks using photographs, videos or masks. Most existing methods focus on spoofing detection in Euclidean space, and ignore the features' manifold structure and interrelationships, thus limiting their capabilities of discrimination and generalization. In this paper, we propose a relativity representation on Riemannian manifold for face spoofing detection. The relativity representation improves generalization capability while ensuring discriminability, at both levels of feature description and classification score. The feature-level relativity representation generalizes information by modeling interrelationships among basic features, and would not depend too much on characteristics of a particular dataset. The score-level relativity representation makes decisions relatively, not absolutely, according to interrelationships (via Riemannian metric) and competitions (via example reweighting) among data samples on Riemannian manifold. The discriminability is ensured by the high-order nature of the feature-level relativity representation as well as Riemannian reweighted discriminative learning of the score-level relativity representation. Moreover, we integrate an attack-sensitive SVM classifier in Euclidean space to improve spoofing detection. Experiments demonstrate the effectiveness of our method on both intra-dataset and cross-dataset testing. Chengtang Yao, Yunde Jia, Huijun Di, Yuwei Wu 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2020 | Confidence-Guided Self Refinement for Action Prediction in Untrimmed VideosabstractMany existing methods formulate the action prediction task as recognizing early parts of actions in trimmed videos. In this paper, we focus on predicting actions from ongoing untrimmed videos where actions might not happen at the very beginning of videos. It is extremely challenging to predict actions in such untrimmed videos due to ambiguous or even no information of actions in the early parts of videos. To address this problem, we propose a prediction confidence that assesses the decision quality of a prediction model. Guided by the confidence, the model continuously refines the prediction results by itself with the increasing observed video frames. Specifically, we build a Self Prediction Refining Network (SPR-Net) which incrementally learns the confidence for action prediction. SPR-Net consists of three modules: a temporal hybrid network, an incremental confidence learner, and a self-refining Gumbel softmax sampler. The temporal hybrid network generates the action category distributions by integrating static scene and dynamic motion information. The incremental confidence learner calculates the confidence in an incremental manner, judging the extent to which the temporal hybrid network should believe its prediction result. The self-refining Gumbel softmax sampler models the mutual relationship between the prediction confidence and the category distribution, which enables them to be jointly learned in an end-to-end fashion. We also present a sparse self-attention mechanism to encode local spatio-temporal features into the frame-level motion representation to further improve the prediction performance. Extensive experiments on five datasets (i.e., UT-Interaction, BIT-Interaction, UCF101, THUMOS14, and ActivityNet) validate the effectiveness of the proposed method. Jingyi Hou, Xinxiao Wu, Jiebo Luo 0001, Yunde Jia |
IEEE Trans. Image Process. | 5 |
| 2020 | Learning Normal Patterns via Adversarial Attention-Based Autoencoder for Abnormal Event Detection in VideosabstractAutomatically detecting anomalies in videos is a challenging problem due to non-deterministic definitions of abnormal events and lack of sufficient training data. To address these issues, we propose an autoencoder coupled with attention model to discover normal patterns in videos via adversarial learning. Abnormal events are detected by diverging them from the normal patterns with the reconstruction error produced by the autoencoder. To this end, we build an end-to-end trainable adversarial attention-based autoencoder network, called Ada-Net, to make the reconstructed frames indistinguishable from original frames. The Ada-Net combines an autoencoder network and a GAN model that is used to benefit enhancing the reconstruction ability of the autoencoder. To further improve the reconstruction performance, we integrate an attention model into the decoder to dynamically select informative parts of encoding features for decoding. The attenion mechanism is helpful to preserving important information for learning intrinsic normal patterns. Evaluations on four challenging datasets, including the Subway, the UCSD Pedestrian, the CUHK Avenue, and the ShanghaiTech datasets, demonstrate the effectiveness of the proposed method. Hao Song 0002, Che Sun, Xinxiao Wu, Yunde Jia |
IEEE Trans. Multim. | 5 |
| 2020 | A Robust Distance Measure for Similarity-Based Classification on the SPD ManifoldabstractThe symmetric positive definite (SPD) matrices, forming a Riemannian manifold, are commonly used as visual representations. The non-Euclidean geometry of the manifold often makes developing learning algorithms (e.g., classifiers) difficult and complicated. The concept of similarity-based learning has been shown to be effective to address various problems on SPD manifolds. This is mainly because the similarity-based algorithms are agnostic to the geometry and purely work based on the notion of similarities/distances. However, existing similarity-based models on SPD manifolds opt for holistic representations, ignoring characteristics of information captured by SPD matrices. To circumvent this limitation, we propose a novel SPD distance measure for the similarity-based algorithm. Specifically, we introduce the concept of point-to-set transformation, which enables us to learn multiple lower dimensional and discriminative SPD manifolds from a higher dimensional one. For lower dimensional SPD manifolds obtained by the point-to-set transformation, we propose a tailored set-to-set distance measure by making use of the family of alpha-beta divergences. We further propose to learn the point-to-set transformation and the set-to-set distance measure jointly, yielding a powerful similarity-based algorithm on SPD manifolds. Our thorough evaluations on several visual recognition tasks (e.g., action classification and face recognition) suggest that our algorithm comfortably outperforms various state-of-the-art algorithms. Zhi Gao 0002, Yuwei Wu 0001, Mehrtash Harandi, Yunde Jia |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2019 | 3D Shape Reconstruction From Images in the Frequency DomainabstractReconstructing the high-resolution volumetric 3D shape from images is challenging due to the cubic growth of computational cost. In this paper, we propose a Fourier-based method that reconstructs a 3D shape from images in a 2D space by predicting slices in the frequency domain. According to the Fourier slice projection theorem, we introduce a thickness map to bridge the domain gap between images in the spatial domain and slices in the frequency domain. The thickness map is the 2D spatial projection of the 3D shape, which is easily predicted from the input image by a general convolutional neural network. Each slice in the frequency domain is the Fourier transform of the corresponding thickness map. All slices constitute a 3D descriptor and the 3D shape is the inverse Fourier transform of the descriptor. Using slices in the frequency domain, our method can transfer the 3D shape reconstruction from the 3D space into the 2D space, which significantly reduces the computational cost. The experiment results on the ShapeNet dataset demonstrate that our method achieves competitive reconstruction accuracy and computational efficiency compared with the state-of-the-art reconstruction methods. Weichao Shen, Yunde Jia, Yuwei Wu 0001 |
CVPR | 2 |
| 2019 | Joint Syntax Representation Learning and Visual Cue Translation for Video CaptioningabstractVideo captioning is a challenging task that involves not only visual perception but also syntax representation learning. Recent progress in video captioning has been achieved through visual perception, but syntax representation learning is still under-explored. We propose a novel video captioning approach that takes into account both visual perception and syntax representation learning to generate accurate descriptions of videos. Specifically, we use sentence templates composed of Part-of-Speech (POS) tags to represent the syntax structure of captions, and accordingly, syntax representation learning is performed by directly inferring POS tags from videos. The visual perception is implemented by a mixture model which translates visual cues into lexical words that are conditional on the learned syntactic structure of sentences. Thus, a video captioning task consists of two sub-tasks: video POS tagging and visual cue translation, which are jointly modeled and trained in an end-to-end fashion. Evaluations on three public benchmark datasets demonstrate that our proposed method achieves substantially better performance than the state-of-the-art methods, which validates the superiority of joint modeling of syntax representation learning and visual perception for video captioning. Jingyi Hou, Xinxiao Wu, Wentian Zhao, Jiebo Luo 0001, Yunde Jia |
ICCV | 5 |
| 2019 | Tracker-Level Decision by Deep Reinforcement Learning for Robust Visual Tracking
Wenju Huang, Yuwei Wu 0001, Yunde Jia |
ICIG (1) | 3 |
| 2019 | Person-following for Telepresence Robots Using Web CamerasabstractMany existing mobile robotic telepresence systems have equipped with two web cameras, one is a forward-facing camera (FF camera) for video communication, and the other is a downward-facing camera (DF camera) for robot navigation. In this paper, we present a new framework of autonomous person-following for telepresence robots using the two web cameras. Based on correlation filters tracking methods, we use the FF camera to track the upper body of a person and the DF camera to localize and track the person's feet. We improve the robustness of feet trackers, consisting of a left foot tracker and a right foot tracker, by making full use of the spatial constraints of the human body parts. We conducted experiments on tracking in different environmental situations and real person-following scenario to evaluate the effectiveness of our method. Xianda Cheng, Yunde Jia, Jingyu Su, Yuwei Wu 0001 |
IROS | 2 |
| 2019 | Temporal Invariant Factor Disentangled Model for Representation Learning
Weichao Shen, Yuwei Wu 0001, Yunde Jia |
PRCV (2) | 3 |
| 2019 | Learning Weighted Video Segments for Temporal Action Localization
Che Sun, Hao Song 0002, Xinxiao Wu, Yunde Jia |
PRCV (1) | 4 |
| 2019 | Unsupervised deep quantization for object instance search
Yuwei Wu 0001, Chenchen Jing, Yunde Jia |
Neurocomputing | 5 |
| 2019 | Diffusion-based kernel matrix model for face liveness detection
Changyong Yu, Chengtang Yao, Mingtao Pei, Yunde Jia |
Image Vis. Comput. | 4 |
| 2019 | Deep convolutional network with locality and sparsity constraints for texture classification
Xingyuan Bu, Yuwei Wu 0001, Zhi Gao 0002, Yunde Jia |
Pattern Recognit. | 4 |
| 2019 | Learning a robust representation via a deep network on symmetric positive definite manifolds
Zhi Gao 0002, Yuwei Wu 0001, Xingyuan Bu, Junsong Yuan 0001, Yunde Jia |
Pattern Recognit. | 6 |
| 2019 | A deep Coarse-to-Fine network for head pose estimation from synthetic data
Wei Liang 0008, Jianbing Shen, Yunde Jia, Lap-Fai Yu |
Pattern Recognit. | 4 |
| 2019 | Robust Distracter-Resistive Tracker via Learning a Multi-Component Discriminative DictionaryabstractDiscriminative dictionary learning (DDL) provides an appealing paradigm for appearance modeling in visual tracking. However, most existing DDL-based trackers cannot handle drastic appearance changes, especially for scenarios with background cluster and/or similar object interference. One reason is that they often suffer from the loss of subtle visual information, which is critical to distinguish an object from distracters. In this paper, we explore the use of activations from the convolutional layer of a convolutional neural network to improve the object representation and then propose a robust distracter-resistive tracker via learning a multi-component discriminative dictionary. The proposed method exploits both the intra-class and inter-class visual information to learn shared atoms and the class-specific atoms. By imposing several constraints into the objective function, the learned dictionary is reconstructive, compressive, and discriminative, and thus can better distinguish an object from the background. In addition, our convolutional features have structural information for object localization and balance the discriminative power and semantic information of the object. Tracking is carried out within a Bayesian inference framework where a joint decision measure is used to construct the observation model. To alleviate the drift problem, the reliable tracking results obtained online are accumulated to update the dictionary. Both the qualitative and quantitative results on the CVPR2013 benchmark, the VOT2015 data set, and the SPOT data set demonstrate that our tracker achieves substantially better overall performance against the state-of-the-art approaches. Weichao Shen, Yuwei Wu 0001, Junsong Yuan 0001, Ling-Yu Duan, Jian Zhang 0002, Yunde Jia |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2019 | Heterogeneous Hashing Network for Face Retrieval Across Image and Video DomainsabstractIn this paper, we present a heterogeneous hashing network to generate effective and compact hash representations of both face images and face videos for face retrieval across image and video domains. The network contains an image branch and a video branch to project face images and videos into a common space, respectively. Then, the non-linear hash functions are learned in the common space to obtain the corresponding binary hash representations. The network is trained with three loss functions: 1) the Fisher loss; 2) the softmax loss; and 3) the triplet ranking loss. The Fisher loss uses the difference form of within-class and between-class scatter and is appropriate for the mini-batch-based optimization method. The Fisher loss together with the softmax loss is exploited to enhance the discriminative power of the common space. The triplet ranking loss is enforced on the final binary hash representations to improve retrieval performance. Experiments on a large-scale face video dataset and two challenging TV-series datasets demonstrate the effectiveness of the proposed method. Chenchen Jing, Zhen Dong 0002, Mingtao Pei, Yunde Jia |
IEEE Trans. Multim. | 4 |
| 2019 | Temporal Action Localization in Untrimmed Videos Using Action Pattern TreesabstractIn this paper, we present a novel framework of automatically localizing action instances based on action pattern trees (AP-Trees) in a long untrimmed video. For localizing action instances in videos with varied temporal lengths, we first split videos into sequential segments and then use the AP-Trees to produce precise temporal boundaries of action instances. The AP-Trees can exploit the temporal information between segments of videos based on the label vectors of segments, by learning the occurrence frequency and order of segments. In AP-Trees, nodes stand for action class labels of segments and edges represent the temporal relationships between two consecutive segments. Thus, we can discover the occurrence frequencies of segments by searching paths of AP-Trees. In order to obtain accurate labels of video segments, we introduce deep neural networks to annotate the segments by simultaneously leveraging the spatio-temporal information and the high-level semantic feature of segments. In the networks, informative action maps are generated by a global average pooling layer to retain the spatio-temporal information of segments. An overlap loss function is employed to further improve the precision of label vectors of segments by considering the temporal overlap between segments and the ground truth. The experiments on THUMOS2014, MSR ActionII, and MPII Cooking datasets demonstrate the effectiveness of the method. Hao Song 0002, Xinxiao Wu, Yuwei Wu 0001, Yunde Jia |
IEEE Trans. Multim. | 6 |
| 2018 | Unsupervised Deep Learning of Mid-Level Video Representation for Action RecognitionabstractCurrent deep learning methods for action recognition rely heavily on large scale labeled video datasets. Manually annotating video datasets is laborious and may introduce unexpected bias to train complex deep models for learning video representation. In this paper, we propose an unsupervised deep learning method which employs unlabeled local spatial-temporal volumes extracted from action videos to learn midlevel video representation for action recognition. Specifically, our method simultaneously discovers mid-level semantic concepts by discriminative clustering and optimizes local spatial-temporal features by two relatively small and simple deep neural networks. The clustering generates semantic visual concepts that guide the training of the deep networks, and the networks in turn guarantee the robustness of the semantic concepts. Experiments on the HMDB51 and the UCF101 datasets demonstrate the superiority of the proposed method, even over several supervised learning methods. Jingyi Hou, Xinxiao Wu, Jin Chen 0009, Jiebo Luo 0001, Yunde Jia |
AAAI | 5 |
| 2018 | Deep Stereo Matching With Explicit Cost Aggregation Sub-ArchitectureabstractDeep neural networks have shown excellent performance for stereo matching. Many efforts focus on the feature extraction and similarity measurement of the matching cost computation step while less attention is paid on cost aggregation which is crucial for stereo matching. In this paper, we present a learning-based cost aggregation method for stereo matching by a novel sub-architecture in the end-to-end trainable pipeline. We reformulate the cost aggregation as a learning process of the generation and selection of cost aggregation proposals which indicate the possible cost aggregation results. The cost aggregation sub-architecture is realized by a two-stream network: one for the generation of cost aggregation proposals, the other for the selection of the proposals. The criterion for the selection is determined by the low-level structure information obtained from a light convolutional network. The two-stream network offers a global view guidance for the cost aggregation to rectify the mismatching value stemming from the limited view of the matching cost computation. The comprehensive experiments on challenge datasets such as KITTI and Scene Flow show that our method outperforms the state-of-the-art methods. Lidong Yu, Yuwei Wu 0001, Yunde Jia |
AAAI | 4 |
| 2018 | Set-to-Set Distance Metric Learning on SPD Manifolds
Zhi Gao 0002, Yuwei Wu 0001, Yunde Jia |
PRCV (3) | 3 |
| 2018 | A discriminative structural model for joint segmentation and recognition of human actions
Cuiwei Liu, Jingyi Hou, Xinxiao Wu, Yunde Jia |
Multim. Tools Appl. | 4 |
| 2018 | Deep CNN based binary hash video representations for face retrieval
Zhen Dong 0002, Chenchen Jing, Mingtao Pei, Yunde Jia |
Pattern Recognit. | 4 |
| 2018 | Depth Super-Resolution on RGB-D Video Sequences With Large Displacement 3D MotionabstractTo enhance the resolution and accuracy of depth data, some video-based depth super-resolution methods have been proposed which utilizes its neighboring depth images in the temporal domain. They often consist of two main stages: motion compensation of temporally neighboring depth images and fusion of compensated depth images. However, large displacement 3D motion often leads to compensation error, and the compensation error is further introduced into the fusion. A video-based depth super-resolution method with novel motion compensation and fusion approaches is proposed in this paper. We claim that, 3D Nearest Neighboring Field (NNF) is a better choice than using positions with true motion displacement for depth enhancements. To handle large displacement 3D motion, the compensation stage utilized 3D NNF instead of true motion used in previous methods. Next, the fusion approach is modeled as a regression problem to predict the super-resolution result efficiently for each depth image by using its compensated depth images. A new deep convolutional neural network architecture is designed for fusion, which is able to employ a large amount of video data for learning the complicated regression function. We comprehensively evaluate our method on various RGB-D video sequences to show its superior performance. Yucheng Wang 0003, Jian Zhang 0002, Zicheng Liu 0001, Qiang Wu 0001, Zhengyou Zhang, Yunde Jia |
IEEE Trans. Image Process. | 6 |
| 2018 | Content-Attention Representation by Factorized Action-Scene Network for Action RecognitionabstractDuring action recognition in videos, irrelevant motions in the background can greatly degrade the performance of recognizing specific actions with which we actually concern ourself here. In this paper, a novel deep neural network, called factorized action-scene network (FASNet), is proposed to encode and fuse the most relevant and informative semantic cues for action recognition. Specifically, we decompose the FASNet into two components. One is a newly designed encoding network, named content attention network (CANet), which encodes local spatial-temporal features to learn the action representations with good robustness to the noise of irrelevant motions. The other is a fusion network, which integrates the pretrained CANet to fuse the encoded spatial-temporal features with contextual scene feature extracted from the same video, for learning more descriptive and discriminative action representations. Moreover, different from the existing deep learning based tasks for generic action recognition, which applies softmax loss function as the training guidance, we formulate two loss functions for guiding the proposed model to accomplish more specific action recognition tasks, i.e., the multilabel correlation loss for multilabel action recognition and the triplet loss for complex event detection. Extensive experiments on the Hollywood2 dataset and the TRECVID MEDTest 14 dataset show that our method achieves superior performance compared with the state-of-the-art methods. Jingyi Hou, Xinxiao Wu, Yuchao Sun, Yunde Jia |
IEEE Trans. Multim. | 4 |
| 2018 | Extracting Key Segments of Videos for Event Detection by Learning From Web SourcesabstractIn this paper, we present a novel approach of extracting the key segments for event detection in unconstrained videos. The key segments are automatically extracted by transferring the knowledge learned from Web images and Web videos to consumer videos. We propose an adaptive latent structural support vector machine model, where the locations of key segments in videos are regarded as latent variables due to the unavailability of the ground truth of key-segment locations in training data. In order to alleviate the time-consuming and labor-expensive manual annotation of huge amounts of training videos, a large number of loosely labeled Web images as well as videos are collected from the Web sources. Additionally, a limited number of labeled consumer videos are utilized to guarantee the precision of the model. Considering the semantic diversity of key segments, we learn a set of concepts as the semantic description of key segments and explore the temporal information of concepts to capture the sequential relations between the segments. The concepts are automatically discovered by using Web images and videos with their associated tags and description sentences. Comprehensive experiments on the Columbia's consumer video and the TRECVID 2014 Multimedia Event Detection datasets demonstrate that our method outperforms the state-of-the-art methods. Hao Song 0002, Xinxiao Wu, Wennan Yu, Yunde Jia |
IEEE Trans. Multim. | 4 |
| 2017 | Heterogeneous Multi-group Adaptation for Event Recognition in Consumer Videos
Mingyu Yao, Xinxiao Wu, Yunde Jia |
ICIG (1) | 4 |
| 2017 | Wide-angle image stitching using multi-homography warpingabstractFor wide-angle images with heavy lens distortion, especially those widely used in multi-camera systems, the single global homography cannot satisfy the required warp due to nonlinear distortions, leading to misalignment and shape distortion. This paper proposes a multi-homography warping to stitch wide-angle images with unknown lens distortion, which integrates multiple local homographies with a global homography for accurate alignment and shape preservation. We suggest a solution by conditional sampling to obtain a larger proportion of inliers for more accurate estimation of local and global projective transformations. We introduce an adaptive weighting scheme to combine these transformations for smoothing our warp over the entire target image from the local homographies to the global homography. The experiments evaluate the alignment accuracy and shape preservation of the proposed method. Bin Xu 0019, Yunde Jia |
ICIP | 2 |
| 2017 | Compact discriminative object representation via weakly supervised learning for real-time visual trackingabstractObject representations are of great importance for robust visual tracking. Although the high‐dimensional representation can effectively encode the input data with more information, exploiting it in a real‐time tracking system would be intractable and infeasible due to the high computational cost and memory requirements. In this study, the authors propose a compact discriminative object representation to achieve both good tracking accuracy and efficiency. An ensemble of weak training sets is generated based on the self‐representative ability of tracking samples, which is applied to learn discriminative functions. Each candidate is represented by the concatenation of project values on all the weak training sets. Tracking is then carried out within a Bayesian inference framework where the classification score of the support vector machine is used to construct the observation model. The evaluations on TB50 benchmark dataset demonstrate that the proposed algorithm is much more computationally efficient than the state‐of‐the‐art methods with comparable accuracy. Weichao Shen, Yuwei Wu 0001, Yunde Jia |
IET Comput. Vis. | 3 |
| 2017 | Heterogeneous domain adaptation method for video annotationabstractIn this study, the authors study the video annotation problem over heterogeneous domains, in which data from the image source domain and the video target domain is represented by heterogeneous features with different dimensions and physical meanings. A novel feature learning method, called heterogeneous discriminative analysis of canonical correlation (HDCC), is proposed to discover a common feature subspace in which heterogeneous features can be compared. The HDCC utilises discriminative information from the source domain as well as topology information from the target domain to learn two different projection matrices. By using these two matrices, heterogeneous data can be projected onto a common subspace and different features can be compared. They additionally design a group weighting learning framework for multi‐domain adaptation to effectively leverage knowledge learned from the source domain. Under this framework, source domain images are organised in groups according to their semantic meanings, and different weights are assigned to these groups according to their relevancies to the target domain videos. Extensive experiments on the Columbia Consumer Video and Kodak datasets demonstrate the effectiveness of their HDCC and group weighting methods. Xinxiao Wu, Yunde Jia |
IET Comput. Vis. | 3 |
| 2017 | Recognizing key segments of videos for video annotation by learning from web image sets
Hao Song 0002, Xinxiao Wu, Wei Liang 0008, Yunde Jia |
Multim. Tools Appl. | 4 |
| 2017 | A Hybrid Data Association Framework for Robust Online Multi-Object TrackingabstractGlobal optimization algorithms have shown impressive performance in data-association-based multi-object tracking, but handling online data remains a difficult hurdle to overcome. In this paper, we present a hybrid data association framework with a min-cost multi-commodity network flow for robust online multi-object tracking. We build local target-specific models interleaved with global optimization of the optimal data association over multiple video frames. More specifically, in the min-cost multi-commodity network flow, the target-specific similarities are online learned to enforce the local consistency for reducing the complexity of the global data association. Meanwhile, the global data association taking multiple video frames into account alleviates irrecoverable errors caused by the local data association between adjacent frames. To ensure the efficiency of online tracking, we give an efficient near-optimal solution to the proposed min-cost multi-commodity flow problem, and provide the empirical proof of its sub-optimality. The comprehensive experiments on real data demonstrate the superior tracking performance of our approach in various challenging situations. Min Yang 0003, Yuwei Wu 0001, Yunde Jia |
IEEE Trans. Image Process. | 3 |
| 2016 | Multimedia event detection via deep spatial-temporal neural networksabstractThis paper proposes a novel method using deep spatial-temporal neural networks based on deep Convolutional Neural Network (CNN) for multimedia event detection. To sufficiently take advantage of the motion and appearance information of events from videos, our networks contain two branches: a temporal neural network and a spatial neural network. The temporal neural network captures motion information by Recurrent Neural Networks with the mutation of gated recurrent unit. The spatial neural network catches object information by using the deep CNN, to encode the CNN features as a bag of semantics with more discriminative representations. Both the temporal and spatial features are beneficial for event detection in a fully coupled way. Finally, we employ the generalized multiple kernel learning method to effectively fuse these two types of heterogeneous and complementary features for action recognition. Experiments on TRECVID MEDTest 14 dataset show that our method achieves better performance than the state of the art. Jingyi Hou, Xinxiao Wu, Feiwu Yu, Yunde Jia |
ICME | 4 |
| 2016 | Attention Estimation for Input Switch in Scalable Multi-display Environments
Xingyuan Bu, Mingtao Pei, Yunde Jia |
ICONIP (4) | 3 |
| 2016 | A low-cost tele-presence wheelchair systemabstractThis paper presents the architecture and implementation of a tele-presence wheelchair system based on tele-presence robot, intelligent wheelchair, and touch screen technologies. The tele-presence wheelchair system consists of a commercial electric wheelchair, an add-on tele-presence interaction module, and a touchable live video image based user interface (called TIUI). The tele-presence interaction module is used to provide video-chatting for an elderly or disabled person with the family members or caregivers, and also captures the live video of an environment for tele-operation and semi-autonomous navigation. The user interface developed in our lab allows an operator to access the system anywhere and directly touch the live video image of the wheelchair to push it as if he/she did it in the presence. This paper also discusses the evaluation of the user experience. Bin Xu 0019, Mingtao Pei, Yunde Jia |
IROS | 4 |
| 2016 | Nonnegative correlation coding for image classification
Zhen Dong 0002, Wei Liang 0008, Yuwei Wu 0001, Mingtao Pei, Yunde Jia |
Sci. China Inf. Sci. | 5 |
| 2016 | Temporal dynamic appearance modeling for online multi-person tracking
Min Yang 0003, Yunde Jia |
Comput. Vis. Image Underst. | 2 |
| 2016 | Multi-group-multi-class domain adaptation for event recognitionabstractIn this study, the authors propose a multi‐group–multi‐class domain adaptation framework to recognise events in consumer videos by leveraging a large number of web videos. The authors’ framework is extended from multi‐class support vector machine by adding a novel data‐dependent regulariser, which can force the event classifier to become consistent in consumer videos. To obtain web videos, they search them using several event‐related keywords and refer the videos returned by one keyword search as a group. They also leverage a video representation which is the average of convolutional neural networks features of the video frames for better performance. Comprehensive experiments on the two real‐world consumer video datasets demonstrate the effectiveness of their method for event recognition in consumer videos. Xinxiao Wu, Yunde Jia |
IET Comput. Vis. | 3 |
| 2016 | A Hierarchical Video Description for Complex Activity Understanding
Cuiwei Liu, Xinxiao Wu, Yunde Jia |
Int. J. Comput. Vis. | 3 |
| 2016 | Heterogeneous discriminant analysis for cross-view action recognition
Wanchen Sui, Xinxiao Wu, Yunde Jia |
Neurocomputing | 4 |
| 2016 | Orthonormal dictionary learning and its application to face recognition
Zhen Dong 0002, Mingtao Pei, Yunde Jia |
Image Vis. Comput. | 3 |
| 2016 | Go-ICP: A Globally Optimal Solution to 3D ICP Point-Set RegistrationabstractThe Iterative Closest Point (ICP) algorithm is one of the most widely used methods for point-set registration. However, being based on local iterative optimization, ICP is known to be susceptible to local minima. Its performance critically relies on the quality of the initialization and only local optimality is guaranteed. This paper presents the first globally optimal algorithm, named Go-ICP, for Euclidean (rigid) registration of two 3D point-sets under the$L_2$error metric defined in ICP. The Go-ICP method is based on a branch-and-bound scheme that searches the entire 3D motion space$SE(3)$. By exploiting the special structure of$SE(3)$geometry, we derive novel upper and lower bounds for the registration error function. Local ICP is integrated into the BnB scheme, which speeds up the new method while guaranteeing global optimality. We also discuss extensions, addressing the issue of outlier robustness. The evaluation demonstrates that the proposed method is able to produce reliable registration results regardless of the initialization. Go-ICP can be applied in scenarios where an optimal solution is desirable or where a good initialization is not always available. Jiaolong Yang, Hongdong Li, Dylan Campbell, Yunde Jia |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2016 | Handling Occlusion and Large Displacement Through Improved RGB-D Scene Flow EstimationabstractThe accuracy of scene flow is restricted by several challenges such as occlusion and large displacement motion. When occlusion happens, the positions inside the occluded regions lose their corresponding counterparts in preceding and succeeding frames. Large displacement motion will increase the complexity of motion modeling and computation. Moreover, occlusion and large displacement motion are highly related problems in scene flow estimation, e.g., large displacement motion often leads to considerably occluded regions in the scene. An improved dense scene flow method based on red-green-blue-depth (RGB-D) data is proposed in this paper. To handle occlusion, we model the occlusion status for each point in our problem formulation, and jointly estimate the scene flow and occluded regions. To deal with large displacement motion, we employ an over-parameterized scene flow representation to model both the rotation and translation components of the scene flow, since large displacement motion cannot be well approximated using translational motion only. Furthermore, we employ a two-stage optimization procedure for this overparameterized scene flow representation. In the first stage, we propose a new RGB-D PatchMatch method, which is mainly applied in the RGB-D image space to reduce the computational complexity introduced by the large displacement motion. According to the quantitative evaluation based on the Middlebury data set, our method outperforms other published methods. The improved performance is also comprehensively confirmed on the real data acquired by Kinect sensor. Yucheng Wang 0003, Jian Zhang 0002, Zicheng Liu 0001, Qiang Wu 0001, Philip A. Chou, Zhengyou Zhang, Yunde Jia |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2016 | Online Discriminative Tracking With Active Example SelectionabstractMost existing discriminative tracking algorithms use a sampling-and-labeling strategy to collect examples and treat the training example collection as a task that is independent of classifier learning. However, the examples collected directly by sampling are neither necessarily informative nor intended to be useful for classifier learning. Updating the classifier with these examples might introduce ambiguity to the tracker. In this paper, we present a novel online discriminative tracking framework that explicitly couples the objectives of example collection and classifier learning. Our method uses Laplacian regularized least squares (LapRLS) to learn a robust classifier that can sufficiently exploit unlabeled data and preserve the local geometrical structure of the feature space. To ensure the high classification confidence of the classifier, we propose an active example selection approach to automatically select the most informative examples for LapRLS. Part of the selected examples that satisfy strict constraints are labeled to enhance the adaptivity of our tracker, which actually provides robust supervisory information to guide semisupervised learning. With active example selection, we are able to avoid the ambiguity introduced by an independent example collection strategy and to alleviate the drift problem caused by misaligned examples. Comparison with the state-of-the-art trackers on the comprehensive benchmark demonstrates that our tracking algorithm is more effective and accurate. Min Yang 0003, Yuwei Wu 0001, Mingtao Pei, Bo Ma 0001, Yunde Jia |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2016 | Transfer Latent SVM for Joint Recognition and Localization of Actions in VideosabstractIn this paper, we develop a novel transfer latent support vector machine for joint recognition and localization of actions by using Web images and weakly annotated training videos. The model takes training videos which are only annotated with action labels as input for alleviating the laborious and time-consuming manual annotations of action locations. Since the ground-truth of action locations in videos are not available, the locations are modeled as latent variables in our method and are inferred during both training and testing phrases. For the purpose of improving the localization accuracy with some prior information of action locations, we collect a number of Web images which are annotated with both action labels and action locations to learn a discriminative model by enforcing the local similarities between videos and Web images. A structural transformation based on randomized clustering forest is used to map the Web images to videos for handling the heterogeneous features of Web images and videos. Experiments on two public action datasets demonstrate the effectiveness of the proposed model for both action localization and action recognition. Cuiwei Liu, Xinxiao Wu, Yunde Jia |
IEEE Trans. Cybern. | 3 |
| 2015 | Discriminative Orthonormal Dictionary Learning for Fast Low-Rank Representation
Zhen Dong 0002, Mingtao Pei, Yunde Jia |
ICONIP (1) | 3 |
| 2015 | Heterogeneous Discriminant Analysis for Cross-View Action Recognition
Wanchen Sui, Xinxiao Wu, Wei Liang 0008, Yunde Jia |
ICONIP (4) | 5 |
| 2015 | Robust Online Multi-object Tracking by Maximum a Posteriori Estimation with Sequential Trajectory Prior
Min Yang 0003, Mingtao Pei, Yunde Jia |
ICONIP (1) | 4 |
| 2015 | Inferring User Preference in Good Abandonment from Eye Movements
Wanxuan Lu, Yunde Jia |
WAIM | 2 |
| 2015 | Learning online structural appearance model for robust object tracking
Min Yang 0003, Mingtao Pei, Yuwei Wu 0001, Yunde Jia |
Sci. China Inf. Sci. | 4 |
| 2015 | Online visual tracking by integrating spatio-temporal cuesabstractThe performance of online visual trackers has improved significantly, but designing an effective appearance‐adaptive model is still a challenging task because of the accumulation of errors during the model updating with newly obtained results, which will cause tracker drift. In this study, the authors propose a novel online tracking algorithm by integrating spatio‐temporal cues to alleviate the drift problem. The authors' goal is to develop a more robust way of updating an adaptive appearance model. The model consists of multiple modules called temporal cues, and these modules are updated in an alternate way which can keep both the historical and current information of the tracked object to handle drastic appearance change. Each module is represented by several fragments called spatial cues. In order to incorporate all the spatial and temporal cues, the authors develop an efficient cue quality evaluation criterion that combines appearance and motion information. Then the tracking results are obtained by a two‐stage dynamic integration mechanism. Both qualitative and quantitative evaluations on challenging video sequences demonstrate that the proposed algorithm performs more favourably against the state‐of‐the‐art methods. Yang He 0004, Mingtao Pei, Min Yang 0003, Yuwei Wu 0001, Yunde Jia |
IET Comput. Vis. | 5 |
| 2015 | Multi-task l0 gradient minimization for visual tracking
Hongwei Hu, Bo Ma 0001, Yunde Jia |
Neurocomputing | 3 |
| 2015 | Cross-domain structural model for video event annotation via web images
Xiabi Liu, Xinxiao Wu, Yunde Jia |
Multim. Tools Appl. | 4 |
| 2015 | Differential tracking with a kernel-based region covariance descriptor
Yuwei Wu 0001, Bo Ma 0001, Yunde Jia |
Pattern Anal. Appl. | 3 |
| 2015 | Robust Match Fusion Using OptimizationabstractIn this paper, we present a novel patch-based match and fusion algorithm by taking account of moving scene in a multiple exposure image sequence using optimization. A uniform iterative approach is developed to match and find the corresponding patches in different exposure images, which are then fused in each iteration. Our approach does not need to align the input multiple exposure images before the fusion process. Considering that the pixel values are affected by various exposure time, we design a new patch-based energy function that will be optimized to improve the matching accuracy. An efficient patch-based exposure fusion approach using the random walker algorithm is developed to preserve the moving objects from the input multiple exposure images. To the best of our knowledge, our algorithm is the first patch-based exposure fusion work to preserve the moving objects of dynamic scenes that does not need the registration process of different exposure images. Experimental results of moving scenes demonstrate that our algorithm achieves visually pleasing fusion results without ghosting artifacts, while the results produced by the state-of-the-art exposure fusion and tone mapping algorithms exhibit different levels of ghosting artifacts. Xiameng Qin, Jianbing Shen, Xiaoyang Mao, Xuelong Li 0001, Yunde Jia |
IEEE Trans. Cybern. | 5 |
| 2015 | Manifold Kernel Sparse Representation of Symmetric Positive-Definite Matrices and Its ApplicationsabstractThe symmetric positive-definite (SPD) matrix, as a connected Riemannian manifold, has become increasingly popular for encoding image information. Most existing sparse models are still primarily developed in the Euclidean space. They do not consider the non-linear geometrical structure of the data space, and thus are not directly applicable to the Riemannian manifold. In this paper, we propose a novel sparse representation method of SPD matrices in the data-dependent manifold kernel space. The graph Laplacian is incorporated into the kernel space to better reflect the underlying geometry of SPD matrices. Under the proposed framework, we design two different positive definite kernel functions that can be readily transformed to the corresponding manifold kernels. The sparse representation obtained has more discriminating power. Extensive experimental results demonstrate good performance of manifold kernel sparse codes in image classification, face recognition, and visual tracking. Yuwei Wu 0001, Yunde Jia, Peihua Li, Jian Zhang 0002, Junsong Yuan 0001 |
IEEE Trans. Image Process. | 2 |
| 2015 | Robust Discriminative Tracking via Landmark-Based Label PropagationabstractThe appearance of an object could be continuously changing during tracking, thereby being not independent identically distributed. A good discriminative tracker often needs a large number of training samples to fit the underlying data distribution, which is impractical for visual tracking. In this paper, we present a new discriminative tracker via landmark-based label propagation (LLP) that is nonparametric and makes no specific assumption about the sample distribution. With an undirected graph representation of samples, the LLP locally approximates the soft label of each sample by a linear combination of labels on its nearby landmarks. It is able to effectively propagate a limited amount of initial labels to a large amount of unlabeled samples. To this end, we introduce a local landmarks approximation method to compute the cross-similarity matrix between the whole data and landmarks. Moreover, a soft label prediction function incorporating the graph Laplacian regularizer is used to diffuse the known labels to all the unlabeled vertices in the graph, which explicitly considers the local geometrical structure of all samples. Tracking is then carried out within a Bayesian inference framework, where the soft label prediction value is used to construct the observation model. Both qualitative and quantitative evaluations on the benchmark data set containing 51 challenging image sequences demonstrate that the proposed algorithm outperforms the state-of-the-art methods. Yuwei Wu 0001, Mingtao Pei, Min Yang 0003, Junsong Yuan 0001, Yunde Jia |
IEEE Trans. Image Process. | 5 |
| 2015 | Cross-View Action Recognition Over Heterogeneous Feature SpacesabstractIn cross-view action recognition, what you saw in one view is different from what you recognize in another view, since the data distribution even the feature space can change from one view to another. In this paper, we address the problem of transferring action models learned in one view (source view) to another different view (target view), where action instances from these two views are represented by heterogeneous features. A novel learning method, called heterogeneous transfer discriminant-analysis of canonical correlations (HTDCC), is proposed to discover a discriminative common feature space for linking source view and target view to transfer knowledge between them. Two projection matrices are learned to, respectively, map data from the source view and the target view into a common feature space via simultaneously minimizing the canonical correlations of interclass training data, maximizing the canonical correlations of intraclass training data, and reducing the data distribution mismatch between the source and target views in the common feature space. In our method, the source view and the target view neither share any common features nor have any corresponding action instances. Moreover, our HTDCC method is capable of handling only a few or even no labeled samples available in the target view, and can also be easily extended to the situation of multiple source views. We additionally propose a weighting learning framework for multiple source views adaptation to effectively leverage action knowledge learned from multiple source views for the recognition task in the target view. Under this framework, different source views are assigned different weights according to their different relevances to the target view. Each weight represents how contributive the corresponding source view is to the target view. Extensive experiments on the IXMAS data set demonstrate the effectiveness of HTDCC on learning the common feature space for heterogeneous cross-view action recognition. In addition, the weighting learning framework can achieve promising results on automatically adapting multiple transferred source-view knowledge to the target view. Xinxiao Wu, Han Wang 0001, Cuiwei Liu, Yunde Jia |
IEEE Trans. Image Process. | 4 |
| 2015 | Vehicle Type Classification Using a Semisupervised Convolutional Neural NetworkabstractIn this paper, we propose a vehicle type classification method using a semisupervised convolutional neural network from vehicle frontal-view images. In order to capture rich and discriminative information of vehicles, we introduce sparse Laplacian filter learning to obtain the filters of the network with large amounts of unlabeled data. Serving as the output layer of the network, the softmax classifier is trained by multitask learning with small amounts of labeled data. For a given vehicle image, the network can provide the probability of each type to which the vehicle belongs. Unlike traditional methods by using handcrafted visual features, our method is able to automatically learn good features for the classification task. The learned features are discriminative enough to work well in complex scenes. We build the challenging BIT-Vehicle dataset, including 9850 high-resolution vehicle frontal-view images. Experimental results on our own dataset and a public dataset demonstrate the effectiveness of the proposed method. Zhen Dong 0002, Yuwei Wu 0001, Mingtao Pei, Yunde Jia |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2015 | Structured-Patch Optimization for Dense CorrespondenceabstractThis paper presents a new method to compute the dense correspondences between two images by using the energy optimization and the structured patches. In terms of the property of the sparse feature and the principle that nearest sub-scenes and neighbors are much more similar, we design a new energy optimization to guide the dense matching process and find the reliable correspondences. The sparse features are also employed to design a new structure to describe the patches. Both transformation and deformation with the structured patches are considered and incorporated into an energy optimization framework. Thus, our algorithm can match the objects robustly in complicated scenes. Finally, a local refinement technique is proposed to solve the perturbation of the matched patches. Experimental results demonstrate that our method outperforms the state-of-the-art matching algorithms. Xiameng Qin, Jianbing Shen, Xiaoyang Mao, Xuelong Li 0001, Yunde Jia |
IEEE Trans. Multim. | 5 |
| 2014 | Weakly Supervised Action Recognition and Localization Using Web Images
Cuiwei Liu, Xinxiao Wu, Yunde Jia |
ACCV (5) | 3 |
| 2014 | Video Annotation by Incremental Learning from Grouped Heterogeneous Sources
Hao Song 0002, Xinxiao Wu, Yunde Jia |
ACCV (5) | 4 |
| 2014 | Landmark-Based Inductive Model for Robust Discriminative Tracking
Yuwei Wu 0001, Mingtao Pei, Min Yang 0003, Yang He 0004, Yunde Jia |
ACCV (5) | 5 |
| 2014 | Coupling Semi-supervised Learning and Example Selection for Online Object Tracking
Min Yang 0003, Yuwei Wu 0001, Mingtao Pei, Bo Ma 0001, Yunde Jia |
ACCV (4) | 5 |
| 2014 | Optimal Essential Matrix Estimation via Inlier-Set Maximization
Jiaolong Yang, Hongdong Li, Yunde Jia |
ECCV (1) | 3 |
| 2014 | A new sparse feature-based patch for dense correspondenceabstractThis paper presents a new method to compute the dense correspondences between two images by using the sparse feature-based patches in an energy optimization framework. Many transformation and deformation cues such as color, scale and rotation should be considered when we finding dense correspondences between images. However, most existing methods only consider part of these transformations, which will introduce the uncorrect correspondence results. In terms of the property of the sparse feature and the principle that nearest sub-scenes and neighbors are much more similar, we design a new energy optimization to guide the dense matching process. Both transformation and deformation are considered in our energy optimization framework since we design the feature-based patches. Thus, our algorithm can match the complicated scenes and objects robustly. At last, a local refinement technique is proposed to solve the perturbation of the matched patches. Experimental results demonstrate that our method outperforms the state-of-the-art algorithms. Xiameng Qin, Jianbing Shen, Xuelong Li 0001, Yunde Jia |
ICME | 4 |
| 2014 | Vehicle Type Classification Using Unsupervised Convolutional Neural NetworkabstractIn this paper, we propose an appearance-based vehicle type classification method from vehicle frontal view images. Unlike other methods using hand-crafted visual features, our method is able to automatically learn good features for vehicle type classification by using a convolutional neural network. In order to capture rich and discriminative information of vehicles, the network is pre-trained by the sparse filtering which is an unsupervised learning method. Besides, the network is with layer-skipping to ensure that final features contain both high-level global and low-level local features. After the final features are obtained, the soft max regression is used to classify vehicle types. We build a challenging vehicle dataset called BIT-Vehicle dataset to evaluate the performance of our method. Experimental results on a public dataset and our own dataset demonstrate that our method is quite effective in classifying vehicle types. Zhen Dong 0002, Mingtao Pei, Yang He 0004, Yanmei Dong, Yunde Jia |
ICPR | 6 |
| 2014 | Visual Tracking Using Multi-stage Random Simple FeaturesabstractIn recent years, deep models offer a promising solution to extract powerful features. Motivated by the effectiveness of the Convolutional Networks (ConvNets) model in image classification and object detection, we present a visual tracking algorithm using the ConvNets model to extract multistage features. The key point of this paper is to show that the multi-stage features extracted by the ConvNets are very proper for visual tracking. In addition, we design a general procedure to generate simple rectangle filters with different complexity, and employ the rectangle filters to construct the ConvNets. The computational cost is reduced by using integral images. The filters are kept constant, thus the update of our tracker would not cost much time. The tracking is formulated as a binary classification problem, and we use an online naive Bayes classifier to build our tracker. The experimental results demonstrate that our tracker achieves comparable results against several state-of-the-art methods. Yang He 0004, Zhen Dong 0002, Min Yang 0003, Lei Chen 0020, Mingtao Pei, Yunde Jia |
ICPR | 6 |
| 2014 | Depth Super-resolution by Fusing Depth Imaging and Stereo Vision with Structural Determinant Information InferenceabstractIn this paper, we present a depth super-resolution framework by fusing depth imaging and stereo vision for high-resolution and high-accuracy depth maps. Depth cameras and stereo vision have their own limitations in some aspects, but their characteristics of range sensing are complementary. Thus, combining both approaches can produce more satisfactory results than either one. Unlike previous fusion methods, we initially taking the noisy depth observation from depth camera as prior information of scene structure. The prior information of scene structure is also utilized to infer structural determinant information, like depth discontinuity and occlusion, which is essential to improve the quality of depth map in the fusion process. In succession, the prior knowledge helps to overcome difficulties of intensity inconsistency in image observation from stereo vision component. Experimental results demonstrate effectiveness and accuracy of the proposed method. Yucheng Wang 0003, Huijun Di, Wei Liang 0008, Jian Zhang 0002, Yunde Jia |
ICPR | 6 |
| 2014 | An Eye-Tracking Study of User Behavior in Web Image Search
Wanxuan Lu, Yunde Jia |
PRICAI | 2 |
| 2014 | Learning a discriminative mid-level feature for action recognition
Cuiwei Liu, Mingtao Pei, Xinxiao Wu, Yu Kong 0001, Yunde Jia |
Sci. China Inf. Sci. | 5 |
| 2014 | Recognising human interaction from videos by a discriminative modelabstractThis study addresses the problem of recognising human interactions between two people. The main difficulties lie in the partial occlusion of body parts and the motion ambiguity in interactions. The authors observed that the interdependencies existing at both the action level and the body part level can greatly help disambiguate similar individual movements and facilitate human interaction recognition. Accordingly, they proposed a novel discriminative method, which model the action of each person by a large‐scale global feature and local body part features, to capture such interdependencies for recognising interaction of two people. A variant of multi‐class Adaboost method is proposed to automatically discover class‐specific discriminative three‐dimensional body parts. The proposed approach is tested on the authors newly introduced BIT‐interaction dataset and the UT‐interaction dataset. The results show that their proposed model is quite effective in recognising human interactions. Yu Kong 0001, Wei Liang 0008, Zhen Dong 0002, Yunde Jia |
IET Comput. Vis. | 4 |
| 2014 | Interactive Phrases: Semantic Descriptionsfor Human Interaction RecognitionabstractThis paper addresses the problem of recognizing human interactions from videos. We propose a novel approach that recognizes human interactions by the learned high-level descriptions, interactive phrases. Interactive phrases describe motion relationships between interacting people. These phrases naturally exploit human knowledge and allow us to construct a more descriptive model for recognizing human interactions. We propose a discriminative model to encode interactive phrases based on the latent SVM formulation. Interactive phrases are treated as latent variables and are used as mid-level features. To complement manually specified interactive phrases, we also discover data-driven phrases from data in order to find potentially useful and discriminative phrases for differentiating human interactions. An information-theoretic approach is employed to learn the data-driven phrases. The interdependencies between interactive phrases are explicitly captured in the model to deal with motion ambiguity and partial occlusion in the interactions. We evaluate our method on the BIT-Interaction data set, UT-Interaction data set, and Collective Activity data set. Experimental results show that our approach achieves superior performance over previous approaches. Yu Kong 0001, Yunde Jia, Yun Fu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2014 | Metric Learning Based Structural Appearance Model for Robust Visual TrackingabstractAppearance modeling is a key issue for the success of a visual tracker. Sparse representation based appearance modeling has received an increasing amount of interest in recent years. However, most of existing work utilizes reconstruction errors to compute the observation likelihood under the generative framework, which may give poor performance, especially for significant appearance variations. In this paper, we advocate an approach to visual tracking that seeks an appropriate metric in the feature space of sparse codes and propose a metric learning based structural appearance model for more accurate matching of different appearances. This structural representation is acquired by performing multiscale max pooling on the weighted local sparse codes of image patches. An online multiple instance metric learning algorithm is proposed that learns a discriminative and adaptive metric, thereby better distinguishing the visual object of interest from the background. The multiple instance setting is able to alleviate the drift problem potentially caused by misaligned training examples. Tracking is then carried out within a Bayesian inference framework, in which the learned metric and the structure object representation are used to construct the observation model. Comprehensive experiments on challenging image sequences demonstrate qualitatively and quantitatively that the proposed algorithm outperforms the state-of-the-art methods. Yuwei Wu 0001, Bo Ma 0001, Min Yang 0003, Jian Zhang 0002, Yunde Jia |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2014 | Video Annotation via Image Groups from the WebabstractSearching desirable events in uncontrolled videos is a challenging task. Current researches mainly focus on obtaining concepts from numerous labeled videos. But it is time consuming and labor expensive to collect a large amount of required labeled videos for training event models under various circumstances. To alleviate this problem, we propose to leverage abundant Web images for videos since Web images contain a rich source of information with many events roughly annotated and taken under various conditions. However, knowledge from the Web is noisy and diverse, brute force knowledge transfer of images may hurt the video annotation performance. Therefore, we propose a novel Group-based Domain Adaptation (GDA) learning framework to leverage different groups of knowledge (source domain) queried from the Web image search engine to consumer videos (target domain). Different from traditional methods using multiple source domains of images, our method organizes the Web images according to their intrinsic semantic relationships instead of their sources. Specifically, two different types of groups (i.e., event-specific groups and concept-specific groups) are exploited to respectively describe the event-level and concept-level semantic meanings of target-domain videos. Under this framework, we assign different weights to different image groups according to the relevances between the source groups and the target domain, and each group weight represents how contributive the corresponding source image group is to the knowledge transferred to the target video. In order to make the group weights and group classifiers mutually beneficial and reciprocal, a joint optimization algorithm is presented for simultaneously learning the weights and classifiers, using two novel data-dependent regularizers. Experimental results on three challenging video datasets (i.e., CCV, Kodak, and YouTube) demonstrate the effectiveness of leveraging grouped knowledge gained from Web images for video annotation. Xinxiao Wu, Yunde Jia |
IEEE Trans. Multim. | 3 |
| 2013 | Discriminatively Trained And-Or Tree Models for Object DetectionabstractThis paper presents a method of learning reconfigurable And-Or Tree (AOT) models discriminatively from weakly annotated data for object detection. To explore the appearance and geometry space of latent structures effectively, we first quantize the image lattice using an over complete set of shape primitives, and then organize them into a directed a cyclic And-Or Graph (AOG) by exploiting their compositional relations. We allow overlaps between child nodes when combining them into a parent node, which is equivalent to introducing an appearance Or-node implicitly for the overlapped portion. The learning of an AOT model consists of three components: (i) Unsupervised sub-category learning (i.e., branches of an object Or-node) with the latent structures in AOG being integrated out. (ii) Weakly supervised part configuration learning (i.e., seeking the globally optimal parse trees in AOG for each sub-category). To search the globally optimal parse tree in AOG efficiently, we propose a dynamic programming (DP) algorithm. (iii) Joint appearance and structural parameters training under latent structural SVM framework. In experiments, our method is tested on PASCAL VOC 2007 and 2010 detection benchmarks of 20 object classes and outperforms comparable state-of-the-art methods. Tianfu Wu 0001, Yunde Jia, Song-Chun Zhu |
CVPR | 3 |
| 2013 | Voice activity detection using convolutive non-negative sparse codingabstractThis paper presents a voice activity detection (VAD) approach using convolutive non-negative sparse coding (CNSC) to improve the detection performance in low signal-to-noise (SNR) conditions. Our idea is to use noise-robust feature for speech signal detection while noise is reduced away. We first use magnitude spectrum as the non-negative and additive low-level representation of audio signals, and learn a speech dictionary from clean speech as well as a noise dictionary from noise samples. Then, the two dictionaries are concatenated to form a global dictionary, and an audio signal is decomposed into coefficient vectors using CNSC on the global dictionary. Only coefficients corresponding to the bases from the speech dictionary are taken as the features for the signal. At last, the activity labels is given by decoding a conditional random field (CRF) which is constructed to model the context of an audio signal for VAD. Experiments demonstrate that our VAD approach has an excellent performance in low SNR conditions. Peng Teng, Yunde Jia |
ICASSP | 2 |
| 2013 | Cross-View Action Recognition over Heterogeneous Feature SpacesabstractIn cross-view action recognition, "what you saw" in one view is different from "what you recognize" in another view. The data distribution even the feature space can change from one view to another due to the appearance and motion of actions drastically vary across different views. In this paper, we address the problem of transferring action models learned in one view (source view) to another different view (target view), where action instances from these two views are represented by heterogeneous features. A novel learning method, called Heterogeneous Transfer Discriminantanalysis of Canonical Correlations (HTDCC), is proposed to learn a discriminative common feature space for linking source and target views to transfer knowledge between them. Two projection matrices that respectively map data from source and target views into the common space are optimized via simultaneously minimizing the canonical correlations of inter-class samples and maximizing the intraclass canonical correlations. Our model is neither restricted to corresponding action instances in the two views nor restricted to the same type of feature, and can handle only a few or even no labeled samples available in the target view. To reduce the data distribution mismatch between the source and target views in the common feature space, a nonparametric criterion is included in the objective function. We additionally propose a joint weight learning method to fuse multiple source-view action classifiers for recognition in the target view. Different combination weights are assigned to different source views, with each weight presenting how contributive the corresponding source view is to the target view. The proposed method is evaluated on the IXMAS multi-view dataset and achieves promising results. Xinxiao Wu, Han Wang 0001, Cuiwei Liu, Yunde Jia |
ICCV | 4 |
| 2013 | Go-ICP: Solving 3D Registration Efficiently and Globally OptimallyabstractRegistration is a fundamental task in computer vision. The Iterative Closest Point (ICP) algorithm is one of the widely-used methods for solving the registration problem. Based on local iteration, ICP is however well-known to suffer from local minima. Its performance critically relies on the quality of initialization, and only local optimality is guaranteed. This paper provides the very first globally optimal solution to Euclidean registration of two 3D point sets or two 3D surfaces under the L2 error. Our method is built upon ICP, but combines it with a branch-and-bound (BnB) scheme which searches the 3D motion space SE(3) efficiently. By exploiting the special structure of the underlying geometry, we derive novel upper and lower bounds for the ICP error function. The integration of local ICP and global BnB enables the new method to run efficiently in practice, and its optimality is exactly guaranteed. We also discuss extensions, addressing the issue of outlier robustness. Jiaolong Yang, Hongdong Li, Yunde Jia |
ICCV | 3 |
| 2013 | Vehicle type classification using distributions of structural and appearance-based featuresabstractClassifying vehicle types from an image is a challenging task due to various light conditions and background interferences. We present a vehicle type classification method using structural and appearance-based features in this paper. The structural feature which characterizes the spatial relative layouts of vehicle parts helps to distinguish a vehicle from the background, and the appearance-based feature is local and robust to the interferences of illumination variation and the background. To obtain compact and discriminative representations of vehicles, the distributions of these two types of features are computed. We further employ Multiple Kernel Learning to combine multiple distributions together for classifying vehicle types. Experimental results demonstrate the effectiveness of our method. Yunde Jia |
ICIP | 2 |
| 2013 | Structural information processing in early vision using Hebbian-based mean shiftabstractThis paper presents a biologically plausible model for the structural information processing in early vision. Our investigation on the frequency spectrum of natural images filtered by the retina shows that the DC component containing much redundancy information and the high frequency components containing much noisy information are reduced, while the middle and low frequency components containing much structural information are enhanced. Simple cells in the primary visual cortex (V1) extract structural primitives from the filtered signals resulting in the emergence of diverse receptive field shapes. We name these structural primitives as structors, and study the neural mechanisms responsible for this diversity of V1 simple cell receptive field shapes. Sparse coding with the L0-norm constraint is reexamined which suggests that the local structure of natural images is determined by few structors regardless of their coefficients. We perform an analysis on the spatial distribution of the input signal and prove that signals in the neighborhood of a special structor has a star shape and peaks at the structor. That is, the structors are the modes of the probability density function of the input signal, and learning the structors can be interpreted as mode detection. Mean sift method is applied to detect modes, and the updating rule for the mean shift appears to be Hebbian. We propose the Hebbian-based mean shift to simulate the emergence of the diversity of simple cell receptive field shapes. The simulation results demonstrate the robustness of the proposed algorithm in producing both Gabor-like and blob-like structors. Jiqian Liu, Yunde Jia |
IJCNN | 3 |
| 2013 | Single-shot extrinsic calibration of a generically configured RGB-D camera rig from scene constraintsabstractWith the increasing use of commodity RGB-D cameras for computer vision, robotics, mixed and augmented reality and other areas, it is of significant practical interest to calibrate the relative pose between a depth (D) camera and an RGB camera in these types of setups. In this paper, we propose a new single-shot, correspondence-free method to extrinsically calibrate a generically configured RGB-D camera rig. We formulate the extrinsic calibration problem as one of geometric 2D-3D registration which exploits scene constraints to achieve single-shot extrinsic calibration. Our method first reconstructs sparse point clouds from a single-view 2D image. These sparse point clouds are then registered with dense point clouds from the depth camera. Finally, we directly optimize the warping quality by evaluating scene constraints in 3D point clouds. Our single-shot extrinsic calibration method does not require correspondences across multiple color images or across different modalities and it is more flexible than existing methods. The scene constraints can be very simple and we demonstrate that a scene containing three sheets of paper is sufficient to obtain reliable calibration and with a lower geometric error than existing methods. Jiaolong Yang, Yuchao Dai, Hongdong Li, Henry J. Gardner, Yunde Jia |
ISMAR | 5 |
| 2013 | Segmentation of the left ventricle in cardiac cine MRI using a shape-constrained snake model
Yuwei Wu 0001, Yuanquan Wang 0001, Yunde Jia |
Comput. Vis. Image Underst. | 3 |
| 2013 | Adaptive diffusion flow active contours for image segmentation
Yuwei Wu 0001, Yuanquan Wang 0001, Yunde Jia |
Comput. Vis. Image Underst. | 3 |
| 2013 | Voice Activity Detection Via Noise Reducing Using Non-Negative Sparse CodingabstractThis letter presents a voice activity detection (VAD) approach using non-negative sparse coding to improve the detection performance in low signal-to-noise ratio (SNR) conditions. The basic idea is to use features extracted from a noise-reduced representation of original audio signals. We decompose the magnitude spectrum of an audio signal on a speech dictionary learned from clean speech and a noise dictionary learned from noise samples. Only coefficients corresponding to the speech dictionary are considered and used as the noise-reduced representation of the signal for feature extraction. A conditional random field (CRF) is used to model the correlation between feature sequences and voice activity labels along audio signals. Then, we assign the voice activity labels for a given audio by decoding the CRF. Experimental results demonstrate that our VAD approach has a good performance in low SNR conditions. Peng Teng, Yunde Jia |
IEEE Signal Process. Lett. | 2 |
| 2013 | Action Recognition Using Multilevel Features and Latent Structural SVMabstractWe first propose a new low-level visual feature, called spatio-temporal context distribution feature of interest points, to describe human actions. Each action video is expressed as a set of relative XYT coordinates between pairwise interest points in a local region. We learn a global Gaussian mixture model (GMM) (referred to as a universal background model) using the relative coordinate features from all the training videos, and then we represent each video as the normalized parameters of a video-specific GMM adapted from the global GMM. In order to capture the spatio-temporal relationships at different levels, multiple GMMs are utilized to describe the context distributions of interest points over multiscale local regions. Motivated by the observation that some actions share similar motion patterns, we additionally propose a novel mid-level class correlation feature to capture the semantic correlations between different action classes. Each input action video is represented by a set of decision values obtained from the pre-learned classifiers of all the action classes, with each decision value measuring the likelihood that the input video belongs to the corresponding action class. Moreover, human actions are often associated with some specific natural environments and also exhibit high correlation with particular scene classes. It is therefore beneficial to utilize the contextual scene information for action recognition. In this paper, we build the high-level co-occurrence relationship between action classes and scene classes to discover the mutual contextual constraints between action and scene. By treating the scene class label as a latent variable, we propose to use the latent structural SVM (LSSVM) model to jointly capture the compatibility between multilevel action features (e.g., low-level visual context distribution feature and the corresponding mid-level class correlation feature) and action classes, the compatibility between multilevel scene features (i.e., SIFT feature and the corresponding class correlation feature) and scene classes, and the contextual relationship between action classes and scene classes. Extensive experiments on UCF Sports, YouTube and UCF50 datasets demonstrate the effectiveness of the proposed multilevel features and action-scene interaction based LSSVM model for human action recognition. Moreover, our method generally achieves higher recognition accuracy than other state-of-the-art methods on these datasets. Xinxiao Wu, Dong Xu 0001, Lixin Duan, Jiebo Luo 0001, Yunde Jia |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2013 | Intrinsic Image Decomposition Using Optimization and User ScribblesabstractIn this paper, we present a novel high-quality intrinsic image recovery approach using optimization and user scribbles. Our approach is based on the assumption of color characteristics in a local window in natural images. Our method adopts a premise that neighboring pixels in a local window having similar intensity values should have similar reflectance values. Thus, the intrinsic image decomposition is formulated by minimizing an energy function with the addition of a weighting constraint to the local image properties. In order to improve the intrinsic image decomposition results, we further specify local constraint cues by integrating the user strokes in our energy formulation, including constant-reflectance, constant-illumination, and fixed-illumination brushes. Our experimental results demonstrate that the proposed approach achieves a better recovery result of intrinsic reflectance and illumination components than the previous approaches. Jianbing Shen, Xiaoshan Yang, Xuelong Li 0001, Yunde Jia |
IEEE Trans. Cybern. | 4 |
| 2012 | On the dimensionality of video bricks under varying illuminationabstractIllumination models of the image set of an object (e.g., human face) under varying lighting conditions have been either empirically or analytically explored. However, the theoretical dimensionality of video bricks of an object under varying illumination is still unknown. In this paper, we focus on this question concretely and give both analytical and empirical results. We derive the theoretical upper bound of the dimensionality of video bricks by investigating the analytical formula of appearance changes due to motion variables of light sources. Theoretical results show in real-world scenes video bricks of an object under varying illumination could be expressed well by a low-dimensional linear subspace. Empirical results of the principal component analysis on the YaleB Face database and our video database are consistent with the theoretical results completely. The application of the low-dimensional linear models of video bricks is demonstrated by the foreground detection task in visual surveillance with drastic illumination changes. Youdong Zhao, Yunde Jia |
CVPR | 3 |
| 2012 | Learning Human Interaction by Interactive Phrases
Yu Kong 0001, Yunde Jia, Yun Fu 0001 |
ECCV (1) | 2 |
| 2012 | View-Invariant Action Recognition Using Latent Kernelized Structural SVM
Xinxiao Wu, Yunde Jia |
ECCV (5) | 2 |
| 2012 | Audio-visual emotion recognition using Boltzmann ZippersabstractThis paper presents a novel approach for automatic audio-visual emotion recognition. The audio and visual channels provide complementary information for human emotional states recognition, and we utilize Boltzmann Zippers as model-level fusion to learn intrinsic correlations between the different modalities. We extract effective audio and visual feature streams with different time scales and feed them to two Boltzmann chains respectively. The hidden units of two chains are interconnected. Second-order methods are applied to Boltzmann Zippers to speed up learning and pruning process. Experimental results on audio-visual emotion data collected in Wizard of Oz scenarios demonstrate our approach is promising and outperforms single modal HMM and conventional coupled HMM methods. Yunde Jia |
ICIP | 2 |
| 2012 | A Hierarchical Model for Human Interaction RecognitionabstractRecognizing human interactions is a challenging task due to partially occluded body parts and motion ambiguities in interactions. We observe that the interdependencies existing at both action level and body part level greatly help disambiguate similar individual movements and facilitate human interaction recognition. In this paper, we propose a novel hierarchical model to capture such interdependencies for recognizing interactions of two persons. We model the action of each person by a large-scale global feature and several body part features. Two types of contextual information are exploited in our model to capture the implicit and complex interdependencies between interaction class, the action classes of two persons and the labels of persons' body parts. We build a challenging human interaction dataset to test our method. Results show that our model is quite effective in recognizing human interactions. Yu Kong 0001, Yunde Jia |
ICME | 2 |
| 2012 | Learning Global and Reconfigurable Part-Based Models for Object DetectionabstractThis paper presents a method of learning global and reconfigurable part-based models (RPM) for object detection. Recently, deformable part-based model (DPM) is widely used. A DPM consists of a root node and a collection of part nodes, which is learned under the latent SVM formulation by treating part nodes as hidden variables. Although the configuration of parts (i.e., the shapes, sizes and locations of parts) plays a major role in improving performance of object detection, it has not been addressed well in the literature. In this paper, we propose RPM to tackle it. A dictionary of part types is defined by enumerating rectangular shapes of different aspect ratios and sizes given the whole lattice (often at twice resolution of the root node), and each part type has a set of part instances when placed in the lattice. So, the configuration space of parts is quantized by the part types and part instances, and then organized into a hierarchical And-Or directed a cyclic graph (AOG). The AOG consists of three types of nodes: terminal nodes (i.e., part instances), And-nodes (representing decompositions of a part instance into two smaller ones) and Or-nodes (representing alternative ways of decompositions). The globally optimal configuration in the AOG is solved using dynamic programming (DP) where the classification error rates of terminal nodes and And-nodes are used as their figures of merit. In experiments, we test our method on the 20 object categories in the PASCAL VOC2007 dataset and obtain comparable performance with state-of-the-art methods. Tianfu Wu 0001, Yi Xie 0006, Yunde Jia |
ICME | 4 |
| 2012 | Probabilistic depth map fusion for real-time multi-view stereo
Mingtao Pei, Yunde Jia |
ICPR | 3 |
| 2012 | Action recognition with discriminative mid-level features
Cuiwei Liu, Yu Kong 0001, Xinxiao Wu, Yunde Jia |
ICPR | 4 |
| 2012 | Audio-visual emotion recognition with boosted coupled HMM
Yunde Jia |
ICPR | 2 |
| 2012 | Annotating videos from the web images
Xinxiao Wu, Yunde Jia |
ICPR | 3 |
| 2012 | Face pose estimation with combined 2D and 3D HOG features
Jiaolong Yang, Wei Liang 0008, Yunde Jia |
ICPR | 3 |
| 2012 | An embedded calibration stereovision systemabstractThis paper describes an embedded calibration stereovision system that is able to be strongly online and auto-calibrated without placing a calibration device in front of the stereo camera, but with the device hidden inside the cavity of the system via a half-mirror. The stereo camera simultaneously observes a scene passing through the half-mirror and the calibration device reflected from the half-mirror, and makes the formation of an embedded calibration stereo pair containing both scene and the calibration device. The features of the calibration patterns are extracted from the embedded calibration stereo pair to estimate the stereo camera's parameters. We use a polyhedral-mirror to generate multiple virtual images of the calibration device to occupy large part of a scene image for accurate estimation. We also use several mirrors to extend the optical path from the calibration device to the stereo camera for depth recovery of distance objects. The system can be easily used in a wide range of applications without considering variation of camera parameters. Yunde Jia, Xiameng Qin |
Intelligent Vehicles Symposium | 1 |
| 2012 | An FPGA-based RGBD imager
Lei Chen 0020, Yunde Jia, Mingxiang Li |
Mach. Vis. Appl. | 2 |
| 2012 | Background modeling by subspace learning on spatio-temporal patches
Youdong Zhao, Haifeng Gong, Yunde Jia, Song-Chun Zhu |
Pattern Recognit. Lett. | 3 |
| 2011 | Intrinsic images using optimizationabstractIn this paper, we present a novel intrinsic image recovery approach using optimization. Our approach is based on the assumption of in a local window in natural images. Our method adopts a premise that neighboring pixels in a local window of a single image having similar intensity values should have similar reflectance values. Thus the intrinsic image decomposition is formulated by optimizing an energy function with adding a weighting constraint to the local image properties. In order to improve the intrinsic image extraction results, we specify local constrain cues by integrating the user strokes in our energy formulation, including constant-reflectance, constant-illumination and fixed-illumination brushes. Our experimental results demonstrate that our approach achieves a better recovery of intrinsic reflectance and illumination components than by previous approaches. Jianbing Shen, Xiaoshan Yang, Yunde Jia, Xuelong Li 0001 |
CVPR | 3 |
| 2011 | Parsing video events with goal inference and intent predictionabstractIn this paper, we present an event parsing algorithm based on Stochastic Context Sensitive Grammar (SCSG) for understanding events, inferring the goal of agents, and predicting their plausible intended actions. The SCSG represents the hierarchical compositions of events and the temporal relations between the sub-events. The alphabets of the SCSG are atomic actions which are defined by the poses of agents and their interactions with objects in the scene. The temporal relations are used to distinguish events with similar structures, interpolate missing portions of events, and are learned from the training data. In comparison with existing methods, our paper makes the following contributions. i) We define atomic actions by a set of relations based on the fluents of agents and their interactions with objects in the scene. ii) Our algorithm handles events insertion and multi-agent events, keeps all possible interpretations of the video to preserve the ambiguities, and achieves the globally optimal parsing solution in a Bayesian framework; iii) The algorithm infers the goal of the agents and predicts their intents by a top-down process; iv) The algorithm improves the detection of atomic actions by event contexts. We show satisfactory results of event recognition and atomic action detection on the data set we captured which contains 12 event categories in both indoor and outdoor videos. Mingtao Pei, Yunde Jia, Song-Chun Zhu |
ICCV | 2 |
| 2011 | Tracking pedestrians with incremental learned intensity and contour templates for PTZ camera visual surveillanceabstractThis paper presents a novel particle-based pedestrian tracking algorithm for PTZ visual surveillance. Most of the state-of-art particle-based tracking algorithms are challenged due to lacking of a reliable moving object detection and drastic scale along with perspective shift of the target. Therefore, pure intensity based algorithms usually miss the target gradually without other features for correcting target location. Our method learns and maintains a contour template of the target besides intensity. Taking into account both the evolution and sudden change of the pedestrian contour, the proposed tracking algorithm maintains several sets of profiles from different perspectives and evolves them incrementally. The effectiveness of our tracking algorithm with extra contour measurement is tested over several surveillance records captured from PTZ camera and estimates the location more robustly than other cutting edge tracking algorithms compared in our experiments. Yi Xie 0006, Mingtao Pei, Guanqun Yu, Yunde Jia |
ICME | 5 |
| 2011 | Dynamic Construction of Multilayer Neural Networks for Classification
Jiqian Liu, Yunde Jia |
ISNN (1) | 2 |
| 2011 | Discriminative structure selection method of Gaussian Mixture Models with its application to handwritten digit recognition
Xiabi Liu, Yunde Jia |
Neurocomputing | 3 |
| 2011 | Multiview Visibility Estimation for Image-Based Modeling
Liuxin Zhang, Ming-Tao Pei, Yunde Jia |
J. Comput. Sci. Technol. | 3 |
| 2011 | Adaptive learning codebook for action recognition
Yu Kong 0001, Xiaoqin Zhang 0002, Weiming Hu 0004, Yunde Jia |
Pattern Recognit. Lett. | 4 |
| 2010 | Pursuing Atomic Video Words by Information Projection
Youdong Zhao, Haifeng Gong, Yunde Jia |
ACCV (2) | 3 |
| 2010 | A High-Performance URL Lookup Engine for URL Filtering SystemsabstractURL filtering systems provide a simple and effective way to prevent people from browsing undesirable or malicious Websites. These systems require a well-designed URL lookup method as the core operation. A high-performance URL lookup engine is proposed in this paper for URL filtering systems. It combines a URL compression algorithm with a multiple string matching based (Wu-Manber-like) matching algorithm. Using this method, the proposed URL lookup engine can achieve high URL lookup performance and efficient memory utilization for storing the ever-increasing URL blacklist with the ability of prefix matching. Experiments with actual URL blacklists and requested URL sets show that our engine can save about 80% memory usage for storing URL blacklists, and reduce 58%-162% URL lookup time compared with the state-of-the-art URL lookup methods. Zhou Zhou 0007, Yunde Jia |
ICC | 3 |
| 2010 | Robust Frame-to-Frame Hybrid MatchingabstractIn this paper, we propose a hybrid approach for addressing feature-based matching problem. We aim to obtain robust and accurate correspondence between features from image frames under unknown and unstructured environments. The approach incorporates image texture analysis, 2-D analytic signal theory and color modeling. It takes advantage of geometric invariant property in texture and monogenic signal information as well as photometric invariant property in HSV color information. The detected features are well localized with high accuracy and the selected matches are robust to changes in scale, blur, viewpoint, and illumination. Experiments conducted on a standard benchmark dataset demonstrate the effectiveness and reliability of our approach. Lei Chen 0020, Yunde Jia |
ICPR | 3 |
| 2010 | Compressive Sampling Recovery for Natural ImagesabstractCompressive sampling (CS) is a novel data collection and coding theory which allows us to recover sparse or compressible signals from a small set of measurements. This paper presents a new model for natural image recovery, in which the smooth l0norm and the approximate total-variation (TV) norm are adopted simultaneously. By using one-order gradient decrease, the speed of algorithm for this new model can be guaranteed. Experimental results demonstrate that the principle of the model is correct and the performance is as good as that based on TV model. The computing speed of the proposed method is two orders of magnitude faster than that of interior point method and two times faster than that of the Nesta optimization based on TV model. Fei Shang, Huiqian Du, Yunde Jia |
ICPR | 3 |
| 2010 | A Discriminative Model for Object Representation and Detection via Sparse FeaturesabstractThis paper proposes a discriminative model that represents an object category with a batch of boosted image patches, motivated by detecting and localizing objects with sparse features. Instead of designing features carefully and category-specifically as in previous work, we extract a massive number of local image patches from the positive object instances and quantize them as weak classifiers. Then we extend the Adaboost algorithm for learning the patch-based model integrating object appearance and structure information. With the learned model, a few features are activated to localize instances in the testing images. In the experiments, we apply the proposed method with several public datasets and achieve advancing performance. Ping Luo 0002, Liang Lin 0004, Yunde Jia |
ICPR | 4 |
| 2010 | Adaptive Diffusion Flow for Parametric Active ContoursabstractThis paper proposes a novel external force for active contours, called adaptive diffusion flow (ADF). We reconsider the generative mechanism of gradient vector flow (GVF) diffusion process from the perspective of image restoration, and exploit a harmonic hyper surface minimal function to substitute smoothness energy term of GVF for alleviating the possible leakage problem. Meanwhile, a ∞- laplacian functional is incorporated in the ADF framework to ensure that the vector flow diffuses mainly along normal direction in homogenous regions of an image. Experiments on synthetic and real images demonstrate the good properties of the ADF snake, including noise robustness, weak edge preserving, and concavity convergence. Yuwei Wu 0001, Yunde Jia, Yuanquan Wang 0001 |
ICPR | 2 |
| 2010 | Tracking Objects with Adaptive Feature Patches for PTZ Camera Visual SurveillanceabstractCompared to the traditional tracking with fixed cameras, the PTZ-camera-based tracking is more challenging due to (i) lacking of reliable background modeling and subtraction; (ii) the appearance and scale of target changing suddenly and drastically. Tackling these problems, this paper proposes a novel tracking algorithm using patch-based object models and demonstrates its advantages with the PTZ-camera in the application of visual surveillance. In our method, the target model is learned and represented by a set of feature patches whose discriminative power is higher than others. The target model is matched and evaluated by both appearance and motion consistency measurements. The homography between frames is also calculated for scale adaptation. The experiment on several surveillance videos shows that our method outperforms the state-of-arts approaches. Yi Xie 0006, Yunde Jia |
ICPR | 3 |
| 2010 | Visibility of Multiple Cameras in a Scene with Unknown GeometryabstractIn this paper, we investigate the problem of determining the visible regions of multiple cameras in a 3D scene without a priori knowledge of the scene geometry. Our approach is based on a variational energy functional where both the unresolved visibility information of multiple cameras and the unknown scene geometry are included. We cast visibility estimation and scene geometry reconstruction as an optimization of the variational energy functional amenable for minimization with the Euler-Lagrange driven evolution. Starting from any initial value, the accurate visibility of multiple cameras as well as the true scene geometry can be obtained at the end of the evolution. Experimental results show the validity of our approach. Liuxin Zhang, Yunde Jia |
ICPR | 2 |
| 2010 | A multimodal labeling interface for wearable computingabstractUnder wearable environments, it is not convenient to label an object with portable keyboards and mice. This paper presents a multimodal labeling interface to solve this problem with natural and efficient operations. Visual and audio modalities cooperate with each other: an object is encircled by visual tracking of a pointing gesture, and meanwhile its name is obtained by speech recognition. In this paper, we propose a concept of virtual touchpad based on stereo vision techniques. With the touchpad, the object encircling task is achieved by drawing a closed curve on a transparent blackboard. The touch events and movements of a pointing gesture are robustly detected for natural gesture interactions. The experimental results demonstrate the efficiency and usability of our multimodal interface. Shanqing Li, Yunde Jia |
IUI | 2 |
| 2010 | Automatic Image Annotation with Cooperation of Concept-Specific and Universal Visual Vocabularies
Xiabi Liu, Yunde Jia |
MMM | 3 |
| 2010 | Layer-Constraint-Based Visibility for Volumetric Multi-view Reconstruction
Yumo Yang, Liuxin Zhang, Yunde Jia |
MMM | 3 |
| 2010 | Surface Reconstruction from Images Using a Variational Formulation
Liuxin Zhang, Yunde Jia |
MMM | 2 |
| 2010 | Discriminative human action recognition in the learned hierarchical manifold space
Xinxiao Wu, Wei Liang 0008, Guangming Hou, Yunde Jia |
Image Vis. Comput. | 5 |
| 2010 | Incremental discriminant-analysis of canonical correlations for action recognition
Xinxiao Wu, Yunde Jia, Wei Liang 0008 |
Pattern Recognit. | 2 |
| 2009 | Learning Group Activity in Soccer Videos from Local Motion
Yu Kong 0001, Weiming Hu 0004, Xiaoqin Zhang 0002, Hanzi Wang, Yunde Jia |
ACCV (1) | 5 |
| 2009 | Convolutional Virtual Electric Field External Force for Active Contours
Yuanquan Wang 0001, Yunde Jia |
ACCV (3) | 2 |
| 2009 | Soft Measure of Visual Token Occurrences for Object Categorization
Xiabi Liu, Yunde Jia |
CAIP | 3 |
| 2009 | Combining evolution strategy and gradient descent method for discriminative learning of bayesian classifiersabstractThe optimization method is one of key issues in discriminative learning of pattern classifiers. This paper proposes a hybrid approach of the Covariance Matrix Adaptation Evolution Strategy (CMA-ES) and the gradient decent method for optimizing Bayesian classifiers under the SOFT target based Max-Min posterior Pseudo-probabilities (Soft-MMP) learning framework. In our hybrid optimization approach, the weighted mean of the parent population in the CMA-ES is adjusted by exploiting the gradient information of objective function, based on which the offspring is generated. As a result, the efficiency and the effectiveness of the CMA-ES are improved. We apply the Soft-MMP with the proposed hybrid optimization approach to handwritten digit recognition. The experiments on the CENPARMI database show that our handwritten digit classifier outperforms other state-of-the-art techniques. Furthermore, our hybrid optimization approach behaved better than not only the single gradient decent method but also the single CMA-ES in the experiments. Xiabi Liu, Yunde Jia |
GECCO | 3 |
| 2009 | Incremental discriminative-analysis of canonical correlations for action recognitionabstractHuman action recognition is a challenging problem due to the large changes of human appearance in the cases of partial occlusions, non-rigid deformations and high irregularities. It is difficult to collect a large set of training samples with the hope of covering all possible variations of an action. In this paper, we propose an online recognition method, namely Incremental Discriminant-Analysis of Canonical Correlations (IDCC), whose discriminative model is incrementally updated to capture the changes of human appearance and thereby facilitates the recognition task in changing environments. As the training sets are acquired sequentially instead of being given completely in advance, our method is able to compute a new discriminant matrix by updating the existing one using the eigenspace merging algorithm. Experimental results on both Weizmann and KTH action data sets show that our method performs better than state-of-the-art methods on both accuracy and efficiency. Moreover, the robustness of our method is demonstrated on the irregular action recognition. Xinxiao Wu, Wei Liang 0008, Yunde Jia |
ICCV | 3 |
| 2009 | Unsupervised Selection and Discriminative Estimation of Orthogonal Gaussian Mixture Models for Handwritten Digit RecognitionabstractThe problem of determining the appropriate number of components is important in finite mixture modeling for pattern classification. This paper considers the application of an unsupervised clustering method called AutoClass to training of orthogonal Gaussian mixture models (OGMM). Actually, the number of components in OGMM of each class is selected based on AutoClass. In this way, the structures of OGMM for difference classes are not necessarily be the same as those in usual modeling scheme, so that the dissimilarity between the data distributions of different classes can be described more exactly. After the model selection is completed, a discriminative learning framework of Bayesian classifiers called max-min posterior pseudo-probabilities (MMP) is employed to estimate component parameters in OGMM of each class. We apply the proposed learning approach of OGMM to handwritten digit recognition. The experimental results on the MNIST database show the effectiveness of our approach. Xiabi Liu, Yunde Jia |
ICDAR | 3 |
| 2009 | Statistical Modeling and Learning for Recognition-Based Handwritten Numeral String SegmentationabstractThis paper proposes a recognition based approach to handwritten numeral string segmentation. We consider two classes: numeral strings segmented correctly or not. The feature vectors containing recognition information for numeral strings segmented correctly are assumed to be of the distribution of Gaussian mixture model (GMM). Based on this modeling, the recognition based segmentation is solved under the max-min posterior pseudo-probabilities (MMP) framework of learning Bayesian classifiers. In the training phase, we use the MMP method to learn a posterior pseudo-probability measure function from positive samples and negative samples of numeral strings segmented correctly. In the process of recognition based segmentation, we generate all possible candidate segmentations of an input string through contour and profile analysis, and then compute the posterior pseudo-probabilities of being the numeral string segmented correctly for all the candidate segmentations. The candidate segmentation with the maximum posterior pseudo-probability is taken as the final result. The effectiveness of our approach is demonstrated by the experiments of numeral string segmentation and recognition on the NIST SD19 database. Xiabi Liu, Yunde Jia |
ICDAR | 3 |
| 2009 | Tracking articulated objects by learning intrinsic structure of motion
Xinxiao Wu, Wei Liang 0008, Yunde Jia |
Pattern Recognit. Lett. | 3 |
| 2009 | Action recognition feedback-based framework for human pose reconstruction from monocular images
Xinxiao Wu, Wei Liang 0008, Yunde Jia |
Pattern Recognit. Lett. | 3 |
| 2008 | Human action recognition using discriminative models in the learned hierarchical manifold spaceabstractA hierarchical learning based approach for human action recognition is proposed in this paper. It consists of hierarchical nonlinear dimensionality reduction based feature extraction and cascade discriminative model based action modeling. Human actions are inferred from human body joint motions and human bodies are decomposed into several physiological body parts according to inherent hierarchy (e.g. right arm, left arm and head all belong to upper body). We explore the underlying hierarchical structures of high-dimensional human pose space using hierarchical Gaussian process latent variable model (HGPLVM) and learn a representative motion pattern set for each body part. In the hierarchical manifold space, the bottom-up cascade conditional random fields (CRFs) are used to predict the corresponding motion pattern in each manifold subspace, and then the final action label is estimated for each observation by a discriminative classifier on the current motion pattern set. Wei Liang 0008, Xinxiao Wu, Yunde Jia |
FG | 4 |
| 2008 | Group action recognition in soccer videosabstractGroup action recognition in soccer videos is a challenging problem due to the difficulties of group action representation and camera motion estimation. This paper presents a novel approach for recognizing group action with a moving camera. In our approach, ego-motion is estimated by the Kanade-Lucas-Tomasi feature sets on successive frames. The optical flow is then computed on compensated frames. Due to the inaccurate ego-motion estimation, the optical flow can not reflect accurate motion of objects. In this paper, we propose a new motion descriptor which treats the optical flow as spatial patterns and extracts accurate global motion from the noisy optical flow. The latent-dynamic conditional random field model is employed to recognize group action. Experimental results show that our approach is promising. Yu Kong 0001, Xiaoqin Zhang 0002, Qingdi Wei, Weiming Hu 0004, Yunde Jia |
ICPR | 5 |
| 2008 | Violence classification based on shape variations from multiple viewsabstractMost existing algorithms for human behavior analysis concentrate on action recognition through assuming that input sequences are well pre-segmented and restricting examples into a small vocabulary. In this paper, we present a novel action violence classification framework which directly evaluates the potential threat based on shape variations. We extract silhouettes as input features, employ the R transform to project binary shapes into the Radon space, and fuse multiple views to classify action violence. Experimental results on the INRIA IXMAS database demonstrate the efficiency and robustness of the proposed method. Fawang Liu, Yunde Jia |
ICPR | 2 |
| 2008 | Spatio-temporal patches for night background modeling by subspace learningabstractIn this paper, a novel background model on spatio-temporal patches is introduced for video surveillance, especially for night outdoor scene, where extreme lighting conditions often cause troubles. The spatio-temporal patch, called brick, is presented to simultaneously capture spatio-temporal information in surveillance video. The set of bricks of a given background patch, under all possible lighting conditions, lies in a low-dimensional subspace, which can be learned by online subspace learning. The proposed method can efficiently model the background and detect the appearance and motion variance caused by foreground. Experimental results on real data show that the proposed method is insensitive to dramatic lighting changes and achieves superior performance to two classical methods. Youdong Zhao, Haifeng Gong, Yunde Jia |
ICPR | 4 |
| 2008 | External Force for Active Contours: Gradient Vector Convolution
Yuanquan Wang 0001, Yunde Jia |
PRICAI | 2 |
| 2008 | A study of regularized Gaussian classifier in high-dimension small sample set case based on MDL principle with application to spectrum recognition
Ping Guo 0002, Yunde Jia, Michael R. Lyu |
Pattern Recognit. | 2 |
| 2008 | Gaussian mixture modeling and learning of neighboring characters for multilingual text extraction in images
Xiabi Liu, Yunde Jia |
Pattern Recognit. | 3 |
| 2007 | Cardiac Motion Estimation from Tagged MRI Using 3D-HARP and NURBS Volumetric Model
Yuanquan Wang 0001, Yunde Jia |
ACCV (1) | 3 |
| 2007 | On the Critical Point of Gradient Vector Flow Snake
Yuanquan Wang 0001, Yunde Jia |
ACCV (2) | 3 |
| 2007 | Discriminant Clustering Embedding for Face Recognition with Image Sets
Youdong Zhao, Yunde Jia |
ACCV (2) | 3 |
| 2007 | A Plane-based Calibration for Multi-camera SystemsabstractIn this paper, a plane-based calibration algorithm with use for calibrating large complexes of linear projective cameras of varying intrinsic and extrinsic parameters is proposed. A planar pattern with known reference points placed at a few different locations is only required as a calibration object. All the cameras do not have to see this planar pattern at all locations simultaneously, and only reasonable overlap between them is necessary. We divide these cameras into groups according to their positions and orientations, to make sure cameras in each group have a common field of view, and estimate the intrinsic parameters of each camera via a plane-based technique. Common views of planes are used to represent each camera in the world coordinate system of its own group and determine the relationships between these world coordinate systems. Rodrigues formula is adopted to deal with the inconsistency when recovering the rigid displacements between cameras. At last, all these cameras can be completely calibrated in a uniform world coordinate system robustly and accurately. Liuxin Zhang, Yunde Jia |
CAD/Graphics | 3 |
| 2007 | Learning Semantic Concepts for Image Retrieval using the Max-Min Posterior Pseudo-ProbabilitiesabstractSemantic gap is the main problem in current content-based image retrieval. This paper proposes an approach which aims to learn semantic concepts from visual features. Each concept is modeled as a posterior pseudo-probability function, and the function parameters are trained from the positive and negative image examples of the concept using the max-min posterior pseudo-probabilities criterion. According to the posterior pseudo-probabilities of the query concept for all images, the image retrieval is realized by classifying all images into two categories: relevant to the query concept and irrelevant. The number of relevant images can be determined automatically. We show the effectiveness and the advantage of our approach through the experiments on Corel database. Xiabi Liu, Yunde Jia |
ICME | 3 |
| 2007 | Facial Expression Analysis on Semantic Neighborhood Preserving Embedding
Yunde Jia, Youdong Zhao |
ISNN (2) | 2 |
| 2007 | A new localization algorithm for planetary roverabstractIn this paper, first of all, the existing localization algorithms of planetary rover are summarized. And then, the RCM (RSSI and Circle centre Mixed localization) Algorithm and its theoretical model, which mixed RSSI (Received Signal Strengt Indicator) Algorithm and the Circle of contact-Center Algorithm, are proposed based on the technology of wireless sensor network. The simulation results show the validity of RCM Algorithm which lowers the localization error down to 10% while the error spreading of RSSI Algorithm is even up to 50%. Hongbin Deng, Yunde Jia |
SMC | 2 |
| 2007 | An airborne image stabilization Method based on the Gaussian Mixture modelabstractIn this paper, an image stabilization method based on the adaptive Gaussian mixture model (GMM) is presented in order to smooth down the airborne image vibration. Firstly, the projection algorithm is adopted for the motion estimation; Secondly, GMM parameter is obtained after analyzing characteristics of the first n images; Finally, a stable image sequence is achieved after GMM motion filter operates on the motion parameter. The experimental results show that the method has the advantage of fast speed and effectively smooth unwanted vibration of image sequences. Hongbin Deng, Yunde Jia, Yihua Xu, Wei Liang 0008 |
SMC | 2 |
| 2007 | Symmetrical null space LDA for face and ear recognition
Xiaoxun Zhang, Yunde Jia |
Neurocomputing | 2 |
| 2007 | A linear discriminant analysis framework based on random subspace for face recognition
Xiaoxun Zhang, Yunde Jia |
Pattern Recognit. | 2 |
| 2006 | Local Steerable Phase (LSP) Feature for Face Representation and RecognitionabstractIn this paper, we propose a novel local steerable phase (LSP) feature extracted from the face image using steerable filter for face representation and recognition. Steerable filter is a kind of oriented filters. It is rotated very efficiently by taking a suitable linear combination of basis filters and allows adaptive control over phase as well as orientation. Phase information provided by steerable filter is locally stable with respect to scale, noise and brightness changes. Furthermore, steerable filter is implemented within a Gaussian pyramid to make use of discriminative power in the scale-space of face images. Each face is represented as multiple "steerablefaces" of different scales and orientations. With simple down-sampling, all the steerablefaces are concatenated to an augmented feature vector for evaluating similarity between face images. A nearest-neighbor classifier based on local weighted phase-correlation is used for final decision rule. Experimental results on FERET and XM2VTS databases demonstrate the performance of the proposed method. Xiaoxun Zhang, Yunde Jia |
CVPR (2) | 2 |
| 2006 | Gaussian Mixture Modeling of Neighbor Characters for Multilingual Text Extraction in ImagesabstractThis paper proposes a new method to extract multilingual text in images through discriminating characters from non-characters based on the Gaussian mixture modeling of neighbor characters. The image is binarized and the morphological closing operation is performed on the binary image, in order that each character in it can be treated as a connected component; the neighborhood of connected components are computed based on the Voronoi partition of the image, and each connected component is labeled as character or non-character according to its neighbors. We applied the proposed text extraction method to Chinese and English text extraction, the effectiveness of which is confirmed by the experimental results. Xiabi Liu, Yunde Jia, Hongbin Deng |
ICIP | 3 |
| 2006 | A Real-Time 3D Human Body Tracking and Modeling SystemabstractIn this paper a real-time system for 3D human upper body tracking and modeling is proposed. The system uses multiple cameras to recover the depth maps in real-time, then integrates both color and depth information to track the human body, head, and hands, and finally recovers the 3D upper body model parameters from the tracking results. Extensive experiments demonstrate that the system can track and rebuild human model in complicated situations. The system makes good tradeoff between the accuracy and system simplicity, and can be widely used in many applications, such as desktop interaction and digital entertainments. Jing-Feng Li, Yihua Xu, Yunde Jia |
ICIP | 4 |
| 2006 | Segmentation of the Left Venctricle from MR Images via Snake Models Incorporating Shape SimilaritiesabstractMagnetic resonance imaging is a noninvasive method to measure the geometry and deformation of the heart during one cardiac cycle. To make thorough use of this anatomical and functional information, it is necessary to segment the endo- and epicardium of the left ventricle. In this study, we present a method based on GVF snake model for this purpose. For endocardium segmentation, the proposed method pays particular attention to papillary muscle and artifacts by adopting a shape energy, with this energy, the snake contour can overcome the spurious edges stemming from artifacts and the final results could depend much less on the initial contour. In order to segment the epicardium, a novel energy based on shape similarity is proposed by assuming that the epicardium resembles the endocardium in shape. In addition, a new strategy is developed to derive the external force. This new force can push the snake contour directly to the epicardium when using the endocardium as initialization. By applying the proposed method to a set of 140 cardiac images, its accuracy and robustness is demonstrated and validated. Yuanquan Wang 0001, Yunde Jia |
ICIP | 2 |
| 2006 | Coding Facial Expression with Oriented Steerable FiltersabstractIn this paper, facial expression is coded by steerable filters which are rotated very efficiently by taking a suitable linear combination of basis filters. Local features extracted by steerable filters are locally stable with respect to scale, noise, and brightness changes, and distinctive enough to capture subtle facial expression cues. Further more, steerable filters are implemented within a Gaussian pyramid to exploit discriminative power in scale-space. Responses of the filters are concatenated to an augmented feature vector to evaluate the similarity between different facial expression images with the nearest-neighbor rule for final decisions. In comparison with Gabor filters, steerable filters save much computational cost and obtain comparable recognition performance with fewer features. Experiments on the JAFFE database demonstrate the effectiveness of steerable filters for coding facial expression. Yunde Jia, Xiaoxun Zhang |
ICIP | 2 |
| 2006 | Maximum-Minimum Similarity Training for Text Extraction
Xiabi Liu, Yunde Jia |
ICONIP (3) | 3 |
| 2006 | Hand-Gesture Based Text Input for Wearable ComputersabstractThis paper proposes a novel text input method based on hand gestures for wearable computers. A character is firstly written by a fingertip and then recognized through a B-splines based character recognition method. The writing procedure is controlled by hand gestures including hand tracking, gesture recognition and fingertip positioning which are performed by an extended CONDENSATION algorithm. We have integrated the proposed text input method into a wearable vision system developed in our lab and tested the resulting text input system on Graffiti 2 alphabet. The experimental result shows that our method is promising for natural text input in wearable computers. Xiabi Liu, Yunde Jia |
ICVS | 3 |
| 2006 | Stereo Vision System on Programmable Chip (SVSoC) for Small Robot NavigationabstractIn this paper, we present a stereo vision system on programmable chip (SVSoC) for dense depth mapping and obstacle detection at video rate. The system is composed of three miniature CMOS cameras with triangular configuration and one FPGA chip for parallel computing. The system algorithm contains nonlinear iteration based cooperative algorithm for high quality dense depth mapping, ground extraction and obstacle location from dense depth maps. With the connection of DSP for moving control, the system is mounted on small hexapod robot for obstacle avoidance and navigation Mingxiang Li, Yunde Jia |
IROS | 2 |
| 2006 | Kernel ICA Feature Extraction for Spectral Recognition of Celestial ObjectsabstractIn the literature of astronomical spectral classification, linear principle component analysis (PCA) was frequently employed to extract features of spectra data. However, the spectral data are too complicated to be well described by a linear model. In this paper, kernel independent component analysis (KICA), which contains a nonlinear kernel mapping component, is adopted to extract features from the spectra of galaxies. Then, a radial basis function neural network is adopted as a classifier to implement the classification. Experiments with real-world spectral data set show that KICA is a very appropriate technique to describe the important features of celestial objects, and the correct classification rate is improved compared with PCA method. Ling Bai, Anbang Xu, Ping Guo 0002, Yunde Jia |
SMC | 4 |
| 2006 | Face recognition with local steerable phase feature
Xiaoxun Zhang, Yunde Jia |
Pattern Recognit. Lett. | 2 |
| 2005 | Simulated Annealing Based Hand Tracking in a Discrete Space
Wei Liang 0008, Yunde Jia, Cheng Ge |
ACII | 2 |
| 2005 | Visual Hand Tracking Using Nonparametric Sequential Belief Propagation
Wei Liang 0008, Yunde Jia, Cheng Ge |
ICIC (1) | 2 |
| 2005 | Hand motion tracking using MDPF methodabstractHand motion tracking is a challenging problem due to the complexity of searching in a high dimensional configuration space for an optimal estimate. This paper represents the hand feasible configurations as a discrete space, which avoids learning to find parameters as general configuration space representations do, meanwhile, it arrange the discrete data on the KD-tree which supports fast nearest neighbor retrieval and it is easy to be modified when new samples are embedded. To track hand motion efficiently, this paper presents a MDPF (Multi-Directional search with Particle Filter) algorithm, in which a 'global' optimization and a 'local' optimization are combined to obtain the best matching configuration. The 'local' method, which is designed to run in multiple processors, could choose more representative samples for global efficiently, and the global method guards the tracking process towards a global minimum. The Experiment results show that this approach is robust and efficient for tracking 3D hand motion. Wei Liang 0008, Cheng Ge, Yunde Jia |
SMC | 3 |
| 2005 | Non-negative matrix factorization framework for face recognitionabstractNon-negative Matrix Factorization (NMF) is a part-based image representation method which adds a non-negativity constraint to matrix factorization. NMF is compatible with the intuitive notion of combining parts to form a whole face. In this paper, we propose a framework of face recognition by adding NMF constraint and classifier constraints to matrix factorization to get both intuitive features and good recognition results. Based on the framework, we present two novel subspace methods: Fisher Non-negative Matrix Factorization (FNMF) and PCA Non-negative Matrix Factorization (PNMF). FNMF adds both the non-negative constraint and the Fisher constraint to matrix factorization. The Fisher constraint maximizes the between-class scatter and minimizes the within-class scatter of face samples. Subsequently, FNMF improves the capability of face recognition. PNMF adds the non-negative constraint and characteristics of PCA, such as maximizing the variance of output coordinates, orthogonal bases, etc. to matrix factorization. Therefore, we can get intuitive features and desirable PCA characteristics. Our experiments show that FNMF and PNMF achieve better face recognition performance than NMF and Local NMF. Yunde Jia, Changbo Hu, Matthew Turk 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2005 | A bottom-up algorithm for finding principal curves with applications to image skeletonization
Xiabi Liu, Yunde Jia |
Pattern Recognit. | 2 |
| 2004 | A robust hand tracking and gesture recognition method for wearable visual interfaces and its applicationsabstractGesture-based interface is one of the most promising modes of human-computer interaction for wearable computers. This paper proposes a robust hand tracking and gesture recognition method for wearable visual interfaces, which is an extension of ICONDENSATION algorithm. The method integrates shape and depth information for robust hand tracking. Gesture recognition is realized through the maximum posterior estimation of several pre-defined gestures. The experimental results show that the proposed method works well in dynamic and complex background. Several promising applications in wearable computers are also discussed. Yunde Jia |
ICIG | 2 |
| 2004 | A fast morphological algorithm for color image multi-scale segmentation using vertex-collapseabstractThis paper proposes a fast morphological algorithm for color image multi-scale segmentation based on the connected sets using the vertex-collapse algorithm. Color morphological operators are defined based on a special full order relation of color vectors in HSV model. Morphological opening-closing operators are implemented using vertex-collapse algorithm to form a nonlinear scale space. Both area and color are considered to calculate the similarity between adjacent regions, and entropy criterion is adopted to control the content contained in the finial scale tree. The experimental result shows that the image is segmented efficiently with perfect shape preserving. Qimin Peng, Yunde Jia |
ICIG | 2 |
| 2004 | Probabilistic classification based image regions labelingabstractImage regions labeling is a crucial step of scene understanding. In this paper, we propose an approach that can solve the image regions labeling problem under the Bayesian probabilistic framework. In our approach, we first extract the efficient low-level visual features from the segmented image regions. Then we represent the knowledge of each label class as a Gaussian mixture model of image region features, and describe the relationship between different region label classes using likelihood function. Finally, the label of each image region is determined by probabilistic classification. The proposed approach is evaluated on the outdoor scene images and the overall recognition accuracy is over 86%. Yunde Jia |
ICIG | 2 |
| 2003 | A Miniature Stereo Vision Machine for Real-Time Dense Depth Mapping
Yunde Jia, Yihua Xu, Wanchun Liu, Yuwen Zhu, Xiaoxun Zhang, Luping An |
ICVS | 1 |
| 1992 | Description and recognition of curved objectsabstractThe author presents an approach to the approximation of curves for the description and recognition of curved objects. One needs only to make a rotation and a scaling respectively to a circular arc to find the longest arc of the approximation of the part of a curve which might contain linear segments. As all points in the curve are independent from each other in computing, it is easy to use a parallel algorithm for fast computing. Experiments on several high resolution tactile images show the algorithm behaves well. This method can also be extended to the recognition of occluded curved objects.> Yunde Jia |
ICPR (3) | 1 |