Kaiqi Huang

dblp:89/7026 · DBLP profile ↗
← Back
234ranked-venue papers
12as first author
66since 2021 · last 2026
0000-0002-2677-9273ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 173 · 5 first-author · 37 since 2021Artificial intelligence and machine learning · 142 · 5 first-author · 47 since 2021Human-computer interaction and ubiquitous computing · 6 · 3 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3Databases, data management, data science and information retrieval · 2 · 1 since 2021Systems, architecture and hardware · 1Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 No-Regret Strategy Solving in Imperfect-Information Games via Pre-Trained Embedding
abstract
High-quality information set abstraction remains a core challenge in solving large-scale imperfect-information extensive-form games (IIEFGs)--such as no-limit Texas Hold’em--where the finite nature of spatial resources hinders solving strategies for the full game. State-of-the-art AI methods rely on pre-trained discrete clustering for abstraction, yet their hard classification irreversibly discards critical information: specifically, the quantifiable subtle differences between information sets--vital for strategy solving--thus compromising the quality of such solving. Inspired by the word embedding paradigm in natural language processing, this paper proposes the Embedding CFR algorithm, a novel approach for solving strategies in IIEFGs within an embedding space. The algorithm pre-trains and embeds the features of individual information sets into an interconnected low-dimensional continuous space, where the resulting vectors more precisely capture both the distinctions and connections between information sets. Embedding CFR introduces a strategy-solving process driven by regret accumulation and strategy updates in this embedding space, with supporting theoretical analysis verifying its ability to reduce cumulative regret. Experiments on poker show that with the same spatial overhead, Embedding CFR achieves significantly faster exploitability convergence compared to cluster-based abstraction algorithms, confirming its effectiveness. Furthermore, to our knowledge, it is the first algorithm in poker AI that pre-trains information set abstractions via low-dimensional embedding for strategy solving.
Yanchang Fu, Shengda Liu, Pei Xu 0003, Kaiqi Huang
AAAI4
2026 CausalStep: A Benchmark for Explicit Stepwise Causal Reasoning in Videos
abstract
Recent advances in large language models (LLMs) have improved reasoning in text and image domains, yet achieving robust video reasoning remains a significant challenge. Existing video benchmarks mainly assess shallow understanding and reasoning and allow models to exploit global context, failing to rigorously evaluate true causal and stepwise reasoning. We present CausalStep, a benchmark designed for explicit stepwise causal reasoning in videos. CausalStep segments videos into causally linked units and enforces a strict stepwise question-answer (QA) protocol, requiring sequential answers and preventing shortcut solutions. Each question includes carefully constructed distractors based on error type taxonomy to ensure diagnostic value. The benchmark features 100 videos across six categories and 1,852 multiple-choice QA pairs. We introduce seven diagnostic metrics for comprehensive evaluation, enabling precise diagnosis of causal reasoning capabilities. Experiments with leading proprietary and open-source models, as well as human baselines, reveal a significant gap between current models and human-level stepwise reasoning. CausalStep provides a rigorous benchmark to drive progress in robust and interpretable video reasoning.
Xuchen Li 0001, Xuzhao Li, Kaiqi Huang, Wentao Zhang 0001
AAAI4
2026 RefRea: Reference-Guided Reasoning with Meta-Cognition for Accurate Language Model Agents
abstract
In recent years, with the rapid development of large language models (LLMs), LLM-based agents have achieved remarkable progress across a wide range of tasks. However, reasoning inconsistencies in LLMs still significantly limit the performance of agents in complex decision-making scenarios. Cognitive science research suggests that individuals can benefit from observing others' explicit thinking processes to improve their strategy-making. Inspired by this mechanism, we propose Reference-guided Reasoning with meta-cognition (RefRea), a novel approach that enhances decision-making by introducing a reference language model to guide and calibrate the reasoning model's actions. RefRea enhances reasoning accuracy and stability by integrating a reference model and a meta-cognition module. The reference model relies solely on validated meta-cognition for consistent guidance, while the reasoning model interacts with the environment using both validated and exploratory meta-cognition. Guidance is provided by comparing the action similarity between the reference and reasoning models. This process is supported by the meta-cognition module, which generates summary knowledge by reflecting on action history and environmental feedback, leading to more adaptive and reliable behavior. We evaluate our algorithm in the text-based reasoning environment ScienceWorld. Experimental results demonstrate that RefRea outperforms state-of-the-art methods. Comprehensive ablation studies further highlight the effectiveness of both the reference model and the meta-cognition module.
Yuxiang Mai, Qiyue Yin, Wancheng Ni, Xiaogang Ouyang, Pei Xu 0003, Kaiqi Huang
AAAI7
2026 ImagerySearch: Adaptive Test-Time Search for Video Generation Beyond Semantic Dependency Constraints
abstract
Video generation models have achieved remarkable progress, particularly excelling in realistic scenarios; however, their performance degrades notably in imaginative scenarios. These prompts often involve rarely co-occurring concepts with long-distance semantic relationships, falling outside training distributions. Existing methods typically apply test-time scaling for improving video quality, but their fixed search spaces and static reward designs limit adaptability to imaginative scenarios. To fill this gap, we propose ImagerySearch, a dynamic test-time scaling law strategy inspired by imagery that adaptively adjusts the inference search space and reward guided by prompts, effectively enhancing generation quality in imaginative scenarios. Furthermore, we introduce LDT-Bench, the first benchmark targeting long-distance semantic prompts, designed to evaluate the creativity of video generation models. It comprises 2,839 challenging concept pairs from diverse recognition datasets and incorporates an automatic evaluation protocol to assess creative capacity. Extensive experiments on LDT-Bench demonstrate that our approach consistently outperforms general generation models and test-time scaling approaches. Additionally, ImagerySearch achieves strong performance on VBench, confirming its effectiveness in improving video generation quality under diverse conditions.
Meiqi Wu, Jiashu Zhu, Xiaokun Feng, Chubin Chen, Bingze Song, Fangyuan Mao, Jiahong Wu 0005, Xiangxiang Chu, Kaiqi Huang
AAAI10
2026 Talk With Your Fingers: A Depth-Aware Benchmark for Air-Writing Recognition
abstract
Air-writing has emerged as a promising communication modality for AR/VR and metaverse environments, enabling quiet, non-contact text input by translating finger movements into natural language. However, existing approaches typically project in-air writing onto a virtual 2D plane and assume characters are formed with a single continuous stroke–an oversimplification that neglects the rich 3D structure inherent in natural handwriting. In this work, we challenge the “single-stroke 2D” paradigm and explore the role of depth cues in enhancing air-writing recognition. To this end, we present DAAWBench, the first large-scale Depth-Aware Air-Writing dataset, featuring 8.8 million RGB-D frames annotated with 3,755 Chinese characters from the GB2312-80 Level-1 set. Our analysis reveals consistent depth variations at stroke boundaries, indicating that stroke segmentation and character recognition can benefit from depth modeling. Based on these insights, we propose DARec, a novel 3D trajectory-based recognition model that effectively leverages depth-aware priors. Extensive experiments across in-domain and out-of-domain settings, including evaluations with vision-language models (e.g., GPT-4o, Qwen-VL) and human baselines, show that DARec significantly outperforms 2D-only counterparts, achieving 87.73% accuracy versus 9.05%. Our findings demonstrate the critical importance of depth modeling in human-computer co-creative interfaces, and we will publicly release our dataset and code at https://github.com/wmeiqi/DAAWBench.
Meiqi Wu, Yuzhong Zhao, Xuchen Li 0001, Yuanqiang Cai, Jiahong Wu 0005, Weiqiang Wang 0001, Kaiqi Huang
IEEE Trans. Circuits Syst. Video Technol.8
2026 COSOS-1k: A Benchmark Dataset and Occlusion-Aware Uncertainty Learning for Multi-View Video Object Detection
abstract
Confined spaces refer to partially or fully enclosed areas, e.g., sewage wells, where working conditions pose significant risks to the workers. The evaluation of COfined Space Operational Safety (COSOS) refers to verifying whether workers are properly equipped with safety equipment before entering a confined space, which is crucial for protecting their safety and health. Due to the crowded nature of such environments and the small size of certain safety equipment, existing methods face significant challenges. Moreover, there is a lack of dedicated datasets to support research in this domain. In this paper, in order to advance research in this challenging task, we present COSOS-1k, an extensive dataset constructed from diverse confined space scenarios. It comprises multi-view videos for each scenario, covers 10 essential safety protective equipments and 6 attributes of worker, and is annotated with expressive object locations, fine-grained attributes, and occlusion status. The COSOS-1k is the first dataset known to date, tailored explicitly for the real-world COSOS scenarios. In addition, we address the challenge of occlusion from three perspectives: instance, video, and view. Firstly, at the instance level, we propose Occlusion-aware Uncertainty Estimation (OUE) method, which leverages box-level occlusion annotations to enable part-level occlusion prediction for objects. Secondly, at the video level, we introduce Cross-Frame Cluster (CFC) attention, which integrates temporal context features from the same object category to mitigate the impact of occlusions in the current frame. Finally, we extend CFC to the view level and form Cross-View Cluster (CVC) attention, where complementary information is mined from another view. Extensive experiments demonstrate the effectiveness of the proposed methods and provide insights into the importance of dataset diversity and expressivity. The COSOS-1k dataset and code are available at https://github.com/deepalchemist/cosos-1k.
Wenjie Yang 0005, Yueying Kao, Yuanlong Yu 0001, Kaiqi Huang
IEEE Trans. Image Process.5
2025 Sequential Preference Optimization: Multi-Dimensional Preference Alignment with Implicit Reward Modeling
abstract
Human preference alignment is critical in building powerful and reliable large language models (LLMs). However, current methods either ignore the multi-dimensionality of human preferences (e.g. helpfulness and harmlessness) or struggle with the complexity of managing multiple reward models. To address these issues, we propose Sequential Preference Optimization (SPO), a method that sequentially fine-tunes LLMs to align with multiple dimensions of human preferences. SPO avoids explicit reward modeling, directly optimizing the models to align with nuanced human preferences. We theoretically derive closed-form optimal SPO policy and loss function. Gradient analysis is conducted to show how SPO manages to fine-tune the LLMs while maintaining alignment on previously optimized dimensions. Empirical results on LLMs of different size and multiple evaluation datasets demonstrate that SPO successfully aligns LLMs across multiple dimensions of human preferences and significantly outperforms the baselines.
Xingzhou Lou, Junge Zhang, Lifeng Liu, Kaiqi Huang
AAAI6
2025 Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues
abstract
Vision-Language Tracking (VLT) aims to localize a target in video sequences using a visual template and language description. While textual cues enhance tracking potential, current datasets typically contain much more image data than text, limiting the ability of VLT methods to align the two modalities effectively. To address this imbalance, we propose a novel plug-and-play method named CTVLT that leverages the strong text-image alignment capabilities of foundation grounding models. CTVLT converts textual cues into interpretable visual heatmaps, which are easier for trackers to process. Specifically, we design a textual cue mapping module that transforms textual cues into target distribution heatmaps, visually representing the location described by the text. Additionally, the heatmap guidance module fuses these heatmaps with the search image to guide tracking more effectively. Extensive experiments on mainstream benchmarks demonstrate the effectiveness of our approach, achieving state-of-the-art performance and validating the utility of our method for enhanced VLT.
Xiaokun Feng, Dailing Zhang, Xuchen Li 0001, Meiqi Wu, Jing Zhang 0110, Xiaotang Chen, Kaiqi Huang
ICASSP8
2025 ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking
abstract
Vision-language tracking aims to locate the target object in the video sequence using a template patch and a language description provided in the initial frame. To achieve robust tracking, especially in complex long-term scenarios that reflect real-world conditions as recently highlighted by MGIT, it is essential not only to characterize the target features but also to utilize the context features related to the target. However, the visual and textual target-context cues derived from the initial prompts generally align only with the initial target state. Due to their dynamic nature, target states are constantly changing, particularly in complex long-term sequences. It is intractable for these cues to continuously guide Vision-Language Trackers (VLTs). Furthermore, for the text prompts with diverse expressions, our experiments reveal that existing VLTs struggle to discern which words pertain to the target or the context, complicating the utilization of textual cues. In this work, we present a novel tracker named ATCTrack, which can obtain multimodal cues Aligned with the dynamic target states through comprehensive Target-Context feature modeling, thereby achieving robust tracking. Specifically, (1) for the visual modality, we propose an effective temporal visual target-context modeling approach that provides the tracker with timely visual cues. (2) For the textual modality, we achieve precise target words identification solely based on textual content, and design an innovative context words calibration method to adaptively utilize auxiliary context words. (3) We conduct extensive experiments on mainstream benchmarks and ATCTrack achieves a new SOTA performance. The code and models will be released at: https://github.com/XiaokunFeng/ATCTrack.
Xiaokun Feng, Xuchen Li 0001, Dailing Zhang, Meiqi Wu, Jing Zhang 0110, Xiaotang Chen, Kaiqi Huang
ICCV8
2025 A Temporal Modeling Framework for Video Pre-Training on Video Instance Segmentation
abstract
Contemporary Video Instance Segmentation (VIS) methods typically adhere to a pre-train then fine-tune regime, where a segmentation model trained on images is fine-tuned on videos. However, the lack of temporal knowledge in the pre-trained model introduces a domain gap which may adversely affect the VIS performance. To effectively bridge this gap, we present a novel "video pre-training" approach to enhance VIS models, especially for videos with intricate instance relationships. Our crucial innovation focuses on reducing disparities between the pre-training and fine-tuning stages. Specifically, we first introduce consistent pseudo-video augmentations to create diverse pseudo-video samples for pre-training while maintaining the instance consistency across frames. Then, we incorporate a multi-scale temporal module to enhance the model’s ability to model temporal relations through self- and cross-attention at short- and long-term temporal spans. Our approach does not set constraints on model architecture and can integrate seamlessly with various VIS methods. Experiment results on commonly adopted VIS benchmarks show that our method consistently outperforms state-of-the-art methods. Our approach achieves a notable 4.0% increase in average precision on the challenging OVIS dataset
Peng-Tao Jiang, Guodong Ding, Kaiqi Huang
ICME6
2025 LLM Data Selection and Utilization via Dynamic Bi-level Optimization
abstract
While large-scale training data is fundamental for developing capable large language models (LLMs), strategically selecting high-quality data has emerged as a critical approach to enhance training efficiency and reduce computational costs. Current data selection methodologies predominantly rely on static, training-agnostic criteria, failing to account for the dynamic model training and data interactions. In this paper, we propose a new Data Weighting Model (DWM) to adjust the weight of selected data within each batch to achieve a dynamic data utilization during LLM training. Specially, to better capture the dynamic data preference of the trained model, a bi-level optimization framework is implemented to update the weighting model. Our experiments demonstrate that DWM enhances the performance of models trained with randomly-selected data, and the learned weighting model can be transferred to enhance other data selection methods and models of different sizes. Moreover, we further analyze how a model’s data preferences evolve throughout training, providing new insights into the data preference of the model during training.
Yang Yu 0056, Kai Han 0002, Yehui Tang 0001, Kaiqi Huang, Yunhe Wang 0001, Dacheng Tao
ICML5
2025 CSTrack: Enhancing RGB-X Tracking via Compact Spatiotemporal Features
abstract
Effectively modeling and utilizing spatiotemporal features from RGB and other modalities (e.g., depth, thermal, and event data, denoted as X) is the core of RGB-X tracker design. Existing methods often employ two parallel branches to separately process the RGB and X input streams, requiring the model to simultaneously handle two dispersed feature spaces, which complicates both the model structure and computation process. More critically, intra-modality spatial modeling within each dispersed space incurs substantial computational overhead, limiting resources for inter-modality spatial modeling and temporal modeling. To address this, we propose a novel tracker, CSTrack, which focuses on modeling Compact Spatiotemporal features to achieve simple yet effective tracking. Specifically, we first introduce an innovative Spatial Compact Module that integrates the RGB-X dual input streams into a compact spatial feature, enabling thorough intra- and inter-modality spatial modeling. Additionally, we design an efficient Temporal Compact Module that compactly represents temporal features by constructing the refined target distribution heatmap. Extensive experiments validate the effectiveness of our compact spatiotemporal modeling method, with CSTrack achieving new SOTA results on mainstream RGB-X benchmarks. The code and models will be released at: https://github.com/XiaokunFeng/CSTrack.
Xiaokun Feng, Dailing Zhang, Xuchen Li 0001, Meiqi Wu, Jing Zhang 0110, Xiaotang Chen, Kaiqi Huang
ICML8
2025 Uncertainty-Aware Opponent Modeling for Deep Reinforcement Learning
Likun Yang, Pei Xu 0003, Shiyue Cao, Xiaotang Chen, Kaiqi Huang
AAMAS6
2025 Constructive Conflict-Driven Multi-Agent Reinforcement Learning for Strategic Diversity
abstract
In recent years, diversity has emerged as a useful mechanism to enhance the efficiency of multi-agent reinforcement learning (MARL). However, existing methods predominantly focus on designing policies based on individual agent characteristics, often neglecting the interplay and mutual influence among agents during policy formation. To address this gap, we propose Competitive Diversity through Constructive Conflict (CoDiCon), a novel approach that incorporates competitive incentives into cooperative scenarios to encourage policy exchange and foster strategic diversity among agents. Drawing inspiration from sociological research, which highlights the benefits of moderate competition and constructive conflict in group decision-making, we design an intrinsic reward mechanism using ranking features to introduce competitive motivations. A centralized intrinsic reward module generates and distributes varying reward values to agents, ensuring an effective balance between competition and cooperation. By optimizing the parameterized centralized reward module to maximize environmental rewards, we reformulate the constrained bilevel optimization problem to align with the original task objectives. We evaluate our algorithm against state-of-the-art methods in the SMAC and GRF environments. Experimental results demonstrate that CoDiCon achieves superior performance, with competitive intrinsic rewards effectively promoting diverse and adaptive strategies among cooperative agents.
Yuxiang Mai, Qiyue Yin, Wancheng Ni, Pei Xu 0003, Kaiqi Huang
IJCAI5
2025 AdaPT: Adaptive Policy Transfer for Multi-agent Reinforcement Learning with Domain Classifiers
abstract
Transfer learning has shown great potential in accelerating Multi-Agent Reinforcement Learning (MARL) training by adapting agent policies to varying input and output dimensions across tasks with different agent numbers. However, directly applying policies from previous tasks often leads to low performance due to domain differences, and policy fine-tuning is inefficient. In this paper, we propose a novel Adaptive Policy Transfer method in MARL (AdaPT), which needs the source domain data and only a small amount of the target domain data to learn a policy in the source domain that works in the target domain. Based on the probabilistic inference view over trajectories, we implement this process by modifying the reward function according to the similarity of source and target domain dynamics. Intuitively, AdaPT achieves policy transfer by encouraging more exploration and exploiting the similar state transitions in the source and target domains, making the policy more like in the target domain rather than the source domain. To enable the policy to perform in tasks with varying numbers of agents, we propose a transformer-based maximum entropy policy model. Besides, we use a novel centralized classifiers module to discriminate whether state transitions are similar. Experimental results on StarCraft II micro-management environment show that our base model outperforms or is comparable to the SOTA non-transfer models. In transfer tasks, our method outperforms the SOTA algorithms and achieves performance comparable to policies trained directly in the target domain.
Yuxiang Mai, Qiyue Yin, Wancheng Ni, Kaiqi Huang
IJCNN4
2025 Unveiling Chain of Step Reasoning for Vision-Language Models with Fine-grained Rewards
abstract
Chain of thought reasoning has demonstrated remarkable success in large language models, yet its adaptation to vision-language reasoning remains an open challenge with unclear best practices. Existing attempts typically employ reasoning chains at a coarse-grained level, which struggles to perform fine-grained structured reasoning and, more importantly, are difficult to evaluate the reward and quality of intermediate reasoning. In this work, we delve into chain of step reasoning for vision-language models, enabling assessing reasoning step quality accurately and leading to effective reinforcement learning and inference-time scaling with fine-grained rewards. We present a simple, effective, and fully transparent framework, including the step-level reasoning data, process reward model (PRM), and reinforcement learning training. With the proposed approaches, our models set strong baselines with consistent improvements on challenging vision-language benchmarks. More importantly, we conduct a thorough empirical analysis and ablation study, unveiling the impact of each component and several intriguing properties of inference-time scaling. We believe this paper serves as a baseline for vision-language models and offers insights into more complex multimodal reasoning. Our dataset, PRM, and code at https://github.com/baaivision/CoS.
Xingzhou Lou, Xiaokun Feng, Kaiqi Huang
NeurIPS4
2025 Relation-Aware Learning for Multitask Multiagent Cooperative Games
abstract
Collaboration among multiple tasks is advantageous for enhancing learning efficiency in multiagent reinforcement learning. To guide agents in cooperating with different teammates in multiple tasks, contemporary approaches encourage agents to exploit common cooperative patterns or identify the learning priorities of multiple tasks. Despite the progress made by these methods, they all assume that all cooperative tasks to be learned are related and desire similar agent policies. This is rarely the case in multiagent cooperation, where minor changes in team composition can lead to significant variations in cooperation, resulting in distinct cooperative strategies compete for limited learning resources. In this article, to tackle the challenge posed by multitask learning in potentially competing cooperative tasks, we propose a novel framework called relation-aware learning (RAL). RAL incorporates a relation awareness module in both task representation and task optimization, aiding in reasoning about task relationships and mitigating negative transfers among dissimilar tasks. To assess the performance of RAL, we conduct a comparative analysis with baseline methods in a multitaskStarCraftenvironment. The results demonstrate the superiority of RAL in multitask cooperative scenarios, particularly in scenarios involving multiple conflicting tasks.
Yang Yu 0056, Likun Yang, Zhourui Guo, Qiyue Yin, Junge Zhang, Kaiqi Huang
IEEE Trans. Games7
2025 Exploration via Embracing Diversity in Reinforcement Learning for Sparse-Reward Procedurally-Generated Tasks
abstract
A key challenge in reinforcement learning is how to guide agents to efficiently explore sparse reward environments. In order to overcome this challenge, the state-of-the-art methods introduce additional intrinsic rewards based on state-related information, such as the novelty of states. Unfortunately, these methods frequently fail in procedurally-generated tasks, where a different environment is generated in each episode so that the agent is not likely to visit the same state more than once. Recently, some exploration methods designed specifically for procedurally-generated tasks have been proposed. However, they still only consider state-related information, which leads to relatively inefficient exploration. In this work, we propose a novel exploration method, which utilizes cross-episode policy-related information and intraepisode state-related information to jointly encourage exploration in procedurally-generated tasks. In term of policy-related information, we first use an imitator-based unbalanced policy diversity to measure the difference between the agent’s current policy and the agent’s previous policies, and then encourage the agent to maximize this difference. In term of state-related information, we encourage the agent to maximize the state diversity within an episode, thereby visiting as many different states as possible in an episode. We show that our method significantly improves sample efficiency over state-of-the-art methods on three challenging benchmarks, including MiniGrid, MiniWorld, and the sparse-reward version of Procgen.
Pei Xu 0003, Hao Chen 0103, Wenjie Yang 0005, Kaiqi Huang
IEEE Trans. Syst. Man Cybern. Syst.4
2024 DDAE: Towards Deep Dynamic Vision BERT Pretraining
abstract
Recently, masked image modeling (MIM) has demonstrated promising prospects in self-supervised representation learning. However, existing MIM frameworks recover all masked patches equivalently, ignoring that the reconstruction difficulty of different patches can vary sharply due to their diverse distance from visible patches. In this paper, we propose a novel deep dynamic supervision to enable MIM methods to dynamically reconstruct patches with different degrees of difficulty at different pretraining phases and depths of the model. Our deep dynamic supervision helps to provide more locality inductive bias for ViTs especially in deep layers, which inherently makes up for the absence of local prior for self-attention mechanism. Built upon the deep dynamic supervision, we propose Deep Dynamic AutoEncoder (DDAE), a simple yet effective MIM framework that utilizes dynamic mechanisms for pixel regression and feature self-distillation simultaneously. Extensive experiments across a variety of vision tasks including ImageNet classification, semantic segmentation on ADE20K and object detection on COCO demonstrate the effectiveness of our approach.
Xiangwen Kong, Xiangyu Zhang 0005, Xin Zhao 0012, Kaiqi Huang
AAAI5
2024 TAPE: Leveraging Agent Topology for Cooperative Multi-Agent Policy Gradient
abstract
Multi-Agent Policy Gradient (MAPG) has made significant progress in recent years. However, centralized critics in state-of-the-art MAPG methods still face the centralized-decentralized mismatch (CDM) issue, which means sub-optimal actions by some agents will affect other agent's policy learning. While using individual critics for policy updates can avoid this issue, they severely limit cooperation among agents. To address this issue, we propose an agent topology framework, which decides whether other agents should be considered in policy gradient and achieves compromise between facilitating cooperation and alleviating the CDM issue. The agent topology allows agents to use coalition utility as learning objective instead of global utility by centralized critics or local utility by individual critics. To constitute the agent topology, various models are studied. We propose Topology-based multi-Agent Policy gradiEnt (TAPE) for both stochastic and deterministic MAPG methods. We prove the policy improvement theorem for stochastic TAPE and give a theoretical explanation for the improved cooperation among agents. Experiment results on several benchmarks show the agent topology is able to facilitate agent cooperation and alleviate CDM issue respectively to improve performance of TAPE. Finally, multiple ablation studies and a heuristic graph search algorithm are devised to show the efficacy of the agent topology.
Xingzhou Lou, Junge Zhang, Timothy J. Norman, Kaiqi Huang, Yali Du 0001
AAAI4
2024 PeLK: Parameter-Efficient Large Kernel ConvNets with Peripheral Convolution
abstract
Recently, some large kernel convnets strike back with appealing performance and efficiency. However, given the square complexity of convolution, scaling up kernels can bring about an enormous amount of parameters and the proliferated parameters can induce severe optimization problem. Due to these issues, current CNNs compromise to scale up to 51 × 51 in the form of stripe convolution (i.e., 51 ×5 + 5 ×51) and start to saturate as the kernel size continues growing. In this paper, we delve into addressing these vital issues and explore whether we can continue scaling up kernels for more performance gains. Inspired by human vision, we propose a human-like peripheral convolution that efficiently reduces over 90% parameter count of dense grid convolution through parameter sharing, and manage to scale up kernel size to extremely large. Our peripheral convolution behaves highly similar to human, reducing the complexity of convolution from O(K2) to O(logK) without backfiring performance. Built on this, we propose Parameter-efficient Large Kernel Network (PeLK). Our PeLK outperforms modern vision Transformers and ConvNet architectures like Swin, ConvNeXt, RepLKNet and SLaK on various vision tasks including ImageNet classification, semantic segmentation on ADE20K and object detection on MS COCO. For the first time, we successfully scale up the kernel size of CNNs to an unprecedented 101 × 101 and demonstrate consistent improvements.
Xiangxiang Chu, Xin Zhao 0012, Kaiqi Huang
CVPR5
2024 Information Bottleneck Based Data Correction in Continual Learning
Mingyi Zhang 0004, Junge Zhang, Kaiqi Huang
ECCV (87)4
2024 Task-Wise Prompt Query Function for Rehearsal-Free Continual Learning
abstract
Continual learning (CL) aims to enable a model to retain knowledge of old tasks while learning new ones. One effective approach to CL is based on data rehearsal method. However, this approach increases the cost of storing data and cannot be used when data from old tasks is unavailable for some reason. Recently, with the emergence of large- scale pre-trained transformer models, prompt-based methods have become an alternative to data rehearsal. These methods rely on a query mechanism to generate prompts and have demonstrated resistance to forgetting in CL scenarios without rehearsal. However, these methods generate prompts in a task-wise way while queries for samples in an instance-wise way, and usually directly use pre-trained models as the encoding function for generating queries. This may lead to data retrieval errors and failure to match the correct prompts. In contrast, we propose building a new task-wise prompt query function that can continuously learn as the task progresses, thereby avoiding the issue of pre-trained models being unable to correctly match appropriate sample-prompt pairs. Our approach improves the effectiveness of the current state-of- the-art methods and has been verified on a series of datasets through our experimental results.
Mingyi Zhang 0004, Junge Zhang, Kaiqi Huang
ICASSP4
2024 Revealing the Dark Secrets of Extremely Large Kernel ConvNets on Robustness
abstract
Robustness is a vital aspect to consider when deploying deep learning models into the wild. Numerous studies have been dedicated to the study of the robustness of vision transformers (ViTs), which have dominated as the mainstream backbone choice for vision tasks since the dawn of 2020s. Recently, some large kernel convnets make a comeback with impressive performance and efficiency. However, it still remains unclear whether large kernel networks are robust and the attribution of their robustness. In this paper, we first conduct a comprehensive evaluation of large kernel convnets’ robustness and their differences from typical small kernel counterparts and ViTs on six diverse robustness benchmark datasets. Then to analyze the underlying factors behind their strong robustness, we design experiments from both quantitative and qualitative perspectives to reveal large kernel convnets’ intriguing properties that are completely different from typical convnets. Our experiments demonstrate for the first time that pure CNNs can achieve exceptional robustness comparable or even superior to that of ViTs. Our analysis on occlusion invariance, kernel attention patterns and frequency characteristics provide novel insights into the source of robustness. Code available at: https://github.com/Lauch1ng/LKRobust.
Xiaokun Feng, Xiangxiang Chu, Kaiqi Huang
ICML5
2024 ADMN: Agent-Driven Modular Network for Dynamic Parameter Sharing in Cooperative Multi-Agent Reinforcement Learning
Yang Yu 0056, Qiyue Yin, Junge Zhang, Pei Xu 0003, Kaiqi Huang
IJCAI5
2024 Population-Based Diverse Exploration for Sparse-Reward Multi-Agent Tasks
Pei Xu 0003, Junge Zhang, Kaiqi Huang
IJCAI3
2024 MemVLT: Vision-Language Tracking with Adaptive Memory-based Prompts
abstract
Vision-language tracking (VLT) enhances traditional visual object tracking by integrating language descriptions, requiring the tracker to flexibly understand complex and diverse text in addition to visual information. However, most existing vision-language trackers still overly rely on initial fixed multimodal prompts, which struggle to provide effective guidance for dynamically changing targets. Fortunately, the Complementary Learning Systems (CLS) theory suggests that the human memory system can dynamically store and utilize multimodal perceptual information, thereby adapting to new scenarios. Inspired by this, (i) we propose a Memory-based Vision-Language Tracker (MemVLT). By incorporating memory modeling to adjust static prompts, our approach can provide adaptive prompts for tracking guidance. (ii) Specifically, the memory storage and memory interaction modules are designed in accordance with CLS theory. These modules facilitate the storage and flexible interaction between short-term and long-term memories, generating prompts that adapt to target variations. (iii) Finally, we conduct extensive experiments on mainstream VLT datasets (e.g., MGIT, TNL2K, LaSOT and LaSOT$_{ext}$). Experimental results show that MemVLT achieves new state-of-the-art performance. Impressively, it achieves 69.4% AUC on the MGIT and 63.3% AUC on the TNL2K, improving the existing best result by 8.4% and 4.7%, respectively.
Xiaokun Feng, Xuchen Li 0001, Dailing Zhang, Meiqi Wu, Jing Zhang 0110, Xiaotang Chen, Kaiqi Huang
NeurIPS8
2024 Beyond Accuracy: Tracking more like Human via Visual Search
abstract
Human visual search ability enables efficient and accurate tracking of an arbitrary moving target, which is a significant research interest in cognitive neuroscience. The recently proposed Central-Peripheral Dichotomy (CPD) theory sheds light on how humans effectively process visual information and track moving targets in complex environments. However, existing visual object tracking algorithms still fall short of matching human performance in maintaining tracking over time, particularly in complex scenarios requiring robust visual search skills. These scenarios often involve Spatio-Temporal Discontinuities (i.e., STDChallenge), prevalent in long-term tracking and global instance tracking. To address this issue, we conduct research from a human-like modeling perspective: (1) Inspired by the CPD, we pro- pose a new tracker named CPDTrack to achieve human-like visual search ability. The central vision of CPDTrack leverages the spatio-temporal continuity of videos to introduce priors and enhance localization precision, while the peripheral vision improves global awareness and detects object movements. (2) To further evaluate and analyze STDChallenge, we create the STDChallenge Benchmark. Besides, by incorporating human subjects, we establish a human baseline, creating a high- quality environment specifically designed to assess trackers’ visual search abilities in videos across STDChallenge. (3) Our extensive experiments demonstrate that the proposed CPDTrack not only achieves state-of-the-art (SOTA) performance in this challenge but also narrows the behavioral differences with humans. Additionally, CPDTrack exhibits strong generalizability across various challenging benchmarks. In summary, our research underscores the importance of human-like modeling and offers strategic insights for advancing intelligent visual target tracking. Code and models are available at https://github.com/ZhangDailing8/CPDTrack.
Dailing Zhang, Xiaokun Feng, Xuchen Li 0001, Meiqi Wu, Jing Zhang 0110, Kaiqi Huang
NeurIPS7
2024 VS-LLM: Visual-Semantic Depression Assessment Based on LLM for Drawing Projection Test
Meiqi Wu, Yaxuan Kang, Xuchen Li 0001, Xiaotang Chen, Yunfeng Kang, Weiqiang Wang 0001, Kaiqi Huang
PRCV (9)8
2024 An Asymmetric Game Theoretic Learning Model
Qiyue Yin, Tongtong Yu, Xueou Feng, Jun Yang 0028, Kaiqi Huang
PRCV (3)5
2024 Redact4Trace: A solution for auditing the data and tracing the users in the redactable blockchain
Jianwei Hu 0002, Kaiqi Huang, Genqing Bian, Yanpeng Cui 0002
Comput. Networks2
2024 SOTVerse: A User-Defined Task Space of Single Object Tracking
Xin Zhao 0012, Kaiqi Huang
Int. J. Comput. Vis.3
2024 Correction: SOTVerse: A User-Defined Task Space of Single Object Tracking
Xin Zhao 0012, Kaiqi Huang
Int. J. Comput. Vis.3
2024 Leveraging Joint-Action Embedding in Multiagent Reinforcement Learning for Cooperative Games
abstract
State-of-the-art multi-agent policy gradient (MAPG) methods have demonstrated convincing capability in many cooperative games. However, the exponentially growing joint-action space severely challenges the critic's value evaluation and hinders performance of MAPG methods. To address this issue, we augment Central-Q policy gradient with a joint-action embedding function and propose Mutual-information Maximization MAPG (M3APG). The joint-action embedding function makes joint-actions contain information of state transitions, which will improve the critic's generalization over the joint-action space by allowing it to infer joint-actions' outcomes. We theoretically prove that with a fixed joint-action embedding function, the convergence of M3APG is guaranteed. Experiment results on the StarCraft Multi-Agent Challenge (SMAC) demonstrate that M3APG gives evaluation results with better accuracy and outperform other MAPG basic models across various maps of multiple difficulty levels. We empirically show that our joint-action embedding model can be extended to value-based multi-agent reinforcement learning methods and state-of-the-art MAPG methods. Finally, we run ablation study to show that the usage of mutual information in our method is necessary and effective.
Xingzhou Lou, Junge Zhang, Yali Du 0001, Chao Yu 0004, Zhaofeng He 0001, Kaiqi Huang
IEEE Trans. Games6
2024 Deep Multitask Multiagent Reinforcement Learning With Knowledge Transfer
abstract
Despite the potential of Multi-Agent Reinforcement Learning (MARL) in addressing numerous complex tasks, training a single team of MARL agents to handle multiple diverse team tasks remains a challenge. In this paper, we introduce a novel Multi-task method based on Knowledge Transfer in cooperative MARL (MKT-MARL). By learning from task-specific teachers, our approach empowers a single team of agents to attain expert-level performance in multiple tasks. MKT-MARL utilizes a knowledge distillation algorithm specifically designed for the multi-agent architecture, which rapidly learns a team control policy incorporating common coordinated knowledge from the experience of task-specific teachers. Additionally, we enhance this training with teacher annealing, gradually shifting the model's learning from distillation towards environmental rewards. This enhancement helps the multi-task model surpass its single-task teachers. We extensively evaluate our algorithm using two commonly-used benchmarks: StarCraft II micro-management and multi-agent particle environment. The experimental results demonstrate that our algorithm outperforms both the single-task teachers and a jointly-trained team of agents. Extensive ablation experiments illustrate the effectiveness of the supervised knowledge transfer and the teacher annealing strategy.
Yuxiang Mai, Yifan Zang 0001, Qiyue Yin, Wancheng Ni, Kaiqi Huang
IEEE Trans. Games5
2024 Finger in Camera Speaks Everything: Unconstrained Air-Writing for Real-World
abstract
Air-writing is a challenging task that combines the fields of computer vision and natural language processing, offering an intuitive and natural approach for human-computer interaction. However, current air-writing solutions face two primary challenges: (1) their dependency on complex sensors (e.g., Radar, EEGs and others) for capturing precise handwritten trajectories, and (2) the absence of a video-based air-writing dataset that covers a comprehensive vocabulary range. These limitations impede their practicality in various real-world scenarios, including the use on devices like iPhones and laptops. To tackle these challenges, we present the groundbreaking air-writing Chinese character video dataset (AWCV-100K), serving as a pioneering benchmark for video-based air-writing. This dataset captures handwritten trajectories in various real-world scenarios using commonly accessible RGB cameras, eliminating the need for complex sensors. AWCV-100K includes 8.8 million video frames, encompassing the complete set of 3,755 characters from the GB2312-80 level-1 set (GB1). Furthermore, we introduce our baseline approach, the video-based character recognizer (VCRec). VCRec adeptly extracts fingertip features from sparse visual cues and employs a spatio-temporal sequence module for analysis. Experimental results showcase the superior performance of VCRec compared to existing models in recognizing air-written characters, both quantitatively and qualitatively. This breakthrough paves the way for enhanced human-computer interaction in real-world contexts. Moreover, our approach leverages affordable RGB cameras, enabling its applicability in a diverse range of scenarios. The code and data examples will be made public at https://github.com/wmeiqi/AWCV.
Meiqi Wu, Kaiqi Huang, Yuanqiang Cai, Yuzhong Zhao, Weiqiang Wang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2024 Contrastive Correlation Preserving Replay for Online Continual Learning
abstract
Online Continual Learning (OCL), as a core step towards achieving human-level intelligence, aims to incrementally learn and accumulate novel concepts from streaming data that can be seen only once, while alleviating catastrophic forgetting on previously acquired knowledge. Under this mode, the model needs to learn new classes or tasks in an online manner, and the data distribution may change over time. Moreover, task boundaries and identities are not available during training and evaluation. To balance the stability and plasticity of networks, in this work, we propose a replay-based framework for OCL, named Contrastive Correlation Preserving Replay (CCPR), which focuses on not only instances but also correlations between multiple instances. Specifically, besides the previous raw samples, the corresponding representations are stored in the memory and used to construct correlations for the past and the current model. To better capture correlation and higher-order dependencies, we maximize the low bound of mutual information between the past correlation and the current correlation by leveraging contrastive objectives. Furthermore, to improve the performance, we propose a new memory update strategy, which simultaneously encourages the balance and diversity of samples within the memory. With limited memory slots, it allows less redundant and more representative samples for later replay. We conduct extensive evaluations on several popular CL datasets, and experiments show that our method consistently outperforms the state-of-the-art methods and can effectively consolidate knowledge to alleviate forgetting.
Mingyi Zhang 0004, Mantian Li, Fusheng Zha, Junge Zhang, Lining Sun, Kaiqi Huang
IEEE Trans. Circuits Syst. Video Technol.7
2023 Subspace-Aware Exploration for Sparse-Reward Multi-Agent Tasks
abstract
Exploration under sparse rewards is a key challenge for multi-agent reinforcement learning problems. One possible solution to this issue is to exploit inherent task structures for an acceleration of exploration. In this paper, we present a novel exploration approach, which encodes a special structural prior on the reward function into exploration, for sparse-reward multi-agent tasks. Specifically, a novel entropic exploration objective which encodes the structural prior is proposed to accelerate the discovery of rewards. By maximizing the lower bound of this objective, we then propose an algorithm with moderate computational cost, which can be applied to practical tasks. Under the sparse-reward setting, we show that the proposed algorithm significantly outperforms the state-of-the-art algorithms in the multiple-particle environment, the Google Research Football and StarCraft II micromanagement tasks. To the best of our knowledge, on some hard tasks (such as 27m_vs_30m}) which have relatively larger number of agents and need non-trivial strategies to defeat enemies, our method is the first to learn winning strategies under the sparse-reward setting.
Pei Xu 0003, Junge Zhang, Qiyue Yin, Chao Yu 0004, Yaodong Yang 0001, Kaiqi Huang
AAAI6
2023 Efficient-VQGAN: Towards High-Resolution Image Generation with Efficient Vision Transformers
abstract
Vector-quantized image modeling has shown great potential in synthesizing high-quality images. However, generating high-resolution images remains a challenging task due to the quadratic computational overhead of the self-attention process. In this study, we seek to explore a more efficient two-stage framework for high-resolution image generation with improvements in the following three aspects. (1) Based on the observation that the first quantization stage has solid local property, we employ a local attention-based quantization model instead of the global attention mechanism used in previous methods, leading to better efficiency and reconstruction quality. (2) We emphasize the importance of multi-grained feature interaction during image generation and introduce an efficient attention mechanism that combines global attention (long-range semantic consistency within the whole image) and local attention (fined-grained details). This approach results in faster generation speed, higher generation fidelity, and improved resolution. (3) We propose a new generation pipeline incorporating autoencoding training and autoregressive generation strategy, demonstrating a better paradigm for image synthesis. Extensive experiments demonstrate the superiority of our approach in high-quality and high-resolution image reconstruction and generation.
Shiyue Cao, Yueqin Yin, Lianghua Huang, Yu Liu 0063, Xin Zhao 0012, Deli Zhao, Kaiqi Huang
ICCV7
2023 Re-parameterizing Your Optimizers rather than Architectures
Xiaohan Ding, Xiangyu Zhang 0005, Kaiqi Huang, Jungong Han, Guiguang Ding
ICLR4
2023 Exploration via Joint Policy Diversity for Sparse-Reward Multi-Agent Tasks
abstract
Exploration under sparse rewards is a key challenge for multi-agent reinforcement learning problems. Previous works argue that complex dynamics between agents and the huge exploration space in MARL scenarios amplify the vulnerability of classical count-based exploration methods when combined with agents parameterized by neural networks, resulting in inefficient exploration. In this paper, we show that introducing constrained joint policy diversity into a classical count-based method can significantly improve exploration when agents are parameterized by neural networks. Specifically, we propose a joint policy diversity to measure the difference between current joint policy and previous joint policies, and then use a filtering-based exploration constraint to further refine the joint policy diversity. Under the sparse-reward setting, we show that the proposed method significantly outperforms the state-of-the-art methods in the multiple-particle environment, the Google Research Football, and StarCraft II micromanagement tasks. To the best of our knowledge, on the hard 3s_vs_5z task which needs non-trivial strategies to defeat enemies, our method is the first to learn winning strategies without domain knowledge under the sparse-reward setting.
Pei Xu 0003, Junge Zhang, Kaiqi Huang
IJCAI3
2023 Underexplored Subspace Mining for Sparse-Reward Cooperative Multi-Agent Reinforcement Learning
abstract
Learning cooperation in sparse-reward multi-agent reinforcement learning is challenging, since agents need to explore in the large joint-state space with sparse feedback. However, in cooperative games, the cooperative target is often related to partial attributes, hence there is no need to treat the whole state space equally. Therefore, we propose Underexplored Subspace Mining (USM), a novel type of intrinsic reward that encourages agents to selectively explore partial attributes instead of wasting time on the whole state space to accelerate learning. Specially, considering that the target-related attributes are varying in different games and hard to predefine, we choose to focus on the underexplored subspace as an alternative, which is an automatic aggregation of the underexplored bottom-level dimensions without any human design or learning parameters. We evaluate our method in cooperative games with discrete and continuous state space separately. Results demonstrate that USM consistently outperforms existing state-of-the-art methods, and becomes the only method that has succeeded in sparse-reward games evaluated with larger state space or more complicated cooperation dynamics.
Yang Yu 0056, Qiyue Yin, Junge Zhang, Hao Chen 0103, Kaiqi Huang
IJCNN5
2023 A Multi-modal Global Instance Tracking Benchmark (MGIT): Better Locating Target in Complex Spatio-temporal and Causal Relationship
abstract
Tracking an arbitrary moving target in a video sequence is the foundation for high-level tasks like video understanding. Although existing visual-based trackers have demonstrated good tracking capabilities in short video sequences, they always perform poorly in complex environments, as represented by the recently proposed global instance tracking task, which consists of longer videos with more complicated narrative content. Recently, several works have introduced natural language into object tracking, desiring to address the limitations of relying only on a single visual modality. However, these selected videos are still short sequences with uncomplicated spatio-temporal and causal relationships, and the provided semantic descriptions are too simple to characterize video content.To address these issues, we (1) first propose a new multi-modal global instance tracking benchmark named MGIT. It consists of 150 long video sequences with a total of 2.03 million frames, aiming to fully represent the complex spatio-temporal and causal relationships coupled in longer narrative content. (2) Each video sequence is annotated with three semantic grains (i.e., action, activity, and story) to model the progressive process of human cognition. We expect this multi-granular annotation strategy can provide a favorable environment for multi-modal object tracking research and long video understanding. (3) Besides, we execute comparative experiments on existing multi-modal object tracking benchmarks, which not only explore the impact of different annotation methods, but also validate that our annotation method is a feasible solution for coupling human understanding into semantic labels. (4) Additionally, we conduct detailed experimental analyses on MGIT, and hope the explored performance bottlenecks of existing algorithms can support further research in multi-modal object tracking. The proposed benchmark, experimental results, and toolkit will be released gradually on http://videocube.aitestunion.com/.
Dailing Zhang, Meiqi Wu, Xiaokun Feng, Xuchen Li 0001, Xin Zhao 0012, Kaiqi Huang
NeurIPS7
2023 A Hierarchical Theme Recognition Model for Sandplay Therapy
Xiaokun Feng, Xiaotang Chen, Kaiqi Huang
PRCV (4)4
2023 EKGRL: Entity-Based Knowledge Graph Representation Learning for Fact-Based Visual Question Answering
Xiaotang Chen, Kaiqi Huang
PRCV (6)3
2023 Global Instance Tracking: Locating Target More Like Humans
abstract
Target tracking, the essential ability of the human visual system, has been simulated by computer vision tasks. However, existing trackers perform well in austere experimental environments but fail in challenges like occlusion and fast motion. The massive gap indicates that researches only measure tracking performance rather than intelligence. How to scientifically judge the intelligence level of trackers? Distinct from decision-making problems, lacking three requirements (a challenging task, a fair environment, and a scientific evaluation procedure) makes it strenuous to answer the question. In this article, we first propose the global instance tracking (GIT) task, which is supposed to search an arbitrary user-specified instance in a video without any assumptions about camera or motion consistency, to model the human visual tracking ability. Whereafter, we construct a high-quality and large-scale benchmark VideoCube to create a challenging environment. Finally, we design a scientific evaluation procedure using human capabilities as the baseline to judge tracking intelligence. Additionally, we provide an online platform with toolkit and an updated leaderboard. Although the experimental results indicate a definite gap between trackers and humans, we expect to take a step forward to generate authentic human-like trackers. The database, toolkit, evaluation server, and baseline results are available at http://videocube.aitestunion.com.
Xin Zhao 0012, Lianghua Huang, Kaiqi Huang
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 QueryProp: Object Query Propagation for High-Performance Video Object Detection
abstract
Video object detection has been an important yet challenging topic in computer vision. Traditional methods mainly focus on designing the image-level or box-level feature propagation strategies to exploit temporal information. This paper argues that with a more effective and efficient feature propagation framework, video object detectors can gain improvement in terms of both accuracy and speed. For this purpose, this paper studies object-level feature propagation, and proposes an object query propagation (QueryProp) framework for high-performance video object detection. The proposed QueryProp contains two propagation strategies: 1) query propagation is performed from sparse key frames to dense non-key frames to reduce the redundant computation on non-key frames; 2) query propagation is performed from previous key frames to the current key frame to improve feature representation by temporal context modeling. To further facilitate query propagation, an adaptive propagation gate is designed to achieve flexible key frame selection. We conduct extensive experiments on the ImageNet VID dataset. QueryProp achieves comparable accuracy with state-of-the-art methods and strikes a decent accuracy/speed trade-off.
Naiyu Gao, Jian Jia, Xin Zhao 0012, Kaiqi Huang
AAAI5
2022 Learning Disentangled Attribute Representations for Robust Pedestrian Attribute Recognition
abstract
Although various methods have been proposed for pedestrian attribute recognition, most studies follow the same feature learning mechanism, \ie, learning a shared pedestrian image feature to classify multiple attributes. However, this mechanism leads to low-confidence predictions and non-robustness of the model in the inference stage. In this paper, we investigate why this is the case. We mathematically discover that the central cause is that the optimal shared feature cannot maintain high similarities with multiple classifiers simultaneously in the context of minimizing classification loss. In addition, this feature learning mechanism ignores the spatial and semantic distinctions between different attributes. To address these limitations, we propose a novel disentangled attribute feature learning (DAFL) framework to learn a disentangled feature for each attribute, which exploits the semantic and spatial characteristics of attributes. The framework mainly consists of learnable semantic queries, a cascaded semantic-spatial cross-attention (SSCA) module, and a group attention merging (GAM) module. Specifically, based on learnable semantic queries, the cascaded SSCA module iteratively enhances the spatial localization of attribute-related regions and aggregates region features into multiple disentangled attribute features, used for classification and updating learnable semantic queries. The GAM module splits attributes into groups based on spatial distribution and utilizes reliable group attention to supervise query attention maps. Experiments on PETA, RAPv1, PA100k, and RAPv2 show that the proposed method performs favorably against state-of-the-art methods.
Jian Jia, Naiyu Gao, Xiaotang Chen, Kaiqi Huang
AAAI5
2022 PanopticDepth: A Unified Framework for Depth-aware Panoptic Segmentation
abstract
This paper presents a unified framework for depth-aware panoptic segmentation (DPS), which aims to reconstruct 3D scene with instance-level semantics from one single image. Prior works address this problem by simply adding a dense depth regression head to panoptic segmentation (PS) networks, resulting in two independent task branches. This neglects the mutually-beneficial relations between these two tasks, thus failing to exploit handy instance-level semantic cues to boost depth accuracy while also producing sub-optimal depth maps. To overcome these limitations, we propose a unified framework for the DPS task by applying a dynamic convolution technique to both the PS and depth prediction tasks. Specifically, instead of predicting depth for all pixels at a time, we generate instance-specific kernels to predict depth and segmentation masks for each instance. Moreover, leveraging the instance-wise depth estimation scheme, we add additional instance-level depth cues to assist with supervising the depth learning via a new depth loss. Extensive experiments on Cityscapes-DPS and SemKITTI-DPS show the effectiveness and promise of our method. We hope our unified solution to DPS can lead a new paradigm in this area. Code is available at https://github.com/NaiyuGao/PanopticDepth.
Naiyu Gao, Jian Jia, Yanhu Shan, Xin Zhao 0012, Kaiqi Huang
CVPR7
2022 Layer-Wisely Supervised Learning For One-Shot Neural Architecture Search
abstract
Neural architecture search aims to automatically discover both efficient and effective neural architectures. Recently, one-shot neural architecture search (one-shot NAS) has drawn great attention due to its high efficiency and competitive performance. One of the most important problems in one-shot NAS is to evaluate the capabilities of architecture candidates. In particular, a pre-trained super-net is served as an evaluator. Due to the large weight-sharing space, current one-shot methods suffer from the ranking disorder issue, that is, the ranking correlation between estimated capabilities and true capabilities of candidates is incorrect. Moreover, the super-net in search is dense thus it is inefficient to train with end-to-end back-propagation. In this paper, we propose to modularize the large weight-sharing space of one-shot NAS into layers by introducing layer-wisely supervised learning. But we discover that greedy layer-wise learning that learns each layer separately with a local objective hurts super-net performance as well as ranking correlation. Instead, we learn each layer by using the gradients propagated from the objective associated with the adjacent upper layer. The simple proposal reduces the representation shift and improves the ranking correlation. In addition, it reduces 47.4% memory footprint and gets a faster convergence of super-net training compared with the strong baseline. Extensive experiments on ImageNet with both supervised and self-supervised objectives demonstrate the effectiveness of our proposal.
Zhourui Guo, Qiyue Yin, Hao Chen 0103, Kaiqi Huang
IJCNN5
2022 Multi-Agent Uncertainty Sharing for Cooperative Multi-Agent Reinforcement Learning
abstract
Cooperative multi-agent reinforcement learning has been considered promising to complete many complex cooperative tasks in the real world such as coordination of robot swarms and self-driving. To promote multi-agent cooperation, Centralized Training with Decentralized Execution emerges as a popular learning paradigm due to partial observability and communication constraints during execution and computational complexity in training. Value decomposition has been known to produce competitive performance to other methods in complex environment within this paradigm such as VDN and QMIX, which approximates the global joint Q-value function with multiple local individual Q-value functions. However, existing works often neglect the uncertainty of multiple agents resulting from the partial observability and very large action space in the multi-agent setting and can only obtain the sub-optimal policy. To alleviate the limitations above, building upon the value decomposition, we propose a novel method called multi-agent uncertainty sharing (MAUS). This method utilizes the Bayesian neural network to explicitly capture the uncertainty of all agents and combines with Thompson sampling to select actions for policy learning. Besides, we impose the uncertainty-sharing mechanism among agents to stabilize training as well as coordinate the behaviors of all the agents for multi-agent cooperation. Extensive experiments on the StarCraft Multi-Agent Challenge (SMAC) environment demonstrate that our approach achieves significant performance to exceed the prior baselines and verify the effectiveness of our method.
Hao Chen 0103, Guangkai Yang, Junge Zhang, Qiyue Yin, Kaiqi Huang
IJCNN5
2022 RACA: Relation-Aware Credit Assignment for Ad-Hoc Cooperation in Multi-Agent Deep Reinforcement Learning
abstract
In recent years, reinforcement learning has faced several challenges in the multi-agent domain, such as the credit assignment issue. Value function factorization emerges as a promising way to handle the credit assignment issue under the centralized training with decentralized execution (CTDE) paradigm. However, existing value function factorization methods cannot deal with ad-hoc cooperation, that is, adapting to new configurations of teammates at test time. Specifically, these methods do not explicitly utilize the relationship between agents and cannot adapt to different sizes of inputs. To address these limitations, we propose a novel method, called Relation-Aware Credit Assignment (RACA), which achieves zero-shot generalization in ad-hoc cooperation scenarios. RACA takes advantage of a graph-based relation encoder to encode the topological structure between agents. Furthermore, RACA utilizes an attention-based observation abstraction mechanism that can generalize to an arbitrary number of teammates with a fixed number of parameters. Experiments demonstrate that our method outperforms baseline methods on the StarCraftII micromanagement benchmark and ad-hoc cooperation scenarios.
Hao Chen 0103, Guangkai Yang, Junge Zhang, Qiyue Yin, Kaiqi Huang
IJCNN5
2022 FGA-NAS: Fast Resource-Constrained Architecture Search by Greedy-ADMM Algorithm
abstract
Differentiable architecture search has demonstrated promising results in automatically designing neural network architectures with desired properties, such as high accuracy and low FLOPs. However, it suffers from a cumbersome training process, and the injection of constraints in the search phase often relies on some hand-crafted heuristic regularizers, the design of which typically requires tremendous human effort. In this paper, to address these critical challenges, we present FGA-NAS, an efficient method for resource-constrained architecture search. First, to reduce the computational cost and improve search flexibility, we propose a novel condensed search space that merges multiple parallel-placed candidates into a single one. Second, to enable the gradient-based optimization for neural architecture search (NAS) under multiple combinatorial constraints, we decompose the constrained NAS into a few simple sub-problems without introducing any heuristics by using the ADMM algorithm [1]. Then, the constrained NAS can be resolved by alternately solving the simple sub-problems. Experimental results on ImageNet show that our method can discover efficient and accurate neural network architectures that achieve the state-of-the-art by only using 0.2 GPU days.
Junge Zhang, Qiaozhe Li, Hao Chen 0103, Kaiqi Huang
IJCNN5
2022 InsPro: Propagating Instance Query and Proposal for Online Video Instance Segmentation
abstract
Video instance segmentation (VIS) aims at segmenting and tracking objects in videos. Prior methods typically generate frame-level or clip-level object instances first and then associate them by either additional tracking heads or complex instance matching algorithms. This explicit instance association approach increases system complexity and fails to fully exploit temporal cues in videos. In this paper, we design a simple, fast and yet effective query-based framework for online VIS. Relying on an instance query and proposal propagation mechanism with several specially developed components, this framework can perform accurate instance association implicitly. Specifically, we generate frame-level object instances based on a set of instance query-proposal pairs propagated from previous frames. This instance query-proposal pair is learned to bind with one specific object across frames through conscientiously developed strategies. When using such a pair to predict an object instance on the current frame, not only the generated instance is automatically associated with its precursors on previous frames, but the model gets a good prior for predicting the same object. In this way, we naturally achieve implicit instance association in parallel with segmentation and elegantly take advantage of temporal clues in videos. To show the effectiveness of our method InsPro, we evaluate it on two popular VIS benchmarks, i.e., YouTube-VIS 2019 and YouTube-VIS 2021. Without bells-and-whistles, our InsPro with ResNet-50 backbone achieves 43.2 AP and 37.6 AP on these two benchmarks respectively, outperforming all other online VIS methods.
Naiyu Gao, Jian Jia, Yanhu Shan, Xin Zhao 0012, Kaiqi Huang
NeurIPS7
2022 Offline reinforcement learning with representations for actions
Xingzhou Lou, Qiyue Yin, Junge Zhang, Chao Yu 0004, Zhaofeng He 0001, Nengjie Cheng, Kaiqi Huang
Inf. Sci.7
2022 Temporal-adaptive sparse feature aggregation for video object detection
Qiaozhe Li, Xin Zhao 0012, Kaiqi Huang
Pattern Recognit.4
2022 Deep Reinforcement Learning With Part-Aware Exploration Bonus in Video Games
abstract
Reinforcement learning algorithms rely on carefully engineering environment rewards that are extrinsic to agents. However, environments with dense rewards are rare, motivating the need for developing reward functions that are intrinsic to agents. Curiosity is a type of successful intrinsic reward function, which uses the prediction error as an reward signal. In prior work, the prediction problem used to generate intrinsic rewards is optimized in the pixel space rather than a learnable feature space to avoid randomness caused by feature changes. However, these methods ignore small but important elements of the states that are often associated with locations of the character, which makes it impossible to generate accurate internal rewards for efficient exploration. In this article, we first demonstrate the effectiveness of introducing prior learned features for existing prediction-based exploration methods. Then, an attention map mechanism is designed to discretize learned features, thereby updating the learned feature and meanwhile reducing the impact of randomness on intrinsic rewards caused by the learning process of features. We verify our method on some video games from the standard reinforcement learning Atari benchmark, achieving clear improvements over random network distillation, which is one of the most advanced exploration methods, in almost all Atari games.
Pei Xu 0003, Qiyue Yin, Junge Zhang, Kaiqi Huang
IEEE Trans. Games4
2022 Bottom-Up Foreground-Aware Feature Fusion for Practical Person Search
abstract
The key to efficient person search is jointly localizing pedestrians and learning discriminative representation for person re-identification (re-ID). Some recently developed models are built with separate detection and re-ID branches on top of shared region feature extraction networks. There are two factors that are detrimental to re-ID feature learning. One is the background information redundancy resulting from the large receptive field of neurons. The other is the body part missing and background clutter caused by inaccurate localization. In this work, a bottom-up fusion (BUF) subnet is proposed to fuse the bounding box features pooled from multiple network stages. With a few parameters introduced, BUF leverages the multi-level features with various sizes of receptive fields to mitigate the background-bias problem. To further suppress the non-pedestrian regions, the newly introduced segmentation head generates a foreground probability map as guidance for the network to focus on the foreground regions. The resulting foreground attention module (FAM) enhances the foreground features. Moreover, for robust feature learning in practical person search, we propose to adaptively smooth the labels of the pedestrian boxes with consideration of the detection quality. Extensive experiments on PRW and CUHK-SYSU validate the effectiveness of the proposals. Our Bottom-Up Foreground-Aware Feature Fusion (BUFF) network with ALS achieves considerable gains over the state-of-the-art on PRW and competitive performance on CUHK-SYSU.
Wenjie Yang 0005, Houjing Huang, Xiaotang Chen, Kaiqi Huang
IEEE Trans. Circuits Syst. Video Technol.4
2021 Learning to Reweight Imaginary Transitions for Model-Based Reinforcement Learning
Wenzhen Huang, Qiyue Yin, Junge Zhang, Kaiqi Huang
AAAI4
2021 Spatial and Semantic Consistency Regularizations for Pedestrian Attribute Recognition
abstract
While recent studies on pedestrian attribute recognition have shown remarkable progress in leveraging complicated networks and attention mechanisms, most of them neglect the inter-image relations and an important prior: spatial consistency and semantic consistency of attributes under surveillance scenarios. The spatial locations of the same attribute should be consistent between different pedestrian images, e.g., the "hat" attribute and the "boots" attribute are always located at the top and bottom of the picture respectively. In addition, the inherent semantic feature of the "hat" attribute should be consistent, whether it is a baseball cap, beret, or helmet. To fully exploit inter-image relations and aggregate human prior in the model learning process, we construct a Spatial and Semantic Consistency (SSC) framework that consists of two complementary regularizations to achieve spatial and semantic consistency for each attribute. Specifically, we first propose a spatial consistency regularization to focus on reliable and stable attribute-related regions. Based on the precise attribute locations, we further propose a semantic consistency regularization to extract intrinsic and discriminative semantic features. We conduct extensive experiments on popular benchmarks including PA100K, RAP, and PETA. Results show that the proposed method performs favorably against state- of-the-art methods without increasing parameters.
Jian Jia, Xiaotang Chen, Kaiqi Huang
ICCV3
2021 CCF-Net: Composite Context Fusion Network with Inter-Slice Correlative Fusion for Multi-Disease Lesion Detection
abstract
Detecting lesions from computed tomography (CT) scans relies on two aspects of the input: intra-slice texture information from the key slice and inter-slice structural context information from the adjacent slices. However, most existing methods ignore the correlation and complementarity between texture and structural information resulting in unexpected loss of performance. In this paper, a novel Composite Context Fusion Network (CCF-Net) is proposed to jointly model intra-slice and inter-slice features so as to prove the effectiveness of the two-steam framework. To extract both texture and structural information, two streams of 2D and 3D convolutional modules are employed in each stage. Moreover, a Composite Fusion architecture equipped with Inter-slice Correlative Fusion (ICF) modules is proposed to achieve stage-by-stage feature fusion in order to excavate and exchange information between texture-aware and context-aware features. Extensive experiments show that the proposed CCF-Net is able to achieve state-of-the-art detection performance on the multi-disease CT lesion detection task and significantly surpass the baseline methods.1
Jiechao Ma, Shu Zhang 0001, Yemin Shi 0001, Junge Zhang, Kaiqi Huang, Yizhou Yu
ICIP6
2021 Can DNN Detectors Compete Against Human Vision in Object Detection Task?
Qiaozhe Li, Xin Zhao 0012, Kaiqi Huang
PRCV (1)4
2021 GOT-10k: A Large High-Diversity Benchmark for Generic Object Tracking in the Wild
abstract
We introduce here a large tracking database that offers an unprecedentedly wide coverage of common moving objects in the wild, called GOT-10k. Specifically, GOT-10k is built upon the backbone of WordNet structure [1] and it populates the majority of over 560 classes of moving objects and 87 motion patterns, magnitudes wider than the most recent similar-scale counterparts [19], [20], [23], [26]. By releasing the large high-diversity database, we aim to provide a unified training and evaluation platform for the development of class-agnostic, generic purposed short-term trackers. The features of GOT-10k and the contributions of this article are summarized in the following. (1) GOT-10k offers over 10,000 video segments with more than 1.5 million manually labeled bounding boxes, enabling unified training and stable evaluation of deep trackers. (2) GOT-10k is by far the first video trajectory dataset that uses the semantic hierarchy of WordNet to guide class population, which ensures a comprehensive and relatively unbiased coverage of diverse moving objects. (3) For the first time, GOT-10k introduces the one-shot protocol for tracker evaluation, where the training and test classes are zero-overlapped. The protocol avoids biased evaluation results towards familiar objects and it promotes generalization in tracker development. (4) GOT-10k offers additional labels such as motion classes and object visible ratios, facilitating the development of motion-aware and occlusion-aware trackers. (5) We conduct extensive tracking experiments with 39 typical tracking algorithms and their variants on GOT-10k and analyze their results in this paper. (6) Finally, we develop a comprehensive platform for the tracking community that offers full-featured evaluation toolkits, an online evaluation server, and a responsive leaderboard. The annotations of GOT-10k's test data are kept private to avoid tuning parameters on it.
Lianghua Huang, Xin Zhao 0012, Kaiqi Huang
IEEE Trans. Pattern Anal. Mach. Intell.3
2021 Universal adversarial perturbations against object detection
Debang Li, Junge Zhang, Kaiqi Huang
Pattern Recognit.3
2021 SSAP: Single-Shot Instance Segmentation With Affinity Pyramid
abstract
Proposal-free instance segmentation methods mainly generate instance-agnostic semantic segmentation labels and instance-aware features to group pixels into different object instances. However, previous methods mostly employ separate modules for these two sub-tasks and require multiple passes for inference. In addition to the lack of efficiency, previous methods also failed to perform as well as proposal-based approaches. To this end, this work proposes a single-shot proposal-free instance segmentation method that requires only one single pass for prediction. Our method is based on learning an affinity pyramid, which computes the probability that two pixels belong to the same instance in a hierarchical manner. Moreover, incorporating with the learned affinity pyramid, a novel cascaded graph partition (CGP) module is presented to fuse the two predictions and segment instances efficiently. As an additional contribution, we conduct an experiment to demonstrate the benefits of proposal-free methods in capturing detailed structures from finely annotated training examples. Our approach is evaluated on the Cityscapes and COCO datasets and achieves state-of-the-art performance.
Naiyu Gao, Yanhu Shan, Yupei Wang, Xin Zhao 0012, Kaiqi Huang
IEEE Trans. Circuits Syst. Video Technol.5
2021 Learning Category- and Instance-Aware Pixel Embedding for Fast Panoptic Segmentation
abstract
Panoptic segmentation (PS) is a complex scene understanding task that requires providing high-quality segmentation for both thing objects and stuff regions. Previous methods handle these two classes with semantic and instance segmentation modules separately, following with heuristic fusion or additional modules to resolve the conflicts between the two outputs. This work simplifies this pipeline of PS by consistently modeling the two classes with a novel PS framework, which extends a detection model with an extra module to predict category- and instance-aware pixel embedding (CIAE). CIAE is a novel pixel-wise embedding feature that encodes both semantic-classification and instance-distinction information. At the inference process, PS results are simply derived by assigning each pixel to a detected instance or a stuff class according to the learned embedding. Our method not only demonstrates fast inference speed but also the first one-stage method to achieve comparable performance to two-stage methods on the challenging COCO benchmark.
Naiyu Gao, Yanhu Shan, Xin Zhao 0012, Kaiqi Huang
IEEE Trans. Image Process.4
2020 Temporal Context Enhanced Feature Aggregation for Video Object Detection
abstract
Video object detection is a challenging task because of the presence of appearance deterioration in certain video frames. One typical solution is to aggregate neighboring features to enhance per-frame appearance features. However, such a method ignores the temporal relations between the aggregated frames, which is critical for improving video recognition accuracy. To handle the appearance deterioration problem, this paper proposes a temporal context enhanced network (TCENet) to exploit temporal context information by temporal aggregation for video object detection. To handle the displacement of the objects in videos, a novel DeformAlign module is proposed to align the spatial features from frame to frame. Instead of adopting a fixed-length window fusion strategy, a temporal stride predictor is proposed to adaptively select video frames for aggregation, which facilitates exploiting variable temporal information and requiring fewer video frames for aggregation to achieve better results. Our TCENet achieves state-of-the-art performance on the ImageNet VID dataset and has a faster runtime. Without bells-and-whistles, our TCENet achieves 80.3% mAP by only aggregating 3 frames.
Naiyu Gao, Qiaozhe Li, Senyao Du, Xin Zhao 0012, Kaiqi Huang
AAAI6
2020 GlobalTrack: A Simple and Strong Baseline for Long-Term Tracking
abstract
A key capability of a long-term tracker is to search for targets in very large areas (typically the entire image) to handle possible target absences or tracking failures. However, currently there is a lack of such a strong baseline for global instance search. In this work, we aim to bridge this gap. Specifically, we propose GlobalTrack, a pure global instance search based tracker that makes no assumption on the temporal consistency of the target's positions and scales. GlobalTrack is developed based on two-stage object detectors, and it is able to perform full-image and multi-scale search of arbitrary instances with only a single query as the guide. We further propose a cross-query loss to improve the robustness of our approach against distractors. With no online learning, no punishment on position or scale changes, no scale smoothing and no trajectory refinement, our pure global instance search based tracker achieves comparable, sometimes much better performance on four large-scale tracking benchmarks (i.e., 52.1% AUC on LaSOT, 63.8% success rate on TLP, 60.3% MaxGM on OxUvA and 75.4% normalized precision on TrackingNet), compared to state-of-the-art approaches that typically require complex post-processing. More importantly, our tracker runs without cumulative errors, i.e., any type of temporary tracking failures will not affect its performance on future frames, making it ideal for long-term tracking. We hope this work will be a strong baseline for long-term tracking and will stimulate future works in this area.
Lianghua Huang, Xin Zhao 0012, Kaiqi Huang
AAAI3
2020 Point Cloud Super Resolution with Adversarial Residual Graph Networks
Huikai Wu, Kaiqi Huang
BMVC2
2020 Composing Good Shots by Exploiting Mutual Relations
abstract
Finding views with a good composition from an input image is a common but challenging problem. There are usually at least dozens of candidates (regions) in an image, and how to evaluate these candidates is subjective. Most existing methods only use the feature corresponding to each candidate to evaluate the quality. However, the mutual relations between the candidates from an image play an essential role in composing a good shot due to the comparative nature of this problem. Motivated by this, we propose a graph-based module with a gated feature update to model the relations between different candidates. The candidate region features are propagated on a graph that models mutual relations between different regions for mining the useful information such that the relation features and region features are adaptively fused. We design a multi-task loss to train the model, especially, a regularization term is adopted to incorporate the prior knowledge about the relations into the graph. A data augmentation method is also developed by mixing nodes from different graphs to improve the model generalization ability. Experimental results show that the proposed model performs favorably against state-of-the-art methods, and comprehensive ablation studies demonstrate the contribution of each module and graph-based inference of the proposed method.
Debang Li, Junge Zhang, Kaiqi Huang, Ming-Hsuan Yang 0001
CVPR3
2020 Learning to Learn Cropping Models for Different Aspect Ratio Requirements
abstract
Image cropping aims at improving the framing of an image by removing its extraneous outer areas, which is widely used in the photography and printing industry. In some cases, the aspect ratio of cropping results is specified depending on some conditions. In this paper, we propose a meta-learning (learning to learn) based aspect ratio specified image cropping method called Mars, which can generate cropping results of different expected aspect ratios. In the proposed method, a base model and two meta-learners are obtained during the training stage. Given an aspect ratio in the test stage, a new model with new parameters can be generated from the base model. Specifically, the two meta-learners predict the parameters of the base model based on the given aspect ratio. The learning process of the proposed method is learning how to learn cropping models for different aspect ratio requirements, which is a typical meta-learning process. In the experiments, the proposed method is evaluated on three datasets and outperforms most state-of-the-art methods in terms of accuracy and speed. In addition, both the intermediate and final results show that the proposed model can predict different cropping windows for an image depending on different aspect ratio requirements.
Debang Li, Junge Zhang, Kaiqi Huang
CVPR3
2020 Human Parsing Based Alignment With Multi-Task Learning For Occluded Person Re-Identification
abstract
Person re-identification (ReID) has obtained great progress in recent years. However, the problem caused by occlusion, which is frequent under surveillance camera, is not sufficiently addressed. When human body is occluded, extracted features are flooded with background noise. Moreover, without knowing location and visibility of parts, directly matching partial images with others will cause misalignment. To tackle the issue, we propose a model named HPNet to extract part-level features and predict visibility of each part, based on human parsing. By extracting features from semantic part regions and perform comparison with consideration of visibility, our method not only reduces background noise but also achieves alignment. Furthermore, ReID and human parsing are learned in a multi-task manner, without the need for an extra part model during testing. In addition to being efficient, the performance of our model surpasses previous methods by a large margin under occlusion scenarios.
Houjing Huang, Xiaotang Chen, Kaiqi Huang
ICME3
2020 Proxy Task Learning For Cross-Domain Person Re-Identification
abstract
Person re-identification (ReID) has achieved rapid improvement recently. However, exploiting the model in a new scene is always faced with huge performance drop. The cause lies in distribution discrepancy between domains, including both low-level (e.g. image quality) and high-level (e.g. pedestrian attribute) variance. To alleviate the problem of domain shift, we propose a novel framework Proxy Task Learning (PTL), which performs body perception tasks on target-domain images while training source-domain ReID, in a multi-task manner. The backbone is shared between tasks and domains, hence both low- and high-level distributions are deeply aligned. We experimentally verify two proxy tasks, i.e. human parsing and attribute recognition, that prominently enhance generalization of the model. When integrating our method into an existing cross-domain pipeline, we achieve state-of-the-art performance on large-scale benchmarks.
Houjing Huang, Xiaotang Chen, Kaiqi Huang
ICME3
2020 Bottom-Up Foreground-Aware Feature Fusion for Person Search
abstract
The key to efficient person search is jointly localizing pedestrians and learning discriminative representation for person re-identification (re-ID). Some recently developed task-joint models are built with separate detection and re-ID branches on top of shared region feature extraction networks, where the large receptive field of neurons leads to background information redundancy for the following re-ID task. Our diagnostic analysis indicates the task-joint model suffers from considerable performance drop when the background is replaced or removed. In this work, we propose a subnet to fuse the bounding box features that pooled from multiple ConvNet stages in a bottom-up manner, termed bottom-up fusion (BUF) network. With a few parameters introduced, BUF leverages the multi-level features with different sizes of receptive fields to mitigate the background-bias problem. Moreover, the newly introduced segmentation head generates a foreground probability map as guidance for the network to focus on the foreground regions. The resulting foreground attention module (FAM) enhances the foreground features. Extensive experiments on PRW and CUHK-SYSU validate the effectiveness of the proposals. Our Bottom-Up Foreground-Aware Feature Fusion (BUFF) network achieves considerable gains over the state-of-the- arts on PRW and competitive performance on CUHK-SYSU.
Wenjie Yang 0005, Dangwei Li, Xiaotang Chen, Kaiqi Huang
ACM Multimedia4
2020 Multi angle optimal pattern-based deep learning for automatic facial expression recognition
Deepak Kumar Jain 0001, Zhang Zhang 0001, Kaiqi Huang
Pattern Recognit. Lett.3
2020 Recurrent Prediction With Spatio-Temporal Attention for Crowd Attribute Recognition
abstract
Crowd attribute recognition is a challenging task for crowd video understanding because a crowd video often contains multiple attributes from various types. Traditional deep learning-based methods directly treat this recognition problem as a multiple binary classification problem and represent the video by vectorizing and fusing the separately learned spatial and temporal features in the fully connected layers. Therefore, the correlations between these attributes may not be well captured. In this paper, a bidirectional recurrent prediction model with a semantic-aware attention mechanism is proposed to explore the spatio-temporal and semantic relations between the attributes for more accurate recognition. The ConvLSTM is introduced for feature representation to capture the spatio-temporal structure of the crowd videos and facilitate the visual attention. The bidirectional recurrent attention module is proposed for sequential attribute prediction by associating each subcategory attributes to corresponding semantic-related regions iteratively. The experiments and evaluations on the challenging WWW crowd video dataset not only show that our approach significantly outperforms the state-of-the-art methods but also verify that our approach can effectively capture the spatio-temporal and semantic relations of the crowd attributes.
Qiaozhe Li, Xin Zhao 0012, Ran He 0001, Kaiqi Huang
IEEE Trans. Circuits Syst. Video Technol.4
2020 Improve Person Re-Identification With Part Awareness Learning
abstract
Person re-identification (ReID) aims to predict whether two images from different cameras belong to the same person. Due to low image quality and variance in view point and body pose, it remains a difficult task. To solve the task, a model is supposed to appropriately capture features that describe body regions for identification. With the simple intuition that explicitly incorporating ReID model with part awareness could be beneficial for learning a more discriminative feature space, we propose part segmentation as an assistant body perception task during the training of a ReID model. Specifically, we add a lightweight segmentation head to the backbone of ReID model during training, which is supervised with part labels. Note that our segmentation head is only introduced during training and that it does not change network input or the way of extracting ReID feature. Experiments show that part segmentation considerably improves the performance of ReID. Through quantitative and qualitative analyses, we further reveal that body part perception helps ReID model to capture a set of more diverse features from the body, with decreased similarity between part features and increased focus on different body regions. We experiment with various representative ReID models and achieve consistent improvement on several large-scale datasets including Market1501, CUHK03, DukeMTMC-reID and MSMT17. E.g. on MSMT17, our method increases Rank-1 Accuracy of GlobalPool-ResNet-50, PCB and MGN by 2.3%, 2.9% and 3.9%, respectively. Incorporated with MGN, our model achieves state-of-the-art performance, with Rank-1 Accuracy 95.8%, 78.8%, 90.0% and 84.0% on four datasets, respectively.
Houjing Huang, Wenjie Yang 0005, Jinbin Lin, Guan Huang 0003, Jiamiao Xu, Xiaotang Chen, Kaiqi Huang
IEEE Trans. Image Process.8
2019 Bootstrap Estimated Uncertainty of the Environment Model for Model-Based Reinforcement Learning
abstract
Model-based reinforcement learning (RL) methods attempt to learn a dynamics model to simulate the real environment and utilize the model to make better decisions. However, the learned environment simulator often has more or less model error which would disturb making decision and reduce performance. We propose a bootstrapped model-based RL method which bootstraps the modules in each depth of the planning tree. This method can quantify the uncertainty of environment model on different state-action pairs and lead the agent to explore the pairs with higher uncertainty to reduce the potential model errors. Moreover, we sample target values from their bootstrap distribution to connect the uncertainties at current and subsequent time-steps and introduce the prior mechanism to improve the exploration efficiency. Experiment results demonstrate that our method efficiently decreases model error and outperforms TreeQN and other stateof-the-art methods on multiple Atari games.
Wenzhen Huang, Junge Zhang, Kaiqi Huang
AAAI3
2019 Visual-Semantic Graph Reasoning for Pedestrian Attribute Recognition
abstract
Pedestrian attribute recognition in surveillance is a challenging task due to poor image quality, significant appearance variations and diverse spatial distribution of different attributes. This paper treats pedestrian attribute recognition as a sequential attribute prediction problem and proposes a novel visual-semantic graph reasoning framework to address this problem. Our framework contains a spatial graph and a directed semantic graph. By performing reasoning using the Graph Convolutional Network (GCN), one graph captures spatial relations between regions and the other learns potential semantic relations between attributes. An end-to-end architecture is presented to perform mutual embedding between these two graphs to guide the relational learning for each other. We verify the proposed framework on three large scale pedestrian attribute datasets including PETA, RAP, and PA100k. Experiments show superiority of the proposed method over state-of-the-art methods and effectiveness of our joint GCN structures for sequential attribute prediction.
Qiaozhe Li, Xin Zhao 0012, Ran He 0001, Kaiqi Huang
AAAI4
2019 3D Object Detection Using Scale Invariant and Feature Reweighting Networks
abstract
3D object detection plays an important role in a large number of real-world applications. It requires us to estimate the localizations and the orientations of 3D objects in real scenes. In this paper, we present a new network architecture which focuses on utilizing the front view images and frustum point clouds to generate 3D detection results. On the one hand, a PointSIFT module is utilized to improve the performance of 3D segmentation. It can capture the information from different orientations in space and the robustness to different scale shapes. On the other hand, our network obtains the useful features and suppresses the features with less information by a SENet module. This module reweights channel features and estimates the 3D bounding boxes more effectively. Our method is evaluated on both KITTI dataset for outdoor scenes and SUN-RGBD dataset for indoor scenes. The experimental results illustrate that our method achieves better performance than the state-of-the-art methods especially when point clouds are highly sparse.
Xin Zhao 0012, Zhe Liu 0033, Ruolan Hu, Kaiqi Huang
AAAI4
2019 Towards Rich Feature Discovery With Class Activation Maps Augmentation for Person Re-Identification
abstract
The fundamental challenge of small inter-person variation requires Person Re-Identification (Re-ID) models to capture sufficient fine-grained information. This paper proposes to discover diverse discriminative visual cues without extra assistance, e.g., pose estimation, human parsing. Specifically, a Class Activation Maps (CAM) augmentation model is proposed to expand the activation scope of baseline Re-ID model to explore rich visual cues, where the backbone network is extended by a series of ordered branches which share the same input but output complementary CAM. A novel Overlapped Activation Penalty is proposed to force the new branch to pay more attention to the image regions less activated by the old ones, such that spatial diverse visual features can be discovered. The proposed model achieves state-of-the-art results on three person Re-ID benchmarks. Moreover, a visualization approach termed ranking activation map (RAM) is proposed to explicitly interpret the ranking results in the test stage, which gives qualitative validations of the proposed method.
Wenjie Yang 0005, Houjing Huang, Zhang Zhang 0001, Xiaotang Chen, Kaiqi Huang, Shu Zhang 0001
CVPR5
2019 SSAP: Single-Shot Instance Segmentation With Affinity Pyramid
abstract
Recently, proposal-free instance segmentation has received increasing attention due to its concise and efficient pipeline. Generally, proposal-free methods generate instance-agnostic semantic segmentation labels and instance-aware features to group pixels into different object instances. However, previous methods mostly employ separate modules for these two sub-tasks and require multiple passes for inference. We argue that treating these two sub-tasks separately is suboptimal. In fact, employing multiple separate modules significantly reduces the potential for application. The mutual benefits between the two complementary sub-tasks are also unexplored. To this end, this work proposes a single-shot proposal-free instance segmentation method that requires only one single pass for prediction. Our method is based on a pixel-pair affinity pyramid, which computes the probability that two pixels belong to the same instance in a hierarchical manner. The affinity pyramid can also be jointly learned with the semantic class labeling and achieve mutual benefits. Moreover, incorporating with the learned affinity pyramid, a novel cascaded graph partition module is presented to sequentially generate instances from coarse to fine. Unlike previous time-consuming graph partition methods, this module achieves 5× speedup and 9% relative improvement on Average-Precision (AP). Our approach achieves new state of the art on the challenging Cityscapes dataset.
Naiyu Gao, Yanhu Shan, Yupei Wang, Xin Zhao 0012, Yinan Yu, Ming Yang 0007, Kaiqi Huang
ICCV7
2019 Bridging the Gap Between Detection and Tracking: A Unified Approach
abstract
Object detection models have been a source of inspiration for many tracking-by-detection algorithms over the past decade. Recent deep trackers borrow designs or modules from the latest object detection methods, such as bounding box regression, RPN and ROI pooling, and can deliver impressive performance. In this paper, instead of redesigning a new tracking-by-detection algorithm, we aim to explore a general framework for building trackers directly upon almost any advanced object detector. To achieve this, three key gaps must be bridged: (1) Object detectors are class-specific, while trackers are class-agnostic. (2) Object detectors do not differentiate intra-class instances, while this is a critical capability of a tracker. (3) Temporal cues are important for stable long-term tracking while they are not considered in still-image detectors. To address the above issues, we first present a simple target-guidance module for guiding the detector to locate target-relevant objects. Then a meta-learner is adopted for the detector to fast learn and adapt a target-distractor classifier online. We further introduce an anchored updating strategy to alleviate the problem of overfitting. The framework is instantiated on SSD and FasterRCNN, the typical one- and two-stage detectors, respectively. Experiments on OTB, UAV123 and NfS have verified our framework and show that our trackers can benefit from deeper backbone networks, as opposed to many recent trackers.
Lianghua Huang, Xin Zhao 0012, Kaiqi Huang
ICCV3
2019 SparseMask: Differentiable Connectivity Learning for Dense Image Prediction
abstract
In this paper, we aim at automatically searching an efficient network architecture for dense image prediction. Particularly, we follow the encoder-decoder style and focus on designing a connectivity structure for the decoder. To achieve that, we design a densely connected network with learnable connections, named Fully Dense Network, which contains a large set of possible final connectivity structures. We then employ gradient descent to search the optimal connectivity from the dense connections. The search process is guided by a novel loss function, which pushes the weight of each connection to be binary and the connections to be sparse. The discovered connectivity achieves competitive results on two segmentation datasets, while runs more than three times faster and requires less than half parameters compared to the state-of-the-art methods. An extensive experiment shows that the discovered connectivity is compatible with various backbones and generalizes well to other dense image prediction tasks.
Huikai Wu, Junge Zhang, Kaiqi Huang
ICCV3
2019 An Effective Adversarial Training Based Spatial-Temporal Network for Abnormal Behavior Detection
abstract
Unsupervised abnormal behavior detection has attracted much attention in recent years. It is a challenging task due to the undefinition and the sparsity of abnormal behaviors, etc. Existing generative model based methods usually perform poorly due to the unknown types of abnormal behaviors and insufficient exploiting of spatial-temporal information. In this paper, we propose a novel adversarial training based spatial-temporal network to tackle these problems. Firstly, we introduce an adversarial training strategy to deal with unknown types of abnormal behaviors. Secondly, to better explore spatial-temporal information, we design an effective two-stream spatial-temporal network, which is identity mapping free and spatial-temporal complementary. Finally, we combine them together to get the final adversarial spatial-temporal network. Our method is evaluated on various challenging public datasets and achieves the state-of-the-art performance.
Zhiyu Yin, Xiaotang Chen, Kaiqi Huang
ICIP3
2019 Pedestrian Attribute Recognition by Joint Visual-semantic Reasoning and Knowledge Distillation
abstract
Pedestrian attribute recognition in surveillance is a challenging task in computer vision due to significant pose variation, viewpoint change and poor image quality. To achieve effective recognition, this paper presents a graph-based global reasoning framework to jointly model potential visual-semantic relations of attributes and distill auxiliary human parsing knowledge to guide the relational learning. The reasoning framework models attribute groups on a graph and learns a projection function to adaptively assign local visual features to the nodes of the graph. After feature projection, graph convolution is utilized to perform global reasoning between the attribute groups to model their mutual dependencies. Then, the learned node features are projected back to visual space to facilitate knowledge transfer. An additional regularization term is proposed by distilling human parsing knowledge from a pre-trained teacher model to enhance feature representations. The proposed framework is verified on three large scale pedestrian attribute datasets including PETA, RAP, and PA-100k. Experiments show that our method achieves state-of-the-art results.
Qiaozhe Li, Xin Zhao 0012, Ran He 0001, Kaiqi Huang
IJCAI4
2019 MVP-Net: Multi-view FPN with Position-Aware Attention for Deep Universal Lesion Detection
Shu Zhang 0001, Junge Zhang, Kaiqi Huang, Yizhou Wang 0001, Yizhou Yu
MICCAI (6)4
2019 GP-GAN: Towards Realistic High-Resolution Image Blending
abstract
It is common but challenging to address high-resolution image blending in the automatic photo editing application. In this paper, we would like to focus on solving the problem of high-resolution image blending, where the composite images are provided. We propose a framework called Gaussian-Poisson Generative Adversarial Network (GP-GAN) to leverage the strengths of the classical gradient-based approach and Generative Adversarial Networks. To the best of our knowledge, it's the first work that explores the capability of GANs in high-resolution image blending task. Concretely, we propose Gaussian-Poisson Equation to formulate the high-resolution image blending problem, which is a joint optimization constrained by the gradient and color information. Inspired by the prior works, we obtain gradient information via applying gradient filters. To generate the color information, we propose a Blending GAN to learn the mapping between the composite images and the well-blended ones. Compared to the alternative methods, our approach can deliver high-resolution, realistic images with fewer bleedings and unpleasant artifacts. Experiments confirm that our approach achieves the state-of-the-art performance on Transient Attributes dataset. A user study on Amazon Mechanical Turk finds that the majority of workers are in favor of the proposed method. The source code is available in \urlhttps://github.com/wuhuikai/GP-GAN, and there's also an online demo in \urlhttp://wuhuikai.me/DeepJS.
Huikai Wu, Shuai Zheng 0001, Junge Zhang, Kaiqi Huang
ACM Multimedia4
2019 Semi-supervised Lesion Detection with Reliable Label Propagation and Missing Label Mining
Shu Zhang 0001, Junge Zhang, Kaiqi Huang
PRCV (2)5
2019 Mixed Supervised Object Detection with Robust Objectness Transfer
abstract
In this paper, we consider the problem of leveraging existing fully labeled categories to improve the weakly supervised detection (WSD) of new object categories, which we refer to as mixed supervised detection (MSD). Different from previous MSD methods that directly transfer the pre-trained object detectors from existing categories to new categories, we propose a more reasonable and robust objectness transfer approach for MSD. In our framework, we first learn domain-invariant objectness knowledge from the existing fully labeled categories. The knowledge is modeled based on invariant features that are robust to the distribution discrepancy between the existing categories and new categories; therefore the resulting knowledge would generalize well to new categories and could assist detection models to reject distractors (e.g., object parts) in weakly labeled images of new categories. Under the guidance of learned objectness knowledge, we utilize multiple instance learning (MIL) to model the concepts of both objects and distractors and to further improve the ability of rejecting distractors in weakly labeled images. Our robust objectness transfer approach outperforms the existing MSD methods, and achieves state-of-the-art results on the challenging ILSVRC2013 detection dataset and the PASCAL VOC datasets.
Yan Li 0043, Junge Zhang, Kaiqi Huang, Jianguo Zhang 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2019 Fast A3RL: Aesthetics-Aware Adversarial Reinforcement Learning for Image Cropping
abstract
Image cropping aims at improving the quality of images by removing unwanted outer areas, which is widely used in the photography and printing industry. Most previous cropping methods that don't need bounding box supervision rely on the sliding window mechanism. The sliding window method results in fixed aspect ratios and limits the shape of the cropping region. Moreover, the sliding window method usually produces lots of candidates on the input image, which is very time-consuming. Motivated by these challenges, we formulate image cropping as a sequential decision-making process and propose a reinforcement learning based framework to address this problem, namely Fast Aesthetics-Aware Adversarial Reinforcement Learning (Fast A3RL). Particularly, the proposed method develops an aesthetics-aware reward function, which is dedicated for image cropping. Similar to human's decisionmaking process, we use a comprehensive state representation including both the current observation and historical experience. We train the agent using the actor-critic architecture in an end-to-end manner. The adversarial learning process is also applied during the training stage. The proposed method is evaluated on several popular cropping datasets, in which the images are unseen during training. Experiment results show that our method achieves state-of-the-art performance with much fewer candidate windows and much less time compared with related methods.
Debang Li, Huikai Wu, Junge Zhang, Kaiqi Huang
IEEE Trans. Image Process.4
2019 A Richly Annotated Pedestrian Dataset for Person Retrieval in Real Surveillance Scenarios
abstract
Retrieving specific persons with various types of queries, e.g., a set of attributes or a portrait photo has great application potential in large-scale intelligent surveillance systems. In this paper, we propose a richly annotated pedestrian (RAP) dataset which serves as a unified benchmark for both attribute-based and image-based person retrieval in real surveillance scenarios. Typically, previous datasets have three improvable aspects, including limited data scale and annotation types, heterogeneous data source, and controlled scenarios. Differently, RAP is a large-scale dataset which contains 84928 images with 72 types of attributes and additional tags of viewpoint, occlusion, body parts, and 2589 person identities. It is collected in the real uncontrolled scene and has complex visual variations in pedestrian samples due to the change of viewpoints, pedestrian postures, and cloth appearance. Towards a high-quality person retrieval benchmark, an amount of state-of-the-art algorithms on pedestrian attribute recognition and person re-identification (ReID), are performed for quantitative analysis with three evaluation tasks, i.e., attribute recognition, attribute-based and image-based person retrieval, where a new instance-based metric is proposed to measure the dependency of the prediction of multiple attributes. Finally, some interesting problems, e.g., the joint feature learning of attribute recognition and ReID, and the problem of cross-day person ReID, are explored to show the challenges and future directions in person retrieval.
Dangwei Li, Zhang Zhang 0001, Xiaotang Chen, Kaiqi Huang
IEEE Trans. Image Process.4
2019 Focal Boundary Guided Salient Object Detection
abstract
The performance of salient object segmentation has been significantly advanced by using deep convolutional networks. However, these networks often produce blob-like saliency maps without accurate object boundaries. This is caused by the limited spatial resolution of their feature maps after multiple pooling operations, and might hinder downstream applications that require precise object shapes. To address this issue, we propose a novel deep model-Focal Boundary Guided (Focal- BG) network. Our model is designed to jointly learn to segment salient object masks and detect salient object boundaries. Our key idea is that additional knowledge about object boundaries can help to precisely identify the shape of the object. Moreover, our model incorporates a refinement pathway to refine the mask prediction, and makes use of the focal loss to facilitate the learning of the hard boundary pixels. To evaluate our model, we conduct extensive experiments. Our Focal-BG network consistently outperforms state-of-the-art methods on five major benchmarks. We provide a detailed analysis of these results and demonstrate that our joint modeling of salient object boundary and mask helps to better capture shape details, especially in the vicinity of object boundaries.
Yupei Wang, Xin Zhao 0012, Xuecai Hu, Yin Li 0003, Kaiqi Huang
IEEE Trans. Image Process.5
2019 Deep Crisp Boundaries: From Boundaries to Higher-Level Tasks
abstract
Edge detection has made significant progress with the help of deep convolutional networks (ConvNet). These ConvNet-based edge detectors have approached human level performance on standard benchmarks. We provide a systematical study of these detectors' outputs. We show that the detection results did not accurately localize edge pixels, which can be adversarial for tasks that require crisp edge inputs. As a remedy, we propose a novel refinement architecture to address the challenging problem of learning a crisp edge detector using ConvNet. Our method leverages a top-down backward refinement pathway, and progressively increases the resolution of feature maps to generate crisp edges. Our results achieve superior performance, surpassing human accuracy when using standard criteria on BSDS500, and largely outperforming the state-of-the-art methods when using more strict criteria. More importantly, we demonstrate the benefit of crisp edge maps for several important applications in computer vision, including optical flow estimation, object proposal generation, and semantic segmentation.
Yupei Wang, Xin Zhao 0012, Yin Li 0003, Kaiqi Huang
IEEE Trans. Image Process.4
2019 ISEE: An Intelligent Scene Exploration and Evaluation Platform for Large-Scale Visual Surveillance
abstract
Intelligent video surveillance (IVS) is always an interesting research topic to utilize visual analysis algorithms for exploring richly structured information from big surveillance data. However, existing IVS systems either struggle to utilize computing resources adequately to improve the efficiency of large-scale video analysis, or present a customized system for specific video analytic functions. It still lacks of a comprehensive computing architecture to enhance efficiency, extensibility and flexibility of IVS system. Moreover, it is also an open problem to study the effect of the combinations of multiple vision modules on the final performance of end applications of IVS system. Motivated by these challenges, we develop an Intelligent Scene Exploration and Evaluation (ISEE) platform based on a heterogeneous CPU-GPU cluster and some distributed computing tools, where Spark Streaming serves as the computing engine for efficient large-scale video processing and Kafka is adopted as a middle-ware message center to decouple different analysis modules flexibly. To validate the efficiency of the ISEE and study the evaluation problem on composable systems, we instantiate the ISEE for an end application on person retrieval with three visual analysis modules, including pedestrian detection with tracking, attribute recognition and re-identification. Extensive experiments are performed on a large-scale surveillance video dataset involving 25 camera scenes, totally 587 hours 720p synchronous videos, where a two-stage question-answering procedure is proposed to measure the performance of execution pipelines composed of multiple visual analysis algorithms based on millions of attribute-based and relationship-based queries. The case study of system-level evaluations may inspire researchers to improve visual analysis algorithms and combining strategies from the view of a scalable and composable system in the future.
Da Li 0003, Zhang Zhang 0001, Kai Yu 0003, Kaiqi Huang, Tieniu Tan
IEEE Trans. Parallel Distributed Syst.4
2018 Deep Semantic Structural Constraints for Zero-Shot Learning
abstract
Zero-shot learning aims to classify unseen image categories by learning a visual-semantic embedding space. In most cases, the traditional methods adopt a separated two-step pipeline that extracts image features are utilized to learn the embedding space. It leads to the lack of specific structural semantic information of image features for zero-shot learning task. In this paper, we propose an end-to-end trainable Deep Semantic Structural Constraints model to address this issue. The proposed model contains the Image Feature Structure constraint and the Semantic Embedding Structure constraint, which aim to learn structure-preserving image features and endue the learned embedding space with stronger generalization ability respectively. With the assistance of semantic structural information, the model gains more auxiliary clues for zero-shot learning. The state-of-the-art performance certifies the effectiveness of our proposed method.
Yan Li 0043, Junge Zhang, Kaiqi Huang, Tieniu Tan
AAAI4
2018 DF2Net: Discriminative Feature Learning and Fusion Network for RGB-D Indoor Scene Classification
abstract
This paper focuses on the task of RGB-D indoor scene classification. It is a very challenging task due to two folds. 1) Learning robust representation for indoor scene is difficult because of various objects and layouts. 2) Fusing the complementary cues in RGB and Depth is nontrivial since there are large semantic gaps between the two modalities. Most existing works learn representation for classification by training a deep network with softmax loss and fuse the two modalities by simply concatenating the features of them. However, these pipelines do not explicitly consider intra-class and inter-class similarity as well as inter-modal intrinsic relationships. To address these problems, this paper proposes a Discriminative Feature Learning and Fusion Network (DF2Net) with two-stage training. In the first stage, to better represent scene in each modality, a deep multi-task network is constructed to simultaneously minimize the structured loss and the softmax loss. In the second stage, we design a novel discriminative fusion network which is able to learn correlative features of multiple modalities and distinctive features of each modality. Extensive analysis and experiments on SUN RGB-D Dataset and NYU Depth Dataset V2 show the superiority of DF2Net over other state-of-the-art methods in RGB-D indoor scene classification task.
Yabei Li, Junge Zhang, Yanhua Cheng, Kaiqi Huang, Tieniu Tan
AAAI4
2018 Adversarially Occluded Samples for Person Re-Identification
abstract
Person re-identification (ReID) is the task of retrieving particular persons across different cameras. Despite its great progress in recent years, it is still confronted with challenges like pose variation, occlusion, and similar appearance among different persons. The large gap between training and testing performance with existing models implies the insufficiency of generalization. Considering this fact, we propose to augment the variation of training data by introducing Adversarially Occluded Samples. These special samples are both a) meaningful in that they resemble real-scene occlusions, and b) effective in that they are tough for the original model and thus provide the momentum to jump out of local optimum. We mine these samples based on a trained ReID model and with the help of network visualization techniques. Extensive experiments show that the proposed samples help the model discover new discriminative clues on the body and generalize much better at test time. Our strategy makes significant improvement over strong baselines on three large-scale ReID datasets, Market1501, CUHK03 and DukeMTMC-reID.
Houjing Huang, Dangwei Li, Zhang Zhang 0001, Xiaotang Chen, Kaiqi Huang
CVPR5
2018 A2-RL: Aesthetics Aware Reinforcement Learning for Image Cropping
abstract
Image cropping aims at improving the aesthetic quality of images by adjusting their composition. Most weakly supervised cropping methods (without bounding box supervision) rely on the sliding window mechanism. The sliding window mechanism requires fixed aspect ratios and limits the cropping region with arbitrary size. Moreover, the sliding window method usually produces tens of thousands of windows on the input image which is very time-consuming. Motivated by these challenges, we firstly formulate the aesthetic image cropping as a sequential decision-making process and propose a weakly supervised Aesthetics Aware Reinforcement Learning (A2-RL) framework to address this problem. Particularly, the proposed method develops an aesthetics aware reward function which especially benefits image cropping. Similar to human's decision making, we use a comprehensive state representation including both the current observation and the historical experience. We train the agent using the actor-critic architecture in an end-to-end manner. The agent is evaluated on several popular unseen cropping datasets. Experiment results show that our method achieves the state-of-the-art performance with much fewer candidate windows and much less time compared with previous weakly supervised methods.
Debang Li, Huikai Wu, Junge Zhang, Kaiqi Huang
CVPR4
2018 Discriminative Learning of Latent Features for Zero-Shot Recognition
abstract
Zero-shot learning (ZSL) aims to recognize unseen image categories by learning an embedding space between image and semantic representations. For years, among existing works, it has been the center task to learn the proper mapping matrices aligning the visual and semantic space, whilst the importance to learn discriminative representations for ZSL is ignored. In this work, we retrospect existing methods and demonstrate the necessity to learn discriminative representations for both visual and semantic instances of ZSL. We propose an end-to-end network that is capable of 1) automatically discovering discriminative regions by a zoom network; and 2) learning discriminative semantic representations in an augmented space introduced for both user-defined and latent attributes. Our proposed method is tested extensively on two challenging ZSL datasets, and the experiment results show that the proposed method significantly outperforms state-of-the-art methods.
Yan Li 0043, Junge Zhang, Jianguo Zhang 0001, Kaiqi Huang
CVPR4
2018 Fast End-to-End Trainable Guided Filter
abstract
Image processing and pixel-wise dense prediction have been advanced by harnessing the capabilities of deep learning. One central issue of deep learning is the limited capacity to handle joint upsampling. We present a deep learning building block for joint upsampling, namely guided filtering layer. This layer aims at efficiently generating the high-resolution output given the corresponding low-resolution one and a high-resolution guidance map. The proposed layer is composed of a guided filter, which is reformulated as a fully differentiable block. To this end, we show that a guided filter can be expressed as a group of spatial varying linear transformation matrices. This layer could be integrated with the convolutional neural networks (CNNs) and jointly optimized through end-to-end training. To further take advantage of end-to-end training, we plug in a trainable transformation function that generates task-specific guidance maps. By integrating the CNNs and the proposed layer, we form deep guided filtering networks. The proposed networks are evaluated on five advanced image processing tasks. Experiments on MIT-Adobe FiveK Dataset demonstrate that the proposed approach runs 10-100× faster and achieves the state-of-the-art performance. We also show that the proposed guided filtering layer helps to improve the performance of multiple pixel-wise dense prediction tasks. The code is available at https://github.com/wuhuikai/DeepGuidedFilter.
Huikai Wu, Shuai Zheng 0001, Junge Zhang, Kaiqi Huang
CVPR4
2018 Pose Guided Deep Model for Pedestrian Attribute Recognition in Surveillance Scenarios
abstract
Recognizing pedestrian attributes, such as gender, backpack, and cloth types, has obtained increasing attention recently due to its great potential in intelligent video surveillance. Existing methods usually solve it with end-to-end multi-label deep neural networks, while the structure knowledge of pedestrian body has been little utilized. Considering that attributes have strong spatial correlations with human structures, e.g. glasses are around the head, in this paper, we introduce pedestrian body structure into this task and propose a Pose Guided Deep Model (PGDM) to improve attribute recognition. The PGDM consists of three main components: 1) coarse pose estimation which distillates the pose knowledge from a pre-trained pose estimation model, 2) body parts localization which adaptively locates informative image regions with only image-level supervision, 3) multiple features fusion which combines the part-based features for attribute recognition. In the inference stage, we fuse the part-based PGDM results with global body based results for final attribute prediction and the performance can be consistently improved. Compared with state-of-the-art models, the performances on three large-scale pedestrian attribute datasets, i.e., PETA, RAP, and PA-100K, demonstrate the effectiveness of the proposed method.
Dangwei Li, Xiaotang Chen, Zhang Zhang 0001, Kaiqi Huang
ICME4
2018 Densely Connected Single-Shot Detector
abstract
One-stage object detection approach which utilizes multi-scale feature maps to predict objects is currently the best real-time detector. However, in this approach, the high-resolution feature maps which are responsible for detecting small objects are harder to learn a proper abstraction of objects than the low-resolution feature maps. The problem is that these feature maps have to transform sufficient low-level information to the next layer while learning high-level abstraction. In this paper, we develop a transformation module which adopts the dense structure to simplify the learning problem of high-resolution feature maps. In addition, we utilize the inception module to enrich the representation power of high-resolution feature maps. Extensive experiments on most object detection datasets clearly demonstrate the effectiveness of our method. In particular, on PASCAL VOC 2007/2012, our method outperforms all the existing one-stage methods. Our model based on the VGG-16 network also achieves competitive result on MS COCO.
Pei Xu 0003, Xin Zhao 0012, Kaiqi Huang
ICPR3
2018 Densely Cascaded Shadow Detection Network via Deeply Supervised Parallel Fusion
abstract
Shadow detection is an important and challenging problem in computer vision. Recently, single image shadow detection had achieved major progress with the development of deep convolutional networks. However, existing methods are still vulnerable to background clutters, and often fail to capture the global context of an input image. These global contextual and semantic cues are essential for accurately localizing the shadow regions. Moreover, rich spatial details are required to segment shadow regions with precise shape. To this end, this paper presents a novel model characterized by a deeply supervised parallel fusion (DSPF) network and a densely cascaded learning scheme. The DSPF network achieves a comprehensive fusion of global semantic cues and local spatial details by multiple stacked parallel fusion branches, which are learned in a deeply supervised manner. Moreover, the densely cascaded learning scheme is employed to refine the spatial details. Our method is evaluated on two widely used shadow detection benchmarks. Experimental results show that our method outperforms state-of-the-arts by a large margin.
Yupei Wang, Xin Zhao 0012, Yin Li 0003, Xuecai Hu, Kaiqi Huang
IJCAI5
2018 Gestalt laws based tracklets analysis for human crowd understanding
Zhang Zhang 0001, Kaiqi Huang
Pattern Recognit.3
2018 Special issue on Video Surveillance-oriented Biometrics
Changxing Ding, Kaiqi Huang, Vishal M. Patel, Brian C. Lovell
Pattern Recognit. Lett.2
2018 Random walk-based feature learning for micro-expression recognition
Deepak Kumar Jain 0001, Zhang Zhang 0001, Kaiqi Huang
Pattern Recognit. Lett.3
2017 A Multi-Task Deep Network for Person Re-Identification
abstract
Person re-identification (ReID) focuses on identifying people across different scenes in video surveillance, which is usually formulated as a binary classification task or a ranking task in current person ReID approaches. In this paper, we take both tasks into account and propose a multi-task deep network (MTDnet) that makes use of their own advantages and jointly optimize the two tasks simultaneously for person ReID. To the best of our knowledge, we are the first to integrate both tasks in one network to solve the person ReID. We show that our proposed architecture significantly boosts the performance. Furthermore, deep architecture in general requires a sufficient dataset for training, which is usually not met in person ReID. To cope with this situation, we further extend the MTDnet and propose a cross-domain architecture that is capable of using an auxiliary set to assist training on small target sets. In the experiments, our approach outperforms most of existing person ReID algorithms on representative datasets including CUHK03, CUHK01, VIPeR, iLIDS and PRID2011, which clearly demonstrates the effectiveness of the proposed approach.
Xiaotang Chen, Jianguo Zhang 0001, Kaiqi Huang
AAAI4
2017 Weakly-supervised Learning of Mid-level Features for Pedestrian Attribute Recognition and Localization
Kai Yu 0003, Biao Leng, Zhang Zhang 0001, Dangwei Li, Kaiqi Huang
BMVC6
2017 Beyond Triplet Loss: A Deep Quadruplet Network for Person Re-identification
abstract
Person re-identification (ReID) is an important task in wide area video surveillance which focuses on identifying people across different cameras. Recently, deep learning networks with a triplet loss become a common framework for person ReID. However, the triplet loss pays main attentions on obtaining correct orders on the training set. It still suffers from a weaker generalization capability from the training set to the testing set, thus resulting in inferior performance. In this paper, we design a quadruplet loss, which can lead to the model output with a larger inter-class variation and a smaller intra-class variation compared to the triplet loss. As a result, our model has a better generalization ability and can achieve a higher performance on the testing set. In particular, a quadruplet deep network using a margin-based online hard negative mining is proposed based on the quadruplet loss for the person ReID. In extensive experiments, the proposed network outperforms most of the state-of-the-art algorithms on representative datasets which clearly demonstrates the effectiveness of our proposed method.
Xiaotang Chen, Jianguo Zhang 0001, Kaiqi Huang
CVPR4
2017 Locality-Sensitive Deconvolution Networks with Gated Fusion for RGB-D Indoor Semantic Segmentation
abstract
This paper focuses on indoor semantic segmentation using RGB-D data. Although the commonly used deconvolution networks (DeconvNet) have achieved impressive results on this task, we find there is still room for improvements in two aspects. One is about the boundary segmentation. DeconvNet aggregates large context to predict the label of each pixel, inherently limiting the segmentation precision of object boundaries. The other is about RGB-D fusion. Recent state-of-the-art methods generally fuse RGB and depth networks with equal-weight score fusion, regardless of the varying contributions of the two modalities on delineating different categories in different scenes. To address the two problems, we first propose a locality-sensitive DeconvNet (LS-DeconvNet) to refine the boundary segmentation over each modality. LS-DeconvNet incorporates locally visual and geometric cues from the raw RGB-D data into each DeconvNet, which is able to learn to upsample the coarse convolutional maps with large context whilst recovering sharp object boundaries. Towards RGB-D fusion, we introduce a gated fusion layer to effectively combine the two LS-DeconvNets. This layer can learn to adjust the contributions of RGB and depth over each pixel for high-performance object recognition. Experiments on the large-scale SUN RGB-D dataset and the popular NYU-Depth v2 dataset show that our approach achieves new state-of-the-art results for RGB-D indoor semantic segmentation.
Yanhua Cheng, Rui Cai 0002, Zhiwei Li 0006, Xin Zhao 0012, Kaiqi Huang
CVPR5
2017 Learning Deep Context-Aware Features over Body and Latent Parts for Person Re-identification
abstract
Person Re-identification (ReID) is to identify the same person across different cameras. It is a challenging task due to the large variations in person pose, occlusion, background clutter, etc. How to extract powerful features is a fundamental problem in ReID and is still an open problem today. In this paper, we design a Multi-Scale Context-Aware Network (MSCAN) to learn powerful features over full body and body parts, which can well capture the local context knowledge by stacking multi-scale convolutions in each layer. Moreover, instead of using predefined rigid parts, we propose to learn and localize deformable pedestrian parts using Spatial Transformer Networks (STN) with novel spatial constraints. The learned body parts can release some difficulties, e.g. pose variations and background clutters, in part-based representation. Finally, we integrate the representation learning processes of full body and body parts into a unified framework for person ReID through multi-class person identification tasks. Extensive evaluations on current challenging large-scale person ReID datasets, including the image-based Market1501, CUHK03 and sequence-based MARS datasets, show that the proposed method achieves the state-of-the-art results.
Dangwei Li, Xiaotang Chen, Zhang Zhang 0001, Kaiqi Huang
CVPR4
2017 Deep Crisp Boundaries
abstract
Edge detection had made significant progress with the help of deep Convolutional Networks (ConvNet). ConvNet based edge detectors approached human level performance on standard benchmarks. We provide a systematical study of these detector outputs, and show that they failed to accurately localize edges, which can be adversarial for tasks that require crisp edge inputs. In addition, we propose a novel refinement architecture to address the challenging problem of learning a crisp edge detector using ConvNet. Our method leverages a top-down backward refinement pathway, and progressively increases the resolution of feature maps to generate crisp edges. Our results achieve promising performance on BSDS500, surpassing human accuracy when using standard criteria, and largely outperforming state-of-the-art methods when using more strict criteria. We further demonstrate the benefit of crisp edge maps for estimating optical flow and generating object proposals.
Yupei Wang, Xin Zhao 0012, Kaiqi Huang
CVPR3
2017 Automatic image cropping with aesthetic map and gradient energy map
abstract
Image cropping is a fundamental task in image editing to enhance the aesthetic quality of images. In this paper, we propose an automatic image cropping technique based on aesthetic map and gradient energy map. Instead of utilizing aesthetic rules in previous methods, we learn the aesthetic map by a deep convolutional neural network with a large-scale dataset for aesthetic quality assessment. The aesthetic map can highlight the discriminative image regions for high (or low) aesthetic quality category. The gradient energy map presents edge spatial distribution of images and is developed to compute the simplicity of images. Then a composition model is learned with the aesthetic map and gradient energy map to evaluate the quality of composition for crops. Moreover, an aesthetic preservation model is developed to compute the aesthetic information remained in crops to avoid cropping out high aesthetic regions. Experiments show that our approach significantly outperforms state-of-the-art cropping methods.
Yueying Kao, Ran He 0001, Kaiqi Huang
ICASSP3
2017 A Self-Paced Category-Aware Approach for Unsupervised Adaptation Networks
abstract
The success of deep neural networks usually relies on a large number of labeled training samples, which unfortunately are not easy to obtain in practice. Unsupervised domain adaptation focuses on the problem where there is no labeled data in the target domain. In this paper, we propose a novel deep unsupervised domain adaptation method that learns transferable features. Different from most existing methods, it attempts to learn a better domain-invariant feature representation by performing a category-wise adaptation to match the conditional distributions of samples with respect to each category. A self-paced learning strategy is used to bring the awareness of label information gradually, which makes the category-wise adaptation feasible even if the labels are unavailable in target domain. Then, we give detailed theoretical analysis to explain how the better performance is obtained. The experimental results show that our method outperforms the current state of the arts on standard domain adaptation datasets.
Wenzhen Huang, Peipei Yang, Kaiqi Huang
ICDM3
2017 Encyclopedia enhanced semantic embedding for zero-shot learning
abstract
There are tremendous object categories in the real world besides those in image datasets. Zero-shot learning aims to recognize image categories which are unseen in the training set. A large number of previous zero-shot learning models use word vectors of the class labels directly as category prototypes in the semantic embedding space. But word vectors cannot obtain the global knowledge of an image category sufficiently. In this paper, we propose a new encyclopedia enhanced semantic embedding model to promote the discriminative capability of word vector prototypes with the global knowledge of each image category. The proposed model extracts the TF-IDF key words from encyclopedia articles to acquire the global knowledge of each category. The convex combination of the key words' word vectors acts as the prototypes of the object categories. The prototypes of seen and unseen classes build up the embedding space where the nearest neighbour search is implemented to recognize the unseen images. The experiments show that the proposed method achieves the state-of-the-art performance on the challenging ImageNet Fall 2011 1k2hop dataset.
Junge Zhang, Kaiqi Huang, Tieniu Tan
ICIP3
2017 Semantics-guided multi-level RGB-D feature fusion for indoor semantic segmentation
abstract
Indoor RGB-D semantic segmentation is a new and challenging problem. Traditional methods usually apply two-stream convolutional neural networks (CNNs) to represent RGB and depth images respectively, and fuse the two streams on a specific layer. In this paper, we explore several fusion strategies based on this two-stream-CNN framework and point out such a single-layer fusion method cannot exploit the complementary RGB and depth cues well for semantic segmentation. To address this problem, we propose a novel Semantics-guided Multi-level feature fusion approach, which first learns deep feature representation from bottom to up, and then gradually fuses the RGB and depth features from high level to low level under the guidance of the semantic cues. Experimental results on SUN RGB-D dataset demonstrate the advantages of the proposed method over the state of the arts.
Yabei Li, Junge Zhang, Yanhua Cheng, Kaiqi Huang, Tieniu Tan
ICIP4
2017 Local structured representation for generic object detection
Junge Zhang, Kaiqi Huang, Tieniu Tan, Zhaoxiang Zhang 0001
Frontiers Comput. Sci.2
2017 GRMA: Generalized Range Move Algorithms for the Efficient Optimization of MRFs
Junge Zhang, Peipei Yang, Stephen J. Maybank, Kaiqi Huang
Int. J. Comput. Vis.5
2017 An Equalized Global Graph Model-Based Approach for Multicamera Object Tracking
abstract
Nonoverlapping multicamera visual object tracking typically consists of two steps: single-camera object tracking (SCT) and inter-camera object tracking (ICT). Most of tracking methods focus on SCT, which happens in the same scene, while for real surveillance scenes, ICT is needed and single-camera tracking methods cannot work effectively. In this paper, we try to improve the overall multicamera object tracking performance by a global graph model with an improved similarity metric. Our method treats the similarities of single-camera tracking and inter-camera tracking differently and obtains the optimization in a global graph model. The results show that our method can work better even in the condition of poor SCT.
Lijun Cao, Xiaotang Chen, Kaiqi Huang
IEEE Trans. Circuits Syst. Video Technol.4
2017 A Semi-Supervised Method for Surveillance-Based Visual Location Recognition
abstract
In this paper, we are devoted to solving the problem of crossing surveillance and mobile phone visual location recognition, especially for the case that the query and reference images are captured by mobile phone and surveillance camera, respectively. Besides, we also study the influence of the environmental condition variations on this problem. To explore that problem, we first build a cross-device location recognition dataset, which includes images of 22 locations taken by mobile phones and surveillance cameras under different time and weather conditions. Then based on careful analysis of the problems existing in the data, we specifically design a method which unifies an unsupervised subspace alignment method and the semi-supervised Laplacian support vector machine. Experiments are performed on our dataset. Compared with several related methods, our method shows to be more efficient on the problem of crossing surveillance and mobile phone visual location recognition. Furthermore, the influence of several factors such as feature, time, and weather is studied.
Pengcheng Liu 0001, Peipei Yang, Kaiqi Huang, Tieniu Tan
IEEE Trans. Cybern.4
2017 ORGM: Occlusion Relational Graphical Model for Human Pose Estimation
abstract
Articulated human pose estimation from monocular image is a challenging problem in computer vision. Occlusion is a main challenge for human pose estimation, which is largely ignored in popular tree structured models. The tree structured model is simple and convenient for exact inference, but short in modeling the occlusion coherence especially in the case of self-occlusion. We propose an occlusion relational graphical model, which is able to model both self-occlusion and occlusion by the other objects simultaneously. The proposed model can encode the interactions between human body parts and objects, and enables it to learn occlusion coherence from data discriminatively. We evaluate our model on several public benchmarks for human pose estimation, including challenging subsets featuring significant occlusion. The experimental results show that our method is superior to the previous state-of-the-arts, and is robust to occlusion for 2D human pose estimation.
Lianrui Fu, Junge Zhang, Kaiqi Huang
IEEE Trans. Image Process.3
2017 Deep Aesthetic Quality Assessment With Semantic Information
abstract
Human beings often assess the aesthetic quality of an image coupled with the identification of the image's semantic content. This paper addresses the correlation issue between automatic aesthetic quality assessment and semantic recognition. We cast the assessment problem as the main task among a multi-task deep model, and argue that semantic recognition task offers the key to address this problem. Based on convolutional neural networks, we employ a single and simple multi-task framework to efficiently utilize the supervision of aesthetic and semantic labels. A correlation item between these two tasks is further introduced to the framework by incorporating the inter-task relationship learning. This item not only provides some useful insight about the correlation but also improves assessment accuracy of the aesthetic task. In particular, an effective strategy is developed to keep a balance between the two tasks, which facilitates to optimize the parameters of the framework. Extensive experiments on the challenging Aesthetic Visual Analysis dataset and Photo.net dataset validate the importance of semantic recognition in aesthetic quality assessment, and demonstrate that multitask deep models can discover an effective aesthetic representation to achieve the state-of-the-art results.
Yueying Kao, Ran He 0001, Kaiqi Huang
IEEE Trans. Image Process.3
2017 Guest Editorial Introduction to the Special Issue on Large-Scale Video Analytics for Enhanced Security: Algorithms and Systems
abstract
Due to the rapid increase of the number of cameras used in the video surveillance and the huge needs of the smart city and public security, video surveillance by human beings is no longer suitable. Hence, since the end of the last century, video analytics for security or visual surveillance has become one of the hottest research topics. Wide-area video surveillance systems can have extremely high data rates and high data volumes. Therefore, the challenge of video analytics is to extract meaningful information efficiently from the huge flow of video data in order to produce high-level semantic descriptions of the activities occurring in the area under surveillance.
Kaiqi Huang, Tieniu Tan, Stephen J. Maybank, Rama Chellappa, Jake Aggarval
IEEE Trans. Syst. Man Cybern. Syst.1
2016 ReD-SFA: Relation Discovery Based Slow Feature Analysis for Trajectory Clustering
abstract
For spectral embedding/clustering, it is still an open problem on how to construct an relation graph to reflect the intrinsic structures in data. In this paper, we proposed an approach, named Relation Discovery based Slow Feature Analysis (ReD-SFA), for feature learning and graph construction simultaneously. Given an initial graph with only a few nearest but most reliable pairwise relations, new reliable relations are discovered by an assumption of reliability preservation, i.e., the reliable relations will preserve their reliabilities in the learnt projection subspace. We formulate the idea as a cross entropy (CE) minimization problem to reduce the discrepancy between two Bernoulli distributions parameterized by the updated distances and the existing relation graph respectively. Furthermore, to overcome the imbalanced distribution of samples, a Boosting-like strategy is proposed to balance the discovered relations over all clusters. To evaluate the proposed method, extensive experiments are performed with various trajectory clustering tasks, including motion segmentation, time series clustering and crowd detection. The results demonstrate that ReDSFA can discover reliable intra-cluster relations with high precision, and competitive clustering performance can be achieved in comparison with state-of-the-art.
Zhang Zhang 0001, Kaiqi Huang, Tieniu Tan, Peipei Yang, Jun Li 0010
CVPR2
2016 Learning temporally correlated representations using lstms for visual tracking
abstract
In this paper, we propose to learn object representations with inference from temporal correlation in videos to achieve effective visual tracking. Unlike traditional methods which perform feature learning either at image level or based on intuitive temporal constraint, we employ the recurrent network with Long Short Term Memory (LSTM) units to directly learn temporally correlated representations of the objects in long sequences. The recurrent network is pre-trained offline with auxiliary data and then online optimized to adapt to the target-specific object. A structured SVM is employed to account for the temporally correlated object appearance as well as distinguish the object from background distraction. Experiment results not only show that the appearance and dynamic patterns of the objects can be characterized via temporally correlated feature learning, but also demonstrate that the proposed tracking algorithm performs favorably against the state-of-the-art methods.
Qiaozhe Li, Xin Zhao 0012, Kaiqi Huang
ICIP3
2016 Joint crowd detection and semantic scene modeling using a Gestalt laws-based similarity
abstract
This paper presents a novel approach to detecting crowd groups and learning semantic regions with a Gestalt laws-based similarity. Different from the existing approaches based on optical flows or complete trajectories, our model adopts tracklets as the original input, because they carry more detailed information. Though those tracklets do not appear in the same duration, they are more robust to noise in crowd scene. According to the Gestalt laws of grouping, we propose three priors to define a unified similarity measure to calculate the affinities of pairs of original tracklets and pairs of representative tracklets in crowd groups. Therefore, the short-term crowd groups and the long-term semantic paths in crowded scene can be detected by a bottom-up hierarchical clustering algorithm simultaneously. Extensive experiments on hundreds of video clips demonstrate that our approach is effective and reliable for crowd detection and semantic scene understanding.
Zhang Zhang 0001, Kaiqi Huang
ICIP3
2016 Semi-Supervised Multimodal Deep Learning for RGB-D Object Recognition
Yanhua Cheng, Xin Zhao 0012, Rui Cai 0002, Zhiwei Li 0006, Kaiqi Huang, Yong Rui
IJCAI5
2016 FastLCD: Fast Label Coordinate Descent for the Efficient Optimization of 2D Label MRFs
Junge Zhang, Peipei Yang, Kaiqi Huang
IJCAI4
2016 Weakly Supervised Large Scale Object Localization with Multiple Instance Learning and Bag Splitting
abstract
Localizing objects of interest in images when provided with only image-level labels is a challenging visual recognition task. Previous efforts have required carefully designed features and have difficulty in handling images with cluttered backgrounds. Up-scaling to large datasets also poses a challenge to applying these methods to real applications. In this paper, we propose an efficient and effective learning framework called MILinear, which is able to learn an object localization model from large-scale data without using bounding box annotations. We integrate rich general prior knowledge into a learning model using a large pre-trained convolutional network. Moreover, to reduce ambiguity in positive images, we present a bag-splitting algorithm that iteratively generates new negative bags from positive ones. We evaluate the proposed approach on the challenging Pascal VOC 2007 dataset, and our method outperforms other state-of-the-art methods by a large margin; some results are even comparable to fully supervised models trained with bounding box annotations. To further demonstrate scalability, we also present detection results on the ILSVRC 2013 detection dataset, and our method outperforms supervised deformable part-based model without using box annotations.
Weiqiang Ren, Kaiqi Huang, Dacheng Tao, Tieniu Tan
IEEE Trans. Pattern Anal. Mach. Intell.2
2016 Hierarchical aesthetic quality assessment using deep convolutional neural networks
Yueying Kao, Kaiqi Huang, Stephen J. Maybank
Signal Process. Image Commun.2
2016 Severely Blurred Object Tracking by Learning Deep Image Representations
abstract
An implicit assumption in many generic object trackers is that the videos are blur free. However, motion blur is very common in real videos. The performance of a generic object tracker may drop significantly when it is applied to videos with severe motion blur. In this paper, we propose a new Tracking-Learning-Data approach to transfer a generic object tracker to a blur-invariant object tracker without deblurring image sequences. Before object tracking, a large set of unlabeled images is used to learn objects' visual prior knowledge, which is then transferred to the appearance model of a specific target. During object tracking, online training samples are collected from the tracking results and the context information. Different blur kernels are involved with the training samples to increase the robustness of the appearance model to severe blur, and the motion parameters of the object are estimated in the particle filter framework. Extensive experimental results demonstrate that the proposed algorithm can robustly track objects not only in severely blurred videos but also in other challenging scenes.
Jianwei Ding, Yongzhen Huang, Wei Liu 0023, Kaiqi Huang
IEEE Trans. Circuits Syst. Video Technol.4
2015 Convolutional Fisher Kernels for RGB-D Object Recognition
abstract
This paper studies the problem of improving object recognition using the novel RGB-D data. To address the problem, a new convolutional Fisher Kernels (CFK) method is proposed to represent RGB-D objects powerfully yet efficiently. The core idea of our approach is to integrate the both advantages of the convolutional neural networks (CNN) and Fisher Kernel encoding (FK): CNN model is flexible to adapt to new data sources, but requires for large amounts of training data with significant computational resources for good generalization, In comparison, FK encoding is able to represent objects powerfully and efficiently with small training data, however, its success highly depends on the well-designed SIFT features in literature, which may not be suitable for the new depth data. CFK can be interpreted as a two-layer feature learning structure to bridge the two models. The first layer employs a single-layer CNN to learn low-level translation ally invariant features for both RGB and depth data efficiently. The second layer aggregates the convolutional responses by FK encoding. Here 2D and 3D spatial pyramids are applied to further improve the Fisher vector representation of each modality. Experiments on RGB-D object recognition benchmarks demonstrate that our approach can achieve the state-of-the-art results.
Yanhua Cheng, Rui Cai 0002, Xin Zhao 0012, Kaiqi Huang
3DV4
2015 Cross-Domain Object Recognition Using Object Alignment
abstract
In this paper, we focus on the problem of cross-domain object recognition [4], which has long been one of the challenging problems in computer vision. This problem typically arises when training (source domain) and test (target domain) samples are drawn from different distributions. In the problem of object recognition, this case is usually caused by the situation that training and test samples are acquired under different sets of background, lighting, view point, resolution conditions, etc. One popular solution to the problem of cross-domain object recognition is minimizing the difference between the source and target distributions. Existing methods are devoted to minimizing that domain difference in a complex image space, which makes the problem hard to solve because of background influence, as shown in Figure 1 (a). Since the object and background are twisted in that image feature space, the discrepancy caused by background is difficult to eliminate, which makes it hard to learn optimal fS and fT for minimizing D( fS(XS), fT (XT )). To discount the influence of the background, we propose to minimize that difference using object alignment. As shown in Figure 1 (b), we minimize the domain difference by transferring to the feature space of aligned objects XS and XT , but not the image feature space having background influence. The key insight of our approach is that the difference between the source and target distributions can be reduced by discounting the influence from the ambiguous background. We define the semantic object as the object that occurs in all the images of one class. To discount the background influence, our primary goal is to automatically localize the semantic object so that the irrelevant background can be eliminated. Then based on the semantic object regions, we can learn an object detector that is robust to the influence of the irrelevant background and makes the crossdomain object recognition much easier than before. In addition, since our detectors are learned in a weakly supervised way, we utilize the classificaSource domain Selective search Object alignment — Topic discovery
Pengcheng Liu 0001, Peipei Yang, Kaiqi Huang, Tieniu Tan
BMVC4
2015 GRSA: Generalized range swap algorithm for the efficient optimization of MRFs
abstract
Markov Random Field (MRF) is an important tool and has been widely used in many vision tasks. Thus, the optimization of MRFs is a problem of fundamental importance. Recently, Veskler and Kumar et. al propose the range move algorithms, which are one of the most successful solvers to this problem. However, two problems have limited the applicability of previous range move algorithms: 1) They are limited in the types of energies they can handle (i.e. only truncated convex functions); 2) These algorithms tend to be very slow compared to other graph-cut based algorithms (e.g. α-expansion and αβ-swap). In this paper, we propose a generalized range swap algorithm (GRSA) for efficient optimization of MRFs. To address the first problem, we extend the GRSA to arbitrary semimetric energies by restricting the chosen labels in each move so that the energy is submodular on the chosen subset. Furthermore, to feasibly choose the labels satisfying the submodular condition, we provide a sufficient condition of the submodularity. For the second problem, unlike previous range move algorithms which execute the set of all possible range moves, we dynamically obtain the iterative moves by solving a set cover problem, which greatly reduces the number of moves during the optimization. Experiments show that the GRSA offers a great speedup over previous range swap algorithms, while it obtains competitive solutions.
Junge Zhang, Peipei Yang, Kaiqi Huang
CVPR4
2015 Query Adaptive Similarity Measure for RGB-D Object Recognition
abstract
This paper studies the problem of improving the top-1 accuracy of RGB-D object recognition. Despite of the impressive top-5 accuracies achieved by existing methods, their top-1 accuracies are not very satisfactory. The reasons are in two-fold: (1) existing similarity measures are sensitive to object pose and scale changes, as well as intra-class variations, and (2) effectively fusing RGB and depth cues is still an open problem. To address these problems, this paper first proposes a new similarity measure based on dense matching, through which objects in comparison are warped and aligned, to better tolerate variations. Towards RGB and depth fusion, we argue that a constant and golden weight doesn't exist. The two modalities have varying contributions when comparing objects from different categories. To capture such a dynamic characteristic, a group of matchers equipped with various fusion weights is constructed, to explore the responses of dense matching under different fusion configurations. All the response scores are finally merged following a learning-to-combination way, which provides quite good generalization ability in practice. The proposed approach win the best results on several public benchmarks, e.g., achieves 92.7% top-1 test accuracy on the Washington RGB-D object dataset, with a 5.1% improvement over the state-of-the-art.
Yanhua Cheng, Rui Cai 0002, Chi Zhang 0069, Zhiwei Li 0006, Xin Zhao 0012, Kaiqi Huang, Yong Rui
ICCV6
2015 Beyond Tree Structure Models: A New Occlusion Aware Graphical Model for Human Pose Estimation
abstract
Occlusion is a main challenge for human pose estimation, which is largely ignored in popular tree structure models. The tree structure model is simple and convenient for exact inference, but short in modeling the occlusion coherence especially in the case of self-occlusion. We propose an occlusion aware graphical model which is able to model both self-occlusion and occlusion by the other objects simultaneously. The proposed model structure can encodes the interactions between human body parts and objects, and hence enables it to learn occlusion coherence from data discriminatively. We evaluate our model on several public benchmarks for human pose estimation including challenging subsets featuring significant occlusion. The experimental results show that our method obtains comparable accuracy with the state-of-the-arts, and is robust to occlusion for 2D human pose estimation.
Lianrui Fu, Junge Zhang, Kaiqi Huang
ICCV3
2015 Context aware model for articulated human pose estimation
abstract
Simple tree model prevails for 2D pose estimation for its simplicity and efficiency. However, the limited kinetic constraints often lead to double-counting and damage the accuracy of leaf parts, and this is largely ignored in previous work. In this paper, we propose a novel enhanced tree model which incorporates both local kinetic constraints and global contextual constraints among non-adjacent parts. By introducing virtual parts, we are able to model richer constraints within a tree structure and dynamic programming can be utilized for efficient inference. Experiments on public benchmarks show that our method is more effective in tackling double counting problem and can improve the localization accuracy, especially for the challenging lower limbs.
Lianrui Fu, Junge Zhang, Kaiqi Huang
ICIP3
2015 Visual aesthetic quality assessment with a regression model
abstract
Aesthetic image analysis has drawn much attention in recent years. However, assessing the aesthetic quality especially aesthetic score prediction is a challenging problem. In this paper, we interpret aesthetic quality assessment as a regression problem and present a new framework by directly training a regression model using a neural network. Firstly, to extract the aesthetic features which are difficult to design manually, we utilize the convolutional network to learn the features. Then, a regression model is trained based on the aesthetic features. Different from classification models which can only predict aesthetic class (high or low) in most existing works, the regression model can predict continuous aesthetic score. Experimental results on a recently published large-scale dataset show that the proposed method can assess the degree of aesthetic quality similar to human visual system effectively and outperforms the state-of-the-art methods.
Yueying Kao, Kaiqi Huang
ICIP3
2015 Learning occlusion patterns using semantic phrases for object detection
abstract
Occlusion inference in image is a classical as well as difficult problem in computer vision. Most approaches model occlusion with different occlusion patterns in complex ways. In this paper, we propose a simple and efficient way to represent occlusion patterns called `occlusion pattern phrase', for example `dog occlude person'. These phrases model the occlusion patterns between two occluded objects, which can be used in applications like object detection and object classification. Here, we focus on using the occlusion pattern phrases in object detection. DPM is used to learn appearance models and a inference procedure is introduced to infer occlusion patterns with structural outputs. Unlike other methods, our method produces not only location boxes for different objects, but also demonstrates their occlusion relations. A new dataset with well annotated occlusion pattern images collected from Pascal VOC2007 and search engines like Bing and Google is introduced in this paper. Experiments show that our method outperforms the baseline and the occlusion pattern phrases can describe the relations between objects as we expect. In the future, we will explore the use of occlusion pattern phrases in scene understanding.
Jinde Liu, Kaiqi Huang, Tieniu Tan
ICIP2
2015 Semi-supervised learning and feature evaluation for RGB-D object recognition
Yanhua Cheng, Xin Zhao 0012, Kaiqi Huang, Tieniu Tan
Comput. Vis. Image Underst.3
2015 Tracking by local structural manifold learning in a new SSIR particle filter
Jianwei Ding, Yunqi Tang, Wei Liu 0023, Yongzhen Huang, Kaiqi Huang
Neurocomputing5
2015 ISEE Smart Home (ISH): Smart video analysis for home security
Junge Zhang, Yanhu Shan, Kaiqi Huang
Neurocomputing3
2015 How to use Bag-of-Words model better for image classification
Kaiqi Huang
Image Vis. Comput.2
2015 VFM: Visual Feedback Model for Robust Object Recognition
Kaiqi Huang
J. Comput. Sci. Technol.2
2015 Large scale crowd analysis based on convolutional neural network
Lijun Cao, Weiqiang Ren, Kaiqi Huang
Pattern Recognit.4
2015 Adaptive Slice Representation for Human Action Classification
abstract
Common action recognition methods describe an action sequence along with its time axis, i.e., first extracting features from the x y plane, and then modeling the dynamic changes along with the time axis. Other than the ordinary x y plane-based representation, other views, e.g., xt slice-based representation, may be more efficient to distinguish different actions. In this paper, we investigate different slicing views of the spatiotemporal volume to organize action sequences and propose an efficient slice representation for human action recognition. First, a minimum average entropy principle is proposed to select the optimal slicing angle for each action sequence adaptively. This allows the foreground pixels to be distributed in the fewest slices so as to reduce more uncertainty caused by the information dispersed in different slices. Then, the obtained slice sequence is transformed into a pair of 1-D signals to describe the distribution of foreground pixels along the time axis. Finally, the mel frequency cepstrum coefficient features are calculated to describe the spectrum characteristics of the 1-D signals over time. Thus, a 3-D spatiotemporal action volume is efficiently transformed into a low-dimensional spectrum features. Extensive experiments on the 2-D human action data sets (the UIUC and the WEIZMANN) as well as the Microsoft Research (MSR) Action3-D depth data set demonstrate the effectiveness of the slice-based representation, where the recognition performance can reach to the state-of-the-art level with high efficiency.
Yanhu Shan, Zhang Zhang 0001, Peipei Yang, Kaiqi Huang
IEEE Trans. Circuits Syst. Video Technol.4
2015 High-Order Topology Modeling of Visual Words for Image Classification
abstract
Modeling relationship between visual words in feature encoding is important in image classification. Recent methods consider this relationship in either image or feature space, and most of them incorporate only pairwise relationship (between visual words). However, in situations involving large variability in images, one cannot capture intrinsic invariance of intra-class images using low-order pairwise relationship. The result is not robust to larger variations in images. In addition, as the number of potential pairings grows exponentially with the number of visual words, the task of learning becomes computationally expensive. To overcome these two limitations, we propose an efficient classification framework that exploits high-order topology of visual words in the feature space, as follows. First, we propose a search algorithm that seeks dependence between the visual words. This dependence is used to construct higher order topology in the feature space. Then, the local features are encoded according to this higher order topology to improve the image classification. Experiments involving four common data sets, namely PASCAL VOC 2007, 15 Scenes, Caltech 101, and UIUC Sport Event, demonstrate that the dependence search significantly improves the efficiency of higher order topological construction, and consequently increases the image classification in all these data sets.
Kaiqi Huang, Dacheng Tao
IEEE Trans. Image Process.1
2015 Large-Scale Weakly Supervised Object Localization via Latent Category Learning
abstract
Localizing objects in cluttered backgrounds is challenging under large-scale weakly supervised conditions. Due to the cluttered image condition, objects usually have large ambiguity with backgrounds. Besides, there is also a lack of effective algorithm for large-scale weakly supervised localization in cluttered backgrounds. However, backgrounds contain useful latent information, e.g., the sky in the aeroplane class. If this latent information can be learned, object-background ambiguity can be largely reduced and background can be suppressed effectively. In this paper, we propose the latent category learning (LCL) in large-scale cluttered conditions. LCL is an unsupervised learning method which requires only image-level class labels. First, we use the latent semantic analysis with semantic object representation to learn the latent categories, which represent objects, object parts or backgrounds. Second, to determine which category contains the target object, we propose a category selection strategy by evaluating each category's discrimination. Finally, we propose the online LCL for use in large-scale conditions. Evaluation on the challenging PASCAL Visual Object Class (VOC) 2007 and the large-scale imagenet large-scale visual recognition challenge 2013 detection data sets shows that the method can improve the annotation precision by 10% over previous methods. More importantly, we achieve the detection precision which outperforms previous results by a large margin and can be competitive to the supervised deformable part model 5.0 baseline on both data sets.
Kaiqi Huang, Weiqiang Ren, Junge Zhang, Stephen J. Maybank
IEEE Trans. Image Process.2
2014 Deformable Object Matching via Deformation Decomposition Based 2D Label MRF
abstract
Deformable object matching, which is also called elastic matching or deformation matching, is an important and challenging problem in computer vision. Although numerous deformation models have been proposed in different matching tasks, not many of them investigate the intrinsic physics underlying deformation. Due to the lack of physical analysis, these models cannot describe the structure changes of deformable objects very well. Motivated by this, we analyze the deformation physically and propose a novel deformation decomposition model to represent various deformations. Based on the physical model, we formulate the matching problem as a two-mensional label Markov Random Field. The MRF energy function is derived from the deformation decomposition model. Furthermore, we propose a two-stage method to optimize the MRF energy function. To provide a quantitative benchmark, we build a deformation matching database with an evaluation criterion. Experimental results show that our method outperforms previous approaches especially on complex deformations.
Junge Zhang, Kaiqi Huang, Tieniu Tan
CVPR3
2014 Weakly Supervised Object Localization with Latent Category Learning
Weiqiang Ren, Kaiqi Huang, Tieniu Tan
ECCV (6)3
2014 A novel solution for multi-camera object tracking
abstract
The traditional multi-camera object tracking contains two steps: single camera object tracking (SCT) and inter-camera object tracking (ICT). The ICT performance strongly relies on the great results of SCT. In practice, most of current SCT methods are unperfect and products much more fragments. In this paper, a novel solution using a global tracklet association is proposed, which can provide a good ICT performance when the SCT results are not perfect. The proposed solution is also available in non-overlapping views through a new tracklet representation and experiments shows the effectiveness of the proposed novel solution in real scene.
Lijun Cao, Xiaotang Chen, Kaiqi Huang
ICIP4
2014 Window mining by clustering mid-level representation for weakly supervised object detection
abstract
Discovering positive detection windows in training images is a challenging problem in weakly supervised object detection. In this paper, we propose a window mining strategy by the simple and efficient k-means clustering. Firstly, a recent segmentation based object proposal is used for its highly semantic candidate windows; secondly, the bag-of-words model is adopted as mid-level object representation for each window. By clustering these windows with k-means, semantic clusters can be generated. Then, to discover the positive windows from these clusters, we further propose a cluster selection method based on each cluster's discrimination, which is evaluated by classification performance given the category label. With the semantic clusters, this selection process is effective and efficient. Evaluation on the challenging PASCAL VOC 2007 dataset shows that the proposed method outperforms all previous weakly supervised approaches.
Weiqiang Ren, Kaiqi Huang
ICIP3
2014 Semi-supervised Learning for RGB-D Object Recognition
abstract
Conventional supervised object recognition methods have been investigated for many years. Despite their successes, there are still two suffering limitations: (1) various information of an object is represented by artificial features only derived from RGB images, (2) lots of manually labeled data is required by supervised learning. To address those limitations, we propose a new semi-supervised learning framework based on RGB and depth (RGB-D) images to improve object recognition. In particular, our framework has two modules: (1) RGB and depth images are represented by convolutional-recursive neural networks to construct high level features, respectively, (2) co-training is exploited to make full use of unlabeled RGB-D instances due to the existing two independent views. Experiments on the standard RGB-D object dataset demonstrate that our method can compete against with other state-of-the-art methods with only 20% labeled data.
Yanhua Cheng, Xin Zhao 0012, Kaiqi Huang, Tieniu Tan
ICPR3
2014 Semi-supervised Learning for Cross-Device Visual Location Recognition
abstract
The aim of this work is to localize a query mobile photograph by utilizing surveillance images, which naturally provide location information. We cast this cross-device visual localization problem as a classification task. By exploiting the surveillance network to collect reference images, the data acquisition process is significantly facilitated. However, the discrepancy between mobile images and surveillance images makes the training samples difficult to be used directly, and the scarcity of training samples caused by the immobility of surveillance cameras further degrades the performance. In contrast to most traditional domain adaptation problems and semi-supervised problems, the scarce labeled data and plentiful unlabeled data exist in different domains. Our location recognition method first exploits the unsupervised subspace alignment to weaken the discrepancy between the two domains, and then adopts the semi-supervised Laplacian SVM to reinforce the discriminant information utilizing the unlabeled mobile images. Experimental results show that our location recognition method significantly outperforms other related methods.
Pengcheng Liu 0001, Peipei Yang, Kaiqi Huang, Tieniu Tan, Hongwei Hao
ICPR3
2014 Improved Optimization Based on Graph Cuts for Discrete Energy Minimization
abstract
Discrete energy optimization is a NP hard problem. Recent years, the graph cuts based algorithms especially the a-expansion and aß-swap, become more and more popular. Both the a-expansion and a/3-swap have been widely used in many applications, and they perform extremely well for the Potts energies. However, since all pixels only have a choice of two labels in one move, both the expansion and swap algorithm get approximate solution by a series of iterations and they do not perform well for more general energies, such as the truncated convex energies [1]. In this paper, we analyze the problems of both the expansion and swap algorithms. The expansion algorithm usually encourages more pixels to get the label fa, since all pixels are only allowed to change their current labels to fa. In contrast, the swap move sometimes cannot swap the labels of pixels reasonably. Based on the analysis, we propose the Interleaved Expansion-Swap Algorithm (IESA) by combining the expansion and swap moves effectively. To prove the effectiveness of the algorithm, we test it on both image restoration and stereo correspondence. The experimental evaluations show that our algorithm gets better optimization compared with both a-expansion and aß-swap.
Junge Zhang, Kaiqi Huang
ICPR3
2014 Learning Convolutional Nonlinear Features for K Nearest Neighbor Image Classification
abstract
Learning low-dimensional feature representations is a crucial task in machine learning and computer vision. Recently the impressive breakthrough in general object recognition made by large scale convolutional networks shows that convolutional networks are able to extract discriminative hierarchical features in large scale object classification task. However, for vision tasks other than end-to-end classification, such as K Nearest Neighbor classification, the learned intermediate features are not necessary optimal for the specific problem. In this paper, we aim to exploit the power of deep convolutional networks and optimize the output feature layer with respect to the task of K Nearest Neighbor (kNN) classification. By directly optimizing the kNN classification error on training data, we in fact learn convolutional nonlinear features in a data-driven and task-driven way. Experimental results on standard image classification benchmarks show that the proposed method is able to learn better feature representations than other general end-to-end classification methods on kNN classification task.
Weiqiang Ren, Yinan Yu, Junge Zhang, Kaiqi Huang
ICPR4
2014 Robust Object Recognition via Visual Pathway Feedback
abstract
Object recognition, which consists of classification and detection, has two important attributes for robustness: (1) Closeness: detection windows should be close to object locations, and (2) Adaptiveness: object matching should be adaptive to object variations in classification. It is difficult to satisfy both attributes by considering classification and detection separately, thus recent studies combine them based on confidence contextualization and foreground modeling. However, these combinations neglect feature saliency and object structure, which are important for recognition. In fact, object recognition originates in the mechanism of "what" and "where" pathways in human visual systems, and more importantly, these pathways have feedback to each other, which provides a probable way to improve closeness and adaptiveness. Inspired by the feedback, we propose a robust object recognition framework by designing a computational model of the feedback mechanism. In the "what" feedback, the feature saliency from classification is exploited to rectify detection windows for better closeness, while in the "where" feedback, object parts from detection are used to model object matching of object structure for better adaptiveness. Experiments show that the "what" and "where" feedback can be effective to improve closeness and adaptiveness for robust object recognition, and encouraging results are obtained on the challenging PASCAL VOC 2007 dataset.
Junge Zhang, Peipei Yang, Kaiqi Huang
ICPR4
2014 Object tracking across non-overlapping views by learning inter-camera transfer models
Xiaotang Chen, Kaiqi Huang, Tieniu Tan
Pattern Recognit.2
2013 Exploring the Power of Kernel in Feature Representation for Object Categorization
Weiqiang Ren, Yinan Yu, Junge Zhang, Kaiqi Huang
ICONIP (3)4
2013 View independent object classification by exploring scene consistency information for traffic scene surveillance
Zhaoxiang Zhang 0001, Kaiqi Huang, Yunhong Wang 0001, Min Li 0022
Neurocomputing2
2013 Practical Camera Calibration From Moving Objects for Traffic Scene Surveillance
abstract
We address the problem of camera calibration for traffic scene surveillance, which supplies a connection between 2-D image features and 3-D measurement. It is helpful to deal with appearance distortion related to view angles, establish multiview correspondences, and make use of 3-D object models as prior information to enhance surveillance performance. A convenient and practical camera calibration method is proposed in this paper. With the camera heightHmeasured as the only user input, we can recover both intrinsic and extrinsic parameters of the camera based on redundant information supplied by moving objects in monocular videos. All cases of traffic scene layouts are considered and corresponding solutions are given to make our method applicable to almost all kinds of traffic scenes in reality. Numerous experiments are conducted in different scenes, and experimental results demonstrate the accuracy and practicability of our approach. It is shown that our approach can be effectively adopted in all kinds of traffic scene surveillance applications.
Zhaoxiang Zhang 0001, Tieniu Tan, Kaiqi Huang, Yunhong Wang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2012 Local Hypersphere Coding Based on Edges between Visual Words
Weiqiang Ren, Yongzhen Huang, Xin Zhao 0012, Kaiqi Huang, Tieniu Tan
ACCV (1)4
2012 Data Decomposition and Spatial Mixture Modeling for Part Based Model
Junge Zhang, Yongzhen Huang, Kaiqi Huang, Zifeng Wu, Tieniu Tan
ACCV (1)3
2012 Tracking Blurred Object with Data-Driven Tracker
abstract
Motion blur is very common in the low quality of image sequences and videos captured by low speed of cameras. Object tracking without accounting for the motion blur would easily fail in these kinds of videos. We propose a new data-driven tracker in the particle filter framework to address this problem without deblurring the image sequences. The motion blur is detected by exploring the property of the blurred input image through Fourier analysis. The appearance model is integrated with a set of motion blur kernels which could reflect different blur effects in real scenes. The motion model is improved to be more robust to sudden motion of the target object. To evaluate the proposed algorithm, several challenging videos with significant motion blur are used in the experiments. The experimental results demonstrate the robustness and accuracy of our algorithm.
Jianwei Ding, Kaiqi Huang, Tieniu Tan
AVSS2
2012 Interest Point Selection with Spatio-temporal Context for Realistic Action Recognition
abstract
Spatio-Temporal Interest Point (STIP) has been widely used for human action recognition. However, the performance of the STIP based methods are still limited in realistic datasets which often include large variations in illuminations, viewpoints and camera motions. One reason of the low performance is that the STIPs only reflect the local change in videos, which is not enough to obtain stable informative features for action representation in realistic scene. To tackle the problem, we proposed an approach to selecting the "stable STIPs" with the spatio-temporal distribution of STIPs in neighbor region. Then, BoW feature is constructed to represent actions with these selected points. The experimental results on KTH dataset and HMDB (the largest realistic human action dataset) demonstrate that the proposed approach has obvious effect on improving the recognition rates of realistic data.
Yanhu Shan, Zhang Zhang 0001, Junge Zhang, Kaiqi Huang, Oh Se Hyun
AVSS4
2012 CLUMOC: Multiple Motion Estimation by Cluster Motion Consensus
abstract
In this paper, we present techniques for robust multiple motions estimation based on dual consensus via clustering in both the image spatial space and the motion parameter space. Starting from traditional Random Samples Consensus algorithm, we novelly propose the CLUster MOtion Consensus (CLUMOC) to extract robust motions. The proposed algorithm has two advantages: (1), instead of random samples, the CLUMOC employs clustering in initial sample selection, which can remove outliers from correct pairs of motion, (2), CLUMOC automatically decides the number of motions, by employing competition among motion and samples, that each motion needs to compete for matching pairs and each pair of matching competes for motions. The experimental results show that the proposed method is effective and efficient under various situations.
Yinan Yu, Weiqiang Ren, Yongzhen Huang, Kaiqi Huang, Tieniu Tan
AVSS4
2012 Feature coding via vector difference for image classification
abstract
An effective image representation is important to an image classification task. The most popular image representation framework utilizes a feature coding algorithm to encode the extracted low-level feature descriptors into a vector representation. In this paper, we analyze the recently developed feature coding methods in a general way. According to their common characteristics, we propose a new coding scheme to perform feature coding based on the vector difference in a high-dimensional space which is obtained by explicit feature maps. As we illustrate, our method has promising results with small codebook sizes and generalizes most existing coding methods in a unified form.
Xin Zhao 0012, Yinan Yu, Yongzhen Huang, Kaiqi Huang, Tieniu Tan
ICIP4
2012 Semantic windows mining in sliding window based object detection
Junge Zhang, Xin Zhao 0012, Yongzhen Huang, Kaiqi Huang, Tieniu Tan
ICPR4
2012 A cascade fusion scheme for gait and cumulative foot pressure image recognition
Shuai Zheng 0001, Kaiqi Huang, Tieniu Tan, Dacheng Tao
Pattern Recognit.2
2012 Cast Shadow Removal in a Hierarchical Manner Using MRF
abstract
In this paper, we present a novel method for shadow removal using Markov random fields (MRF). In our method, we first construct the shadow model in a hierarchical manner. At the pixel level, we use the Gaussian mixture model to model the behavior of cast shadows for every pixel in the HSV color space. The samples which are used to update the shadow model should satisfy a pre-classifier. This pre-classifier indicates the color feature of shadow in current frame. At the global level, we exploit the statistical features of shadow in the whole scene over several consecutive frames to make this pre-classifier accurate and adaptive to the change of shadow. Then, based on the shadow model, an MRF model is constructed for shadow removal. The main contribution of this paper is twofold. First, although our method is a chroma-based method, we make the pre-classifier accurate and adaptive to the change of shadow by using the statistical features of shadow at the global level. Moreover, tracking information can make this global-level statistical information more robust. Second, we construct an MRF model to represent the dependencies between the label of a pixel and the shadow models of its neighbors. Experimental results show that the proposed method is efficient and robust.
Kaiqi Huang, Tieniu Tan
IEEE Trans. Circuits Syst. Video Technol.2
2012 A Discriminative Model of Motion and Cross Ratio for View-Invariant Action Recognition
abstract
Action recognition is very important for many applications such as video surveillance, human-computer interaction, and so on; view-invariant action recognition is hot and difficult as well in this field. In this paper, a new discriminative model is proposed for video-based view-invariant action recognition. In the discriminative model, motion pattern and view invariants are perfectly fused together to make a better combination of invariance and distinctiveness. We address a series of issues, including interest point detection in image sequence, motion feature extraction and description, and view-invariant calculation. First, motion detection is used to extract motion information from videos, which is much more efficient than traditional background modeling and tracking-based methods. Second, as for feature representation, we exact variety of statistical information from motion and view-invariant feature based on cross ratio. Last, in the action modeling, we apply a discriminative probabilistic model-hidden conditional random field to model motion patterns and view invariants, by which we could fuse the statistics of motion and projective invariability of cross ratio in one framework. Experimental results demonstrate that our method can improve the ability to distinguish different categories of actions with high robustness to view change in real circumstances.
Kaiqi Huang, Yeying Zhang, Tieniu Tan
IEEE Trans. Image Process.1
2012 Efficient Object Tracking by Incremental Self-Tuning Particle Filtering on the Affine Group
abstract
We propose an incremental self-tuning particle filtering (ISPF) framework for visual tracking on the affine group, which can find the optimal state in a chainlike way with a very small number of particles. Unlike traditional particle filtering, which only relies on random sampling for state optimization, ISPF incrementally draws particles and utilizes an online-learned pose estimator (PE) to iteratively tune them to their neighboring best states according to some feedback appearance-similarity scores. Sampling is terminated if the maximum similarity of all tuned particles satisfies a target-patch similarity distribution modeled online or if the permitted maximum number of particles is reached. With the help of the learned PE and some appearance-similarity feedback scores, particles in ISPF become "smart" and can automatically move toward the correct directions; thus, sparse sampling is possible. The optimal state can be efficiently found in a step-by-step way in which some particles serve as bridge nodes to help others to reach the optimal state. In addition to the single-target scenario, the "smart" particle idea is also extended into a multitarget tracking problem. Experimental results demonstrate that our ISPF can achieve great robustness and very high accuracy with only a very small number of particles.
Min Li 0022, Tieniu Tan, Wei Chen 0012, Kaiqi Huang
IEEE Trans. Image Process.4
2012 Foreground Object Detection Using Top-Down Information Based on EM Framework
abstract
In this paper, we present a novel foreground object detection scheme that integrates the top-down information based on the expectation maximization (EM) framework. In this generalized EM framework, the top-down information is incorporated in an object model. Based on the object model and the state of each target, a foreground model is constructed. This foreground model can augment the foreground detection for the camouflage problem. Thus, an object's state-specific Markov random field (MRF) model is constructed for detection based on the foreground model and the background model. This MRF model depends on the latent variables that describe each object's state. The maximization of the MRF model is the M-step in the EM framework. Besides fusing spatial information, this MRF model can also adjust the contribution of the top-down information for detection. To obtain detection result using this MRF model, sampling importance resampling is used to sample the latent variable and the EM framework refines the detection iteratively. Besides the proposed generalized EM framework, our method does not need any prior information of the moving object, because we use the detection result of moving object to incorporate the domain knowledge of the object shapes into the construction of top-down information. Moreover, in our method, a kernel density estimation (KDE)-Gaussian mixture model (GMM) hybrid model is proposed to construct the probability density function of background and moving object model. For the background model, it has some advantages over GMM- and KDE-based methods. Experimental results demonstrate the capability of our method, particularly in handling the camouflage problem.
Kaiqi Huang, Tieniu Tan
IEEE Trans. Image Process.2
2012 A Novel Algorithm for View and Illumination Invariant Image Matching
abstract
The challenges in local-feature-based image matching are variations of view and illumination. Many methods have been recently proposed to address these problems by using invariant feature detectors and distinctive descriptors. However, the matching performance is still unstable and inaccurate, particularly when large variation in view or illumination occurs. In this paper, we propose a view and illumination invariant image-matching method. We iteratively estimate the relationship of the relative view and illumination of the images, transform the view of one image to the other, and normalize their illumination for accurate matching. Our method does not aim to increase the invariance of the detector but to improve the accuracy, stability, and reliability of the matching results. The performance of matching is significantly improved and is not affected by the changes of view and illumination in a valid range. The proposed method would fail when the initial view and illumination method fails, which gives us a new sight to evaluate the traditional detectors. We propose two novel indicators for detector evaluation, namely, valid angle and valid illumination, which reflect the maximum allowable change in view and illumination, respectively. Extensive experimental results show that our method improves the traditional detector significantly, even in large variations, and the two indicators are much more distinctive.
Yinan Yu, Kaiqi Huang, Wei Chen 0012, Tieniu Tan
IEEE Trans. Image Process.2
2012 Three-Dimensional Deformable-Model-Based Localization and Recognition of Road Vehicles
abstract
We address the problem of model-based object recognition. Our aim is to localize and recognize road vehicles from monocular images or videos in calibrated traffic scenes. A 3-D deformable vehicle model with 12 shape parameters is set up as prior information, and its pose is determined by three parameters, which are its position on the ground plane and its orientation about the vertical axis under ground-plane constraints. An efficient local gradient-based method is proposed to evaluate the fitness between the projection of the vehicle model and image data, which is combined into a novel evolutionary computing framework to estimate the 12 shape parameters and three pose parameters by iterative evolution. The recovery of pose parameters achieves vehicle localization, whereas the shape parameters are used for vehicle recognition. Numerous experiments are conducted in this paper to demonstrate the performance of our approach. It is shown that the local gradient-based method can evaluate accurately and efficiently the fitness between the projection of the vehicle model and the image data. The evolutionary computing framework is effective for vehicles of different types and poses is robust to all kinds of occlusion.
Zhaoxiang Zhang 0001, Tieniu Tan, Kaiqi Huang, Yunhong Wang 0001
IEEE Trans. Image Process.3
2011 Exploring relations of visual codes for image classification
abstract
The classic Bag-of-Features (BOF) model and its extensional work use a single value to represent a visual code. This strategy ignores the relation of visual codes. In this paper, we explore this relation and propose a new algorithm for image classification. It consists of two main parts: 1) construct the codebook graph wherein a visual code is linked with other codes; 2) describe each local feature using a pair of related codes, corresponding to an edge of the graph. Our approach contains richer information than previous BOF models. Moreover, we demonstrate that these models are special cases of ours. Various coding and pooling algorithms can be embedded into our framework to obtain better performance. Experiments on different kinds of image classification databases demonstrate that our approach can stably achieve excellent performance compared with various BOF models.
Yongzhen Huang, Kaiqi Huang, Tieniu Tan
CVPR2
2011 Salient coding for image classification
abstract
The codebook based (bag-of-words) model is a widely applied model for image classification. We analyze recent coding strategies in this model, and find that saliency is the fundamental characteristic of coding. The saliency in coding means that if a visual code is much closer to a descriptor than other codes, it will obtain a very strong response. The salient representation under maximum pooling operation leads to the state-of-the-art performance on many databases and competitions. However, most current coding schemes do not recognize the role of salient representation, so that they may lead to large deviations in representing local descriptors. In this paper, we propose “salient coding”, which employs the ratio between descriptors' nearest code and other codes to describe descriptors. This approach can guarantee salient representation without deviations. We study salient coding on two sets of image classification databases (15-Scenes and PASCAL VOC2007). The experimental results demonstrate that our approach outperforms all other coding methods in image classification.
Yongzhen Huang, Kaiqi Huang, Yinan Yu, Tieniu Tan
CVPR2
2011 Boosted local structured HOG-LBP for object localization
abstract
Object localization is a challenging problem due to variations in object's structure and illumination. Although existing part based models have achieved impressive progress in the past several years, their improvement is still limited by low-level feature representation. Therefore, this paper mainly studies the description of object structure from both feature level and topology level. Following the bottom-up paradigm, we propose a boosted Local Structured HOG-LBP based object detector. Firstly, at feature level, we propose Local Structured Descriptor to capture the object's local structure, and develop the descriptors from shape and texture information, respectively. Secondly, at topology level, we present a boosted feature selection and fusion scheme for part based object detector. All experiments are conducted on the challenging PASCAL VOC2007 datasets. Experimental results show that our method achieves the state-of-the-art performance.
Junge Zhang, Kaiqi Huang, Yinan Yu, Tieniu Tan
CVPR2
2011 Direction-based stochastic matching for pedestrian recognition in non-overlapping cameras
abstract
Pedestrian recognition is a challenging problem in non-overlapping multi-camera object tracking. In this paper, we present a novel approach for matching pedestrians across non-overlapping multiple cameras without the need of a training phase or spatio-temporal cues across cameras. To deal with viewpoint changes, we introduce the concept of directional angles estimated using the spatio-temporal continuity in the single camera tracking. To deal with pose changes, a stochastic matching strategy is performed, where the similarity of two blobs belonging to different viewpoints is calculated by a novel similarity measurement algorithm. The experiments are performed on different multi-view datasets. Experimental results demonstrate the effectiveness and robustness of the proposed method.
Xiaotang Chen, Kaiqi Huang, Tieniu Tan
ICIP2
2011 Partial Least Squares based subwindow search for pedestrian detection
abstract
In this paper, we propose a Partial Least Squares based sub- window search method for pedestrian detection, by which the detection speed can be improved effectively while maintaining high detection accuracy. Firstly, a sparse search is implemented to find all the possible locations containing parts of a pedestrian. Then a pre-learned Partial Least Squares regression model is applied to estimate the displacements of the subwindows to guide them towards the approximate locations of the pedestrians. Finally, we conduct a dense search around the approximate locations to obtain the exact locations of the pedestrians. Experiments on the INRIA dataset demonstrate that our method greatly reduces the number of search windows, which leads to much fewer feature extraction in the detection phase. Thus, it is about 10 times faster than the sliding window method with a jump step of 8 x 8.
Jinchen Wu, Wei Chen 0012, Kaiqi Huang, Tieniu Tan
ICIP3
2011 Evaluation framework on translation-invariant representation for cumulative foot pressure image
abstract
Ground reaction force can be used to distinguish different human gait like limb movement. Cumulative foot pressure image is a 2-D data that recorded the spatial and temporal change of ground reaction force during one gait cycle. However, when putting it into practice as a new biometric for gait recognition, it suffers from the problem of large translation variation within class caused by wearing different shoes and walking in different speed. In this paper, an evaluation framework is proposed to address the problem. The framework consists of a database containing cumulative foot pressure images and a well designed benchmark. The data are collected from 118 subjects. A locality-constrained sparse coding scheme is developed to be compared with the benchmark PCA approach. Experimental results show the potential of evaluation framework for evaluating the translation-invariant power of image representation algorithms.
Shuai Zheng 0001, Kaiqi Huang, Tieniu Tan
ICIP2
2011 Robust view transformation model for gait recognition
abstract
Recent gait recognition systems often suffer from the challenges including viewing angle variation and large intra-class variations. In order to address these challenges, this paper presents a robust View Transformation Model for gait recognition. Based on the gait energy image, the proposed method establishes a robust view transformation model via robust principal component analysis. Partial least square is used as feature selection method. Compared with the existing methods, the proposed method finds out a shared linear correlated low rank subspace, which brings the advantages that the view transformation model is robust to viewing angle variation, clothing and carrying condition changes. Conducted on the CASIA gait dataset, experimental results show that the proposed method outperforms the other existing methods.
Shuai Zheng 0001, Junge Zhang, Kaiqi Huang, Ran He 0001, Tieniu Tan
ICIP3
2011 Multi-view Pedestrian Recognition Using Shared Dictionary Learning with Group Sparsity
Shuai Zheng 0001, Bo Xie 0002, Kaiqi Huang, Dacheng Tao
ICONIP (3)3
2011 An Extended Grammar System for Learning and Recognizing Complex Visual Events
abstract
For a grammar-based approach to the recognition of visual events, there are two major limitations that prevent it from real application. One is that the event rules are predefined by domain experts, which means huge manual cost. The other is that the commonly used grammar can only handle sequential relations between subevents, which is inadequate to recognize more complex events involving parallel subevents. To solve these problems, we propose an extended grammar approach to modeling and recognizing complex visual events. First, motion trajectories as original features are transformed into a set of basic motion patterns of a single moving object, namely, primitives (terminals) in the grammar system. Then, a Minimum Description Length (MDL) based rule induction algorithm is performed to discover the hidden temporal structures in primitive stream, where Stochastic Context-Free Grammar (SCFG) is extended by Allen's temporal logic to model the complex temporal relations between subevents. Finally, a Multithread Parsing (MTP) algorithm is adopted to recognize interesting complex events in a given primitive stream, where a Viterbi-like error recovery strategy is also proposed to handle large-scale errors, e.g., insertion and deletion errors. Extensive experiments, including gymnastic exercises, traffic light events, and multi-agent interactions, have been executed to validate the effectiveness of the proposed approach.
Zhang Zhang 0001, Tieniu Tan, Kaiqi Huang
IEEE Trans. Pattern Anal. Mach. Intell.3
2011 Enhanced Biologically Inspired Model for Object Recognition
abstract
The biologically inspired model (BIM) proposed by Serre presents a promising solution to object categorization. It emulates the process of object recognition in primates' visual cortex by constructing a set of scale- and position-tolerant features whose properties are similar to those of the cells along the ventral stream of visual cortex. However, BIM has potential to be further improved in two aspects: mismatch by dense input and randomly feature selection due to the feedforward framework. To solve or alleviate these limitations, we develop an enhanced BIM (EBIM) in terms of the following two aspects: 1) removing uninformative inputs by imposing sparsity constraints, 2) apply a feedback loop to middle level feature selection. Each aspect is motivated by relevant psychophysical research findings. To show the effectiveness of the EBIM, we apply it to object categorization and conduct empirical studies on four computer vision data sets. Experimental results demonstrate that the EBIM outperforms the BIM and is comparable to state-of-the-art approaches in terms of accuracy. Moreover, the new system is about 20 times faster than the BIM.
Yongzhen Huang, Kaiqi Huang, Dacheng Tao, Tieniu Tan, Xuelong Li 0001
IEEE Trans. Syst. Man Cybern. Part B2
2011 Biologically Inspired Features for Scene Classification in Video Surveillance
abstract
Inspired by human visual cognition mechanism, this paper first presents a scene classification method based on an improved standard model feature. Compared with state-of-the-art efforts in scene classification, the newly proposed method is more robust, more selective , and of lower complexity. These advantages are demonstrated by two sets of experiments on both our own database and standard public ones. Furthermore, occlusion and disorder problems in scene classification in video surveillance are also first studied in this paper.
Kaiqi Huang, Dacheng Tao, Yuan Yan Tang, Xuelong Li 0001, Tieniu Tan
IEEE Trans. Syst. Man Cybern. Part B1
2010 Modeling Complex Scenes for Accurate Moving Objects Segmentation
Jianwei Ding, Min Li 0022, Kaiqi Huang, Tieniu Tan
ACCV (2)3
2010 A Heuristic Deformable Pedestrian Detection Method
Yongzhen Huang, Kaiqi Huang, Tieniu Tan
ACCV (2)2
2010 Multi-Target Tracking by Learning Class-Specific and Instance-Specific Cues
Min Li 0022, Wei Chen 0012, Kaiqi Huang, Tieniu Tan
ACCV (2)3
2010 Visual tracking via incremental self-tuning particle filtering on the affine group
abstract
We propose an incremental self-tuning particle filtering (ISPF) framework for visual tracking on the affine group. SIFT (Scale Invariant Feature Transform) like descriptors are used as basic features, and IPCA (Incremental Principle Component Analysis) is utilized to learn an adaptive appearance subspace for similarity measurement. ISPF tries to find the optimal target position in a step-by-step way: particles are incrementally drawn and intelligently tuned to their best states by an online LWPR (Local Weighted Projection Regression) pose estimator; searching is terminated if the maximum similarity of all tuned particles satisfies a target similarity distribution (TSD) modeled online or the permitted maximum number of particles is reached. Experimental results demonstrate that our ISPF can achieve great robustness and very high accuracy with only a very small number of random particles.
Min Li 0022, Wei Chen 0012, Kaiqi Huang, Tieniu Tan
CVPR3
2010 Recovering the Topology of Multiple Cameras by Finding Continuous Paths in a Trellis
abstract
In this paper, we propose an unsupervised method for recovering the topology of multiple cameras with non-overlapping fields of view. The nodes in the topology graph are defined as entry/exit zones in each camera while the connectivity between nodes is inferred through finding continuous paths in a trellis where appearance information and temporal information of moving objects are encoded. Unlike previous methods which assume a single mode transition distribution between nodes, our method is capable of dealing with multi-modal transition situations when both cars and pedestrians are in the scene. Results on simulated and real-life datasets demonstrate the effectiveness of the proposed method.
Yinghao Cai, Kaiqi Huang, Tieniu Tan, Matti Pietikäinen
ICPR2
2010 3D Model Based Vehicle Tracking Using Gradient Based Fitness Evaluation under Particle Filter Framework
abstract
We address the problem of 3D model based vehicle tracking from monocular videos of calibrated traffic scenes. A 3D wire-frame model is set up as prior information and an efficient fitness evaluation method based on image gradients is introduced to estimate the fitness score between the projection of vehicle model and image data, which is then combined into a particle filter based framework for robust vehicle tracking. Numerous experiments are conducted and experimental results demonstrate the effectiveness of our approach for accurate vehicle tracking and robustness to noise and occlusions.
Zhaoxiang Zhang 0001, Kaiqi Huang, Tieniu Tan, Yunhong Wang 0001
ICPR2
2010 Vs-star: A visual interpretation system for visual surveillance
Kaiqi Huang, Tieniu Tan
Pattern Recognit. Lett.1
2010 Discriminative Orthogonal Neighborhood-Preserving Projections for Classification
abstract
Orthogonal neighborhood-preserving projection (ONPP) is a recently developed orthogonal linear algorithm for overcoming the out-of-sample problem existing in the well-known manifold learning algorithm, i.e., locally linear embedding. It has been shown that ONPP is a strong analyzer of high-dimensional data. However, when applied to classification problems in a supervised setting, ONPP only focuses on the intraclass geometrical information while ignores the interaction of samples from different classes. To enhance the performance of ONPP in classification, a new algorithm termed discriminative ONPP (DONPP) is proposed in this paper. DONPP 1) takes into account both intraclass and interclass geometries; 2) considers the neighborhood information of interclass relationships; and 3) follows the orthogonality property of ONPP. Furthermore, DONPP is extended to the semisupervised case, i.e., semisupervised DONPP (SDONPP). This uses unlabeled samples to improve the classification accuracy of the original DONPP. Empirical studies demonstrate the effectiveness of both DONPP and SDONPP.
Tianhao Zhang 0002, Kaiqi Huang, Xuelong Li 0001, Jie Yang 0002, Dacheng Tao
IEEE Trans. Syst. Man Cybern. Part B2
2009 A Novel Visual Organization Based on Topological Perception
Yongzhen Huang, Kaiqi Huang, Tieniu Tan, Dacheng Tao
ACCV (1)2
2009 A Harris-Like Scale Invariant Feature Detector
Yinan Yu, Kaiqi Huang, Tieniu Tan
ACCV (2)2
2009 A convergent solution to two dimensional linear discriminant analysis
abstract
The matrix based data representation has been recognized to be effective for face recognition because it can deal with the undersampled problem. One of the most popular algorithms, the two dimensional linear discriminant analysis (2DLDA), has been identified to be effective to encode the discriminative information for training matrix represented samples. However, 2DLDA does not converge in the training stage. This paper presents an evolutionary computation based solution, referred to as E-2DLDA, to provide a convergent training stage for 2DLDA. In E-2DLDA, every randomly generated candidate projection matrices are first normalized. The evolutionary computation method optimizes the projection matrices to best separate different classes. Experimental results show E-2DLDA is convergent and outperforms 2DLDA.
Wei Chen 0012, Kaiqi Huang, Tieniu Tan, Dacheng Tao
ICIP2
2009 Computational primitives of visual perception
abstract
Great stride has been made in psychological research about primitives of visual perception, which is important to computer vision and image processing. In this paper, we propose a computational model to imitate the primitives of visual perception based on the pyschological theory of topological perceptual organization. First, we adopt geodesic distance based descriptor to describe an independent topological structure. Then, we consider the spatial relationship of two independent structures. Experiments on structures classification demonstrates that the propose model is consistent with the psychological theory. Further experiments on patches clustering prove that our approach can be used to enhance other algorithms.
Yongzhen Huang, Kaiqi Huang, Tieniu Tan
ICIP2
2009 Rapid and robust human detection and tracking based on omega-shape features
abstract
This paper proposes a novel method for rapid and robust human detection and tracking based on the omega-shape features of people's head-shoulder parts. There are two modules in this method. In the first module, a Viola-Jones type classifier and a local HOG (Histograms of Oriented Gradients) feature based AdaBoost classifier are combined to detect head-shoulders rapidly and effectively. Then, in the second module, each detected head-shoulder is tracked by a particle filter tracker using local HOG features to model target's appearance, which shows great robustness in scenarios of crowding, background distractors and partial occlusions. Experimental results demonstrate the effectiveness and efficiency of the proposed approach.
Min Li 0022, Zhaoxiang Zhang 0001, Kaiqi Huang, Tieniu Tan
ICIP3
2009 Robust visual tracking based on simplified biologically inspired features
abstract
We address the problem of robust appearance-based visual tracking. First, a set of simplified biologically inspired features (SBIF) is proposed for object representation and the Bhattacharyya coefficient is used to measure the similarity between the target model and candidate targets. Then, the proposed appearance model is combined into a Bayesian state inference tracking framework utilizing the SIR (sampling importance resampling) particle filter to propagate sample distributions over time. Numerous experiments are conducted and experimental results demonstrate that our algorithm is robust to partial occlusions and variations of illumination and pose, resistent to nearby distractors, as well as possesses the state-of-the-art tracking accuracy.
Min Li 0022, Zhaoxiang Zhang 0001, Kaiqi Huang, Tieniu Tan
ICIP3
2009 Object detection and tracking for night surveillance based on salient contrast analysis
abstract
Night surveillance is a challenging task because of low brightness, low contrast, low Signal to Noise Ratio (SNR) and low appearance information. Most existing models for night surveillance share the following problems: a lack of adaptability for different scenes and separation between detection and tracking. To solve these problems we propose a model based on Salient Contrast Change (SCC) feature, which applies learning process to enhance adaptability and analyzes trajectories to improve the effectiveness of detection. Empirical studies on several real night videos show that the proposed model is more effective than the original CC model and other traditional models.
Liangsheng Wang, Kaiqi Huang, Yongzhen Huang, Tieniu Tan
ICIP2
2009 A compact optical flowbased motion representation for real-time action recognition in surveillance scenes
abstract
We address the problem of action recognition. Our aim is to recognize single person activities in surveillance scenes. To meet the requirements of real scene action recognition, we present a compact motion representation for human activity recognition. With the employment of efficient features extracted from optical flow as the main part, together with global information, our motion representation is compact and discriminative. We also build a novel human action dataset(CASIA) in surveillance scene with three vertically different viewpoints and distant people. Experiments on CASIA dataset and WEIZMANN dataset show that our method can achieve satisfying recognition performance with low computational cost as well as robustness against both horizontal(panning) and vertical(tilting) viewpoint changes.
Shiquan Wang, Kaiqi Huang, Tieniu Tan
ICIP2
2009 View-invariant action recognition using cross ratios across frames
abstract
We present a new method of computing invariants in videos captured from different views to achieve view-invariant action recognition. To avoid the constraints of collinearity or coplanarity of image points for constructing invariants, we consider several neighboring frames to compute cross ratios, namely cross ratios across frames (CRAF), as our invariant representation of action. For every five points sampled with different intervals from the trajectories of action, we construct a pair of cross ratios (CRs). Afterwards, we transform the CRs to histograms as the feature vectors for classification. Experimental results demonstrate that the proposed method outperforms the state-of-the-art methods in effectiveness and stability.
Yeying Zhang, Kaiqi Huang, Yongzhen Huang, Tieniu Tan
ICIP2
2009 Visual information analysis for security
Dacheng Tao, Yuan Yuan 0001, Jialie Shen 0001, Kaiqi Huang, Xuelong Li 0001
Signal Process.4
2009 Human Behavior Analysis Based on a New Motion Descriptor
abstract
Human behavior analysis is an important area of research in computer vision and is also driven by a wide spectrum of applications, such as smart video surveillance and human-computer interface. In this paper, we present a novel approach for human behavior analysis. Two research challenges, motion representation and behavior recognition, are addressed. A novel motion descriptor, which is an improved feature based on optical flow, is proposed for motion representation. Optical flow is improved with a motion filter, and feature fusion with the shape and trajectory information. To recognize the behavior, the support vector machine is employed to train the classifier where the concatenation of histograms is formed as the input features. Experimental results on the Weizmann behavior database and the Institute of Automation, Chinese Academy of Science real-world multiview behavior database demonstrate the robustness and effectiveness of our method.
Kaiqi Huang, Shiquan Wang, Tieniu Tan, Stephen J. Maybank
IEEE Trans. Circuits Syst. Video Technol.1
2009 A Study on Gait-Based Gender Classification
abstract
Gender is an important cue in social activities. In this correspondence, we present a study and analysis of gender classification based on human gait. Psychological experiments were carried out. These experiments showed that humans can recognize gender based on gait information, and that contributions of different body components vary. The prior knowledge extracted from the psychological experiments can be combined with an automatic method to further improve classification accuracy. The proposed method which combines human knowledge achieves higher performance than some other methods, and is even more accurate than human observers. We also present a numerical analysis of the contributions of different human components, which shows that head and hair, back, chest and thigh are more discriminative than other components. We also did challenging cross-race experiments that used Asian gait data to classify the gender of Europeans, and vice versa. Encouraging results were obtained. All the above prove that gait-based gender classification is feasible in controlled environments. In real applications, it still suffers from many difficulties, such as view variation, clothing and shoes changes, or carrying objects. We analyze the difficulties and suggest some possible solutions.
Shiqi Yu 0001, Tieniu Tan, Kaiqi Huang, Kui Jia, Xinyu Wu 0001
IEEE Trans. Image Process.3
2009 View-Independent Behavior Analysis
abstract
The motion analysis of the human body is an important topic of research in computer vision devoted to detecting, tracking, and understanding people's physical behavior. This strong interest is driven by a wide spectrum of applications in various areas such as smart video surveillance. Most research in behavior (or gesture) representation focusses on view-dependent representation, and some research on view invariance considers only information from 3-D models, which is effective under considerable changes of viewpoint. This paper introduces a view-independent behavior-analysis framework based on decision fusion in which distance and view angle factors are analyzed. This is a first effort to tackle the problem of behaviors under significant changes in view angle, and a first corresponding video database is built.
Kaiqi Huang, Dacheng Tao, Yuan Yuan 0001, Xuelong Li 0001, Tieniu Tan
IEEE Trans. Syst. Man Cybern. Part B1
2008 Enhanced biologically inspired model
abstract
It has been demonstrated by Serre et al. that the biologically inspired model (BIM) is effective for object recognition. It outperforms many state-of-the-art methods in challenging databases. However, BIM has the following three problems: a very heavy computational cost due to dense input, a disputable pooling operation in modeling relations of the visual cortex, and blind feature selection in a feed-forward framework. To solve these problems, we develop an enhanced BIM (EBIM), which removes uninformative input by imposing sparsity constraints, utilizes a novel local weighted pooling operation with stronger physiological motivations, and applies a feedback procedure that selects effective features for combination. Empirical studies on the CalTech5 database and CalTech101 database show that EBIM is more effective and efficient than BIM. We also apply EBIM to the MIT-CBCL street scene database to show it achieves comparable performance in comparison with the current best performance. Moreover, the new system can process images with resolution 128 times 128 at a rate of 50 frames per second and enhances the speed 20 times at least in comparison with BIM in common applications.
Yongzhen Huang, Kaiqi Huang, Liangsheng Wang, Dacheng Tao, Tieniu Tan, Xuelong Li 0001
CVPR2
2008 Practical camera auto-calibration based on object appearance and motion for traffic scene visual surveillance
abstract
Camera calibration, as a fundamental issue in computer vision, is indispensable in many visual surveillance applications. Firstly, calibrated camera can help to deal with perspective distortion of object appearance on image plane. Secondly, calibrated camera makes it possible to recover metrics from images which are robust to scene or view an gle changes. In addition, with calibrated cameras, we can make use of prior information of 3D models to estimate 3D pose of objects and make object detection or tracking more robust to noise and occlusions. In this paper, we propose an automatic method to recover camera models from traffic scene surveillance videos. With only the camera height H measured, we can completely recover both intrinsic and extrinsic parameters of cameras based on appearance and motion of objects in videos. Experiments are conducted in different scenes and experimental results demonstrate the effectiveness and practicability of our approach, which can be adopted in many traffic scene surveillance applications.
Zhaoxiang Zhang 0001, Min Li 0022, Kaiqi Huang, Tieniu Tan
CVPR3
2008 Multi-thread Parsing for Recognizing Complex Events in Videos
Zhang Zhang 0001, Kaiqi Huang, Tieniu Tan
ECCV (3)2
2008 Matching tracking sequences across widely separated cameras
abstract
In this paper, we present a new solution to the problem of matching tracking sequences across different cameras. Unlike snapshot-based appearance matching which matches objects by a single image, we focus on sequence matching to alleviate the uncertainties brought by segmentation errors and partial occlusions. By incorporating multiple snapshots of the same object, the influence of the variation is alleviated. At the training stage, given the sequence of a queried person under one camera, the appearance model is formulated by concatenating feature vectors with the majority of votes over the sequence. At the testing stage, Bayesian inference is incorporated into the identification framework to accumulate the temporal information in the sequence. Experimental results demonstrate the effectiveness of the proposed method.
Yinghao Cai, Kaiqi Huang, Tieniu Tan
ICIP2
2008 Robust automated ground plane rectification based on moving vehicles for traffic scene surveillance
abstract
Most outdoor visual surveillance scenes involve objects of interest moving on the ground plane. However, perspective distortion introduces many difficulties to various applications like object classification and activity recognition. In this paper, we propose a robust automated method for both affine and metric rectification of the ground plane based on appearance and motion of vehicles in traffic scene surveillance videos. This rectification enables normalization of object properties like size, length and velocity. Various useful applications are presented and experimental results demonstrate the effectiveness and robustness of the proposed method.
Zhaoxiang Zhang 0001, Min Li 0022, Kaiqi Huang, Tieniu Tan
ICIP3
2008 Human appearance matching across multiple non-overlapping cameras
abstract
In this paper, we present a new solution to the problem of appearance matching across multiple non-overlapping cameras. Objects of interest, pedestrians are represented by a set of region signatures centered at points sampled from edges. The problem of frame-to-frame appearance matching is formulated as finding corresponding points in two images as minimization of a cost function over the space of correspondence. The correspondence problem is solved under integer optimization framework where the cost function is determined by similarity of region signatures as well as geometric constraints between points. Experimental results demonstrate the effectiveness of the proposed method.
Yinghao Cai, Kaiqi Huang, Tieniu Tan
ICPR2
2008 Estimating the number of people in crowded scenes by MID based foreground segmentation and head-shoulder detection
abstract
This paper proposes a novel method to address the problem of estimating the number of people in surveillance scenes with people gathering and waiting. The proposed method combines a MID (mosaic image difference) based foreground segmentation algorithm and a HOG (histograms of oriented gradients) based head-shoulder detection algorithm to provide an accurate estimation of people counts in the observed area. In our framework, the MID-based foreground segmentation module provides active areas for the head-shoulder detection module to detect heads and count the number of people. Numerous experiments are conducted and convincing results demonstrate the effectiveness of our method.
Min Li 0022, Zhaoxiang Zhang 0001, Kaiqi Huang, Tieniu Tan
ICPR3
2008 Boosting local feature descriptors for automatic objects classification in traffic scene surveillance
abstract
We address the problem of automatic object classification for traffic scene surveillance, which is very challenging for the low resolution videos, large intra-class variations and real-time requirement. In this paper, we propose a new strategy for object classification by boosting different local feature descriptors in motion blobs. We not only evaluate the performance of each local feature descriptor, but also fuse these descriptors to achieve better performance. Numerous experiments are conducted and experimental results demonstrate the effectiveness and efficiency of our approach with robustness to noise and variance of view angles, lighting conditions and environments.
Zhaoxiang Zhang 0001, Min Li 0022, Kaiqi Huang, Tieniu Tan
ICPR3
2008 3D model based vehicle localization by optimizing local gradient based fitness evaluation
abstract
We address the problem of 3D model based vehicle localization in calibrated traffic scenes. A wire-frame vehicle model is set up as prior information and an efficient local gradient based method is proposed to evaluate the fitness between the projection of 3D model and image data, which illustrates smooth optimization surface and more conspicuous peak with low computational cost. Gradient decent is then applied to optimize the evaluation score for localization. Experimental results demonstrate the accuracy, efficiency and robustness of the proposed method for model based vehicle localization.
Zhaoxiang Zhang 0001, Min Li 0022, Kaiqi Huang, Tieniu Tan
ICPR3
2008 A real-time object detecting and tracking system for outdoor night surveillance
Kaiqi Huang, Liangsheng Wang, Tieniu Tan, Stephen J. Maybank
Pattern Recognit.1
2007 Continuously Tracking Objects Across Multiple Widely Separated Cameras
Yinghao Cai, Wei Chen 0012, Kaiqi Huang, Tieniu Tan
ACCV (1)3
2007 Multi-view Gymnastic Activity Recognition with Fused HMM
Ying Wang 0003, Kaiqi Huang, Tieniu Tan
ACCV (1)2
2007 Cast Shadow Removal Combining Local and Global Features
abstract
In this paper, we present a method using pixel-level information, local region-level information and global-level information to remove shadow. At the pixel-level, we employ GMM to model the behavior of cast shadow for every pixel in the HSV color space, as it can deal with complex illumination conditions. However, unlike the GMM for background which can obtain sample every frame, this model for shadow needs more frames to get the same number of sample, because shadow may not appear at the same pixel for each frame. Therefore, it will take a long time to converge. To overcome this drawback, we use the local region-level information to get more samples and global-level information to improve a preclassifier and then, by using it, we get samples which are more likely to be shadow. Also, at the local region-level, we use Markov random fields to represent dependencies between the label of single pixel and labels of its neighborhood. Moreover, to make global level information more robust, tracking information is used. Experimental results show that the proposed method is efficient and robust.
Kaiqi Huang, Tieniu Tan, Liangsheng Wang
CVPR2
2007 Recognizing Night Walkers Based on One Pseudoshape Representation of Gait
abstract
Gait is a promising biometric cue which can facilitate the recognition of human beings, particularly when other biometrics are unavailable. Existing work for gait recognition, however, lays more emphasis on the problem of daytime walker recognition and overlooks the significance of walker recognition at night. This paper deals with the problem of recognizing nighttime walkers. We take advantage of infrared gait patterns to accomplish this task: 1) Walker detection is improved using intensity compensation-based background subtraction; 2) pseudoshape-based features are proposed to describe gait patterns; 3) the dimension of gait features is reduced through the principal component analysis (PCA) and linear discriminant analysis (LDA) techniques; 4) temporal cues are exploited in the form of the relevant component analysis (RCA) learning; 5) the nearest neighbor classifier is used to recognize unknown gait. Experimental results justify the effectiveness of our method and show that our method has an encouraging potential for the application in surveillance systems.
Daoliang Tan, Kaiqi Huang, Shiqi Yu 0001, Tieniu Tan
CVPR2
2007 Human Activity Recognition Based on R Transform
abstract
This paper addresses human activity recognition based on a new feature descriptor. For a binary human silhouette, an extended radon transform, R transform, is employed to represent low-level features. The advantage of the R transform lies in its low computational complexity and geometric invariance. Then a set of HMMs based on the extracted features are trained to recognize activities. Compared with other commonly-used feature descriptors, R transform is robust to frame loss in video, disjoint silhouettes and holes in the shape, and thus achieves better performance in recognizing similar activities. Rich experiments have proved the efficiency of the proposed method.
Ying Wang 0003, Kaiqi Huang, Tieniu Tan
CVPR2
2007 EDA Approach for Model Based Localization and Recognition of Vehicles
abstract
We address the problem of model based recognition. Our aim is to localize and recognize road vehicles from monocular images in calibrated scenes. A deformable 3D geometric vehicle model with 12 parameters is set up as prior information and Bayesian Classification Error is adopted for evaluation of fitness between the model and images. Using a novel evolutionary computing method called EDA (Estimation of Distribution Algorithm), we can not only determine the 3D pose of the vehicle, but also obtain a 12 dimensional vector which corresponds to the 12 shape parameters of the model. By clustering obtained vectors in the parameter space, we can recognize different types of vehicles. Experimental results demonstrate the effectiveness of the approach to vehicles of different types and poses. Thanks to EDA, we can not only localize and recognize vehicles, but also show the whole evolution procedure of the deformable model which gradually fits the image better and better.
Zhaoxiang Zhang 0001, Weishan Dong, Kaiqi Huang, Tieniu Tan
CVPR3
2007 Trajectory Series Analysis based Event Rule Induction for Visual Surveillance
abstract
In this paper, a generic rule induction framework based on trajectory series analysis is proposed to learn the event rules. First the trajectories acquired by a tracking system are mapped into a set of primitive events that represent some basic motion patterns of moving object. Then a minimum description length (MDL) principle based grammar induction algorithm is adopted to infer the meaningful rules from the primitive event series. Compared with previous grammar rule based work on event recognition where the rules are all defined manually, our work aims to learn the event rules automatically. Experiments in a traffic crossroad have demonstrated the effectiveness of our methods. Shown in the experimental results, most of the grammar rules obtained by our algorithm are consistent with the actual traffic events in the crossroad. Furthermore the traffic lights rule in the crossroad can also be leaned correctly with the help of eliminating the irrelevant trajectories.
Zhang Zhang 0001, Kaiqi Huang, Tieniu Tan, Liangsheng Wang
CVPR2
2007 Orthogonal Diagonal Projections for Gait Recognition
abstract
Gait has received much attention from researchers in the vision field due to its utility in walker identification. One of the key issues in gait recognition is how to extract discriminative shape features from 2D human silhouette images. This paper deals with the problem of gait-based walker recognition using statistical shape features. First, we normalize walkers' silhouettes (to facilitate gait feature comparison) into a square form and use the orthogonal projections in the positive and negative diagonal directions to draw personal signatures contained in gait patterns. Then principal component analysis (PCA) and linear discriminant analysis (LDA) are applied to reduce the dimensionality of original gait features and to improve the topological structure in the feature space. Finally, this paper accomplishes the recognition of unknown gait features based on the nearest neighbor rule, with the discussion of the effect of distance metrics and scales on discriminating performance. Experimental results justify the potential of our method.
Daoliang Tan, Kaiqi Huang, Shiqi Yu 0001, Tieniu Tan
ICIP (1)2
2007 Abnormal Activity Recognition in Office Based on R Transform
abstract
This paper introduces an abnormal activity recognition method based on a new feature descriptor for human silhouette. For a binary human silhouette, an extended radon transform, R transform, is employed to represent low-level features. The information that the initial silhouette carries is transformed in a compact way preserving important spatial information of the activities. Then a set of HMMs based on the features extracted by our method are trained to recognize abnormal activities. Experiments have proved the accuracy and efficiency of the proposed method, and the comparison with Fourier descriptor illustrates its robustness to disjoint shapes and shapes with holes.
Ying Wang 0003, Kaiqi Huang, Tieniu Tan
ICIP (1)2
2007 Group Activity Recognition Based on ARMA Shape Sequence Modeling
abstract
In this paper, we propose a system identification approach for group activity recognition in traffic surveillance. Statistical shape theory is used to extract features, and then ARMA (autoregressive and moving average) is adopted for feature learning and activity identification. Here only a few points, instead of the complete trajectory of each object are used to describe the dynamic information of group activity. And ARMA is employed to learn activity sequences. The performance of the proposed method is proved by experiments on 570 video sequences, with the average recognition rate of 88% (compared with 81% of HMM). The extracted features are invariant to zoom, pan and tilt, which is also proved in the experiments.
Ying Wang 0003, Kaiqi Huang, Tieniu Tan
ICIP (3)2
2007 Real-Time Moving Object Classification with Automatic Scene Division
abstract
We address the problem of moving object classification. Our aim is to classify moving objects of traffic scene videos into pedestrians, bicycles and vehicles. Instead of supervised learning and manual labeling of large training samples, our classifiers are initialized and refined online automatically. With efficient features extracted and organized, the approach can be real-time and achieve high classification accuracy. Once the view or scene changes detected, the algorithm can automatically refine the classifiers and adapt them to new environments. Experimental results demonstrate the effectiveness and robustness of the proposed approach.
Zhaoxiang Zhang 0001, Yinghao Cai, Kaiqi Huang, Tieniu Tan
ICIP (5)3
2006 Detecting and Tracking Distant Objects at Night Based on Human Visual System
Kaiqi Huang, Liangsheng Wang, Tieniu Tan
ACCV (2)1
2006 Complex Activity Representation and Recognition by Extended Stochastic Grammar
Zhang Zhang 0001, Kaiqi Huang, Tieniu Tan
ACCV (1)2
2006 Natural color image enhancement and evaluation algorithm based on human visual system
Kaiqi Huang, Zhenyang Wu
Comput. Vis. Image Underst.1
2005 Image enhancement based on the statistics of visual representation
Kaiqi Huang, Zhenyang Wu
Image Vis. Comput.1
2004 Color image enhancement and evaluation algorithm based on human visual system
abstract
We propose a novel color image enhancement method, the human visual system controlled color image enhancement and evaluation (HCCIEE) algorithm, and apply it to color images by considering natural image quality metrics. This HCCIEE algorithm is based on a multiscale representation of pattern, luminance and color processing in the human visual system. Experiments illustrate that the HCCIEE algorithm can produce distinguishing details while avoiding artifacts, which often occur in conventional multiscale enhancement methods, as well as producing images that appear as similar as possible to the viewer's perception of actual scenes.
Kaiqi Huang, Zhenyang Wu
ICASSP (3)1