EDBT 2026 Demo / reviewers in the wild / expert
Gang Hua 0001
dblp:75/5209
· DBLP profile ↗
270ranked-venue papers
18as first author
113since 2021 · last 2026
0000-0001-9522-6157ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 210 · 16 first-author · 91 since 2021Graphics, computer vision, multimedia, augmented reality and games · 192 · 10 first-author · 69 since 2021Computer networks · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Systems, architecture and hardware · 2Security and privacy · 2Applied, interdisciplinary, general and emerging computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RayD3D: Distilling Depth Knowledge Along the Ray for Robust Multi-View 3D Object DetectionabstractMulti-view 3D detection with bird’s eye view (BEV) is crucial for autonomous driving and robotics, but its robustness in real-world is limited as it struggles to predict accurate depth values. A mainstream solution, cross-modal distillation, transfers depth information from LiDAR to camera models but also unintentionally transfers depth-irrelevant information (e.g. LiDAR density). To mitigate this issue, we propose RayD3D, which transfers crucial depth knowledge along the ray: a line projecting from the camera to true location of an object. It is based on the fundamental imaging principle that predicted location of this object can only vary along this ray, which is finally determined by predicted depth value. Therefore, distilling along the ray enables more effective depth information transfer. More specifically, we design two ray-based distillation modules. Ray-based Contrastive Distillation (RCD) incorporates contrastive learning into distillation by sampling along the ray to learn how LiDAR accurately locates objects. Ray-based Weighted Distillation (RWD) adaptively adjusts distillation weight based on the ray to minimize the interference of depth-irrelevant information in LiDAR. For validation, we widely apply RayD3D into three representative types of BEV-based models, including BEVDet, BEVDepth4D, and BEVFormer. Our method is trained on clean NuScenes, and tested on both clean NuScenes and RoboBEV with a variety types of data corruptions. Our method significantly improves the robustness of all the three base models in all scenarios without increasing inference costs, and achieves the best when compared to recently released multi-view and distillation models. Zhaonian Kuang, Zongwei Zhou, Meng Yang 0002, Xinhu Zheng, Gang Hua 0001 |
AAAI | 6 |
| 2026 | Efficient LLM-Jailbreaking via Multimodal-LLM JailbreakabstractThis paper focuses on jailbreaking attacks against large language models (LLMs), eliciting them to generate objectionable content in response to harmful user queries. Unlike previous LLM-jailbreak methods that directly orient to LLMs, our approach begins by constructing a multimodal large language model (MLLM) built upon the target LLM. Subsequently, we perform an efficient MLLM jailbreak and obtain a jailbreaking embedding. Finally, we convert the embedding into a textual jailbreaking suffix to carry out the jailbreak of target LLM. Compared to the direct LLM-jailbreak methods, our indirect jailbreaking approach is more efficient, as MLLMs are more vulnerable to jailbreak than pure LLMs. Additionally, to improve the attack success rate of jailbreak, we propose an image-text semantic matching scheme to identify a suitable initial input. Extensive experiments demonstrate that our approach surpasses current state-of-the-art jailbreak methods in terms of both efficiency and effectiveness. Moreover, our approach exhibits superior cross-class generalization abilities. Haoxuan Ji, Zhenxing Niu, Xinbo Gao 0001, Gang Hua 0001 |
AAAI | 5 |
| 2026 | Object-Scene-Camera Decomposition and Recomposition for Data Efficient Monocular 3D Object Detection
Zhaonian Kuang, Meng Yang 0002, Xinhu Zheng, Gang Hua 0001 |
Int. J. Comput. Vis. | 5 |
| 2026 | Self-Supervised Video Representation Learning in a Heuristic Decoupled Perspective
Zeen Song, Wenwen Qiang, Changwen Zheng, Hui Xiong 0001, Gang Hua 0001 |
Int. J. Comput. Vis. | 5 |
| 2026 | Sparse Trajectory PredictionabstractPedestrian trajectory prediction is crucial for ensuring safe decision-making in intelligent robotic systems. While this task demands real-time performance, previous works have primarily focused on improving prediction accuracy, often neglecting efficiency. Dense predictions with time-consuming post-clustering steps and global interactions with quadratic computational complexity result in a trade-off between accuracy and speed. In this paper, we propose a novel Sparse Trajectory Prediction (STP) model that aims to achieve both high accuracy and real-time speed by following an efficient principle: leveraging sparse structures to achieve global effects. STP instantiates this principle within a transformer-style encoder-decoder framework. In the encoder, STP introduces irregular interaction, which builds sparse interactions with dynamic interactive positions, reducing computational complexity to linearithmic/linear while maintaining global interaction. In the decoder, STP applies an early-sparsity strategy to generate sparse motion modes that represent global motion behaviors. These modes are shared across all predictions, eliminating redundant computations. By harnessing the expressive power of transformers, STP maps these sparse motion modes into multimodal future trajectories, significantly improving prediction speed while ensuring accuracy. Experimental results on four commonly used datasets demonstrate that STP maximizes both accuracy and prediction speed, achieving state-of-the-art performance and significantly improving prediction speed by about $100 \times$100× - $150 \times$150× to satisfy the real-time demand. Liushuai Shi, Le Wang 0003, Sanping Zhou, Wei Tang 0016, Gang Hua 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2026 | Unifying Multi-Modal Hair Editing via Proxy Feature BlendingabstractHair editing is a long-standing problem in computer vision that demands both fine-grained local control and intuitive user interactions across diverse modalities. Despite the remarkable progress of GANs and diffusion models, existing methods still lack a unified framework that simultaneously supports arbitrary interaction modes (e.g., text, sketch, mask, and reference image) while ensuring precise editing and faithful preservation of irrelevant attributes. In this work, we introduce a novel paradigm that reformulates hair editing as proxy-based hair transfer. Specifically, we leverage the dense and semantically disentangled latent space of StyleGAN for precise manipulation and exploit its feature space for disentangled attribute preservation, thereby decoupling the objectives of editing and preservation. Our framework unifies different modalities by converting editing conditions into distinct transfer proxies, whose features are seamlessly blended to achieve global or local edits. Beyond 2D, we extend our paradigm to 3D-aware settings by incorporating EG3D and PanoHead, where we propose a multi-view boosted hair feature localization strategy together with 3D-tailored proxy generation methods that exploit the inherent properties of 3D-aware generative models. Extensive experiments demonstrate that our method consistently outperforms prior approaches in editing effects, attribute preservation, visual naturalness, and multi-view consistency, while offering unprecedented support for multimodal and mixed-modal interactions. Tianyi Wei, Dongdong Chen 0001, Wenbo Zhou 0004, Jing Liao 0001, Can Wang 0007, Weiming Zhang 0001, Gang Hua 0001, Nenghai Yu |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2026 | Data-Driven Bidirectional Spatial-Adaptive Network for Weakly Supervised Object Detection in Remote Sensing ImagesabstractWeakly-supervised object detection (WSOD) learns detectors with only image-level classification annotations. Without precise instance-level labels, most previous WSOD methods in remote sensing images (RSIs) select the highest-scoring proposals as the final detection results, which are confronted by two major challenges: (1) instances with small scale or rare poses are easily neglected; (2) optimizing network by the top-scoring region inevitably overlooks many valuable candidate proposals. To mitigate the above-mentioned challenges, we propose a data-driven bidirectional spatial-adaptive network (BSANet). It contains a forward-reverse spatial dropout (FRSD) module to reduce instance ambiguity induced from extreme scales and poses, as well as crowded scene, and to better excavate the entire instances. From attention learning perspective, the proposed FRSD is conceptually similar to a data-driven hard attention mechanism, which adaptively samples and reconstructs the spatially related regions for mining more latent feature responses. Meanwhile, our FRSD effectively alleviates the inherent problem that non-parametric hard attention learning fashion cannot adapt to different datasets. In addition, we build a soft attention branch to simultaneously model soft pixel-level and hard region-level attention information for exploring the complementary benefit between soft and hard attention learning. We evaluate our BSANet on the challenging NWPU VHR-10.v2 and DIOR datasets. Experimental results demonstrate that our method sets a new state-of-the-art. Zebin Wu 0001, Shangdong Zheng, Yang Xu 0006, Le Wang 0003, Zhihui Wei, Gang Hua 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2026 | Deployment Prior Injection for Run-Time Re-Biasable Object Detection
Yiding Yang, Vishal M. Patel, Gang Hua 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2026 | RAMP: Iterative Refinement and Adaptive Multi-granularity Perception for embodied dialog localization
Le Wang 0003, Sanping Zhou, Gang Hua 0001, Wei Tang 0016 |
Pattern Recognit. | 5 |
| 2026 | Multi-Modal Decouple and Recouple Network for Robust 3D Object DetectionabstractMulti-modal 3D object detection with bird’s eye view (BEV) has achieved desired advances on benchmarks. Nonetheless, the accuracy may drop significantly in the real world due to data corruption such as sensor configurations for LiDAR and scene conditions for camera. One design bottleneck of previous models resides in the tightly coupling of multi-modal BEV features during fusion, which may degrade the overall system performance if one modality or both is corrupted. To mitigate, we propose a Multi-Modal Decouple and Recouple Network for robust 3D object detection under data corruption. Different modalities commonly share some high-level invariant features. We observe that these invariant features across modalities do not always fail simultaneously, because different types of data corruption affect each modality in distinct ways. These invariant features can be recovered across modalities for robust fusion under data corruption. To this end, we explicitly decouple Camera/LiDAR BEV features into modality-invariant and modality-specific parts. It allows invariant features to compensate each other while mitigates the negative impact of a corrupted modality on the other. We then recouple these features into three experts to handle different types of data corruption, respectively, i.e., LiDAR, camera, and both. For each expert, we use modality-invariant features as robust information, while modality-specific features serve as a complement. Finally, we adaptively fuse the three experts to exact robust features for 3D object detection. For validation, we collect a benchmark with a large quantity of data corruption for LiDAR, camera, and both based on nuScenes. Our model is trained on clean nuScenes and tested on all types of data corruption. Our model consistently achieves the best accuracy on both corrupted and clean data compared to recent models. Zhaonian Kuang, Yuzhe Ji, Meng Yang 0002, Xinhu Zheng, Gang Hua 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | Pedestrian Trajectory Prediction via Hierarchical Dynamics DecompositionabstractPredicting human future trajectories is crucial for various intelligent systems and applications. Previous approaches typically adopt a direct prediction strategy, which decodes trajectory features directly into future coordinates. However, they overlook different hierarchical high-order velocities, which have stronger representational abilities in dynamics. In this paper, we introduce HDDNet, a novel trajectory prediction framework that follows dynamical principles and employs a hierarchical dynamics decomposition strategy. Specifically, HDDNet models future trajectories by progressively transferring trajectory coordinates into velocity, acceleration, and jerk, up to the highest-order velocity, which sequentially represent a broader receptive field and a more compact representation of motion dynamics. Furthermore, we design a hierarchical dynamics decomposition decoder with a corresponding dynamics loss, which predicts future trajectories by sequentially refining human motions from the highest-order velocity down to the final coordinates. Compared to the traditional direct prediction strategy, our approach makes better use of dynamic information at different levels. Extensive experiments and ablation studies on the ETH-UCY, SDD and GigaTraj datasets demonstrate that our method outperforms existing state-of-the-art approaches. Yonghao Dong, Le Wang 0003, Sanping Zhou, Gang Hua 0001, Changyin Sun 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | RSRNav: Reasoning Spatial Relationship for Image-Goal NavigationabstractRecent image-goal navigation (ImageNav) methods learn a perception-action policy by separately capturing semantic features of the goal and egocentric images, then passing them to a policy network. However, challenges remain: (1) Semantic features often fail to provide accurate directional information, leading to superfluous actions, and (2) performance drops significantly when viewpoint inconsistencies arise between training and application. To address these challenges, we propose RSRNav, a simple yet effective method that reasons spatial relationships between the goal and current observations as navigation guidance. Specifically, we model the spatial relationship by constructing correlations between the goal and current observations, which are then passed to the policy network for action prediction. These correlations are progressively refined using fine-grained cross-correlation and direction-aware correlation for more precise navigation. Extensive evaluation of RSRNav on three benchmark datasets demonstrates superior navigation performance, particularly in the "user-matched goal" setting, highlighting its potential for real-world applications. Code: https://github.com/ qinzheng2000/RSRNav.git. Le Wang 0003, Sanping Zhou, Gang Hua 0001, Wei Tang 0016 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Diversifying Query: Region-Guided Transformer for Temporal Sentence GroundingabstractTemporal sentence grounding is a challenging task that aims to localize the moment spans relevant to a language description. Although recent DETR-based models have achieved notable progress by leveraging multiple learnable moment queries, they suffer from overlapped and redundant proposals, leading to inaccurate predictions. We attribute this limitation to the lack of task-related guidance for the learnable queries to serve a specific mode. Furthermore, the complex solution space generated by variable and open-vocabulary language descriptions complicates optimization, making it harder for learnable queries to adaptively distinguish each other, leading to more severe overlapped proposals. To address this limitation, we present the Region-Guided TRansformer (RGTR) for temporal sentence grounding, which introduces regional guidance to increase query diversity and eliminate overlapped proposals. Instead of using learnable queries, RGTR adopts a set of anchor pairs as moment queries to introduce explicit regional guidance. Each moment query takes charge of moment prediction for a specific temporal region, which reduces the optimization difficulty and ensures the diversity of the proposals. In addition, we design an IoU-aware scoring head to improve proposal quality. Extensive experiments demonstrate the effectiveness of RGTR, outperforming state-of-the-art methods on three public benchmarks and exhibiting good generalization and robustness on out-of-distribution splits. Xiaolong Sun, Liushuai Shi, Le Wang 0003, Sanping Zhou, Gang Hua 0001 |
AAAI | 7 |
| 2025 | Boosting Point-Supervised Temporal Action Localization through Integrating Query Reformation and Optimal TransportabstractPoint-supervised Temporal Action Localization poses significant challenges due to the difficulty of identifying complete actions with a single-point annotation per action. Existing methods typically employ Multiple Instance Learning, which struggles to capture global temporal context and requires heuristic post-processing. In research on fully-supervised tasks, DETR-based structures have effectively addressed these limitations. However, it is nontrivial to merely adapt DETR to this task, encountering two major bottlenecks. (1) How to integrate point label information into the model and (2) How to select optimal decoder proposals for training in the absence of complete action segment annotations. To address this issue, we introduce an end-to-end framework by integrating Query Reformation and Optimal Transport (QROT). Specifically, we encode point labels through a set of semantic consensus queries, enabling effective focus on action-relevant snippets. Furthermore, we integrate an optimal transport mechanism to generate high-quality pseudo labels. These pseudo-labels facilitate precise proposals selection based on the Hungarian algorithm, significantly enhancing localization accuracy in point-supervised settings. Extensive experiments on the THUMOS14 and ActivityNet-v1.3 datasets demonstrate that our method outperforms existing MIL-based approaches, offering more stable and accurate temporal action localization in point-level supervision. Mengnan Liu 0001, Le Wang 0003, Sanping Zhou, Xiaolong Sun, Gang Hua 0001 |
CVPR | 6 |
| 2025 | Towards Precise Embodied Dialogue Localization via Causality Guided DiffusionabstractEmbodied localization based on vision and natural language dialogues presents a persistent challenge in embodied intelligence. Existing methods often approach this task as an image translation problem, leveraging encoder-decoder architectures to predict heatmaps. However, these methods frequently experience a deficiency in accuracy, largely due to their heavy reliance on resolution. To address this issue, we introduce CGD, a novel framework that utilizes causality guided diffusion model to directly model coordinate distributions. Specifically, CGD employs a denoising network to regress coordinates, while integrating causal learning modules, namely back-door adjustment (BDA) and front-door adjustment (FDA) to mitigate confounders during the diffusion process. This approach reduces the dependency on high resolution for improving accuracy, while effectively minimizing spurious correlations, thereby promoting unbiased learning. By guiding the denoising process with causal adjustments, CGD offers flexible control over intensity, ensuring seamless integration with diffusion models. Experimental results demonstrate that CGD outperforms state-of-the-art methods across all metrics. Additionally, we also evaluate CGD in a multi-shot setting, achieving consistently high accuracy. Le Wang 0003, Sanping Zhou, Jingyi Tian, Gang Hua 0001, Wei Tang 0016 |
CVPR | 7 |
| 2025 | Jigsaw++: Imagining Complete Shape Priors for Object ReassemblyabstractThe automatic assembly problem has attracted increasing interest due to its complex challenges that involve 3D representation. This paper introduces Jigsaw++, a novel generative method designed to tackle the multifaceted challenges of reconstructing complete shape for the reassembly problem. Existing approach focusing primarily on piecewise information for both part and fracture assembly, often overlooking the integration of complete object prior. Jigsaw++ distinguishes itself by learning a shape prior of complete objects. It employs the proposed "retargeting" strategy that effectively leverages the output of any existing assembly method to generate complete shape reconstructions. This capability allows it to function orthogonally to the current methods. Through extensive evaluations on Breaking Bad dataset and PartNet, Jigsaw++ has demonstrated its effectiveness, reducing reconstruction errors and enhancing the precision of shape reconstruction, which sets a new direction for future reassembly model developments. Jiaxin Lu 0001, Gang Hua 0001, Qixing Huang |
ICCV | 2 |
| 2025 | Moment Quantization for Video Temporal GroundingabstractVideo temporal grounding is a critical video understanding task, which aims to localize moments relevant to a language description. The challenge of this task lies in distinguishing relevant and irrelevant moments. Previous methods focused on learning continuous features exhibit weak differentiation between foreground and background features. In this paper, we propose a novel Moment-Quantization based Video Temporal Grounding method (MQVTG), which quantizes the input video into various discrete vectors to enhance the discrimination between relevant and irrelevant moments. Specifically, MQVTG maintains a learnable moment codebook, where each video moment matches a codeword. Considering the visual diversity, i.e., various visual expressions for the same moment, MQVTG treats moment-codeword matching as a clustering process without using discrete vectors, avoiding the loss of useful information from direct hard quantization. Additionally, we employ effective prior-initialization and joint-projection strategies to enhance the maintained moment codebook. With its simple implementation, the proposed method can be integrated into existing temporal grounding models as a plug-and-play component. Extensive experiments on six popular benchmarks demonstrate the effectiveness and generalizability of MQVTG, significantly outperforming state-of-the-art methods. Further qualitative analysis shows that our method effectively groups relevant features and separates irrelevant ones, aligning with our goal of enhancing discrimination. Xiaolong Sun, Le Wang 0003, Sanping Zhou, Liushuai Shi, Mengnan Liu 0001, Gang Hua 0001 |
ICCV | 8 |
| 2025 | Token Coordinated Prompt Attention is Needed for Visual PromptingabstractVisual prompting techniques are widely used to efficiently fine-tune pretrained Vision Transformers (ViT) by learning a small set of shared prompts for all tokens. However, existing methods overlook the unique roles of different tokens in conveying discriminative information and interact with all tokens using the same prompts, thereby limiting the representational capacity of ViT. This often leads to indistinguishable and biased prompt-extracted features, hindering performance. To address this issue, we propose a plug-and-play Token Coordinated Prompt Attention (TCPA) module, which assigns specific coordinated prompts to different tokens for attention-based interactions. Firstly, recognizing the distinct functions of CLS and image tokens-global information aggregation and local feature extraction, we disentangle the prompts into CLS Prompts and Image Prompts, which interact exclusively with CLS tokens and image tokens through attention mechanisms. This enhances their respective discriminative abilities. Furthermore, as different image tokens correspond to distinct image patches and contain diverse information, we employ a matching function to automatically assign coordinated prompts to individual tokens. This enables more precise attention interactions, improving the diversity and representational capacity of the extracted features. Extensive experiments across various benchmarks demonstrate that TCPA significantly enhances the diversity and discriminative power of the extracted features. Xu Zou 0002, Gang Hua 0001, Jiahuan Zhou |
ICML | 3 |
| 2025 | Componential Prompt-Knowledge Alignment for Domain Incremental LearningabstractDomain Incremental Learning (DIL) aims to learn from non-stationary data streams across domains while retaining and utilizing past knowledge. Although prompt-based methods effectively store multi-domain knowledge in prompt parameters and obtain advanced performance through cross-domain prompt fusion, we reveal an intrinsic limitation: component-wise misalignment between domain-specific prompts leads to conflicting knowledge integration and degraded predictions. This arises from the random positioning of knowledge components within prompts, where irrelevant component fusion introduces interference. To address this, we propose Componential Prompt-Knowledge Alignment (KA-Prompt), a novel prompt-based DIL method that introduces component-aware prompt-knowledge alignment during training, significantly improving both the learning and inference capacity of the model. KA-Prompt operates in two phases: (1) Initial Componential Structure Configuring, where a set of old prompts containing knowledge relevant to the new domain are mined via greedy search, which is then exploited to initialize new prompts to achieve reusable knowledge transfer and establish intrinsic alignment between new and old prompts. (2) Online Alignment Preservation, which dynamically identifies the target old prompts and applies adaptive componential consistency constraints as new prompts evolve. Extensive experiments on DIL benchmarks demonstrate the effectiveness of our KA-Prompt. Our source code is available at https://github.com/zhoujiahuan1991/ICML2025-KA-Prompt. Kunlun Xu, Xu Zou 0002, Gang Hua 0001, Jiahuan Zhou |
ICML | 3 |
| 2025 | DynaRend: Learning 3D Dynamics via Masked Future Rendering for Robotic ManipulationabstractLearning generalizable robotic manipulation policies remains a key challenge due to the scarcity of diverse real-world training data. While recent approaches have attempted to mitigate this through self-supervised representation learning, most either rely on 2D vision pretraining paradigms such as masked image modeling, which primarily focus on static semantics or scene geometry, or utilize large-scale video prediction models that emphasize 2D dynamics, thus failing to jointly learn the geometry, semantics, and dynamics required for effective manipulation. In this paper, we present DynaRend, a representation learning framework that learns 3D-aware and dynamics-informed triplane features via masked reconstruction and future prediction using differentiable volumetric rendering. By pretraining on multi-view RGB-D video data, DynaRend jointly captures spatial geometry, future dynamics, and task semantics in a unified triplane representation. The learned representations can be effectively transferred to downstream robotic manipulation tasks via action value map prediction. We evaluate DynaRend on two challenging benchmarks, RLBench and Colosseum, as well as in real-world robotic experiments, demonstrating substantial improvements in policy success rate, generalization to environmental perturbations, and real-world applicability across diverse manipulation tasks. Jingyi Tian, Le Wang 0003, Sanping Zhou, Gang Hua 0001 |
NeurIPS | 6 |
| 2025 | SAMPO: Scale-wise Autoregression with Motion Prompt for Generative World ModelsabstractWorld models allow agents to simulate the consequences of actions in imagined environments for planning, control, and long-horizon decision-making. However, existing autoregressive world models struggle with visually coherent predictions due to disrupted spatial structure, inefficient decoding, and inadequate motion modeling. In response, we propose Scale-wise Autoregression with Motion PrOmpt (SAMPO), a hybrid framework that combines visual autoregressive modeling for intra-frame generation with causal modeling for next-frame generation. Specifically, SAMPO integrates temporal causal decoding with bidirectional spatial attention, which preserves spatial locality and supports parallel decoding within each scale. This design significantly enhances both temporal consistency and rollout efficiency. To further improve dynamic scene understanding, we devise an asymmetric multi-scale tokenizer that preserves spatial details in observed frames and extracts compact dynamic representations for future frames, optimizing both memory usage and model performance. Additionally, we introduce a trajectory-aware motion prompt module that injects spatiotemporal cues about object and robot trajectories, focusing attention on dynamic regions and improving temporal consistency and physical realism. Extensive experiments show that SAMPO achieves competitive performance in action-conditioned video prediction and model-based control, improving generation quality with 4.4× faster inference. We also evaluate SAMPO's zero-shot generalization and scaling behavior, demonstrating its ability to generalize to unseen tasks and benefit from larger model sizes. Jingyi Tian, Le Wang 0003, Zhimin Liao, Huaiyi Dong, Sanping Zhou, Wei Tang 0016, Gang Hua 0001 |
NeurIPS | 10 |
| 2025 | State Space Prompting via Gathering and Spreading Spatio-Temporal Information for Video UnderstandingabstractRecently, pre-trained state space models have shown great potential for video classification, which sequentially compresses visual tokens in videos with linear complexity, thereby improving the processing efficiency of video data while maintaining high performance. To apply powerful pre-trained models to downstream tasks, prompt learning is proposed to achieve efficient downstream task adaptation with only a small number of fine-tuned parameters. However, the sequentially compressed visual prompt tokens fail to capture the spatial and temporal contextual information in the video, thus limiting the effective propagation of spatial information within a video frame and temporal information between frames in the state compression model and the extraction of discriminative information. To tackle the above issue, we proposed a State Space Prompting (SSP) method for video understanding, which combines intra-frame and inter-frame prompts to aggregate and propagate key spatiotemporal information in the video. Specifically, an Intra-Frame Gathering (IFG) module is designed to aggregate spatial key information within each frame. Besides, an Inter-Frame Spreading (IFS) module is designed to spread discriminative spatio-temporal information across different frames. By adaptively balancing and compressing key spatio-temporal information within and between frames, our SSP effectively propagates discriminative information in videos in a complementary manner. Extensive experiments on four video benchmark datasets verify that our SSP significantly outperforms existing SOTA methods by 2.76\% on average while reducing the overhead of fine-tuning parameters. Jiahuan Zhou, Zhenyu Cui, Xu Zou 0002, Gang Hua 0001 |
NeurIPS | 6 |
| 2025 | Class-aware Domain Knowledge Fusion and Fission for Continual Test-Time AdaptationabstractContinual Test-Time Adaptation (CTTA) aims to quickly fine-tune the model during the test phase so that it can adapt to multiple unknown downstream domain distributions without pre-acquiring downstream domain data.
To this end, existing advanced CTTA methods mainly reduce the catastrophic forgetting of historical knowledge caused by irregular switching of downstream domain data by restoring the initial model or reusing historical models. However, these methods are usually accompanied by serious insufficient learning of new knowledge and interference from potentially harmful historical knowledge, resulting in severe performance degradation. To this end, we propose a class-aware domain Knowledge Fusion and Fission method for continual test-time adaptation, called KFF, which adaptively expands and merges class-aware domain knowledge in old and new domains according to the test-time data from different domains, where discriminative historical knowledge can be dynamically accumulated. Specifically, considering the huge domain gap within streaming data, a domain Knowledge FIssion (KFI) module is designed to adaptively separate new domain knowledge from a paired class-aware domain prompt pool, alleviating the impact of negative knowledge brought by old domains that are distinct from the current domain. Besides, to avoid the cumulative computation and storage overheads from continuously fissioning new knowledge, a domain Knowledge FUsion (KFU) module is further designed to merge the fissioned new knowledge into the existing knowledge pool with minimal cost, where a greedy knowledge dynamic merging strategy is designed to improve the compatibility of new and old knowledge while keeping the computational efficiency. Jiahuan Zhou, Zhenyu Cui, Xu Zou 0002, Gang Hua 0001 |
NeurIPS | 6 |
| 2025 | Editorial: Introduction to the Special Section on Best of CVPR'2022
Kristin J. Dana, Gang Hua 0001, Stefan Roth 0001, Dimitris Samaras |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | AFC-RNN: Adaptive Forgetting-Controlled Recurrent Neural Network for Pedestrian Trajectory PredictionabstractPedestrian trajectory prediction plays a crucial and fundamental role in many computer vision tasks. Most existing works utilize recurrent neural networks to extract temporal features from trajectories because their recursive structure is inherently well-suited for time series data. However, previous methods overlook the forgetting characteristics of pedestrians when modeling historical trajectories, which may cause the model to focus on the wrong positions of historical information. In this paper, we propose a simple yet effective Adaptive Forgetting-Controlled Recurrent Neural Network (AFC-RNN) for pedestrian trajectory prediction. The core idea of AFC-RNN is a novel Adaptive Forgetting Controller (AFC), which controls the forgetting degree of the historical information at each time step explicitly and adaptively. Specifically, AFC first learns memory factors for each time step based on the temporal correlation of observed trajectories using the self-attention mechanism. Then, AFC-RNN applies these memory factors to regulate the forgetting degree of observed features at each time step from RNN. Extensive experiments and ablation studies on ETH, UCY, SDD, and NBA datasets demonstrate that our method outperforms existing state-of-the-art approaches. Additionally, we provide a mathematical analysis to demonstrate the superiority of our adaptive forgetting strategy in the AFC-RNN over traditional RNNs for trajectory forgetting modeling. Yonghao Dong, Le Wang 0003, Sanping Zhou, Wei Tang 0016, Gang Hua 0001, Changyin Sun 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | Glissando-Net: Deep Single View Category Level Pose Estimation and 3D ReconstructionabstractWe present a deep learning model, dubbed Glissando-Net, to simultaneously estimate the pose and reconstruct the 3D shape of objects at the category level from a single RGB image. Previous works predominantly focused on either estimating poses (often at the instance level), or reconstructing shapes, but not both. Glissando-Net is composed of two auto-encoders that are jointly trained, one for RGB images and the other for point clouds. We embrace two key design choices in Glissando-Net to achieve a more accurate prediction of the 3D shape and pose of the object given a single RGB image as input. First, we augment the feature maps of the point cloud encoder and decoder with transformed feature maps from the image decoder, enabling effective 2D-3D interaction in both training and prediction. Second, we predict both the 3D shape and pose of the object in the decoder stage. This way, we better utilize the information in the 3D point clouds presented only in the training stage to train the network for more accurate prediction. We jointly train the two encoder-decoders for RGB and point cloud data to learn how to pass latent features to the point cloud decoder during inference. In testing, the encoder of the 3D point cloud is discarded. The design of Glissando-Net is inspired by codeSLAM. Unlike codeSLAM, which targets 3D reconstruction of scenes, we focus on pose estimation and shape reconstruction of objects, and directly predict the object pose and a pose invariant 3D reconstruction without the need of the code optimization step. Extensive experiments, involving both ablation studies and comparison with competing methods, demonstrate the efficacy of our proposed method, and compare favorably with the state-of-the-art. Hao Kang, Philippos Mordohai, Gang Hua 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | Scale Propagation Network for Generalizable Depth CompletionabstractDepth completion, inferring dense depth maps from sparse measurements, is crucial for robust 3D perception. Although deep learning based methods have made tremendous progress in this problem, these models cannot generalize well across different scenes that are unobserved in training, posing a fundamental limitation that yet to be overcome. A careful analysis of existing deep neural network architectures for depth completion, which are largely borrowing from successful backbones for image analysis tasks, reveals that a key design bottleneck actually resides in the conventional normalization layers. These normalization layers are designed, on one hand, to make training more stable, on the other hand, to build more visual invariance across scene scales. However, in depth completion, the scale is actually what we want to robustly estimate in order to better generalize to unseen scenes. To mitigate, we propose a novel scale propagation normalization (SP-Norm) method to propagate scales from input to output, and simultaneously preserve the normalization operator for easy convergence. More specifically, we rescale the input using learned features of a single-layer perceptron from the normalized input, rather than directly normalizing the input as conventional normalization layers. We then develop a new network architecture based on SP-Norm and the ConvNeXt V2 backbone. We explore the composition of various basic blocks and architectures to achieve superior performance and efficient inference for generalizable depth completion. Extensive experiments are conducted on six unseen datasets with various types of sparse depth maps, i.e., randomly sampled 0.1%/1%/10% valid pixels, 4/8/16/32/64-line LiDAR points, and holes from Structured-Light. Our model consistently achieves the best accuracy with faster speed and lower memory when compared to state-of-the-art methods. Haotian Wang 0009, Meng Yang 0002, Xinhu Zheng, Gang Hua 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Recurrent Aligned Network for Generalized Pedestrian Trajectory PredictionabstractPedestrian trajectory prediction is a crucial component in computer vision and robotics, but remains challenging due to the domain shift problem. Previous studies have tried to tackle this problem by leveraging a portion of trajectory data from the target domain to fine-tune the model. However, such domain adaptation methods are impractical in real-world scenarios, as it is infeasible to collect trajectory data from all potential target domains. In this paper, we study a new task named generalized pedestrian trajectory prediction, with the aim of generalizing the model to unseen domains without accessing their trajectories. To tackle this task, we further introduce a Recurrent Aligned Network (RAN) to minimize the domain gap through domain alignment. Specifically, we devise a recurrent alignment module to effectively align the trajectory feature spaces at both time-state and time-sequence levels by the recurrent alignment strategy. Furthermore, we introduce a pre-aligned representation module to combine social interactions with the recurrent alignment strategy, which aims to consider social interactions during the alignment process instead of just target trajectories. We extensively evaluate our method and compare it with state-of-the-art methods on three widely used benchmarks. The experimental results demonstrate the superior generalization capability of our method. Our work not only fills the gap in the generalization setting for practical pedestrian trajectory prediction, but also sets strong baselines in this field. Yonghao Dong, Le Wang 0003, Sanping Zhou, Gang Hua 0001, Changyin Sun 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Pluralistic Salient Object DetectionabstractWe introduce pluralistic salient object detection (PSOD), a novel task aimed at generating multiple plausible salient segmentation results for a given input image. Unlike conventional SOD methods that produce a single segmentation mask for salient objects, this new setting recognizes the inherent complexity of real-world images, comprising multiple objects, and the ambiguity in defining salient objects due to different user intentions. To study this task, we present two new SOD datasets "DUTS-MM" and "DUTS-MQ", along with newly designed evaluation metrics. DUTS-MM builds upon the DUTS dataset but enriches the ground-truth mask annotations from three aspects which 1) improves the mask quality especially for boundary and fine-grained structures; 2) alleviates the annotation inconsistency issue; and 3) provides multiple ground-truth masks for images with saliency ambiguity. DUTS-MQ consists of approximately 100K image-mask pairs with human-annotated preference scores, enabling the learning of real human preferences in measuring mask quality. Building upon these two datasets, we propose a simple yet effective pluralistic SOD baseline based on a Mixture-of-Experts (MOE) design. Equipped with two prediction heads, it simultaneously predicts multiple masks using different query prompts and predicts human preference scores for each mask candidate. Extensive experiments and analyses underscore the significance of our proposed datasets and affirm the effectiveness of our PSOD framework. Xuelu Feng, Yunsheng Li, Dongdong Chen 0001, Chunming Qiao, Junsong Yuan 0001, Lu Yuan 0001, Gang Hua 0001 |
IEEE Trans. Image Process. | 7 |
| 2025 | Meta Pairwise Relationship Distillation for Unsupervised Person Re-IdentificationabstractUnsupervised person re-identification (Re-ID) is challenging due to the lack of ground-truth labels. Most existing methods rely on pseudo labels estimated via iterative clustering and thus are highly susceptible to performance penalties incurred by the inaccurate estimated number of clusters. Alternatively, we utilize the sample pairs with pairwise pseudo labels to guide the feature learning to avoid the dilemma of determining cluster numbers. In this article, we propose a meta pairwise relationship distillation (MPRD) method that incorporates a graph convolutional network (GCN) to provide high-fidelity pairwise relationships to supervise the model training. A small amount of metadata with very-confidence pairwise relationships and the unlabeled pairs with the provided pseudo pairwise relationships participate in the GCN training. Besides, we introduce a hard sample deduction (HSD) module to timely mine the sample pairs with error-prone pairwise pseudo labels to mitigate the misled optimization by noisy labels. Furthermore, since the features of each positive pair represent the same person, we design a positive pair alignment (PPA) module to reduce the redundant information in each feature, which is achieved by minimizing the difference between each positive pair's feature distributions. Extensive experiments on the Market-1501, DukeMTMC-reID, and MSMT17 datasets show that our method outperforms the state-of-the-art unsupervised methods. Haoxuanye Ji, Le Wang 0003, Sanping Zhou, Wei Tang 0016, Nanning Zheng 0001, Gang Hua 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2024 | Temporal Correlation Vision Transformer for Video Person Re-IdentificationabstractVideo Person Re-Identification (Re-ID) is a task of retrieving persons from multi-camera surveillance systems. Despite the progress made in leveraging spatio-temporal information in videos, occlusion in dense crowds still hinders further progress. To address this issue, we propose a Temporal Correlation Vision Transformer (TCViT) for video person Re-ID. TCViT consists of a Temporal Correlation Attention (TCA) module and a Learnable Temporal Aggregation (LTA) module. The TCA module is designed to reduce the impact of non-target persons by relative state, while the LTA module is used to aggregate frame-level features based on their completeness. Specifically, TCA is a parameter-free module that first aligns frame-level features to restore semantic coherence in videos and then enhances the features of the target person according to temporal correlation. Additionally, unlike previous methods that treat each frame equally with a pooling layer, LTA introduces a lightweight learnable module to weigh and aggregate frame-level features under the guidance of a classification score. Extensive experiments on four prevalent benchmarks demonstrate that our method achieves state-of-the-art performance in video Re-ID. Le Wang 0003, Sanping Zhou, Gang Hua 0001, Changyin Sun 0001 |
AAAI | 4 |
| 2024 | Evidential Active Recognition: Intelligent and Prudent Open-World Embodied PerceptionabstractActive recognition enables robots to intelligently explore novel observations, thereby acquiring more information while circumventing undesired viewing conditions. Recent approaches favor learning policies from simulated or collected data, wherein appropriate actions are more frequently selected when the recognition is accurate. However, most recognition modules are developed under the closed-world assumption, which makes them ill-equipped to handle unexpected inputs, such as the absence of the target object in the current observation. To address this issue, we propose treating active recognition as a sequential evidence-gathering process, providing by-step uncertainty quantification and reliable prediction under the evidence combination theory. Additionally, the reward function developed in this paper effectively characterizes the merit of actions when operating in open-world environments. To evaluate the performance, we collect a dataset from an indoor simulator, encompassing various recognition challenges such as distance, occlusion levels, and visibility. Through a series of experiments on recognition and robustness analysis, we demonstrate the necessity of introducing uncertainties to active recognition and the superior performance of the proposed method. Lei Fan 0005, Mingfu Liang, Yunxuan Li, Gang Hua 0001, Ying Wu 0001 |
CVPR | 4 |
| 2024 | DL3DV-10K: A Large-Scale Scene Dataset for Deep Learning-based 3D VisionabstractWe have witnessed significant progress in deep learning-based 3D vision, ranging from neural radiance field (NeRF) based 3D representation learning to applications in novel view synthesis (NVS). However, existing scene-level datasets for deep learning-based 3D vision, limited to ei-ther synthetic environments or a narrow selection of real-world scenes, are quite insufficient. This insufficiency not only hinders a comprehensive benchmark of existing methods but also caps what could be explored in deep learning-based 3D analysis. To address this critical gap, we present DL3DV-10K, a large-scale scene dataset, featuring 51.2 million frames from 10,510 videos captured from 65 types of point- of-interest (POI) locations, covering both bounded and unbounded scenes, with different levels of reflection, transparency, and lighting. We conducted a comprehensive benchmark of recent NVS methods on DL3DV-10K, which revealed valuable insights for future research in NVS. In addition, we have obtained encouraging results in a pilot study to learn generalizable NeRF from DL3DV-10K, which manifests the necessity of a large-scale scene-level dataset to forge a path toward a foundation model for learning 3D representation. Our DL3DV-10K dataset, benchmark results, and models will be publicly accessible. Lu Ling, Yichen Sheng, Zhi Tu, Wentian Zhao, Cheng Xin, Kun Wan 0001, Lantao Yu, Zixun Yu, Yawen Lu, Xuanmao Li, Xingpeng Sun, Rohan Ashok, Aniruddha Mukherjee, Hao Kang, Xiangrui Kong, Gang Hua 0001, Tianyi Zhang 0001, Bedrich Benes, Aniket Bera |
CVPR | 17 |
| 2024 | Towards Generalizable Multi-Object TrackingabstractMulti-Object Tracking (MOT) encompasses various tracking scenarios, each characterized by unique traits. Ef-fective trackers should demonstrate a high degree of gen-eralizability across diverse scenarios. However, existing trackers struggle to accommodate all aspects or necessi-tate hypothesis and experimentation to customize the asso-ciation information (motion and/or appearance) for a given scenario, leading to narrowly tailored solutions with limited generalizability. In this paper, we investigate the factors that influence trackers' generalization to different scenar-ios and concretize them into a set of tracking scenario at-tributes to guide the design of more generalizable trackers. Furthermore, we propose a “point-wise to instance-wise relation” framework for MOT, i.e., GeneralTrack, which can generalize across diverse scenarios while eliminating the need to balance motion and appearance. Thanks to its supe-rior generalizability, our proposed GeneralTrack achieves state-of-the-art performance on multiple benchmarks and demonstrates the potential for domain generalization. Le Wang 0003, Sanping Zhou, Panpan Fu, Gang Hua 0001, Wei Tang 0016 |
CVPR | 5 |
| 2024 | Stepwise Multi-grained Boundary Detector for Point-Supervised Temporal Action Localization
Mengnan Liu 0001, Le Wang 0003, Sanping Zhou, Qilin Zhang 0004, Gang Hua 0001 |
ECCV (7) | 7 |
| 2024 | UGG: Unified Generative Grasping
Jiaxin Lu 0001, Hao Kang, Bo Liu 0043, Yiding Yang, Qixing Huang, Gang Hua 0001 |
ECCV (67) | 7 |
| 2024 | Learning Anomalies with Normality Prior for Unsupervised Video Anomaly Detection
Haoyue Shi 0002, Le Wang 0003, Sanping Zhou, Gang Hua 0001, Wei Tang 0016 |
ECCV (6) | 4 |
| 2024 | Exploring Pre-trained Text-to-Video Diffusion Models for Referring Video Object Segmentation
Zixin Zhu, Xuelu Feng, Dongdong Chen 0001, Junsong Yuan 0001, Chunming Qiao, Gang Hua 0001 |
ECCV (12) | 6 |
| 2024 | Enhancing Implicit Shape Generators Using Topological RegularizationsabstractA fundamental problem in learning 3D shapes generative models is that when the generative model is simply fitted to the training data, the resulting synthetic 3D models can present various artifacts. Many of these artifacts are topological in nature, e.g., broken legs, unrealistic thin structures, and small holes. In this paper, we introduce a principled approach that utilizes topological regularization losses on an implicit shape generator to rectify topological artifacts. The objectives are two-fold. The first is to align the persistent diagram (PD) distribution of the training shapes with that of synthetic shapes. The second ensures that the PDs are smooth among adjacent synthetic shapes. We show how to achieve these two objectives using two simple but effective formulations. Specifically, distribution alignment is achieved to learn a generative model of PDs and align this generator with PDs of synthetic shapes. We show how to handle discrete and continuous variabilities of PDs by using a shape-regularization term when performing PD alignment. Moreover, we enforce the smoothness of the PDs using a smoothness loss on the PD generator, which further improves the behavior of PD distribution alignment. Experimental results on ShapeNet show that our approach leads to much better generalization behavior than state-of-the-art implicit shape generators. Yang Li 0104, Lohit Anirudh Jagarapu, Hao Kang, Gang Hua 0001, Qixing Huang |
ICML | 7 |
| 2024 | Multimodal LLM Enhanced Cross-lingual Cross-modal RetrievalabstractCross-lingual cross-modal retrieval (CCR) aims to retrieve visually relevant content based on non-English queries, without relying on human-labeled cross-modal data pairs during training. One popular approach involves utilizing machine translation (MT) to create pseudo-parallel data pairs, establishing correspondence between visual and non-English textual data. However, aligning their representations poses challenges due to the significant semantic gap between vision and text, as well as the lower quality of non-English representations caused by pre-trained encoders and data noise. To overcome these challenges, we propose LECCR, a novel solution that incorporates the multi-modal large language model (MLLM) to improve the alignment between visual and non-English representations. Specifically, we first employ MLLM to generate detailed visual content descriptions and aggregate them into multi-view semantic slots that encapsulate different semantics. Then, we take these semantic slots as internal features and leverage them to interact with the visual features. By doing so, we enhance the semantic information within the visual features, narrowing the semantic gap between modalities and generating local visual semantics for subsequent multi-level matching. Additionally, to further enhance the alignment between visual and non-English features, we introduce softened matching under English guidance. This approach provides more comprehensive and reliable inter-modal correspondences between visual and non-English features. Extensive experiments on four CCR benchmarks, i.e., Multi30K, MSCOCO, VATEX, and MSR-VTT-CN, demonstrate the effectiveness of our proposed method. Code: https://github.com/LiJiaBei-7/leccr. Le Wang 0003, Qiang Zhou 0001, Zhibin Wang 0004, Hao Li 0030, Gang Hua 0001, Wei Tang 0016 |
ACM Multimedia | 6 |
| 2024 | End-to-end pedestrian trajectory prediction via Efficient Multi-modal Predictors
Sanping Zhou, Le Wang 0003, Liushuai Shi, Yonghao Dong, Gang Hua 0001 |
Comput. Vis. Image Underst. | 6 |
| 2024 | Residual feature learning with hierarchical calibration for gaze estimation
Zhengdan Yin, Sanping Zhou, Le Wang 0003, Gang Hua 0001, Nanning Zheng 0001 |
Mach. Vis. Appl. | 5 |
| 2024 | Transformer Based Pluralistic Image Completion With Reduced Information LossabstractTransformer based methods have achieved great success in image inpainting recently. However, we find that these solutions regard each pixel as a token, thus suffering from an information loss issue from two aspects: 1) They downsample the input image into much lower resolutions for efficiency consideration. 2) They quantize 2563RGB values to a small number (such as 512) of quantized color values. The indices of quantized pixels are used as tokens for the inputs and prediction targets of the transformer. To mitigate these issues, we propose a new transformer based framework called “PUT”. Specifically, to avoid input downsampling while maintaining computation efficiency, we design a patch-based auto-encoder P-VQVAE. The encoder converts the masked image into non-overlapped patch tokens and the decoder recovers the masked regions from the inpainted tokens while keeping the unmasked regions unchanged. To eliminate the information loss caused by input quantization, an Un-quantized Transformer is applied. It directly takes features from the P-VQVAE encoder as input without any quantization and only regards the quantized tokens as prediction targets.Furthermore, to make the inpainting process more controllable, we introduce semantic and structural conditions as extra guidance. Extensive experiments show that our method greatly outperforms existing transformer based methods on image fidelity and achieves much higher diversity and better fidelity than state-of-the-art pluralistic inpainting methods on complex large-scale datasets (e.g., ImageNet). Codes are available athttps://github.com/liuqk3/PUT. Qiankun Liu 0001, Zhentao Tan, Dongdong Chen 0001, Ying Fu 0001, Qi Chu 0001, Gang Hua 0001, Nenghai Yu |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2024 | Towards Unified Robustness Against Both Backdoor and Adversarial AttacksabstractDeep Neural Networks (DNNs) are known to be vulnerable to both backdoor and adversarial attacks. In the literature, these two types of attacks are commonly treated as distinct robustness problems and solved separately, since they belong to training-time and inference-time attacks respectively. However, this paper revealed that there is an intriguing connection between them: (1) planting a backdoor into a model will significantly affect the model's adversarial examples and (2) for an infected model, its adversarial examples have similar features as the triggered images. Based on these observations, a novel Progressive Unified Defense (PUD) algorithm is proposed to defend against backdoor and adversarial attacks simultaneously. Specifically, our PUD has a progressive model purification scheme to jointly erase backdoors and enhance the model's adversarial robustness. At the early stage, the adversarial examples of infected models are utilized to erase backdoors. With the backdoor gradually erased, our model purification can naturally turn into a stage to boost the model's robustness against adversarial attacks. Besides, our PUD algorithm can effectively identify poisoned images, which allows the initial extra dataset not to be completely clean. Extensive experimental results show that, our discovered connection between backdoor and adversarial attacks is ubiquitous, no matter what type of backdoor attack. The proposed PUD outperforms the state-of-the-art backdoor defense, including the model repairing-based and data filtering-based methods. Besides, it also has the ability to compete with the most advanced adversarial defense methods. The code is available here. Zhenxing Niu, Yuyao Sun, Qiguang Miao, Rong Jin 0001, Gang Hua 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | Deep Image Matting With Sparse User InteractionsabstractImage matting is a fundamental and challenging problem in computer vision and graphics. Most existing matting methods leverage a user-supplied trimap as an auxiliary input to produce good alpha matte. However, obtaining high-quality trimap itself is arduous. Recently, some hint-free methods have emerged, however, the matting quality is still far behind the trimap-based methods. The main reason is that, some hints for removing semantic ambiguity and improving matting quality are essential. Apparently, there is a trade-off between interaction cost and matting quality. To balance performance and user-friendliness, we propose an improved deep image matting framework which is trimap-free and only needs sparse user click or scribble interaction to minimize the needed auxiliary constraints while still allowing interactivity. Moreover, we introduce uncertainty estimation that predicts which parts need polishing and conduct uncertainty-guided refinement. To trade off runtime against refinement quality, users can also choose different refinement modes. Experimental results show that our method performs better than existing trimap-free methods and comparably to state-of-the-art trimap-based methods with minimal user effort. Finally, we demonstrate the extensibility of our framework to video human matting without any structure modification, by adding optical flow-based sparse hint propagation and temporal consistency regularization imposed on the single frame. Tianyi Wei, Dongdong Chen 0001, Wenbo Zhou 0004, Jing Liao 0001, Weiming Zhang 0001, Gang Hua 0001, Nenghai Yu |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2024 | Robust Model Watermarking for Image Processing Networks via Structure ConsistencyabstractThe intellectual property of deep networks can be easily "stolen" by surrogate model attack. There has been significant progress in protecting the model IP in classification tasks. However, little attention has been devoted to the protection of image processing models. By utilizing consistent invisible spatial watermarks, the work (Zhang et al. 2020) first considered model watermarking for deep image processing networks and demonstrated its efficacy in many downstream tasks. Its success depends on the hypothesis that if a consistent watermark exists in all prediction outputs, that watermark will be learned into the attacker's surrogate model. However, when the attacker uses common data augmentation attacks (e.g., rotate, crop, and resize) during surrogate model training, it will fail because the underlying watermark consistency is destroyed. To mitigate this issue, we propose a new watermarking methodology, "structure consistency", based on which a new deep structure-aligned model watermarking algorithm is designed. Specifically, the embedded watermarks are designed to be aligned with physically consistent image structures, such as edges or semantic regions. Experiments demonstrate that our method is more robust than the baseline in resisting data augmentation attacks. Besides that, we test the generalization ability and robustness of our method to a broader range of adaptive attacks. Jie Zhang 0073, Dongdong Chen 0001, Jing Liao 0001, Zehua Ma, Han Fang 0004, Weiming Zhang 0001, Huamin Feng, Gang Hua 0001, Nenghai Yu |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2024 | Adversarial Attack and Defense in Deep RankingabstractDeep Neural Network classifiers are vulnerable to adversarial attacks, where an imperceptible perturbation could result in misclassification. However, the vulnerability of DNN-based image ranking systems remains under-explored. In this paper, we propose two attacks against deep ranking systems, i.e., Candidate Attack and Query Attack, that can raise or lower the rank of chosen candidates by adversarial perturbations. Specifically, the expected ranking order is first represented as a set of inequalities. Then a triplet-like objective function is designed to obtain the optimal perturbation. Conversely, an anti-collapse triplet defense is proposed to improve the ranking model robustness against all proposed attacks, where the model learns to prevent the adversarial attack from pulling the positive and negative samples close to each other. To comprehensively measure the empirical adversarial robustness of a ranking model with our defense, we propose an empirical robustness score, which involves a set of representative attacks against ranking models. Our adversarial ranking attacks and defenses are evaluated on MNIST, Fashion-MNIST, CUB200-2011, CARS196, and Stanford Online Products datasets. Experimental results demonstrate that our attacks can effectively compromise a typical deep ranking system. Nevertheless, our defense can significantly improve the ranking system's robustness and simultaneously mitigate a wide range of attacks. Le Wang 0003, Zhenxing Niu, Qilin Zhang 0004, Nanning Zheng 0001, Gang Hua 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2024 | Transfer easy to hard: Adversarial contrastive feature learning for unsupervised person re-identification
Haoxuanye Ji, Le Wang 0003, Sanping Zhou, Wei Tang 0016, Nanning Zheng 0001, Gang Hua 0001 |
Pattern Recognit. | 6 |
| 2024 | Toward High Quality Multi-Object Tracking and Segmentation Without Mask SupervisionabstractRecently studies have shown the potential of weakly supervised multi-object tracking and segmentation, but the drawbacks of coarse pseudo mask label and limited utilization of temporal information remain to be unresolved. To address these issues, we present a framework that directly uses box label to supervise the segmentation network without resorting to pseudo mask label. In addition, we propose to fully exploit the temporal information from two perspectives. Firstly, we integrate optical flow-based pairwise consistency to ensure mask consistency across frames, thereby improving mask quality for segmentation. Secondly, we propose a temporally adjacent pair-based sampling strategy to adapt instance embedding learning for data association in tracking. We combine these techniques into an end-to-end deep model, named BoxMOTS, which requires only box annotation without mask supervision. Extensive experiments demonstrate that our model surpasses current state-of-the-art by a large margin, and produces promising results on KITTI MOTS and BDD100K MOTS. The source code is available at https://github.com/Spritea/BoxMOTS. Wensheng Cheng, Zhenyu Wu 0002, Haibin Ling, Gang Hua 0001 |
IEEE Trans. Image Process. | 5 |
| 2024 | PointCAT: Contrastive Adversarial Training for Robust Point Cloud RecognitionabstractNotwithstanding the prominent performance shown in various applications, point cloud recognition models have often suffered from natural corruptions and adversarial perturbations. In this paper, we delve into boosting the general robustness of point cloud recognition, proposing Point-Cloud Contrastive Adversarial Training (PointCAT). The main intuition of PointCAT is encouraging the target recognition model to narrow the decision gap between clean point clouds and corrupted point clouds by devising feature-level constraints rather than logit-level constraints. Specifically, we leverage a supervised contrastive loss to facilitate the alignment and the uniformity of hypersphere representations, and design a pair of centralizing losses with dynamic prototype guidance to prevent features from deviating outside their belonging category clusters. To generate more challenging corrupted point clouds, we adversarially train a noise generator concurrently with the recognition model from the scratch. This differs from previous adversarial training methods that utilized gradient-based attacks as the inner loop. Comprehensive experiments show that the proposed PointCAT outperforms the baseline methods, significantly enhancing the robustness of diverse point cloud recognition models under various corruptions, including isotropic point noises, the LiDAR simulated noises, random point dropping, and adversarial perturbations. Our code is available at: https://github.com/shikiw/PointCAT. Qidong Huang, Xiaoyi Dong, Dongdong Chen 0001, Hang Zhou 0007, Weiming Zhang 0001, Gang Hua 0001, Yueqiang Cheng, Nenghai Yu |
IEEE Trans. Image Process. | 7 |
| 2024 | Disentangled Sample Guidance Learning for Unsupervised Person Re-IdentificationabstractUnsupervised person re-identification (Re-ID) is challenging due to the lack of ground truth labels. Most existing methods employ iterative clustering to generate pseudo labels for unlabeled training data to guide the learning process. However, how to select samples that are both associated with high-confidence pseudo labels and hard (discriminative) enough remains a critical problem. To address this issue, a disentangled sample guidance learning (DSGL) method is proposed for unsupervised Re-ID. The method consists of disentangled sample mining (DSM) and discriminative feature learning (DFL). DSM disentangles (unlabeled) person images into identity-relevant and identity-irrelevant factors, which are used to construct disentangled positive/negative groups that contain discriminative enough information. DFL incorporates the mined disentangled sample groups into model training by a surrogate disentangled learning loss and a disentangled second-order similarity regularization, to help the model better distinguish the characteristics of different persons. By using the DSGL training strategy, the mAP on Market-1501 and MSMT17 increases by 6.6% and 10.1% when applying the ResNet50 framework, and by 0.6% and 6.9% with the vision transformer (VIT) framework, respectively, validating the effectiveness of the DSGL method. Moreover, DSGL surpasses previous state-of-the-art methods by achieving higher Top-1 accuracy and mAP on the Market-1501, MSMT17, PersonX, and VeRi-776 datasets. The source code for this paper is available at https://github.com/jihaoxuanye/DiseSGL. Haoxuanye Ji, Le Wang 0003, Sanping Zhou, Wei Tang 0016, Gang Hua 0001 |
IEEE Trans. Image Process. | 5 |
| 2024 | Sparse Pedestrian Character Learning for Trajectory PredictionabstractPedestrian trajectory prediction in a first-person view has recently attracted much attention due to its importance in autonomous driving. Recent work utilizes pedestrian character information, i.e., action and appearance, to improve the learned trajectory embedding and achieves state-of-the-art performance. However, it neglects the invalid and negative pedestrian character information, which is harmful to trajectory representation and thus leads to performance degradation. To address this issue, we present a two-stream sparse-character-based network (TSNet) for pedestrian trajectory prediction. Specifically, TSNet learns the negative-removed characters in the sparse character representation stream to improve the trajectory embedding obtained in the trajectory representation stream. Moreover, to model the negative-removed characters, we propose a novel sparse character graph, including the sparse category and sparse temporal character graphs, to learn the different effects of various characters in category and temporal dimensions, respectively. Extensive experiments on two first-person view datasets, PIE and JAAD, show that our method outperforms existing state-of-the-art methods. In addition, ablation studies demonstrate different effects of various characters and prove that TSNet outperforms approaches without eliminating negative characters. Yonghao Dong, Le Wang 0003, Sanping Zhou, Gang Hua 0001, Changyin Sun 0001 |
IEEE Trans. Multim. | 4 |
| 2024 | Abnormal Ratios Guided Multi-Phase Self-Training for Weakly-Supervised Video Anomaly DetectionabstractWeakly-supervised Video Anomaly Detection (W-VAD) aims to detect abnormal events in videos given only video-level labels for training. Recent methods relying on multiple instance learning (MIL) and self-training achieve good performance, but they tend to focus on learning easy abnormal patterns while ignoring hard ones, e.g., unusual driving trajectory or over-speeding driving. How to detect hard anomalies is a critical but largely ignored problem in W-VAD. To tackle this challenge, we propose a novel framework, termed Abnormal Ratios guided Multi-phase Self-training (ARMS), for W-VAD. It includes a new abnormal ratio-based MIL (AR-MIL) loss and a new multi-phase self-training paradigm. The AR-MIL loss guides the learning of hard anomalies by enforcing a minimum ratio of abnormal snippets in an abnormal video and no abnormal snippets in a normal video. Our multi-phase self-training paradigm sequentially performs bootstrapping, hard anomalies mining, and adaptive self-training so as to address pseudo labeling on easy anomalies, detect hard anomalies, and setting adaptive abnormal ratios for different videos in a unified framework. Experimental results on three benchmark datasets, i.e., ShanghaiTech, UCF-Crime, and XD-Violence, show that ARMS outperforms all previous state-of-the-art methods and has a great advantage in detecting hard anomalies. Haoyue Shi 0002, Le Wang 0003, Sanping Zhou, Gang Hua 0001, Wei Tang 0016 |
IEEE Trans. Multim. | 4 |
| 2023 | Weakly-Guided Self-Supervised Pretraining for Temporal Activity DetectionabstractTemporal Activity Detection aims to predict activity classes per frame, in contrast to video-level predictions in Activity Classification (i.e., Activity Recognition). Due to the expensive frame-level annotations required for detection, the scale of detection datasets is limited. Thus, commonly, previous work on temporal activity detection resorts to fine-tuning a classification model pretrained on large-scale classification datasets (e.g., Kinetics-400). However, such pretrained models are not ideal for downstream detection, due to the disparity between the pretraining and the downstream fine-tuning tasks. In this work, we propose a novel weakly-guided self-supervised pretraining method for detection. We leverage weak labels (classification) to introduce a self-supervised pretext task (detection) by generating frame-level pseudo labels, multi-action frames, and action segments. Simply put, we design a detection task similar to downstream, on large-scale classification data, without extra annotations. We show that the models pretrained with the proposed weakly-guided self-supervised detection task outperform prior work on multiple challenging activity detection benchmarks, including Charades and MultiTHUMOS. Our extensive ablations further provide insights on when and how to use the proposed models for activity detection. Code is available at github.com/kkahatapitiya/SSDet. Kumara Kahatapitiya, Zhou Ren, Zhenyu Wu 0002, Michael S. Ryoo, Gang Hua 0001 |
AAAI | 6 |
| 2023 | Multi-Stream Representation Learning for Pedestrian Trajectory PredictionabstractForecasting the future trajectory of pedestrians is an important task in computer vision with a range of applications, from security cameras to autonomous driving. It is very challenging because pedestrians not only move individually across time but also interact spatially, and the spatial and temporal information is deeply coupled with one another in a multi-agent scenario. Learning such complex spatio-temporal correlation is a fundamental issue in pedestrian trajectory prediction. Inspired by the procedure that the hippocampus processes and integrates spatio-temporal information to form memories, we propose a novel multi-stream representation learning module to learn complex spatio-temporal features of pedestrian trajectory. Specifically, we learn temporal, spatial and cross spatio-temporal correlation features in three respective pathways and then adaptively integrate these features with learnable weights by a gated network. Besides, we leverage the sparse attention gate to select informative interactions and correlations brought by complex spatio-temporal modeling and reduce complexity of our model. We evaluate our proposed method on two commonly used datasets, i.e. ETH-UCY and SDD, and the experimental results demonstrate our method achieves the state-of-the-art performance. Code: https://github.com/YuxuanIAIR/MSRL-master Le Wang 0003, Sanping Zhou, Jinghai Duan, Gang Hua 0001, Wei Tang 0016 |
AAAI | 5 |
| 2023 | Boosted Dynamic Neural NetworksabstractEarly-exiting dynamic neural networks (EDNN), as one type of dynamic neural networks, has been widely studied recently. A typical EDNN has multiple prediction heads at different layers of the network backbone. During inference, the model will exit at either the last prediction head or an intermediate prediction head where the prediction confidence is higher than a predefined threshold. To optimize the model, these prediction heads together with the network backbone are trained on every batch of training data. This brings a train-test mismatch problem that all the prediction heads are optimized on all types of data in training phase while the deeper heads will only see difficult inputs in testing phase. Treating training and testing inputs differently at the two phases will cause the mismatch between training and testing data distributions. To mitigate this problem, we formulate an EDNN as an additive model inspired by gradient boosting, and propose multiple training techniques to optimize the model effectively. We name our method BoostNet. Our experiments show it achieves the state-of-the-art performance on CIFAR100 and ImageNet datasets in both anytime and budgeted-batch prediction modes. Our code is released at https://github.com/SHI-Labs/Boosted-Dynamic-Networks. Haichao Yu, Gang Hua 0001, Gao Huang 0001, Humphrey Shi |
AAAI | 3 |
| 2023 | Regularizing Second-Order Influences for Continual LearningabstractContinual learning aims to learn on non-stationary data streams without catastrophically forgetting previous knowledge. Prevalent replay-based methods address this challenge by rehearsing on a small buffer holding the seen data, for which a delicate sample selection strategy is required. However, existing selection schemes typically seek only to maximize the utility of the ongoing selection, overlooking the interference between successive rounds of selection. Motivated by this, we dissect the interaction of sequential selection steps within a framework built on influence functions. We manage to identify a new class of second-order influences that will gradually amplify incidental bias in the replay buffer and compromise the selection process. To regularize the second-order effects, a novel selection objective is proposed, which also has clear connections to two widely adopted criteria. Furthermore, we present an efficient implementation for optimizing the proposed criterion. Experiments on multiple continual learning benchmarks demonstrate the advantage of our approach over state-of-the-art methods. Code is available at https://github.com/feifeiobama/InfluenceCL. Zhicheng Sun 0001, Yadong Mu, Gang Hua 0001 |
CVPR | 3 |
| 2023 | Diversity-Aware Meta Visual PromptingabstractWe present Diversity-Aware Meta Visual Prompting (DAM-VP), an efficient and effective prompting method for transferring pre-trained models to downstream tasks with frozen backbone. A challenging issue in visual prompting is that image datasets sometimes have a large data diversity whereas a per-dataset generic prompt can hardly handle the complex distribution shift toward the original pretraining data distribution properly. To address this issue, we propose a dataset Diversity-Aware prompting strategy whose initialization is realized by a Meta-prompt. Specifically, we cluster the downstream dataset into small homogeneity subsets in a diversity-adaptive way, with each subset has its own prompt optimized separately. Such a divide-and-conquer design reduces the optimization difficulty greatly and significantly boosts the prompting performance. Furthermore, all the prompts are initialized with a meta-prompt, which is learned across several datasets. It is a bootstrapped paradigm, with the key observation that the prompting knowledge learned from previous datasets could help the prompt to converge faster and perform better on a new dataset. During inference, we dynamically select a proper prompt for each input, based on the feature distance between the input and each subset. Through extensive experiments, our DAM-VP demonstrates superior efficiency and effectiveness, clearly surpassing previous prompting methods in a series of downstream datasets for different pretraining models. Our code is available at: https://github.com/shikiw/DAM-VP. Qidong Huang, Xiaoyi Dong, Dongdong Chen 0001, Weiming Zhang 0001, Gang Hua 0001, Nenghai Yu |
CVPR | 6 |
| 2023 | Progressive Backdoor Erasing via connecting Backdoor and Adversarial AttacksabstractDeep neural networks (DNNs) are known to be vulnera-ble to both backdoor attacks as well as adversarial attacks. In the literature, these two types of attacks are commonly treated as distinct problems and solved separately, since they belong to training-time and inference-time attacks respectively. However, in this paper we find an intriguing connection between them: for a model planted with backdoors, we observe that its adversarial examples have similar behaviors as its triggered images, i.e., both activate the same subset of DNN neurons. It indicates that planting a back-door into a model will significantly affect the model's adversarial examples. Based on these observations, a novel Progressive Backdoor Erasing (PBE) algorithm is proposed to progressively purify the infected model by leveraging un-targeted adversarial attacks. Different from previous back-door defense methods, one significant advantage of our approach is that it can erase backdoor even when the clean extra dataset is unavailable. We empirically show that, against 5 state-of-the-art backdoor attacks, our PBE can effectively erase the backdoor without obvious performance degradation on clean samples and outperforms existing de-fense methods. Bingxu Mu, Zhenxing Niu, Le Wang 0003, Xue Wang 0010, Qiguang Miao, Rong Jin 0001, Gang Hua 0001 |
CVPR | 7 |
| 2023 | MotionTrack: Learning Robust Short-Term and Long-Term Motions for Multi-Object TrackingabstractThe main challenge of Multi-Object Tracking (MOT) lies in maintaining a continuous trajectory for each target. Existing methods often learn reliable motion patterns to match the same target between adjacent frames and discriminative appearance features to re-identify the lost targets after a long period. However, the reliability of motion prediction and the discriminability of appearances can be easily hurt by dense crowds and extreme occlusions in the tracking process. In this paper, we propose a simple yet effective multi-object tracker, i.e., MotionTrack, which learns robust short-term and long-term motions in a unified framework to associate trajectories from a short to long range. For dense crowds, we design a novel Interaction Module to learn interaction-aware motions from short-term trajectories, which can estimate the complex movement of each target. For extreme occlusions, we build a novel Refind Module to learn reliable long-term motions from the target's history trajectory, which can link the interrupted trajectory with its corresponding detection. Our Interaction Module and Refind Module are embedded in the well-known tracking-by-detection paradigm, which can work in tandem to maintain superior performance. Extensive experimental results on MOT17 and MOT20 datasets demonstrate the superiority of our approach in challenging scenarios, and it achieves state-of-the-art performances at various MOT metrics. Code is available at https://github.com/qwomeng/MotionTrack. Sanping Zhou, Le Wang 0003, Jinghai Duan, Gang Hua 0001, Wei Tang 0016 |
CVPR | 5 |
| 2023 | SOAR: Scene-debiasing Open-set Action RecognitionabstractDeep learning models have a risk of utilizing spurious clues to make predictions, such as recognizing actions based on the background scene. This issue can severely degrade the open-set action recognition performance when the testing samples have different scene distributions from the training samples. To mitigate this problem, we propose a novel method, called Scene-debiasing Open-set Action Recognition (SOAR), which features an adversarial scene reconstruction module and an adaptive adversarial scene classification module. The former prevents the decoder from reconstructing the video background given video features, and thus helps reduce the background information in feature learning. The latter aims to confuse scene type classification given video features, with a specific emphasis on the action foreground, and helps to learn scene-invariant information. In addition, we design an experiment to quantify the scene bias. The results indicate that the current open-set action recognizers are biased toward the scene, and our proposed SOAR method better mitigates such bias. Furthermore, our extensive experiments demonstrate that our method outperforms state-of-the-art methods, and the ablation studies confirm the effectiveness of our proposed modules. Yuanhao Zhai 0001, Ziyi Liu 0001, Zhenyu Wu 0002, Chunluan Zhou, David S. Doermann, Junsong Yuan 0001, Gang Hua 0001 |
ICCV | 8 |
| 2023 | Sparse Instance Conditioned Multimodal Trajectory PredictionabstractPedestrian trajectory prediction is critical in many vision tasks but challenging due to the multimodality of the future trajectory. Most existing methods predict multi-modal trajectories conditioned by goals (future endpoints) or instances (all future points). However, goal-conditioned methods ignore the intermediate process and instance-conditioned methods ignore the stochasticity of pedestrian motions. In this paper, we propose a simple yet effective Sparse Instance Conditioned Network (SICNet), which gives a balanced solution between goal-conditioned and instance-conditioned methods. Specifically, SICNet learns comprehensive sparse instances, i.e., representative points of the future trajectory, through a mask generated by a long short-term memory encoder and uses the memory mechanism to store and retrieve such sparse instances. Hence SICNet can decode the observed trajectory into the future prediction conditioned on the stored sparse instance. Moreover, we design a memory refinement module that refines the retrieved sparse instances from the memory to reduce memory recall errors. Extensive experiments on ETH-UCY and SDD datasets show that our method outperforms existing state-of-the-art methods. In addition, ablation studies demonstrate the superiority of our method compared with goal-conditioned and instance-conditioned approaches. Yonghao Dong, Le Wang 0003, Sanping Zhou, Gang Hua 0001 |
ICCV | 4 |
| 2023 | Flexible Visual Recognition by Evidential Modeling of Confusion and IgnoranceabstractIn real-world scenarios, typical visual recognition systems could fail under two major causes, i.e., the misclassification between known classes and the excusable misbehavior on unknown-class images. To tackle these deficiencies, flexible visual recognition should dynamically predict multiple classes when they are unconfident between choices and reject making predictions when the input is entirely out of the training distribution. Two challenges emerge along with this novel task. First, prediction uncertainty should be separately quantified as confusion depicting inter-class uncertainties and ignorance identifying out-of-distribution samples. Second, both confusion and ignorance should be comparable between samples to enable effective decision-making. In this paper, we propose to model these two sources of uncertainty explicitly with the theory of Subjective Logic. Regarding recognition as an evidence-collecting process, confusion is then defined as conflicting evidence, while ignorance is the absence of evidence. By predicting Dirichlet concentration parameters for singletons, comprehensive subjective opinions, including confusion and ignorance, could be achieved via further evidence combinations. Through a series of experiments on synthetic data analysis, visual recognition, and open-set detection, we demonstrate the effectiveness of our methods in quantifying two sources of uncertainties and dealing with flexible recognition. Lei Fan 0005, Bo Liu 0043, Ying Wu 0001, Gang Hua 0001 |
ICCV | 5 |
| 2023 | Improving Adversarial Robustness of Masked Autoencoders via Test-time Frequency-domain PromptingabstractIn this paper, we investigate the adversarial robustness of vision transformers that are equipped with BERT pretraining (e.g., BEiT, MAE). A surprising observation is that MAE has significantly worse adversarial robustness than other BERT pretraining methods. This observation drives us to rethink the basic differences between these BERT pretraining methods and how these differences affect the robustness against adversarial perturbations. Our empirical analysis reveals that the adversarial robustness of BERT pretraining is highly related to the reconstruction target, i.e., predicting the raw pixels of masked image patches will degrade more adversarial robustness of the model than predicting the semantic context, since it guides the model to concentrate more on medium-/high-frequency components of images. Based on our analysis, we provide a simple yet effective way to boost the adversarial robustness of MAE. The basic idea is using the dataset-extracted domain knowledge to occupy the medium-/high-frequency of images, thus narrowing the optimization space of adversarial perturbations. Specifically, we group the distribution of pretraining data and optimize a set of cluster-specific visual prompts on frequency domain. These prompts are incorporated with input images through prototype-based prompt selection during test period. Extensive evaluation shows that our method clearly boost MAE’s adversarial robustness while maintaining its clean performance on ImageNet-1k classification. Our code is available at: https://github.com/shikiw/RobustMAE. Qidong Huang, Xiaoyi Dong, Dongdong Chen 0001, Yinpeng Chen, Lu Yuan 0001, Gang Hua 0001, Weiming Zhang 0001, Nenghai Yu |
ICCV | 6 |
| 2023 | Parallel Attention Interaction Network for Few-Shot Skeleton-based Action RecognitionabstractLearning discriminative features from very few labeled samples to identify novel classes has received increasing attention in skeleton-based action recognition. Existing works aim to learn action-specific embeddings by exploiting either intra-skeleton or inter-skeleton spatial associations, which may lead to less discriminative representations. To address these issues, we propose a novel Parallel Attention Interaction Network (PAINet) that incorporates two complementary branches to strengthen the match by inter-skeleton and intraskeleton correlation. Specifically, a topology encoding module utilizing topology and physical information is proposed to enhance the modeling of interactive parts and joint pairs in both branches. In the Cross Spatial Alignment branch, we employ a spatial cross-attention module to establish joint associations across sequences, and a directional Average Symmetric Surface Metric is introduced to locate the closest temporal similarity. In parallel, the Cross Temporal Alignment branch incorporates a spatial self-attention module to aggregate spatial context within sequences as well as applies the temporal cross-attention network to correct misalignment temporally and calculate similarity. Extensive experiments on three skeleton benchmarks, namely NTU-T, NTU-S, and Kinetics, demonstrate the superiority of our framework and consistently outperform state-of-the-art methods. Sanping Zhou, Le Wang 0003, Gang Hua 0001 |
ICCV | 4 |
| 2023 | Trajectory Unified Transformer for Pedestrian Trajectory PredictionabstractPedestrian trajectory prediction is an essential link to understanding human behavior. Recent work achieves state-of-the-art performance gained from hand-designed post-processing, e.g., clustering. However, this post-processing suffers from expensive inference time and neglects the probability that the predicted trajectory disturbs downstream safety decisions. In this paper, we present Trajectory Unified TRansformer, called TUTR, which unifies the trajectory prediction components, social interaction, and multimodal trajectory prediction, into a transformer encoder-decoder architecture to effectively remove the need for post-processing. Specifically, TUTR parses the relationships across various motion modes using an explicit global prediction and an implicit mode-level transformer encoder. Then, TUTR attends to the social interactions with neighbors by a social-level transformer decoder. Finally, a dual prediction forecasts diverse trajectories and corresponding probabilities in parallel without post-processing. TUTR achieves state-of-the-art accuracy performance and improvements in inference speed of about 10× - 40× compared to previous well-tuned state-of-the-art methods using post-processing. Liushuai Shi, Le Wang 0003, Sanping Zhou, Gang Hua 0001 |
ICCV | 4 |
| 2023 | HairCLIPv2: Unifying Hair Editing via Proxy Feature BlendingabstractHair editing has made tremendous progress in recent years. Early hair editing methods use well-drawn sketches or masks to specify the editing conditions. Even though they can enable very fine-grained local control, such interaction modes are inefficient for the editing conditions that can be easily specified by language descriptions or reference images. Thanks to the recent breakthrough of cross-modal models (e.g., CLIP), HairCLIP is the first work that enables hair editing based on text descriptions or reference images. However, such text-driven and reference-driven interaction modes make HairCLIP unable to support fine-grained controls specified by sketch or mask. In this paper, we propose HairCLIPv2, aiming to support all the aforementioned interactions with one unified framework. Simultaneously, it improves upon HairCLIP with better irrelevant attributes (e.g., identity, background) preservation and unseen text descriptions support. The key idea is to convert all the hair editing tasks into hair transfer tasks, with editing conditions converted into different proxies accordingly. The editing effects are added upon the input image by blending the corresponding proxy features within the hairstyle or hair color feature spaces. Besides the unprecedented user interaction mode support, quantitative and qualitative experiments demonstrate the superiority of HairCLIPv2 in terms of editing effects, irrelevant attribute preservation and visual naturalness. Our code is available at https://github.com/wty-ustc/HairCLIPv2. Tianyi Wei, Dongdong Chen 0001, Wenbo Zhou 0004, Jing Liao 0001, Weiming Zhang 0001, Gang Hua 0001, Nenghai Yu |
ICCV | 6 |
| 2023 | Learning from Noisy Pseudo Labels for Semi-Supervised Temporal Action LocalizationabstractSemi-Supervised Temporal Action Localization (SS-TAL) aims to improve the generalization ability of action detectors with large-scale unlabeled videos. Albeit the recent advancement, one of the major challenges still remains: noisy pseudo labels hinder efficient learning on abundant unlabeled videos, embodied as location biases and category errors. In this paper, we dive deep into such an important but understudied dilemma. To this end, we propose a unified framework, termed Noisy Pseudo-Label Learning, to handle both location biases and category errors. Specifically, our method is featured with (1) Noisy Label Ranking to rank pseudo labels based on the semantic confidence and boundary reliability, (2) Noisy Label Filtering to address the class-imbalance problem of pseudo labels caused by category errors, (3) Noisy Label Learning to penalize in-consistent boundary predictions to achieve noise-tolerant learning for heavy location biases. As a result, our method could effectively handle the label noise problem and improve the utilization of a large amount of unlabeled videos. Extensive experiments on THUMOS14 and ActivityNet v1.3 demonstrate the effectiveness of our method. The code is available at github.com/kunnxia/NPL. Le Wang 0003, Sanping Zhou, Gang Hua 0001, Wei Tang 0016 |
ICCV | 4 |
| 2023 | Implicit Autoencoder for Point-Cloud Self-Supervised Representation LearningabstractThis paper advocates the use of implicit surface representation in autoencoder-based self-supervised 3D representation learning. The most popular and accessible 3D representation, i.e., point clouds, involves discrete samples of the underlying continuous 3D surface. This discretization process introduces sampling variations on the 3D shape, making it challenging to develop transferable knowledge of the true 3D geometry. In the standard autoencoding paradigm, the encoder is compelled to encode not only the 3D geometry but also information on the specific discrete sampling of the 3D shape into the latent code. This is because the point cloud reconstructed by the decoder is considered unacceptable unless there is a perfect mapping between the original and the reconstructed point clouds. This paper introduces the Implicit AutoEncoder (IAE), a simple yet effective method that addresses the sampling variation issue by replacing the commonly-used point-cloud decoder with an implicit decoder. The implicit decoder reconstructs a continuous representation of the 3D shape, independent of the imperfections in the discrete samples. Extensive experiments demonstrate that the proposed IAE achieves state-of-the-art performance across various self-supervised learning benchmarks. Our code is available at https://github.com/SimingYan/IAE. Siming Yan, Zhenpei Yang, Hao Kang, Gang Hua 0001, Qixing Huang |
ICCV | 7 |
| 2023 | Representing Multimodal Behaviors With Mean Location for Pedestrian Trajectory PredictionabstractRepresenting multimodal behaviors is a critical challenge for pedestrian trajectory prediction. Previous methods commonly represent this multimodality with multiple latent variables repeatedly sampled from a latent space, encountering difficulties in interpretable trajectory prediction. Moreover, the latent space is usually built by encoding global interaction into future trajectory, which inevitably introduces superfluous interactions and thus leads to performance reduction. To tackle these issues, we propose a novel Interpretable Multimodality Predictor (IMP) for pedestrian trajectory prediction, whose core is to represent a specific mode by its mean location. We model the distribution of mean location as a Gaussian Mixture Model (GMM) conditioned on sparse spatio-temporal features, and sample multiple mean locations from the decoupled components of GMM to encourage multimodality. Our IMP brings four-fold benefits: 1) Interpretable prediction to provide semantics about the motion behavior of a specific mode; 2) Friendly visualization to present multimodal behaviors; 3) Well theoretical feasibility to estimate the distribution of mean locations supported by the central-limit theorem; 4) Effective sparse spatio-temporal features to reduce superfluous interactions and model temporal continuity of interaction. Extensive experiments validate that our IMP not only outperforms state-of-the-art methods but also can achieve a controllable prediction by customizing the corresponding mean location. Liushuai Shi, Le Wang 0003, Chengjiang Long, Sanping Zhou, Wei Tang 0016, Nanning Zheng 0001, Gang Hua 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2023 | Semantic Probability Distribution Modeling for Diverse Semantic Image SynthesisabstractSemantic image synthesis, translating semantic layouts to photo-realistic images, is a one-to-many mapping problem. Though impressive progress has been recently made, diverse semantic synthesis that can efficiently produce semantic-level or even instance-level multimodal results, still remains a challenge. In this article, we propose a novel diverse semantic image synthesis framework from the perspective of semantic class distributions, which naturally supports diverse generation at both semantics and instance level. We achieve this by modeling class-level conditional modulation parameters as continuous probability distributions instead of discrete values, and sampling per-instance modulation parameters through instance-adaptive stochastic sampling that is consistent across the network. Moreover, we propose prior noise remapping, through linear perturbation parameters encoded from paired references, to facilitate supervised training and exemplar-based instance style control at test time. To further extend the user interaction function of the proposed method, we also introduce sketches into the network. In addition, specially designed generator modules, Progressive Growing Module and Multi-Scale Refinement Module, can be used as a general module to improve the performance of complex scene generation. Extensive experiments on multiple datasets show that our method can achieve superior diversity and comparable quality compared to state-of-the-art methods. Codes are available at https://github.com/tzt101/INADE.git. Zhentao Tan, Qi Chu 0001, Menglei Chai, Dongdong Chen 0001, Jing Liao 0001, Qiankun Liu 0001, Bin Liu 0016, Gang Hua 0001, Nenghai Yu |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2023 | Adaptive Two-Stream Consensus Network for Weakly-Supervised Temporal Action LocalizationabstractWeakly-supervised temporal action localization (W-TAL) aims to classify and localize all action instances in untrimmed videos under only video-level supervision. Without frame-level annotations, it is challenging for W-TAL methods to clearly distinguish actions and background, which severely degrades the action boundary localization and action proposal scoring. In this paper, we present an adaptive two-stream consensus network (A-TSCN) to address this problem. Our A-TSCN features an iterative refinement training scheme: a frame-level pseudo ground truth is generated and iteratively updated from a late-fusion activation sequence, and used to provide frame-level supervision for improved model training. Besides, we introduce an adaptive attention normalization loss, which adaptively selects action and background snippets according to video attention distribution. By differentiating the attention values of the selected action snippets and background snippets, it forces the predicted attention to act as a binary selection and promotes the precise localization of action boundaries. Furthermore, we propose a video-level and a snippet-level uncertainty estimator, and they can mitigate the adverse effect caused by learning from noisy pseudo ground truth. Experiments conducted on the THUMOS14, ActivityNet v1.2, ActivityNet v1.3, and HACS datasets show that our A-TSCN outperforms current state-of-the-art methods, and even achieves comparable performance with several fully-supervised methods. Yuanhao Zhai 0001, Le Wang 0003, Wei Tang 0016, Qilin Zhang 0004, Nanning Zheng 0001, David S. Doermann, Junsong Yuan 0001, Gang Hua 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2023 | ContextLoc++: A Unified Context Model for Temporal Action LocalizationabstractEffectively tackling the problem of temporal action localization (TAL) necessitates a visual representation that jointly pursues two confounding goals, i.e., fine-grained discrimination for temporal localization and sufficient visual invariance for action classification. We address this challenge by enriching the local, global and multi-scale contexts in the popular two-stage temporal localization framework. Our proposed model, dubbed ContextLoc++, can be divided into three sub-networks: L-Net, G-Net, and M-Net. L-Net enriches the local context via fine-grained modeling of snippet-level features, which is formulated as a query-and-retrieval process. Furthermore, the spatial and temporal snippet-level features, functioning as keys and values, are fused by temporal gating. G-Net enriches the global context via higher-level modeling of the video-level representation. In addition, we introduce a novel context adaptation module to adapt the global context to different proposals. M-Net further fuses the local and global contexts with multi-scale proposal features. Specially, proposal-level features from multi-scale video snippets can focus on different action characteristics. Short-term snippets with fewer frames pay attention to the action details while long-term snippets with more frames focus on the action variations. Experiments on the THUMOS14 and ActivityNet v1.3 datasets validate the efficacy of our method against existing state-of-the-art TAL algorithms. Zixin Zhu, Le Wang 0003, Wei Tang 0016, Nanning Zheng 0001, Gang Hua 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Memory-augmented appearance-motion network for video anomaly detection
Le Wang 0003, Junwen Tian, Sanping Zhou, Haoyue Shi 0002, Gang Hua 0001 |
Pattern Recognit. | 5 |
| 2023 | Tolerating Annotation Displacement in Dense Object Counting via Point Annotation Probability MapabstractCounting objects in crowded scenes remains a challenge to computer vision. The current deep learning based approach often formulate it as a Gaussian density regression problem. Such a brute-force regression, though effective, may not consider the annotation displacement properly which arises from the human annotation process and may lead to different distributions. We conjecture that it would be beneficial to consider the annotation displacement in the dense object counting task. To obtain strong robustness against annotation displacement, generalized Gaussian distribution (GGD) function with a tunable bandwidth and shape parameter is exploited to form the learning target point annotation probability map, PAPM. Specifically, we first present a hand-designed PAPM method (HD-PAPM), in which we design a function based on GGD to tolerate the annotation displacement. For end-to-end training, the hand-designed PAPM may not be optimal for the particular network and dataset. An adaptively learned PAPM method (AL-PAPM) is proposed. To improve the robustness to annotation displacement, we design an effective transport cost function based on GGD. The proposed PAPM is capable of integration with other methods. We also combine PAPM with P2PNet through modifying the matching cost matrix, forming P2P-PAPM. This could also improve the robustness to annotation displacement of P2PNet. Extensive experiments show the superiority of our proposed methods. Yuehai Chen, Jing Yang 0014, Badong Chen, Shaoyi Du, Gang Hua 0001 |
IEEE Trans. Image Process. | 5 |
| 2023 | Instance Motion Tendency Learning for Video Panoptic SegmentationabstractVideo panoptic segmentation is an important but challenging task in computer vision. It not only performs panoptic segmentation of each frame, but also associates the same instance across adjacent frames. Due to the lack of temporal coherence modeling, most existing approaches often generate identity switches during instance association, and they cannot handle ambiguous segmentation boundaries caused by motion blur. To address these difficult issues, we introduce a simple yet effective Instance Motion Tendency Network (IMTNet) for video panoptic segmentation. It learns a global motion tendency map for instance association, and a hierarchical classifier for motion boundary refinement. Specifically, a Global Motion Tendency Module (GMTM) is designed to learn robust motion features from optical flows, which can directly associate each instance in the previous frame to the corresponding instance in the current frame. In addition, we propose a Motion Boundary Refinement Module (MBRM) to learn a hierarchical classifier to handle the boundary pixels of moving targets, which can effectively revise the inaccurate segmentation predictions. Experimental results on both Cityscapes and Cityscapes-VPS datasets show that our IMTNet outperforms most state-of-the-art approaches. Le Wang 0003, Hongzhen Liu, Sanping Zhou, Wei Tang 0016, Gang Hua 0001 |
IEEE Trans. Image Process. | 5 |
| 2023 | Exploring Action Centers for Temporal Action LocalizationabstractTemporal action localization aims at detecting the temporal intervals of human actions in untrimmed videos. Most previous methods rely on locating and matching the start and end times of actions. However, action boundaries are ambiguous and uncertain in nature, which leads to inaccurate action localization and a lot of false positives. In this paper, we introduce a new framework for temporal action localization. It explicitly models temporal action centers to reduce unreliable action detection results caused by ambiguous action boundaries. Since action centers are highly related to semantic actions, they can be detected more reliably than the conventional action boundaries. As a result, our framework can exclude false positives and promote high-quality proposals. Based on action centers, we propose a triplet feature fusion mechanism. It performs neural message passing among the boundaries and the center as well as contextual regions outside of the proposal to enrich its representation. In addition, we introduce a centerness scoring method to suppress proposals deviating from the centers of action instances. Consequently, our network can retrieve high-quality action proposals and locate actions more precisely. Experimental results show our method outperforms state-of-the-art methods on the THUMOS14 and ActivityNet v1.3 datasets. Le Wang 0003, Sanping Zhou, Gang Hua 0001, Wei Tang 0016 |
IEEE Trans. Multim. | 5 |
| 2023 | Perceptual Hashing of Deep Convolutional Neural Networks for Model Copy DetectionabstractIn recent years, many model intellectual property (IP) proof methods for IP protection have been proposed, such as model watermarking and model fingerprinting. However, with the increasing number of models transmitted and deployed on the Internet, quickly finding the suspect model among thousands of models on model-sharing platforms such as GitHub is in great demand, which concurrently triggers the new security problem of model copy detection for IP protection. As an important part of the model IP protection system, the model copy detection task has not received enough attention. Due to the high computational complexity, both model watermarking and model fingerprinting lack the capability to efficiently find suspected infringing models among tens of millions of models. In this article, inspired by the hash-based image retrieval methods, we introduce a novel model copy detection mechanism: perceptual hashing for convolutional neural networks (CNNs). The proposed perceptual hashing algorithm can convert the weights of CNNs to fixed-length binary hash codes so that the lightly modified version has the similar hash code as the original model. By comparing the similarity of a pair of hash codes between a query model and a test model in the model library, similar versions of a query model can be retrieved efficiently. To the best of our knowledge, this is the first perceptual hashing algorithm for deep neural network models. Specifically, we first select the important model weights based on the model compression theory, then calculate the normal test statistics (NTS) on the segments of important weights, and finally encode the NTS features into hash codes. The experiment performed on a model library containing 3,565 models indicates that our perceptual hashing scheme has a superior copy detection performance. Hang Zhou 0007, Jie Zhang 0073, Dongdong Chen 0001, Weiming Zhang 0001, Kejiang Chen, Gang Hua 0001, Nenghai Yu |
ACM Trans. Multim. Comput. Commun. Appl. | 7 |
| 2022 | Complementary Attention Gated Network for Pedestrian Trajectory PredictionabstractPedestrian trajectory prediction is crucial in many practical applications due to the diversity of pedestrian movements, such as social interactions and individual motion behaviors. With similar observable trajectories and social environments, different pedestrians may make completely different future decisions. However, most existing methods only focus on the frequent modal of the trajectory and thus are difficult to generalize to the peculiar scenario, which leads to the decline of the multimodal fitting ability when facing similar scenarios. In this paper, we propose a complementary attention gated network (CAGN) for pedestrian trajectory prediction, in which a dual-path architecture including normal and inverse attention is proposed to capture both frequent and peculiar modals in spatial and temporal patterns, respectively. Specifically, a complementary block is proposed to guide normal and inverse attention, which are then be summed with learnable weights to get attention features by a gated network. Finally, multiple trajectory distributions are estimated based on the fused spatio-temporal attention features due to the multimodality of future trajectory. Experimental results on benchmark datasets, i.e., the ETH, and the UCY, demonstrate that our method outperforms state-of-the-art methods by 13.8% in Average Displacement Error (ADE) and 10.4% in Final Displacement Error (FDE). Code will be available at https://github.com/jinghaiD/CAGN Jinghai Duan, Le Wang 0003, Chengjiang Long, Sanping Zhou, Fang Zheng 0009, Liushuai Shi, Gang Hua 0001 |
AAAI | 7 |
| 2022 | Social Interpretable Tree for Pedestrian Trajectory PredictionabstractUnderstanding the multiple socially-acceptable future behaviors is an essential task for many vision applications. In this paper, we propose a tree-based method, termed as Social Interpretable Tree (SIT), to address this multi-modal prediction task, where a hand-crafted tree is built depending on the prior information of observed trajectory to model multiple future trajectories. Specifically, a path in the tree from the root to leaf represents an individual possible future trajectory. SIT employs a coarse-to-fine optimization strategy, in which the tree is first built by high-order velocity to balance the complexity and coverage of the tree and then optimized greedily to encourage multimodality. Finally, a teacher-forcing refining operation is used to predict the final fine trajectory. Compared with prior methods which leverage implicit latent variables to represent possible future trajectories, the path in the tree can explicitly explain the rough moving behaviors (e.g., go straight and then turn right), and thus provides better interpretability. Despite the hand-crafted tree, the experimental results on ETH-UCY and Stanford Drone datasets demonstrate that our method is capable of matching or exceeding the performance of state-of-the-art methods. Interestingly, the experiments show that the raw built tree without training outperforms many prior deep neural network based approaches. Meanwhile, our method presents sufficient flexibility in long-term prediction and different best-of-K predictions. Liushuai Shi, Le Wang 0003, Chengjiang Long, Sanping Zhou, Fang Zheng 0009, Nanning Zheng 0001, Gang Hua 0001 |
AAAI | 7 |
| 2022 | Learning Disentangled Classification and Localization Representations for Temporal Action LocalizationabstractA common approach to Temporal Action Localization (TAL) is to generate action proposals and then perform action classification and localization on them. For each proposal, existing methods universally use a shared proposal-level representation for both tasks. However, our analysis indicates that this shared representation focuses on the most discriminative frames for classification, e.g., ``take-offs" rather than ``run-ups" in distinguishing ``high jump" and ``long jump", while frames most relevant to localization, such as the start and end frames of an action, are largely ignored. In other words, such a shared representation can not simultaneously handle both classification and localization tasks well, and it makes precise TAL difficult. To address this challenge, this paper disentangles the shared representation into classification and localization representations. The disentangled classification representation focuses on the most discriminative frames, and the disentangled localization representation focuses on the action phase as well as the action start and end. Our model could be divided into two sub-networks, i.e., the disentanglement network and the context-based aggregation network. The disentanglement network is an autoencoder to learn orthogonal hidden variables of classification and localization. The context-based aggregation network aggregates the classification and localization representations by modeling local and global contexts. We evaluate our proposed method on two popular benchmarks for TAL, which outperforms all state-of-the-art methods. Zixin Zhu, Le Wang 0003, Wei Tang 0016, Ziyi Liu 0001, Nanning Zheng 0001, Gang Hua 0001 |
AAAI | 6 |
| 2022 | Uncertainty-Based Spatial-Temporal Attention for Online Action Detection
Hongji Guo, Zhou Ren, Gang Hua 0001 |
ECCV (4) | 4 |
| 2022 | Breadcrumbs: Adversarial Class-Balanced Sampling for Long-Tailed Recognition
Bo Liu 0043, Hao Kang, Gang Hua 0001, Nuno Vasconcelos |
ECCV (24) | 4 |
| 2022 | TxVAD: Improved Video Action Detection by TransformersabstractVideo action detection aims to localize persons in both space and time from video sequences and recognize their actions. Most existing methods are composed of many specialized components, e.g., pretrained person/object detectors, region proposal networks (RPN), memory banks, and so on. This paper proposes a conceptually simple paradigm for video action detection using Transformers, which effectively removes the need for specialized components and achieves superior performance. Our proposed Transformer-based Video Action Detector (TxVAD) utilizes two Transformers to capture scene context information and long-range spatio-temporal context information, for person localization and action classification, respectively. Through extensive experiments on four public datasets, AVA, AVA-Kinetics, JHMDB-21, and UCF101-24, we show that our conceptually simple paradigm has achieved state-of-the-art performance for video action detection task, without using pre-trained person/object detectors, RPN, or memory bank. Zhenyu Wu 0002, Zhou Ren, Zhangyang Wang, Gang Hua 0001 |
ACM Multimedia | 5 |
| 2022 | Weakly Supervised Temporal Action Localization Through Contrast Based Evaluation NetworksabstractGiven only video-level action categorical labels during training, weakly-supervised temporal action localization (WS-TAL) learns to detect action instances and locates their temporal boundaries in untrimmed videos. Compared to its fully supervised counterpart, WS-TAL is more cost-effective in data labeling and thus favorable in practical applications. However, the coarse video-level supervision inevitably incurs ambiguities in action localization, especially in untrimmed videos containing multiple action instances. To overcome this challenge, we observe that significant temporal contrasts among video snippets, e.g., caused by temporal discontinuities and sudden changes, often occur around true action boundaries. This motivates us to introduce a Contrast-based Localization EvaluAtioN Network (CleanNet), whose core is a new temporal action proposal evaluator, which provides fine-grained pseudo supervision by leveraging the temporal contrasts among snippet-level classification predictions. As a result, the uncertainty in locating action instances can be resolved via evaluating their temporal contrast scores. Moreover, the new action localization module is an integral part of CleanNet which enables end-to-end training. This is in contrast to many existing WS-TAL methods where action localization is merely a post-processing step. Besides, we also explore the usage of temporal contrast on temporal action proposal (TAP) generation task, which we believe is the first attempt with the weak supervision setting. Experiments on the THUMOS14, ActivityNet v1.2 and v1.3 datasets validate the efficacy of our method against existing state-of-the-art WS-TAL algorithms. Ziyi Liu 0001, Le Wang 0003, Qilin Zhang 0004, Wei Tang 0016, Nanning Zheng 0001, Gang Hua 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2022 | Efficient Semantic Image Synthesis via Class-Adaptive NormalizationabstractSpatially-adaptive normalization (SPADE) is remarkably successful recently in conditional semantic image synthesis in T. Park et al. 2019 which modulates the normalized activation with spatially-varying transformations learned from semantic layouts, to prevent the semantic information from being washed away. Despite its impressive performance, a more thorough understanding of the advantages inside the box is still highly demanded to help reduce the significant computation and parameter overhead introduced by this novel structure. In this paper, from a return-on-investment point of view, we conduct an in-depth analysis of the effectiveness of this spatially-adaptive normalization and observe that its modulation parameters benefit more from semantic-awareness rather than spatial-adaptiveness, especially for high-resolution input masks. Inspired by this observation, we propose class-adaptive normalization (CLADE), a lightweight but equally-effective variant that is only adaptive to semantic class. In order to further improve spatial-adaptiveness, we introduce intra-class positional map encoding calculated from semantic layouts to modulate the normalization parameters of CLADE and propose a truly spatially-adaptive variant of CLADE, namely CLADE-ICPE. Through extensive experiments on multiple challenging datasets, we demonstrate that the proposed CLADE can be generalized to different SPADE-based methods while achieving comparable generation quality compared to SPADE, but it is much more efficient with fewer extra parameters and lower computational cost. The code and pretrained models are available at https://github.com/tzt101/CLADE.git. Zhentao Tan, Dongdong Chen 0001, Qi Chu 0001, Menglei Chai, Jing Liao 0001, Mingming He, Lu Yuan 0001, Gang Hua 0001, Nenghai Yu |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2022 | Deep Model Intellectual Property Protection via Deep WatermarkingabstractDespite the tremendous success, deep neural networks are exposed to serious IP infringement risks. Given a target deep model, if the attacker knows its full information, it can be easily stolen by fine-tuning. Even if only its output is accessible, a surrogate model can be trained through student-teacher learning by generating many input-output training pairs. Therefore, deep model IP protection is important and necessary. However, it is still seriously under-researched. In this work, we propose a new model watermarking framework for protecting deep networks trained for low-level computer vision or image processing tasks. Specifically, a special task-agnostic barrier is added after the target model, which embeds a unified and invisible watermark into its outputs. When the attacker trains one surrogate model by using the input-output pairs of the barrier target model, the hidden watermark will be learned and extracted afterwards. To enable watermarks from binary bits to high-resolution images, a deep invisible watermarking mechanism is designed. By jointly training the target model and watermark embedding, the extra barrier can even be absorbed into the target model. Through extensive experiments, we demonstrate the robustness of the proposed framework, which can resist attacks with different network structures and objective functions. Jie Zhang 0073, Dongdong Chen 0001, Jing Liao 0001, Weiming Zhang 0001, Huamin Feng, Gang Hua 0001, Nenghai Yu |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2022 | Loss functions for pose guided person image generation
Haoyue Shi 0002, Le Wang 0003, Nanning Zheng 0001, Gang Hua 0001, Wei Tang 0016 |
Pattern Recognit. | 4 |
| 2022 | Dual relation network for temporal action localization
Le Wang 0003, Sanping Zhou, Gang Hua 0001, Wei Tang 0016 |
Pattern Recognit. | 4 |
| 2022 | Local to Global Feature Learning for Salient Object Detection
Xuelu Feng, Sanping Zhou, Zixin Zhu, Le Wang 0003, Gang Hua 0001 |
Pattern Recognit. Lett. | 5 |
| 2022 | E2Style: Improve the Efficiency and Effectiveness of StyleGAN InversionabstractThis paper studies the problem of StyleGAN inversion, which plays an essential role in enabling the pretrained StyleGAN to be used for real image editing tasks. The goal of StyleGAN inversion is to find the exact latent code of the given image in the latent space of StyleGAN. This problem has a high demand for quality and efficiency. Existing optimization-based methods can produce high-quality results, but the optimization often takes a long time. On the contrary, forward-based methods are usually faster but the quality of their results is inferior. In this paper, we present a new feed-forward network "E2Style" for StyleGAN inversion, with significant improvement in terms of efficiency and effectiveness. In our inversion network, we introduce: 1) a shallower backbone with multiple efficient heads across scales; 2) multi-layer identity loss and multi-layer face parsing loss to the loss function; and 3) multi-stage refinement. Combining these designs together forms an effective and efficient method that exploits all benefits of optimization-based and forward-based methods. Quantitative and qualitative results show that our E2Style performs better than existing forward-based methods and comparably to state-of-the-art optimization-based methods while maintaining the high efficiency as well as forward-based methods. Moreover, a number of real image editing applications demonstrate the efficacy of our E2Style. Our code is available at https://github.com/wty-ustc/e2style. Tianyi Wei, Dongdong Chen 0001, Wenbo Zhou 0004, Jing Liao 0001, Weiming Zhang 0001, Lu Yuan 0001, Gang Hua 0001, Nenghai Yu |
IEEE Trans. Image Process. | 7 |
| 2022 | Poison Ink: Robust and Invisible Backdoor AttackabstractRecent research shows deep neural networks are vulnerable to different types of attacks, such as adversarial attacks, data poisoning attacks, and backdoor attacks. Among them, backdoor attacks are the most cunning and can occur in almost every stage of the deep learning pipeline. Backdoor attacks have attracted lots of interest from both academia and industry. However, most existing backdoor attack methods are visible or fragile to some effortless pre-processing such as common data transformations. To address these limitations, we propose a robust and invisible backdoor attack called "Poison Ink". Concretely, we first leverage the image structures as target poisoning areas and fill them with poison ink (information) to generate the trigger pattern. As the image structure can keep its semantic meaning during the data transformation, such a trigger pattern is inherently robust to data transformations. Then we leverage a deep injection network to embed such input-aware trigger pattern into the cover image to achieve stealthiness. Compared to existing popular backdoor attack methods, Poison Ink outperforms both in stealthiness and robustness. Through extensive experiments, we demonstrate that Poison Ink is not only general to different datasets and network architectures but also flexible for different attack scenarios. Besides, it also has very strong resistance against many state-of-the-art defense techniques. Jie Zhang 0073, Dongdong Chen 0001, Qidong Huang, Jing Liao 0001, Weiming Zhang 0001, Huamin Feng, Gang Hua 0001, Nenghai Yu |
IEEE Trans. Image Process. | 7 |
| 2022 | Action Coherence Network for Weakly-Supervised Temporal Action LocalizationabstractWeakly-supervised Temporal Action Localization (W-TAL) aims at simultaneously classifying and locating all action instances with only video-level supervision. However, current W-TAL methods have two limitations. First, they ignore the difference in video representations between an action instance and its surrounding background when generating and scoring action proposals. Second, the unique characteristics of the RGB frames and optical flow are largely ignored when fusing these two modalities. To address these problems, an Action Coherence Network (ACN) is proposed in this paper. Its core is a new coherence loss which exploits both classification predictions and video content representations to supervise action boundary regression and thus leads to more accurate action localization results. Besides, the proposed ACN explicitly takes into account the specific characteristics of RGB frames and optical flow by training two separate sub-networks, each of which is able to generate modality-specific action proposals independently. Finally, to take advantage of the complementary action proposals generated by two streams, a novel fusion module is introduced to reconcile them and obtain the final action localization results. Experiments on the THUMOS14 and ActivityNet datasets show that our ACN outperforms the state-of-the-art W-TAL methods, and is even comparable to some recent fully-supervised methods. Particularly, ACN achieves a mean average precision of 26.4% on the THUMOS14 dataset under the IoU threshold 0.5. Yuanhao Zhai 0001, Le Wang 0003, Wei Tang 0016, Qilin Zhang 0004, Nanning Zheng 0001, Gang Hua 0001 |
IEEE Trans. Multim. | 6 |
| 2021 | Weakly Supervised Temporal Action Localization Through Learning Explicit Subspaces for Action and ContextabstractWeakly-supervised Temporal Action Localization (WS-TAL) methods learn to localize temporal starts and ends of action instances in a video under only video-level supervision. Existing WS-TAL methods rely on deep features learned for action recognition. However, due to the mismatch between classification and localization, these features cannot distinguish the frequently co-occurring contextual background, i.e., the context, and the actual action instances. We term this challenge action-context confusion, and it will adversely affect the action localization accuracy. To address this challenge, we introduce a framework that learns two feature subspaces respectively for actions and their context. By explicitly accounting for action visual elements, the action instances can be localized more precisely without the distraction from the context. To facilitate the learning of these two feature subspaces with only video-level categorical labels, we leverage the predictions from both spatial and temporal streams for snippets grouping. In addition, an unsupervised learning task is introduced to make the proposed module focus on mining temporal information. The proposed approach outperforms state-of-the-art WS-TAL methods on three benchmarks, i.e., THUMOS14, ActivityNet v1.2 and v1.3 datasets. Ziyi Liu 0001, Le Wang 0003, Wei Tang 0016, Junsong Yuan 0001, Nanning Zheng 0001, Gang Hua 0001 |
AAAI | 6 |
| 2021 | ACSNet: Action-Context Separation Network for Weakly Supervised Temporal Action LocalizationabstractThe object of Weakly-supervised Temporal Action Localization (WS-TAL) is to localize all action instances in an untrimmed video with only video-level supervision. Due to the lack of frame-level annotations during training, current WS-TAL methods rely on attention mechanisms to localize the foreground snippets or frames that contribute to the video-level classification task. This strategy frequently confuse context with the actual action, in the localization result. Separating action and context is a core problem for precise WS-TAL, but it is very challenging and has been largely ignored in the literature. In this paper, we introduce an Action-Context Separation Network (ACSNet) that explicitly takes into account context for accurate action localization. It consists of two branches (i.e., the Foreground-Background branch and the Action-Context branch). The Foreground-Background branch first distinguishes foreground from background within the entire video while the Action-Context branch further separates the foreground as action and context. We associate video snippets with two latent components (i.e., a positive component and a negative component), and their different combinations can effectively characterize foreground, action and context. Furthermore, we introduce extended labels with auxiliary context categories to facilitate the learning of action-context separation. Experiments on THUMOS14 and ActivityNet v1.2/v1.3 datasets demonstrate the ACSNet outperforms existing state-of-the-art WS-TAL methods by a large margin. Ziyi Liu 0001, Le Wang 0003, Qilin Zhang 0004, Wei Tang 0016, Junsong Yuan 0001, Nanning Zheng 0001, Gang Hua 0001 |
AAAI | 7 |
| 2021 | Any-Precision Deep Neural NetworksabstractWe present any-precision deep neural networks (DNNs), which are trained with a new method that allows the learned DNNs to be flexible in numerical precision during inference. The same model in runtime can be flexibly and directly set to different bit-widths, by truncating the least significant bits, to support dynamic speed and accuracy trade-off. When all layers are set to low-bits, we show that the model achieved accuracy comparable to dedicated models trained at the same precision. This nice property facilitates flexible deployment of deep learning models in real-world applications, where in practice trade-offs between model accuracy and runtime efficiency are often sought. Previous literature presents solutions to train models at each individual fixed efficiency/accuracy trade-off point. But how to produce a model flexible in runtime precision is largely unexplored. When the demand of efficiency/accuracy trade-off varies from time to time or even dynamically changes in runtime, it is infeasible to re-train models accordingly, and the storage budget may forbid keeping multiple models. Our proposed framework achieves this flexibility without performance degradation. More importantly, we demonstrate that this achievement is agnostic to model architectures and applicable to multiple vision tasks. Our code is released at https://github.com/SHI-Labs/Any-Precision-DNNs. Haichao Yu, Humphrey Shi, Thomas S. Huang, Gang Hua 0001 |
AAAI | 5 |
| 2021 | SGCN: Sparse Graph Convolution Network for Pedestrian Trajectory PredictionabstractPedestrian trajectory prediction is a key technology in autopilot, which remains to be very challenging due to complex interactions between pedestrians. However, previous works based on dense undirected interaction suffer from modeling superfluous interactions and neglect of trajectory motion tendency, and thus inevitably result in a considerable deviance from the reality. To cope with these issues, we present a Sparse Graph Convolution Network (SGCN) for pedestrian trajectory prediction. Specifically, the SGCN explicitly models the sparse directed interaction with a sparse directed spatial graph to capture adaptive interaction pedestrians. Meanwhile, we use a sparse directed temporal graph to model the motion tendency, thus to facilitate the prediction based on the observed direction. Finally, parameters of a bi-Gaussian distribution for trajectory prediction are estimated by fusing the above two sparse graphs. We evaluate our proposed method on the ETH and UCY datasets, and the experimental results show our method outperforms comparative state-of-the-art methods by 9% in Average Displacement Error (ADE) and 13% in Final Displacement Error (FDE). Notably, visualizations indicate that our method can capture adaptive interactions between pedestrians and their effective motion tendencies. Liushuai Shi, Le Wang 0003, Chengjiang Long, Sanping Zhou, Zhenxing Niu, Gang Hua 0001 |
CVPR | 7 |
| 2021 | Learning View Selection for 3D ScenesabstractEfficient 3D space sampling to represent an underlying 3D object/scene is essential for 3D vision, robotics, and beyond. A standard approach is to explicitly sample a dense collection of views and formulate it as a view selection problem, or, more generally, a set cover problem. In this paper, we introduce a novel approach that avoids dense view sampling. The key idea is to learn a view prediction network and a trainable aggregation module that takes the predicted views as input and outputs an approximation of their generic scores (e.g., surface coverage, viewing angle from surface normals). This methodology allows us to turn the set cover problem (or multi-view representation optimization) into a continuous optimization problem. We then explain how to effectively solve the induced optimization problem using continuation, i.e., aggregating a hierarchy of smoothed scoring modules. Experimental results show that our approach arrives at similar or better solutions with about 10 x speed up in running time, comparing with the standard methods. Yifan Sun 0007, Qixing Huang, Dun-Yu Hsiao, Gang Hua 0001 |
CVPR | 5 |
| 2021 | Diverse Semantic Image Synthesis via Probability Distribution ModelingabstractSemantic image synthesis, translating semantic layouts to photo-realistic images, is a one-to-many mapping problem. Though impressive progress has been recently made, diverse semantic synthesis that can efficiently produce semantic-level multimodal results, still remains a challenge. In this paper, we propose a novel diverse semantic image synthesis framework from the perspective of semantic class distributions, which naturally supports diverse generation at semantic or even instance level. We achieve this by modeling class-level conditional modulation parameters as continuous probability distributions instead of discrete values, and sampling per-instance modulation parameters through instance-adaptive stochastic sampling that is consistent across the network. Moreover, we propose prior noise remapping, through linear perturbation parameters encoded from paired references, to facilitate supervised training and exemplar-based instance style control at test time. Extensive experiments on multiple datasets show that our method can achieve superior diversity and comparable quality compared to state-of-the-art methods. Code will be available at https://github.com/tzt101/INADE.git Zhentao Tan, Menglei Chai, Dongdong Chen 0001, Jing Liao 0001, Qi Chu 0001, Bin Liu 0016, Gang Hua 0001, Nenghai Yu |
CVPR | 7 |
| 2021 | Learning Dynamics via Graph Neural Networks for Human Pose Estimation and TrackingabstractMulti-person pose estimation and tracking serve as crucial steps for video understanding. Most state-of-the-art approaches rely on first estimating poses in each frame and only then implementing data association and refinement. Despite the promising results achieved, such a strategy is inevitably prone to missed detections especially in heavily-cluttered scenes, since this tracking-by-detection paradigm is, by nature, largely dependent on visual evidences that are absent in the case of occlusion. In this paper, we propose a novel online approach to learning the pose dynamics, which are independent of pose detections in current fame, and hence may serve as a robust estimation even in challenging scenarios including occlusion. Specifically, we derive this prediction of dynamics through a graph neural network (GNN) that explicitly accounts for both spatial-temporal and visual information. It takes as input the historical pose tracklets and directly predicts the corresponding poses in the following frame for each tracklet. The predicted poses will then be aggregated with the detected poses, if any, at the same frame so as to produce the final pose, potentially recovering the occluded joints missed by the estimator. Experiments on PoseTrack 2017 and Pose-Track 2018 datasets demonstrate that the proposed method achieves results superior to the state of the art on both human pose estimation and tracking tasks. Yiding Yang, Zhou Ren, Chunluan Zhou, Xinchao Wang, Gang Hua 0001 |
CVPR | 6 |
| 2021 | Meta Pairwise Relationship Distillation for Unsupervised Person Re-identificationabstractUnsupervised person re-identification (Re-ID) remains challenging due to the lack of ground-truth labels. Existing methods often rely on estimated pseudo labels via iterative clustering and classification, and they are unfortunately highly susceptible to performance penalties incurred by the inaccurate estimated number of clusters. Alternatively, we propose the Meta Pairwise Relationship Distillation (MPRD) method to estimate the pseudo labels of sample pairs for unsupervised person Re-ID. Specifically, it consists of a Convolutional Neural Network (CNN) and Graph Convolutional Network (GCN), in which the GCN estimates the pseudo labels of sample pairs based on the current features extracted by CNN, and the CNN learns better features by involving high-fidelity positive and negative sample pairs imposed by GCN. To achieve this goal, a small amount of labeled samples are used to guide GCN training, which can distill meta knowledge to judge the difference in the neighborhood structure between positive and negative sample pairs. Extensive experiments on Market-1501, DukeMTMC-reID and MSMT17 datasets show that our method outperforms the state-of-the-art approaches. Haoxuanye Ji, Le Wang 0003, Sanping Zhou, Wei Tang 0016, Nanning Zheng 0001, Gang Hua 0001 |
ICCV | 6 |
| 2021 | GistNet: a Geometric Structure Transfer Network for Long-Tailed RecognitionabstractThe problem of long-tailed recognition, where the number of examples per class is highly unbalanced, is considered. It is hypothesized that the well known tendency of standard classifier training to overfit to popular classes can be exploited for effective transfer learning. Rather than eliminating this overfitting, e.g. by adopting popular class-balanced sampling methods, the learning algorithm should instead leverage this overfitting to transfer geometric information from popular to low-shot classes. A new classifier architecture, GistNet, is proposed to support this goal, using constellations of classifier parameters to encode the class geometry. A new learning algorithm is then proposed for GeometrIc Structure Transfer (GIST), with resort to a combination of loss functions that combine class-balanced and random sampling to guarantee that, while overfitting to the popular classes is restricted to geometric parameters, it is leveraged to transfer class geometry from popular to few-shot classes. This enables better generalization for few-shot classes without the need for the manual specification of class weights, or even the explicit grouping of classes into different types. Experiments on two popular long-tailed recognition datasets show that GistNet outperforms existing solutions to this problem. Bo Liu 0043, Hao Kang, Gang Hua 0001, Nuno Vasconcelos |
ICCV | 4 |
| 2021 | Unlimited Neighborhood Interaction for Heterogeneous Trajectory PredictionabstractUnderstanding complex social interactions among agents is a key challenge for trajectory prediction. Most existing methods consider the interactions between pairwise traffic agents or in a local area, while the nature of interactions is unlimited, involving an uncertain number of agents and non-local areas simultaneously. Besides, they treat heterogeneous traffic agents the same, namely those among agents of different categories, while neglecting people’s diverse reaction patterns toward traffic agents in different categories. To address these problems, we propose a simple yet effective Unlimited Neighborhood Interaction Network (UNIN), which predicts trajectories of heterogeneous agents in multiple categories. Specifically, the proposed unlimited neighborhood interaction module generates the fused-features of all agents involved in an interaction simultaneously, which is adaptive to any number of agents and any range of interaction area. Meanwhile, a hierarchical graph attention module is proposed to obtain category-to-category interaction and agent-to-agent interaction. Finally, parameters of a Gaussian Mixture Model are estimated for generating the future trajectories. Extensive experimental results on benchmark datasets demonstrate a significant performance improvement of our method over the state-of-the-art methods. Fang Zheng 0009, Le Wang 0003, Sanping Zhou, Wei Tang 0016, Zhenxing Niu, Nanning Zheng 0001, Gang Hua 0001 |
ICCV | 7 |
| 2021 | Practical Relative Order Attack in Deep RankingabstractRecent studies unveil the vulnerabilities of deep ranking models, where an imperceptible perturbation can trigger dramatic changes in the ranking result. While previous attempts focus on manipulating absolute ranks of certain candidates, the possibility of adjusting their relative order remains under-explored. In this paper, we formulate a new adversarial attack against deep ranking systems, i.e., the Order Attack, which covertly alters the relative order among a selected set of candidates according to an attacker-specified permutation, with limited interference to other unrelated candidates. Specifically, it is formulated as a triplet-style loss imposing an inequality chain reflecting the specified permutation. However, direct optimization of such white-box objective is infeasible in a real-world attack scenario due to various black-box limitations. To cope with them, we propose a Short-range Ranking Correlation metric as a surrogate objective for black-box Order Attack to approximate the white-box method. The Order Attack is evaluated on the Fashion-MNIST and Stanford-Online-Products datasets under both white-box and black-box threat models. The black-box attack is also successfully implemented on a major e-commerce platform. Comprehensive experimental evaluations demonstrate the effectiveness of the proposed methods, revealing a new type of ranking model vulnerability. Le Wang 0003, Zhenxing Niu, Qilin Zhang 0004, Nanning Zheng 0001, Gang Hua 0001 |
ICCV | 7 |
| 2021 | Enriching Local and Global Contexts for Temporal Action LocalizationabstractEffectively tackling the problem of temporal action localization (TAL) necessitates a visual representation that jointly pursues two confounding goals, i.e., fine-grained discrimination for temporal localization and sufficient visual invariance for action classification. We address this challenge by enriching both the local and global contexts in the popular two-stage temporal localization framework, where action proposals are first generated followed by action classification and temporal boundary regression. Our proposed model, dubbed ContextLoc, can be divided into three sub-networks: L-Net, G-Net and P-Net. L-Net enriches the local context via fine-grained modeling of snippet-level features, which is formulated as a query-and-retrieval process. G-Net enriches the global context via higher-level modeling of the video-level representation. In addition, we introduce a novel context adaptation module to adapt the global context to different proposals. P-Net further models the context-aware inter-proposal relations. We explore two existing models to be the P-Net in our experiments. The efficacy of our proposed method is validated by experimental results on the THUMOS14 (54.3% at [email protected]) and ActivityNet v1.3 (56.01% at [email protected]) datasets, which outperforms recent states of the art. Code is available at https://github.com/buxiangzhiren/ContextLoc. Zixin Zhu, Wei Tang 0016, Le Wang 0003, Nanning Zheng 0001, Gang Hua 0001 |
ICCV | 5 |
| 2021 | Robust Pose Estimation in Crowded Scenes with Direct Pose-Level InferenceabstractMulti-person pose estimation in crowded scenes is challenging because overlapping and occlusions make it difficult to detect person bounding boxes and infer pose cues from individual keypoints. To address those issues, this paper proposes a direct pose-level inference strategy that is free of bounding box detection and keypoint grouping. Instead of inferring individual keypoints, the Pose-level Inference Network (PINet) directly infers the complete pose cues for a person from his/her visible body parts. PINet first applies the Part-based Pose Generation (PPG) to infer multiple coarse poses for each person from his/her body parts. Those coarse poses are refined by the Pose Refinement module through incorporating pose priors, and finally are fused in the Pose Fusion module. PINet relies on discriminative body parts to differentiate overlapped persons, and applies visual body cues to infer the global pose cues. Experiments on several crowded scenes pose estimation benchmarks demonstrate the superiority of PINet. For instance, it achieves 59.8% AP on the OCHuman dataset, outperforming the recent works by a large margin. Dongkai Wang, Shiliang Zhang, Gang Hua 0001 |
NeurIPS | 3 |
| 2021 | Graph-based temporal action co-localization from an untrimmed video
Le Wang 0003, Changbo Zhai, Qilin Zhang 0004, Wei Tang 0016, Nanning Zheng 0001, Gang Hua 0001 |
Neurocomputing | 6 |
| 2021 | Explicit Filterbank Learning for Neural Image Style Transfer and Image ProcessingabstractImage style transfer is to re-render the content of one image with the style of another. Most existing methods couple content and style information in their network structures and hyper-parameters, and learn it as a black-box. For better understanding, this paper aims to provide a new explicit decoupled perspective. Specifically, we propose StyleBank, which is composed of multiple convolution filter banks and each filter bank explicitly represents one style. To transfer an image to a specific style, the corresponding filter bank is operated on the intermediate feature produced by a single auto-encoder. The StyleBank and the auto-encoder are jointly learnt in such a way that the auto-encoder does not encode any style information. This explicit representation also enables us to conduct incremental learning to add a new style and fuse styles at not only the image level, but also the region level. Our method is the first style transfer network that links back to traditional texton mapping methods, and provides new understanding on neural style transfer. We further apply this general filterbank learning idea to two different multi-parameter image processing tasks: edge-aware image smoothing and denoising. Experiments demonstrate that it can achieve comparable results to its single parameter setting counterparts. Dongdong Chen 0001, Lu Yuan 0001, Jing Liao 0001, Nenghai Yu, Gang Hua 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2021 | A General Decoupled Learning Framework for Parameterized Image OperatorsabstractMany different deep networks have been used to approximate, accelerate or improve traditional image operators. Among these traditional operators, many contain parameters which need to be tweaked to obtain the satisfactory results, which we refer to as "parameterized image operators". However, most existing deep networks trained for these operators are only designed for one specific parameter configuration, which does not meet the needs of real scenarios that usually require flexible parameters settings. To overcome this limitation, we propose a new decoupled learning algorithm to learn from the operator parameters to dynamically adjust the weights of a deep network for image operators, denoted as the base network. The learned algorithm is formed as another network, namely the weight learning network, which can be end-to-end jointly trained with the base network. Experiments demonstrate that the proposed framework can be successfully applied to many traditional parameterized image operators. To accelerate the parameter tuning for practical scenarios, the proposed framework can be further extended to dynamically change the weights of only one single layer of the base network while sharing most computation cost. We demonstrate that this cheap parameter-tuning extension of the proposed decoupled learning framework even outperforms the state-of-the-art alternative approaches. Qingnan Fan, Dongdong Chen 0001, Lu Yuan 0001, Gang Hua 0001, Nenghai Yu, Baoquan Chen |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2021 | Editorial: Introduction to the Special Section on CVPR2019 Best Papers
Gang Hua 0001, Derek Hoiem, Abhinav Gupta 0001, Zhuowen Tu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | Usability Studies of an Egocentric Vision-Based Robotic WheelchairabstractMotivated by the need to improve the quality of life for the elderly and disabled individuals who rely on wheelchairs for mobility, and who may have limited or no hand functionality at all, we propose an egocentric computer vision based co-robot wheelchair to enhance their mobility without hand usage. The robot is built using a commercially available powered wheelchair modified to be controlled by head motion. Head motion is measured by tracking an egocentric camera mounted on the user’s head and faces outward. Compared with previous approaches to hands-free mobility, our system provides a more natural human robot interface because it enables the user to control the speed and direction of motion in a continuous fashion, as opposed to providing a small number of discrete commands. This article presents three usability studies, which were conducted on 37 subjects. The first two usability studies focus on comparing the proposed control method with existing solutions while the third study was conducted to assess the effectiveness of training subjects to operate the wheelchair over several sessions. A limitation of our studies is that they have been conducted with healthy participants. Our findings, however, pave the way for further studies with subjects with disabilities. Mohammed Kutbi, Xiaoxue Du, Yizhe Chang, Nikolaos Agadakos, Gang Hua 0001, Philippos Mordohai |
ACM Trans. Hum. Robot Interact. | 7 |
| 2021 | Giant Panda IdentificationabstractThe lack of automatic tools to identify giant panda makes it hard to keep track of and manage giant pandas in wildlife conservation missions. In this paper, we introduce a new Giant Panda Identification (GPID) task, which aims to identify each individual panda based on an image. Though related to the human re-identification and animal classification problem, GPID is extraordinarily challenging due to subtle visual differences between pandas and cluttered global information. In this paper, we propose a new benchmark dataset iPanda-50 for GPID. The iPanda-50 consists of 6, 874 images from 50 giant panda individuals, and is collected from panda streaming videos. We also introduce a new Feature-Fusion Network with Patch Detector (FFN-PD) for GPID. The proposed FFN-PD exploits the patch detector to detect discriminative local patches without using any part annotations or extra location sub-networks, and builds a hierarchical representation by fusing both global and local features to enhance the inter-layer patch feature interactions. Specifically, an attentional cross-channel pooling is embedded in the proposed FFN-PD to improve the identify-specific patch detectors. Experiments performed on the iPanda-50 datasets demonstrate the proposed FFN-PD significantly outperforms competing methods. Besides, experiments on other fine-grained recognition datasets (i.e., CUB-200-2011, Stanford Cars, and FGVC-Aircraft) demonstrate that the proposed FFN-PD outperforms existing state-of-the-art methods. Le Wang 0003, Rizhi Ding, Yuanhao Zhai 0001, Qilin Zhang 0004, Wei Tang 0016, Nanning Zheng 0001, Gang Hua 0001 |
IEEE Trans. Image Process. | 7 |
| 2021 | Object Cosegmentation in Noisy Videos With Multilevel HypergraphabstractWith the target of simultaneously segmenting semantically related videos to identify the common objects, video object cosegmentation has attracted the attention of researchers in recent years. Existing methods are primarily based on pair-wise relations between adjacent pixels and regions, which are susceptible to performance degradation from object entries/exists or occlusions. Specifically, we refer these video frames without the common objects present as the “empty” frames. In this paper, we propose a multilevel hypergraph-based full Video object CoSegmentation (VCS) method, which incorporates high-level semantics and low-level appearance/motion/saliency to construct the hyperedge among multiple spatially and temporally adjacent regions. Specifically, the high-level semantic model fuses multiple object proposals from each frame instead of relying on a single object proposal per frame. A hypergraph cut is subsequently utilized to calculate the object cosegmentation. Experiments on four video object segmentation/cosegmentation datasets against state-of-the-art methods with both objective and subjective results manifest the effectiveness of the proposed VCS method, including the SegTrack and VCoSeg datasets without “empty” frames, the XJTU-Stevens dataset with 3.7% “empty” frames, and the Noisy-ViCoSeg dataset proposed together with our method with 30.3% “empty” frames. Le Wang 0003, Qilin Zhang 0004, Zhenxing Niu, Nanning Zheng 0001, Gang Hua 0001 |
IEEE Trans. Multim. | 6 |
| 2020 | Ladder Loss for Coherent Visual-Semantic EmbeddingabstractFor visual-semantic embedding, the existing methods normally treat the relevance between queries and candidates in a bipolar way – relevant or irrelevant, and all “irrelevant” candidates are uniformly pushed away from the query by an equal margin in the embedding space, regardless of their various proximity to the query. This practice disregards relatively discriminative information and could lead to suboptimal ranking in the retrieval results and poorer user experience, especially in the long-tail query scenario where a matching candidate may not necessarily exist. In this paper, we introduce a continuous variable to model the relevance degree between queries and multiple candidates, and propose to learn a coherent embedding space, where candidates with higher relevance degrees are mapped closer to the query than those with lower relevance degrees. In particular, the new ladder loss is proposed by extending the triplet loss inequality to a more general inequality chain, which implements variable push-away margins according to respective relevance degrees. In addition, a proper Coherent Score metric is proposed to better measure the ranking results including those “irrelevant” candidates. Extensive experiments on multiple datasets validate the efficacy of our proposed method, which achieves significant improvement over existing state-of-the-art methods. Zhenxing Niu, Le Wang 0003, Zhanning Gao, Qilin Zhang 0004, Gang Hua 0001 |
AAAI | 6 |
| 2020 | Loss Functions for Person Image Generation
Haoyue Shi 0002, Le Wang 0003, Wei Tang 0016, Nanning Zheng 0001, Gang Hua 0001 |
BMVC | 5 |
| 2020 | Self-Robust 3D Point Recognition via Gather-Vector GuidanceabstractIn this paper, we look into the problem of 3D adversary attack, and propose to leverage the internal properties of the point clouds and the adversarial examples to design a new self-robust deep neural network (DNN) based 3D recognition systems. As a matter of fact, on one hand, point clouds are highly structured. Hence for each local part of clean point clouds, it is possible to learn what is it (``part of a bottle") and its relative position (``upper part of a bottle") to the global object center. On the other hand, with the visual quality constraint, 3D adversarial samples often only produce small local perturbations, thus they will roughly keep the original global center but may cause incorrect local relative position estimation. Motivated by these two properties, we use relative position (dubbed as ``gather-vector") as the adversarial indicator and propose a new robust gather module. Equipped with this module, we further propose a new self-robust 3D point recognition network. Through extensive experiments, we demonstrate that the proposed method can improve the robustness of the target attack under the white-box setting significantly. For I-FGSM based attack, our method reduces the attack success rate from 94.37 \% to 75.69 \%. For C\&W based attack, our method reduces the attack success rate more than 40.00 \%. Moreover, our method is complementary to other types of defense methods to achieve better defense results. Xiaoyi Dong, Dongdong Chen 0001, Hang Zhou 0007, Gang Hua 0001, Weiming Zhang 0001, Nenghai Yu |
CVPR | 4 |
| 2020 | gDLS*: Generalized Pose-and-Scale Estimation Given Scale and Gravity PriorsabstractMany real-world applications in augmented reality (AR), 3D mapping, and robotics require both fast and accurate estimation of camera poses and scales from multiple images captured by multiple cameras or a single moving camera. Achieving high speed and maintaining high accuracy in a pose-and-scale estimator are often conflicting goals. To simultaneously achieve both, we exploit a priori knowledge about the solution space. We present gDLS*, a generalized-camera-model pose-and-scale estimator that utilizes rotation and scale priors. gDLS* allows an application to flexibly weigh the contribution of each prior, which is important since priors often come from noisy sensors. Compared to state-of-the-art generalized-pose-and-scale estimators (e.g., gDLS), our experiments on both synthetic and real data consistently demonstrate that gDLS* accelerates the estimation process and improves scale and pose accuracy. Victor Fragoso, Joseph DeGol, Gang Hua 0001 |
CVPR | 3 |
| 2020 | SaccadeNet: A Fast and Accurate Object DetectorabstractObject detection is an essential step towards holistic scene understanding. Most existing object detection algorithms attend to certain object areas once and then predict the object locations. However, scientists have revealed that human do not look at the scene in fixed steadiness. Instead, human eyes move around, locating informative parts to understand the object location. This active perceiving movement process is called saccade. In this paper, inspired by such mechanism, we propose a fast and accurate object detector called SaccadeNet. It contains four main modules, the Center Attentive Module, the Corner Attentive Module, the Attention Transitive Module, and the Aggregation Attentive Module, which allows it to attend to different informative object keypoints actively, and predict object locations from coarse to fine. The Corner Attentive Module is used only during training to extract more informative corner features which brings free-lunch performance boost. On the MS COCO dataset, we achieve the performance of 40.4% mAP at 28 FPS and 30.5% mAP at 118 FPS. Among all the real-time object detectors, our SaccadeNet achieves the best detection performance, which demonstrates the effectiveness of the proposed detection mechanism. Shiyi Lan, Zhou Ren, Larry Davis 0001, Gang Hua 0001 |
CVPR | 5 |
| 2020 | Few-Shot Open-Set Recognition Using Meta-LearningabstractThe problem of open-set recognition is considered. While previous approaches only consider this problem in the context of large-scale classifier training, we seek a unified solution for this and the low-shot classification setting. It is argued that the classic softmax classifier is a poor solution for open-set recognition, since it tends to overfit on the training classes. Randomization is then proposed as a solution to this problem. This suggests the use of meta-learning techniques, commonly used for few-shot classification, for the solution of open-set recognition. A new oPen sEt mEta LEaRning (PEELER) algorithm is then introduced. This combines the random selection of a set of novel classes per episode, a loss that maximizes the posterior entropy for examples of those classes, and a new metric learning formulation based on the Mahalanobis distance. Experimental results show that PEELER achieves state of the art open set recognition performance for both few-shot and large-scale recognition. On CIFAR and miniImageNet, it achieves substantial gains in seen/unseen class detection AUROC for a given seen-class classification accuracy. Bo Liu 0043, Hao Kang, Gang Hua 0001, Nuno Vasconcelos |
CVPR | 4 |
| 2020 | LG-GAN: Label Guided Adversarial Network for Flexible Targeted Attack of Point Cloud Based Deep NetworksabstractDeep neural networks have made tremendous progress in 3D point-cloud recognition. Recent works have shown that these 3D recognition networks are also vulnerable to adversarial samples produced from various attack methods, including optimization-based 3D Carlini-Wagner attack, gradient-based iterative fast gradient method, and skeleton-detach based point-dropping. However, after a careful analysis, these methods are either extremely slow because of the optimization/iterative scheme, or not flexible to support targeted attack of a specific category. To overcome these shortcomings, this paper proposes a novel label guided adversarial network (LG-GAN) for real-time flexible targeted point cloud attack. To the best of our knowledge, this is the first generation based 3D point cloud attack method. By feeding the original point clouds and target attack label into LG-GAN, it can learn how to deform the point clouds to mislead the recognition network into the specific label only with a single forward pass. In detail, LG-GAN first leverages one multi-branch adversarial network to extract hierarchical features of the input point clouds, then incorporates the specified label information into multiple intermediate features using the label encoder. Finally, the encoded features will be fed into the coordinate reconstruction decoder to generate the target adversarial sample. By evaluating different point-cloud recognition models (e.g., PointNet, PointNet++ and DGCNN), we demonstrate that the proposed LG-GAN can support flexible targeted attack on the fly while guaranteeing good attack performance and higher efficiency simultaneously. Hang Zhou 0007, Dongdong Chen 0001, Jing Liao 0001, Kejiang Chen, Xiaoyi Dong, Kunlin Liu, Weiming Zhang 0001, Gang Hua 0001, Nenghai Yu |
CVPR | 8 |
| 2020 | Two-Stream Consensus Network for Weakly-Supervised Temporal Action Localization
Yuanhao Zhai 0001, Le Wang 0003, Wei Tang 0016, Qilin Zhang 0004, Junsong Yuan 0001, Gang Hua 0001 |
ECCV (6) | 6 |
| 2020 | Adversarial Ranking Attack and Defense
Zhenxing Niu, Le Wang 0003, Qilin Zhang 0004, Gang Hua 0001 |
ECCV (14) | 5 |
| 2020 | Temporal Keypoint Matching and Refinement Network for Pose Estimation and Tracking
Chunluan Zhou, Zhou Ren, Gang Hua 0001 |
ECCV (22) | 3 |
| 2020 | Fine-Grained Giant Panda IdentificationabstractThe image-based fine-grained identification of individual giant pandas (Ailuropoda melanoleuca) is an emerging technology, and it is extraordinarily challenging due to the extremely subtle visual differences between individual giant pandas and limited annotated training data. To address these challenges, we propose the Feature-Fusion Convolutional Neural Network with Patch Detector (FFCNN-PD) algorithm, which exploits the discriminative local patches and builds a hierarchical representation generated by fusing both global and local features. Specifically, an attentional cross-channel pooling is embedded in the FFCNN-PD to improve the class- specific patch detectors. In addition, we propose a new giant panda identification dataset (iPanda-30) to establish a benchmark. Experiments on the proposed iPanda-30 dataset and other fine-grained recognition datasets demonstrate the effectiveness of the FFCNN-PD algorithm against the existing state-of-the-arts. Rizhi Ding, Le Wang 0003, Qilin Zhang 0004, Zhenxing Niu, Nanning Zheng 0001, Gang Hua 0001 |
ICASSP | 6 |
| 2020 | Action Co-localization in an Untrimmed Video by Graph Neural Networks
Changbo Zhai, Le Wang 0003, Qilin Zhang 0004, Zhanning Gao, Zhenxing Niu, Nanning Zheng 0001, Gang Hua 0001 |
MMM (1) | 7 |
| 2020 | Passport-aware Normalization for Deep Model ProtectionabstractDespite tremendous success in many application scenarios, deep learning faces serious intellectual property (IP) infringement threats. Considering the cost of designing and training a good model, infringements will significantly infringe the interests of the original model owner. Recently, many impressive works have emerged for deep model IP protection. However, they either are vulnerable to ambiguity attacks, or require changes in the target network structure by replacing its original normalization layers and hence cause significant performance drops. To this end, we propose a new passport-aware normalization formulation, which is generally applicable to most existing normalization layers and only needs to add another passport-aware branch for IP protection. This new branch is jointly trained with the target model but discarded in the inference stage. Therefore it causes no structure change in the target model. Only when the model IP is suspected to be stolen by someone, the private passport-aware branch is added back for ownership verification. Through extensive experiments, we verify its effectiveness in both image and 3D point recognition models. It is demonstrated to be robust not only to common attack techniques like fine-tuning and model compression, but also to ambiguity attacks. By further combining it with trigger-set based methods, both black-box and white-box verification can be achieved for enhanced security of deep learning models deployed in real systems. Jie Zhang 0073, Dongdong Chen 0001, Jing Liao 0001, Weiming Zhang 0001, Gang Hua 0001, Nenghai Yu |
NeurIPS | 5 |
| 2020 | Calibrated Domain-Invariant Learning for Highly Generalizable Large Scale Re-IdentificationabstractMany real-world applications, such as city scale traffic monitoring and control, requires large scale re-identification. However, previous ReID methods often failed to address two limitations in existing ReID benchmarks, i.e., low spatiotem-poral coverage and sample imbalance. Notwithstanding their demonstrated success in every single benchmark, they have difficulties in generalizing to unseen environments. As a result, these methods are less applicable in a large scale setting due to poor generalization. In seek for a highly generalizable large-scale ReID method, we present an adversarial domain-invariant feature learning framework (ADIN) that explicitly learns to separate identity-related features from challenging variations, where for the first time "free" annotations in ReID data such as video timestamp and camera index are utilized. Furthermore, we find that the imbalance of nuisance classes jeopardizes the adversarial training, and for mitigation we propose a calibrated adversarial loss that is attentive to nuisance distribution. Experiments on existing large-scale person/vehicle ReID datasets demonstrate that ADIN learns more robust and generalizable representations, as evidenced by its outstanding direct transfer performance across datasets, which is a criterion that can better measure the generalizability of large scale Re-ID methods. Ye Yuan 0012, Wuyang Chen 0001, Tianlong Chen 0001, Zhou Ren, Zhangyang Wang, Gang Hua 0001 |
WACV | 7 |
| 2020 | Semi-online Multi-people Tracking by Re-identification
Long Lan, Xinchao Wang, Gang Hua 0001, Thomas S. Huang, Dacheng Tao |
Int. J. Comput. Vis. | 3 |
| 2020 | Controllable Image Processing via Adaptive FilterBank PyramidabstractTraditional image processing operators often provide some control parameters to tweak the final results. Recently, different convolutional neural networks have been used to approximate or improve these operators. However, in those methods, one single model can only handle one operator of a specific parameter value and does not support parameter tuning. In this paper, we propose a new plugin module, “Adaptive Filterbank Pyramid”, which can be inserted into a backbone network to support multiple operators and continuous parameter tuning. Our module explicitly represents one operator with one filterbank pyramid. To generate the results of a specific operator, the corresponding filterbank pyramid is convolved with the intermediate feature pyramid produced by the backbone network. The weights of the filterbank pyramid are directly regressed by another sub-network, which is jointly trained with the backbone network and adapted to the input parameter, thus enabling continuous parameter tuning. We applied the proposed module for a large variety of image processing tasks, including image smoothing, image denoising, image deblocking, image enhancement and neural style transfer. Experiments show that our method is generalized to different types of image processing tasks and different backbone network structures. Compared to the single-operator-single-parameter baseline, our method can produce comparable results but is significantly more efficient in both training and testing. Dongdong Chen 0001, Qingnan Fan, Jing Liao 0001, Angelica I. Avilés-Rivero, Lu Yuan 0001, Nenghai Yu, Gang Hua 0001 |
IEEE Trans. Image Process. | 7 |
| 2020 | Improving Person Re-Identification With Iterative Impression AggregationabstractOur impression about one person often updates after we see more aspects of him/her and this process keeps iterating given more meetings. We formulate such an intuition into the problem of person re-identification (re-ID), where the representation of a query (probe) image is iteratively updated with new information from the candidates in the gallery. Specifically, we propose a simple attentional aggregation formulation to instantiate this idea and showcase that such a pipeline achieves competitive performance on standard benchmarks including CUHK03, Market-1501 and DukeMTMC. Not only does such a simple method improve the performance of the baseline models, it also achieves comparable performance with latest advanced re-ranking methods. Another advantage of this proposal is its flexibility to incorporate different representations and similarity metrics. By utilizing stronger representations and metrics, we further demonstrate state-of-the-art person re-ID performance, which also validates the general applicability of the proposed method. Dengpan Fu, Bo Xin, Jingdong Wang 0001, Dongdong Chen 0001, Jianmin Bao, Gang Hua 0001, Houqiang Li |
IEEE Trans. Image Process. | 6 |
| 2019 | Video Imprint Segmentation for Temporal Action Detection in Untrimmed VideosabstractWe propose a temporal action detection by spatial segmentation framework, which simultaneously categorize actions and temporally localize action instances in untrimmed videos. The core idea is the conversion of temporal detection task into a spatial semantic segmentation task. Firstly, the video imprint representation is employed to capture the spatial/temporal interdependences within/among frames and represent them as spatial proximity in a feature space. Subsequently, the obtained imprint representation is spatially segmented by a fully convolutional network. With such segmentation labels projected back to the video space, both temporal action boundary localization and per-frame spatial annotation can be obtained simultaneously. The proposed framework is robust to variable lengths of untrimmed videos, due to the underlying fixed-size imprint representations. The efficacy of the framework is validated in two public action detection datasets. Zhanning Gao, Le Wang 0003, Qilin Zhang 0004, Zhenxing Niu, Nanning Zheng 0001, Gang Hua 0001 |
AAAI | 6 |
| 2019 | Object Affordances Graph Network for Action Recognition
Haoliang Tan, Le Wang 0003, Qilin Zhang 0004, Zhanning Gao, Nanning Zheng 0001, Gang Hua 0001 |
BMVC | 6 |
| 2019 | Capturing Piecewise SVBRDFs with Content Aware Lighting
Xiao Li 0030, Peiran Ren, Yue Dong 0001, Gang Hua 0001, Xin Tong 0001, Baining Guo |
CGI | 4 |
| 2019 | Weakly Supervised Temporal Action Localization Through Contrast Based Evaluation NetworksabstractWeakly-supervised temporal action localization (WS-TAL) is a promising but challenging task with only video-level action categorical labels available during training. Without requiring temporal action boundary annotations in training data, WS-TAL could possibly exploit automatically retrieved video tags as video-level labels. However, such coarse video-level supervision inevitably incurs confusions, especially in untrimmed videos containing multiple action instances. To address this challenge, we propose the Contrast-based Localization EvaluAtioN Network (CleanNet) with our new action proposal evaluator, which provides pseudo-supervision by leveraging the temporal contrast in snippet-level action classification predictions. Essentially, the new action proposal evaluator enforces an additional temporal contrast constraint so that high-evaluation-score action proposals are more likely to coincide with true action instances. Moreover, the new action localization module is an integral part of CleanNet which enables end-to-end training. This is in contrast to many existing WS-TAL methods where action localization is merely a post-processing step. Experiments on THUMOS14 and ActivityNet datasets validate the efficacy of CleanNet against existing state-ofthe- art WS-TAL algorithms. Ziyi Liu 0001, Le Wang 0003, Qilin Zhang 0004, Zhanning Gao, Zhenxing Niu, Nanning Zheng 0001, Gang Hua 0001 |
ICCV | 7 |
| 2019 | Action Coherence Network for Weakly Supervised Temporal Action LocalizationabstractMost prominent temporal action localization methods are of the fully-supervised type, which rely heavily on frame-level labels, which could be prohibitively expensive to annotate. Thanks to recent developments on the Weakly-supervised Temporal Action Localization (W-TAL), this alternative paradigm requires only video-level labels in training, alleviating such annotation efforts. Specifically, we present Action Coherence Network (ACN) for W-TAL, which features a new coherence loss that better supervises action boundary learning and facilitate proposal regression. In addition, a purpose-built fusion module is proposed for localization inference based on features extracted by two streams of convolutional neural network. Overall, the proposed ACN achieves state-of-the-art W-TAL performance on two challenging datasets (THU-MOS14 and ActivityNet1.2, particularly ACN attains mAP of 24.2% on THUMOS14 under IoU threshold 0.5), which is approaching some recent fully-supervised TAL methods. Yuanhao Zhai 0001, Le Wang 0003, Ziyi Liu 0001, Qilin Zhang 0004, Gang Hua 0001, Nanning Zheng 0001 |
ICIP | 5 |
| 2019 | Priming Deep Pedestrian Detection with Geometric ContextabstractWe investigate the role of geometric context in deep neural networks to establish better pedestrian detectors that are more robust to occlusions. Notwithstanding their demonstrated successes, deep object detectors under-perform in crowded scenes with high intra-category occlusions. One brute-force solution is to collect a large number of labeled training samples under occlusion, but the combinatorial increase in the labeling effort makes it an unaffordable solution. We argue that a promising and complementary direction to solve this problem is to bring geometric context to modulate feature learning in a DNN. We identify that an effective way to leverage geometric context is to induce it in two steps - through early fusion, by guiding region proposal generation to focus on occluded regions, and through late fusion, by penalizing misalignments of bounding boxes in both 2D and 3D. Our experiments on multiple state-of-the-art DNN detectors and detection benchmarks clearly demonstrates that our proposed method outperforms strong baselines by an average of 5%. Ishani Chakraborty, Gang Hua 0001 |
ICRA | 2 |
| 2019 | Gated Context Aggregation Network for Image Dehazing and DerainingabstractImage dehazing aims to recover the uncorrupted content from a hazy image. Instead of leveraging traditional low-level or handcrafted image priors as the restoration constraints, e.g., dark channels and increased contrast, we propose an end-to-end gated context aggregation network to directly restore the final haze-free image. In this network, we adopt the latest smoothed dilation technique to help remove the gridding artifacts caused by the widely-used dilated convolution with negligible extra parameters, and leverage a gated sub-network to fuse the features from different levels. Extensive experiments demonstrate that our method can surpass previous state-of-the-art methods by a large margin both quantitatively and qualitatively. In addition, to demonstrate the generality of the proposed method, we further apply it to the image deraining task, which also achieves the state-of-the-art performance. Dongdong Chen 0001, Mingming He, Qingnan Fan, Jing Liao 0001, Liheng Zhang, Dongdong Hou, Lu Yuan 0001, Gang Hua 0001 |
WACV | 8 |
| 2019 | Video ImprintabstractA new unified video analytics framework (ER3) is proposed for complex event retrieval, recognition and recounting, based on the proposed video imprint representation, which exploits temporal correlations among image features across video frames. With the video imprint representation, it is convenient to reverse map back to both temporal and spatial locations in video frames, allowing for both key frame identification and key areas localization within each frame. In the proposed framework, a dedicated feature alignment module is incorporated for redundancy removal across frames to produce the tensor representation, i.e., the video imprint. Subsequently, the video imprint is individually fed into both a reasoning network and a feature aggregation module, for event recognition/recounting and event retrieval tasks, respectively. Thanks to its attention mechanism inspired by the memory networks used in language modeling, the proposed reasoning network is capable of simultaneous event category recognition and localization of the key pieces of evidence for event recounting. In addition, the latent structure in our reasoning network highlights the areas of the video imprint, which can be directly used for event recounting. With the event retrieval task, the compact video representation aggregated from the video imprint contributes to better retrieval results than existing state-of-the-art methods. Zhanning Gao, Le Wang 0003, Nebojsa Jojic, Zhenxing Niu, Nanning Zheng 0001, Gang Hua 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2019 | Order-Preserving Optimal Transport for Distances between SequencesabstractWe present new distance measures between sequences that can tackle local temporal distortion and periodic sequences with arbitrary starting points. Through viewing the instances of each sequence as empirical samples of an unknown distribution, we cast the calculations of distances between sequences as optimal transport problems. To preserve the inherent temporal relationships of the instances in sequences, we propose two methods through incorporating the temporal information into the spatial ground metric and concentrating the transport with two novel temporal regularization terms, respectively. The inverse difference moment regularization enforces local homogeneous structures in the transport, and the KL-divergence with a prior distribution regularization prevents transport between instances with far temporal positions. We show that the resulting problems can be efficiently solved by the matrix scaling algorithm. Extensive experiments on eight datasets with different classifiers and performance measures show the effectiveness and generality of the proposed distances. Bing Su 0001, Gang Hua 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2018 | Towards Open-Set Identity Preserving Face SynthesisabstractWe propose a framework based on Generative Adversarial Networks to disentangle the identity and attributes of faces, such that we can conveniently recombine different identities and attributes for identity preserving face synthesis in open domains. Previous identity preserving face synthesis processes are largely confined to synthesizing faces with known identities that are already in the training dataset. To synthesize a face with identity outside the training dataset, our framework requires one input image of that subject to produce an identity vector, and any other input face image to extract an attribute vector capturing, e.g., pose, emotion, illumination, and even the background. We then recombine the identity vector and the attribute vector to synthesize a new face of the subject with the extracted attribute. Our proposed framework does not need to annotate the attributes of faces in any way. It is trained with an asymmetric loss function to better preserve the identity and stabilize the training process. It can also effectively leverage large amounts of unlabeled training face images to further improve the fidelity of the synthesized faces for subjects that are not presented in the labeled training face dataset. Our experiments demonstrate the efficacy of the proposed framework. We also present its usage in a much broader set of applications including face frontalization, face attribute morphing, and face adversarial example detection. Jianmin Bao, Dong Chen 0003, Fang Wen 0001, Houqiang Li, Gang Hua 0001 |
CVPR | 5 |
| 2018 | Stereoscopic Neural Style TransferabstractThis paper presents the first attempt at stereoscopic neural style transfer, which responds to the emerging demand for 3D movies or AR/VR. We start with a careful examination of applying existing monocular style transfer methods to left and right views of stereoscopic images separately. This reveals that the original disparity consistency cannot be well preserved in the final stylization results, which causes 3D fatigue to the viewers. To address this issue, we incorporate a new disparity loss into the widely adopted style loss function by enforcing the bidirectional disparity constraint in non-occluded regions. For a practical realtime solution, we propose the first feed-forward network by jointly training a stylization sub-network and a disparity sub-network, and integrate them in a feature level middle domain. Our disparity sub-network is also the first end-to-end network for simultaneous bidirectional disparity and occlusion mask estimation. Finally, our network is effectively extended to stereoscopic videos, by considering both temporal coherence and disparity consistency. We will show that the proposed method clearly outperforms the baseline algorithms both quantitatively and qualitatively. Dongdong Chen 0001, Lu Yuan 0001, Jing Liao 0001, Nenghai Yu, Gang Hua 0001 |
CVPR | 5 |
| 2018 | Revisiting Deep Intrinsic Image DecompositionsabstractWhile invaluable for many computer vision applications, decomposing a natural image into intrinsic reflectance and shading layers represents a challenging, underdetermined inverse problem. As opposed to strict reliance on conventional optimization or filtering solutions with strong prior assumptions, deep learning based approaches have also been proposed to compute intrinsic image decompositions when granted access to sufficient labeled training data. The downside is that current data sources are quite limited, and broadly speaking fall into one of two categories: either dense fully-labeled images in synthetic/narrow settings, or weakly-labeled data from relatively diverse natural scenes. In contrast to many previous learning-based approaches, which are often tailored to the structure of a particular dataset (and may not work well on others), we adopt core network structures that universally reflect loose prior knowledge regarding the intrinsic image formation process and can be largely shared across datasets. We then apply flexibly supervised loss layers that are customized for each source of ground truth labels. The resulting deep architecture achieves state-of-the-art results on all of the major intrinsic image benchmarks, and runs considerably faster than most at test time. Qingnan Fan, Jiaolong Yang, Gang Hua 0001, Baoquan Chen, David P. Wipf |
CVPR | 3 |
| 2018 | Semi-supervised FusedGAN for Conditional Image Generation
Navaneeth Bodla, Gang Hua 0001, Rama Chellappa |
ECCV (5) | 2 |
| 2018 | Decouple Learning for Parameterized Image Operators
Qingnan Fan, Dongdong Chen 0001, Lu Yuan 0001, Gang Hua 0001, Nenghai Yu, Baoquan Chen |
ECCV (13) | 4 |
| 2018 | Stacked Cross Attention for Image-Text Matching
Kuang-Huei Lee, Gang Hua 0001, Houdong Hu, Xiaodong He 0001 |
ECCV (4) | 3 |
| 2018 | LQ-Nets: Learned Quantization for Highly Accurate and Compact Deep Neural Networks
Dongqing Zhang, Jiaolong Yang, Dongqiangzi Ye, Gang Hua 0001 |
ECCV (8) | 4 |
| 2018 | Joint Spatio-Temporal Action Localization in Untrimmed Videos with Per-Frame SegmentationabstractInspired by the recent spatio-temporal action localization efforts with tubelets (sequences of bounding boxes), we present a new spatio-temporal action detector Segment-tube, which consists of sequences of per-frame segmentation masks. The proposed Segment-tube detector can temporally pinpoint the starting/ending frame of each action class in the presence of preceding/subsequent interference actions in untrimmed videos. Simultaneously, the Segment-tube detector produces per-frame segmentation masks instead of bounding boxes, offering superior spatial accuracy to tubelets. This is achieved by alternating iterative optimization between temporal action localization and spatial action segmentation. Experimental results on multiple datasets validate the efficacy of the proposed detector. Xuhuan Duan, Le Wang 0003, Changbo Zhai, Nanning Zheng 0001, Qilin Zhang 0004, Zhenxing Niu, Gang Hua 0001 |
ICIP | 7 |
| 2018 | Video Object Co-Segmentation from Noisy Videos by a Multi-Level Hypergraph ModelabstractDefined as simultaneously segmenting a set of related videos to identify the common objects, video co-segmentation has attracted the attention of researchers in recent years. Existing methods are primarily based on pair-wise relations between adjacent pixels/regions, which are susceptible to performance degradation from “empty” video frames (e.g., due to transient/intermittent common objects). In this paper, a new multilevel hypergraph based method, termed the full Video object Co-Segmentation method (VCS), is proposed, which incorporates both a high-level semantics object model and a low-level appearance/motion/saliency object model to construct the hyperedge among multiple spatially and temporally adjacent regions. Specifically, the high-level semantic model fuses multiple object proposals from each frame instead of relying on a single object proposal per frame. A hypergraph cut is subsequently utilized to calculate the object co-segmentation. Experiments on three datasets demonstrate the efficacy of the proposed VCS method. Le Wang 0003, Qilin Zhang 0004, Nanning Zheng 0001, Gang Hua 0001 |
ICIP | 5 |
| 2018 | Guest Editorial
Lamberto Ballan, Shih-Fu Chang, Gang Hua 0001, Thomas Mensink, Greg Mori, Rahul Sukthankar |
Comput. Vis. Image Underst. | 3 |
| 2018 | Connections with Robust PCA and the Role of Emergent Sparsity in Variational Autoencoder ModelsabstractVariational autoencoders (VAE) represent a popular, flexible form of deep generative model that can be stochastically fit to samples from a given random process using an information-theoretic variational bound on the true underlying distribution. Once so-obtained, the model can be putatively used to generate new samples from this distribution, or to provide a low-dimensional latent representation of existing samples. While quite effective in numerous application domains, certain important mechanisms which govern the behavior of the VAE are obfuscated by the intractable integrals and resulting stochastic approximations involved. Moreover, as a highly non-convex model, it remains unclear exactly how minima of the underlying energy relate to original design purposes. We attempt to better quantify these issues by analyzing a series of tractable special cases of increasing complexity. In doing so, we unveil interesting connections with more traditional dimensionality reduction models, as well as an intrinsic yet underappreciated propensity for robustly dismissing sparse outliers when estimating latent manifolds. With respect to the latter, we demonstrate that the VAE can be viewed as the natural evolution of recent robust PCA models, capable of learning nonlinear manifolds of unknown dimension obscured by gross corruptions. Bin Dai 0008, Yu Wang 0060, John A. D. Aston, Gang Hua 0001, David P. Wipf |
J. Mach. Learn. Res. | 4 |
| 2018 | Collaborative Active Visual Recognition from Crowds: A Distributed Ensemble ApproachabstractActive learning is an effective way of engaging users to interactively train models for visual recognition more efficiently. The vast majority of previous works focused on active learning with a single human oracle. The problem of active learning with multiple oracles in a collaborative setting has not been well explored. We present a collaborative computational model for active learning with multiple human oracles, the input from whom may possess different levels of noises. It leads to not only an ensemble kernel machine that is robust to label noises, but also a principled label quality measure to online detect irresponsible labelers. Instead of running independent active learning processes for each individual human oracle, our model captures the inherent correlations among the labelers through shared data among them. Our experiments with both simulated and real crowd-sourced noisy labels demonstrate the efficacy of our model. Gang Hua 0001, Chengjiang Long, Ming Yang 0007, Yan Gao 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2018 | Probabilistic Elastic Part Model: A Pose-Invariant Representation for Real-World Face VerificationabstractPose variation remains to be a major challenge for real-world face recognition. We approach this problem through a probabilistic elastic part model. We extract local descriptors (e.g., LBP or SIFT) from densely sampled multi-scale image patches. By augmenting each descriptor with its location, a Gaussian mixture model (GMM) is trained to capture the spatial-appearance distribution of the face parts of all face images in the training corpus, namely the probabilistic elastic part (PEP) model. Each mixture component of the GMM is confined to be a spherical Gaussian to balance the influence of the appearance and the location terms, which naturally defines a part. Given one or multiple face images of the same subject, the PEP-model builds its PEP representation by sequentially concatenating descriptors identified by each Gaussian component in a maximum likelihood sense. We further propose a joint Bayesian adaptation algorithm to adapt the universally trained GMM to better model the pose variations between the target pair of faces/face tracks, which consistently improves face verification accuracy. Our experiments show that we achieve state-of-the-art face verification accuracy with the proposed representations on the Labeled Face in the Wild (LFW) dataset, the YouTube video face database, and the CMU MultiPIE dataset. Gang Hua 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2018 | Joint Video Object Discovery and Segmentation by Coupled Dynamic Markov NetworksabstractIt is a challenging task to extract segmentation mask of a target from a single noisy video, which involves object discovery coupled with segmentation. To solve this challenge, we present a method to jointly discover and segment an object from a noisy video, where the target disappears intermittently throughout the video. Previous methods either only fulfill video object discovery, or video object segmentation presuming the existence of the object in each frame. We argue that jointly conducting the two tasks in a unified way will be beneficial. In other words, video object discovery and video object segmentation tasks can facilitate each other. To validate this hypothesis, we propose a principled probabilistic model, where two dynamic Markov networks are coupled-one for discovery and the other for segmentation. When conducting the Bayesian inference on this model using belief propagation, the bi-directional message passing reveals a clear collaboration between these two inference tasks. We validated our proposed method in five data sets. The first three video data sets, i.e., the SegTrack data set, the YouTube-objects data set, and the Davis data set, are not noisy, where all video frames contain the objects. The two noisy data sets, i.e., the XJTU-Stevens data set, and the Noisy-ViDiSeg data set, newly introduced in this paper, both have many frames that do not contain the objects. When compared with state of the art, it is shown that although our method produces inferior results on video data sets without noisy frames, we are able to obtain better results on video data sets with noisy frames. Ziyi Liu 0001, Le Wang 0003, Gang Hua 0001, Qilin Zhang 0004, Zhenxing Niu, Ying Wu 0001, Nanning Zheng 0001 |
IEEE Trans. Image Process. | 3 |
| 2018 | Knowledge-Based Topic Model for Unsupervised Object Discovery and LocalizationabstractUnsupervised object discovery and localization is to discover some dominant object classes and localize all of object instances from a given image collection without any supervision. Previous work has attempted to tackle this problem with vanilla topic models, such as latent Dirichlet allocation (LDA). However, in those methods no prior knowledge for the given image collection is exploited to facilitate object discovery. On the other hand, the topic models used in those methods suffer from the topic coherence issue-some inferred topics do not have clear meaning, which limits the final performance of object discovery. In this paper, prior knowledge in terms of the so-called must-links are exploited from Web images on the Internet. Furthermore, a novel knowledge-based topic model, called LDA with mixture of Dirichlet trees, is proposed to incorporate the must-links into topic modeling for object discovery. In particular, to better deal with the polysemy phenomenon of visual words, the must-link is re-defined as that one must-link only constrains one or some topic(s) instead of all topics, which leads to significantly improved topic coherence. Moreover, the must-links are built and grouped with respect to specific object classes, thus the must-links in our approach are semantic-specific, which allows to more efficiently exploit discriminative prior knowledge from Web images. Extensive experiments validated the efficiency of our proposed approach on several data sets. It is shown that our method significantly improves topic coherence and outperforms the unsupervised methods for object discovery and localization. In addition, compared with discriminative methods, the naturally existing object classes in the given image collection can be subtly discovered, which makes our approach well suited for realistic applications of unsupervised object discovery. Zhenxing Niu, Gang Hua 0001, Le Wang 0003, Xinbo Gao 0001 |
IEEE Trans. Image Process. | 2 |
| 2017 | How to Train a Compact Binary Neural Network with High Accuracy?abstractHow to train a binary neural network (BinaryNet) with both high compression rate and high accuracy on large scale dataset? We answer this question through a careful analysis of previous work on BinaryNets, in terms of training strategies, regularization, and activation approximation. Our findings first reveal that a low learning rate is highly preferred to avoid frequent sign changes of the weights, which often makes the learning of BinaryNets unstable. Secondly, we propose to use PReLU instead of ReLU in a BinaryNet to conveniently absorb the scale factor for weights to the activation function, which enjoys high computation efficiency for binarized layers while maintains high approximation accuracy. Thirdly, we reveal that instead of imposing L2 regularization, driving all weights to zero which contradicts with the setting of BinaryNets, we introduce a regularization term that encourages the weights to be bipolar. Fourthly, we discover that the failure of binarizing the last layer, which is essential for high compression rate, is due to the improper output range. We propose to use a scale layer to bring it to normal. Last but not least, we propose multiple binarizations to improve the approximation of the activations. The composition of all these enables us to train BinaryNets with both high compression rate and high accuracy, which is strongly supported by our extensive empirical study. Wei Tang 0016, Gang Hua 0001, Liang Wang 0001 |
AAAI | 2 |
| 2017 | StyleBank: An Explicit Representation for Neural Image Style TransferabstractWe propose StyleBank, which is composed of multiple convolution filter banks and each filter bank explicitly represents one style, for neural image style transfer. To transfer an image to a specific style, the corresponding filter bank is operated on top of the intermediate feature embedding produced by a single auto-encoder. The StyleBank and the auto-encoder are jointly learnt, where the learning is conducted in such a way that the auto-encoder does not encode any style information thanks to the flexibility introduced by the explicit filter bank representation. It also enables us to conduct incremental learning to add a new image style by learning a new filter bank while holding the auto-encoder fixed. The explicit style representation along with the flexible network design enables us to fuse styles at not only the image level, but also the region level. Our method is the first style transfer network that links back to traditional texton mapping methods, and hence provides new understanding on neural style transfer. Our method is easy to train, runs in real-time, and produces results that qualitatively better or at least comparable to existing methods. Dongdong Chen 0001, Lu Yuan 0001, Jing Liao 0001, Nenghai Yu, Gang Hua 0001 |
CVPR | 5 |
| 2017 | ER3: A Unified Framework for Event Retrieval, Recognition and RecountingabstractWe develop a unified framework for complex event retrieval, recognition and recounting. The framework is based on a compact video representation that exploits the temporal correlations in image features. Our feature alignment procedure identifies and removes the feature redundancies across frames and outputs an intermediate tensor representation we call video imprint. The video imprint is then fed into a reasoning network, whose attention mechanism parallels that of memory networks used in language modeling. The reasoning network simultaneously recognizes the event category and locates the key pieces of evidence for event recounting. In event retrieval tasks, we show that the compact video representation aggregated from the video imprint achieves significantly better retrieval accuracy compared with existing methods. We also set new state of the art results in event recognition tasks with an additional benefit: The latent structure in our reasoning network highlights the areas of the video imprint and can be directly used for event recounting. As video imprint maps back to locations in the video frames, the network allows not only the identification of key frames but also specific areas inside each frame which are most influential to the decision process. Zhanning Gao, Gang Hua 0001, Dongqing Zhang, Nebojsa Jojic, Le Wang 0003, Jianru Xue, Nanning Zheng 0001 |
CVPR | 2 |
| 2017 | Collaborative Deep Reinforcement Learning for Joint Object SearchabstractWe examine the problem of joint top-down active search of multiple objects under interaction, e.g., person riding a bicycle, cups held by the table, etc. Such objects under interaction often can provide contextual cues to each other to facilitate more efficient search. By treating each detector as an agent, we present the first collaborative multi-agent deep reinforcement learning algorithm to learn the optimal policy for joint active object localization, which effectively exploits such beneficial contextual information. We learn inter-agent communication through cross connections with gates between the Q-networks, which is facilitated by a novel multi-agent deep Q-learning algorithm with joint exploitation sampling. We verify our proposed method on multiple object detection benchmarks. Not only does our model help to improve the performance of state-of-the-art active localization models, it also reveals interesting co-detection patterns that are intuitively interpretable. Bo Xin, Yizhou Wang 0001, Gang Hua 0001 |
CVPR | 4 |
| 2017 | Correlational Gaussian Processes for Cross-Domain Visual RecognitionabstractWe present a probabilistic model that captures higher order co-occurrence statistics for joint visual recognition in a collection of images and across multiple domains. More importantly, we predict the structured output across multiple domains by correlating outputs from the multi-classes Gaussian process classifiers in each individual domain. A set of correlational tensors is adopted to model the relationship within a single domain as well as across multiple domains. This renders it possible to explore a high-order relational model instead of using just a set of pairwise relational models. Such tensor relations are based on both the positive and negative co-occurrences of different categories of visual instances across multi-domains. This is in contrast to most previous models where only pair-wise relationships are explored. We conduct experiments on four challenging image collections. The experimental results clearly demonstrate the efficacy of our proposed model. Chengjiang Long, Gang Hua 0001 |
CVPR | 2 |
| 2017 | Order-Preserving Wasserstein Distance for Sequence MatchingabstractWe present a new distance measure between sequences that can tackle local temporal distortion and periodic sequences with arbitrary starting points. Through viewing the instances of sequences as empirical samples of an unknown distribution, we cast the calculation of the distance between sequences as the optimal transport problem. To preserve the inherent temporal relationships of the instances in sequences, we smooth the optimal transport problem with two novel temporal regularization terms. The inverse difference moment regularization enforces transport with local homogeneous structures, and the KL-divergence with a prior distribution regularization prevents transport between instances with far temporal positions. We show that this problem can be efficiently optimized through the matrix scaling algorithm. Extensive experiments on different datasets with different classifiers show that the proposed distance outperforms the traditional DTW variants and the smoothed optimal transport distance without temporal regularization. Bing Su 0001, Gang Hua 0001 |
CVPR | 2 |
| 2017 | Neural Aggregation Network for Video Face RecognitionabstractThis paper presents a Neural Aggregation Network (NAN) for video face recognition. The network takes a face video or face image set of a person with a variable number of face images as its input, and produces a compact, fixed-dimension feature representation for recognition. The whole network is composed of two modules. The feature embedding module is a deep Convolutional Neural Network (CNN) which maps each face image to a feature vector. The aggregation module consists of two attention blocks which adaptively aggregate the feature vectors to form a single feature inside the convex hull spanned by them. Due to the attention mechanism, the aggregation is invariant to the image order. Our NAN is trained with a standard classification or verification loss without any extra supervision signal, and we found that it automatically learns to advocate high-quality face images while repelling low-quality ones such as blurred, occluded and improperly exposed faces. The experiments on IJB-A, YouTube Face, Celebrity-1000 video face recognition benchmarks show that it consistently outperforms naive aggregation methods and achieves the state-of-the-art accuracy. Jiaolong Yang, Peiran Ren, Dongqing Zhang, Dong Chen 0003, Fang Wen 0001, Hongdong Li, Gang Hua 0001 |
CVPR | 7 |
| 2017 | CVAE-GAN: Fine-Grained Image Generation through Asymmetric TrainingabstractWe present variational generative adversarial networks, a general learning framework that combines a variational auto-encoder with a generative adversarial network, for synthesizing images in fine-grained categories, such as faces of a specific person or objects in a category. Our approach models an image as a composition of label and latent attributes in a probabilistic model. By varying the fine-grained category label fed into the resulting generative model, we can generate images in a specific category with randomly drawn values on a latent attribute vector. Our approach has two novel aspects. First, we adopt a cross entropy loss for the discriminative and classifier network, but a mean discrepancy objective for the generative network. This kind of asymmetric loss function makes the GAN training more stable. Second, we adopt an encoder network to learn the relationship between the latent space and the real image space, and use pairwise feature matching to keep the structure of generated images. We experiment with natural images of faces, flowers, and birds, and demonstrate that the proposed models are capable of generating realistic and diverse samples with fine-grained category labels. We further show that our models can be applied to other tasks, such as image inpainting, super-resolution, and data augmentation for training better face recognition models. Jianmin Bao, Dong Chen 0003, Fang Wen 0001, Houqiang Li, Gang Hua 0001 |
ICCV | 5 |
| 2017 | Coherent Online Video Style TransferabstractTraining a feed-forward network for the fast neural style transfer of images has proven successful, but the naive extension of processing videos frame by frame is prone to producing flickering results. We propose the first end-to-end network for online video style transfer, which generates temporally coherent stylized video sequences in near realtime. Two key ideas include an efficient network by incorporating short-term coherence, and propagating short-term coherence to long-term, which ensures consistency over a longer period of time. Our network can incorporate different image stylization networks and clearly outperforms the per-frame baseline both qualitatively and quantitatively. Moreover, it can achieve visually comparable coherence to optimization-based video style transfer, but is three orders of magnitude faster. Dongdong Chen 0001, Jing Liao 0001, Lu Yuan 0001, Nenghai Yu, Gang Hua 0001 |
ICCV | 5 |
| 2017 | A Generic Deep Architecture for Single Image Reflection Removal and Image SmoothingabstractThis paper proposes a deep neural network structure that exploits edge information in addressing representative low-level vision tasks such as layer separation and image filtering. Unlike most other deep learning strategies applied in this context, our approach tackles these challenging problems by estimating edges and reconstructing images using only cascaded convolutional layers arranged such that no handcrafted or application-specific image-processing components are required. We apply the resulting transferrable pipeline to two different problem domains that are both sensitive to edges, namely, single image reflection removal and image smoothing. For the former, using a mild reflection smoothness assumption and a novel synthetic data generation method that acts as a type of weak supervision, our network is able to solve much more difficult reflection cases that cannot be handled by previous methods. For the latter, we also exceed the state-of-the-art quantitative and qualitative results by wide margins. In all cases, the proposed framework is simple, fast, and easy to transfer across disparate domains. Qingnan Fan, Jiaolong Yang, Gang Hua 0001, Baoquan Chen, David P. Wipf |
ICCV | 3 |
| 2017 | Hierarchical Multimodal LSTM for Dense Visual-Semantic EmbeddingabstractWe address the problem of dense visual-semantic embedding that maps not only full sentences and whole images but also phrases within sentences and salient regions within images into a multimodal embedding space. Such dense embeddings, when applied to the task of image captioning, enable us to produce several region-oriented and detailed phrases rather than just an overview sentence to describe an image. Specifically, we present a hierarchical structured recurrent neural network (RNN), namely Hierarchical Multimodal LSTM (HM-LSTM). Compared with chain structured RNN, our proposed model exploits the hierarchical relations between sentences and phrases, and between whole images and image regions, to jointly establish their representations. Without the need of any supervised labels, our proposed model automatically learns the fine-grained correspondences between phrases and image regions towards the dense embedding. Extensive experiments on several datasets validate the efficacy of our method, which compares favorably with the state-of-the-art methods. Zhenxing Niu, Le Wang 0003, Xinbo Gao 0001, Gang Hua 0001 |
ICCV | 5 |
| 2017 | Green Generative Modeling: Recycling Dirty Data using Recurrent Variational Autoencoders
Yu Wang 0060, Bin Dai 0008, Gang Hua 0001, John A. D. Aston, David P. Wipf |
UAI | 3 |
| 2017 | Exemplar-Guided Similarity Learning on Polynomial Kernel Feature Map for Person Re-identification
Dapeng Chen, Zejian Yuan, Jingdong Wang 0001, Badong Chen, Gang Hua 0001, Nanning Zheng 0001 |
Int. J. Comput. Vis. | 5 |
| 2017 | Multi-Timescale Collaborative TrackingabstractWe present the multi-timescale collaborative tracker for single object tracking. The tracker simultaneously utilizes different types of "forces", namely attraction, repulsion and support, to take advantage of their complementary strengths. We model the three forces via three components that are learned from the sample sets with different timescales. The long-term descriptive component attracts the target sample, while the medium-term discriminative component repulses the target from the background. They are collaborated in the appearance model to benefit each other. The short-term regressive component combines the votes of the auxiliary samples to predict the target's position, forming the context-aware motion model. The appearance model and the motion model collaboratively determine the target state, and the optimal state is estimated by a novel coarse-to-fine search strategy. We have conducted an extensive set of experiments on the standard 50 video benchmark. The results confirm the effectiveness of each component and their collaboration, outperforming current state-of-the-art methods. Dapeng Chen, Zejian Yuan, Gang Hua 0001, Jingdong Wang 0001, Nanning Zheng 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2017 | Video Object Discovery and Co-Segmentation with Extremely Weak SupervisionabstractWe present a spatio-temporal energy minimization formulation for simultaneous video object discovery and co-segmentation across multiple videos containing irrelevant frames. Our approach overcomes a limitation that most existing video co-segmentation methods possess, i.e., they perform poorly when dealing with practical videos in which the target objects are not present in many frames. Our formulation incorporates a spatio-temporal auto-context model, which is combined with appearance modeling for superpixel labeling. The superpixel-level labels are propagated to the frame level through a multiple instance boosting algorithm with spatial reasoning, based on which frames containing the target object are identified. Our method only needs to be bootstrapped with the frame-level labels for a few video frames (e.g., usually 1 to 3) to indicate if they contain the target objects or not. Extensive experiments on four datasets validate the efficacy of our proposed method: 1) object segmentation from a single video on the SegTrack dataset, 2) object co-segmentation from multiple videos on a video co-segmentation dataset, and 3) joint object discovery and co-segmentation from multiple videos containing irrelevant frames on the MOViCS dataset and XJTU-Stevens, a new dataset that we introduce in this paper. The proposed method compares favorably with the state-of-the-art in all of these experiments. Le Wang 0003, Gang Hua 0001, Rahul Sukthankar, Jianru Xue, Zhenxing Niu, Nanning Zheng 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2017 | Visual Tracking via Joint Discriminative Appearance LearningabstractIn this paper, we present a discriminative tracking method based on dictionary learning and support vector machine (SVM) classification, where the dictionary and the classifier are jointly learned within a unified objective function. A discriminative differential tracking method is proposed, which estimates the motion parameters iteratively by the gradient-based method to maximize the SVM classification score, leading the bounding box to move purposively. As the target appearance may change across frames, an online update scheme is exploited, which not only reserves the discriminative information, but also adaptively accounts for the appearance changes in the dynamic scenes. We examine the proposed method on the benchmark challenging image sequences, including heavy occlusion, pose change, illumination variation, and so on. Extensive evaluations demonstrate that the proposed tracker performs favorably against other state-of-the-art algorithms. Fu Li 0003, Huchuan Lu, Gang Hua 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2017 | Visual attribute transfer through deep image analogyabstractWe propose a new technique for visual attribute transfer across images that may have very different appearance but have perceptually similar semantic structure. By visual attribute transfer, we mean transfer of visual information (such as color, tone, texture, and style) from one image to another. For example, one image could be that of a painting or a sketch while the other is a photo of a real scene, and both depict the same type of scene. Our technique finds semantically-meaningful dense correspondences between two input images. To accomplish this, it adapts the notion of "image analogy" [Hertzmann et al. 2001] with features extracted from a Deep Convolutional Neutral Network for matching; we call our technique deep image analogy. A coarse-to-fine strategy is used to compute the nearest-neighbor field for generating the results. We validate the effectiveness of our proposed method in a variety of cases, including style/texture transfer, color/style swap, sketch/painting to photo, and time lapse. Jing Liao 0001, Lu Yuan 0001, Gang Hua 0001, Sing Bing Kang |
ACM Trans. Graph. | 4 |
| 2016 | A Multi-level Contextual Model for Person Recognition in Photo AlbumsabstractIn this work, we present a new framework for person recognition in photo albums that exploits contextual cues at multiple levels, spanning individual persons, individual photos, and photo groups. Through experiments, we show that the information available at each of these distinct contextual levels provides complementary cues as to person identities. At the person level, we leverage clothing and body appearance in addition to facial appearance, and to compensate for instances where the faces are not visible. At the photo level we leverage a learned prior on the joint distribution of identities on the same photo to guide the identity assignments. Going beyond a single photo, we are able to infer natural groupings of photos with shared context in an unsupervised manner. By exploiting this shared contextual information, we are able to reduce the identity search space and exploit higher intra-personal appearance consistency within photo groups. Our new framework enables efficient use of these complementary multi-level contextual cues to improve overall recognition rates on the photo album person recognition task, as demonstrated through state-of-theart results on a challenging public dataset. Our results outperform competing methods by a significant margin, while being computationally efficient and practical in a real world application. Jonathan Brandt, Zhe Lin 0001, Xiaohui Shen, Gang Hua 0001 |
CVPR | 5 |
| 2016 | Ordinal Regression with Multiple Output CNN for Age EstimationabstractTo address the non-stationary property of aging patterns, age estimation can be cast as an ordinal regression problem. However, the processes of extracting features and learning a regression model are often separated and optimized independently in previous work. In this paper, we propose an End-to-End learning approach to address ordinal regression problems using deep Convolutional Neural Network, which could simultaneously conduct feature learning and regression modeling. In particular, an ordinal regression problem is transformed into a series of binary classification sub-problems. And we propose a multiple output CNN learning algorithm to collectively solve these classification sub-problems, so that the correlation between these tasks could be explored. In addition, we publish an Asian Face Age Dataset (AFAD) containing more than 160K facial images with precise age ground-truths, which is the largest public age dataset to date. To the best of our knowledge, this is the first work to address ordinal regression problems by using CNN, and achieves the state-of-the-art performance on both the MORPH and AFAD datasets. Zhenxing Niu, Le Wang 0003, Xinbo Gao 0001, Gang Hua 0001 |
CVPR | 5 |
| 2016 | Supervised Transformer Network for Efficient Face Detection
Dong Chen 0003, Gang Hua 0001, Fang Wen 0001, Jian Sun 0001 |
ECCV (5) | 2 |
| 2016 | Supervised Matrix Factorization for Cross-Modality Hashing
Hong Liu 0009, Rongrong Ji, Yongjian Wu 0001, Gang Hua 0001 |
IJCAI | 4 |
| 2016 | An egocentric computer vision based co-robot wheelchairabstractMotivated by the emerging needs to improve the quality of life for the elderly and disabled individuals who rely on wheelchairs for mobility, and who might have limited or no hand functionality at all, we propose an egocentric computer vision based co-robot wheelchair to enhance their mobility without hand usage. The co-robot wheelchair is built upon a typical commercial power wheelchair. The user can access 360 degrees of motion direction as well as a continuous range of speed without the use of hands via the egocentric computer vision based control we developed. The user wears an egocentric camera and collaborates with the robotic wheelchair by conveying the motion commands with head motions. Compared with previous sip-n-puff, chin-control and tongue-operated solutions to hands-free mobility, this egocentric computer vision based control system provides a more natural human robot interface. Our experiments show that this design is of higher usability and users can quickly learn to control and operate the wheelchair. Besides its convenience in manual navigation, the egocentric camera also supports novel user-robot interaction modes by enabling autonomous navigation towards a detected person or object of interest. User studies demonstrate the usability and efficiency of the proposed egocentric computer vision co-robot wheelchair. Mohammed Kutbi, Changjiang Cai, Philippos Mordohai, Gang Hua 0001 |
IROS | 6 |
| 2016 | A Joint Gaussian Process Model for Active Visual Recognition with Expertise Estimation in Crowdsourcing
Chengjiang Long, Gang Hua 0001, Ashish Kapoor |
Int. J. Comput. Vis. | 2 |
| 2016 | Introduction of New Associate EditorsabstractPresents a listing of the new Associate Editors for this issue of the publication. Nikolaos V. Boulgouris, David Bull 0001, Marco Cagnazzo, Andrea Cavallaro, Gene Cheung, Amit K. Roy-Chowdhury, Pedro Comesaña Alfaro, Sarp Ertürk, Markus Flierl, Gian Luca Foresti, Gang Hua 0001, Zhu Li 0001, Weisi Lin, Siwei Ma 0001, Pramod Kumar Meher, Debargha Mukherjee, Aleksandra Pizurica, Andrea Prati 0001, Paolo Remagnino, Arun Ross, Shin'ichi Satoh 0001, Andreas E. Savakis, Heiko Schwarz, Ling Shao 0001, Shervin Shirmohammadi, Giuseppe Valenzise, Meng Wang 0001, Zhou Wang 0001, Yonggang Wen 0001, Dong Xu 0001, Junsong Yuan 0001, Yuan Yuan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 12 |
| 2016 | Three-Dimensional Traffic Scenes Simulation From Road Image SequencesabstractIn this paper, we present a novel framework to allow users to tour simulated traffic scenes from the first-person view. Constructing 3-D scenes from road image sequences is in general difficult, due to the intrinsic complexity of dynamic road scenes, which are composed of a drastically moving background, not to mention numerous other surrounding vehicles. With the definitions of the traffic scene models, we first introduce the construction process of the simple traffic scenes. After the detection of road boundaries by a semantic fast two-cycle (FTC) level set method, we generate the control points on road sides to construct the “floor-wall” background scene that is subsequently propagated to each frame. Furthermore, we approach the cluttered traffic scenes through a three-component processing pipeline as follows: 1) traffic elements segmentation; 2) background images inpainting; and 3) traffic scenes construction. The traffic elements in the cluttered images are segmented by the semantic FTC level set method first. A Gaussian mixture model is then employed to inpaint the occluded background utilizing the optical flows. The cluttered traffic scenes can be constructed after the segmentation and inpainting components. The foreground polygons such as vehicles and traffic signs are then modeled. Users can change their viewpoints according to their own interpretations. We present the evaluations of each technical component, followed by our findings from comprehensive user studies, which well demonstrate the effectiveness of the proposed framework in delivering good touring experience to users. Yaochen Li, Yuehu Liu, Yuanqi Su, Gang Hua 0001, Nanning Zheng 0001 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2015 | Development of an Effective Method and a Portable Device to Evaluate the Pupillary ReflexabstractExamination of the pupil offers an objective evaluation of visual function as well as the vegetative pathways to the eye. In spite of technological advances and substantial progress in many areas, the routine pupil examination with a conventional light source has undergone no significant changes in the last 100 years. This work proposes the development of an effective and objective method and a portable device to test the direct and consensual pupillary reflex. The first results demonstrate the success of a new device construction and new methodology to record the reflex with differ stimulus, in a situation of complete seal of light. Hedenir Machado Pinheiro, Ronaldo Martins da Costa, Leandro Luís Galdino de Oliveira, Eduardo Nery Rossi Camilo, Gang Hua 0001 |
CBMS | 5 |
| 2015 | Similarity learning on an explicit polynomial kernel feature map for person re-identificationabstractIn this paper, we address the person re-identification problem, discovering the correct matches for a probe person image from a set of gallery person images. We follow the learning-to-rank methodology and learn a similarity function to maximize the difference between the similarity scores of matched and unmatched images for a same person. We introduce at least three contributions to person re-identification. First, we present an explicit polynomial kernel feature map, which is capable of characterizing the similarity information of all pairs of patches between two images, called soft-patch-matching, instead of greedily keeping only the best matched patch, and thus more robust. Second, we introduce a mixture of linear similarity functions that is able to discover different soft-patch-matching patterns. Last, we introduce a negative semi-definite regularization over a subset of the weights in the similarity function, which is motivated by the connection between explicit polynomial kernel feature map and the Mahalanobis distance, as well as the sparsity constraint over the parameters to avoid over-fitting. Experimental results over three public benchmarks demonstrate the superiority of our approach. Dapeng Chen, Zejian Yuan, Gang Hua 0001, Nanning Zheng 0001, Jingdong Wang 0001 |
CVPR | 3 |
| 2015 | Hierarchical-PEP model for real-world face recognitionabstractPose variation remains one of the major factors adversely affect the accuracy of real-world face recognition systems. Inspired by the recently proposed probabilistic elastic part (PEP) model and the success of the deep hierarchical architecture in a number of visual tasks, we propose the Hierarchical-PEP model to approach the unconstrained face recognition problem. We apply the PEP model hierarchically to decompose a face image into face parts at different levels of details to build pose-invariant part-based face representations. Following the hierarchy from bottom-up, we stack the face part representations at each layer, discriminatively reduce its dimensionality, and hence aggregate the face part representations layer-by-layer to build a compact and invariant face representation. The Hierarchical-PEP model exploits the fine-grained structures of the face parts at different levels of details to address the pose variations. It is also guided by supervised information in constructing the face part/face representations. We empirically verify the Hierarchical-PEP model on two public benchmarks (i.e., the LFW and YouTube Faces) and a face recognition challenge (i.e., the PaSC grand challenge) for image-based and video-based face verification. The state-of-the-art performance demonstrates the potential of our method. Gang Hua 0001 |
CVPR | 2 |
| 2015 | A convolutional neural network cascade for face detectionabstractIn real-world face detection, large visual variations, such as those due to pose, expression, and lighting, demand an advanced discriminative model to accurately differentiate faces from the backgrounds. Consequently, effective models for the problem tend to be computationally prohibitive. To address these two conflicting challenges, we propose a cascade architecture built on convolutional neural networks (CNNs) with very powerful discriminative capability, while maintaining high performance. The proposed CNN cascade operates at multiple resolutions, quickly rejects the background regions in the fast low resolution stages, and carefully evaluates a small number of challenging candidates in the last high resolution stage. To improve localization effectiveness, and reduce the number of candidates at later stages, we introduce a CNN-based calibration stage after each of the detection stages in the cascade. The output of each calibration stage is used to adjust the detection window position for input to the subsequent stage. The proposed method runs at 14 FPS on a single CPU core for VGA-resolution images and 100 FPS using a GPU, and achieves state-of-the-art detection performance on two public face detection benchmarks. Zhe Lin 0001, Xiaohui Shen, Jonathan Brandt, Gang Hua 0001 |
CVPR | 5 |
| 2015 | Multi-class Multi-annotator Active Learning with Robust Gaussian Process for Visual RecognitionabstractActive learning is an effective way to relieve the tedious work of manual annotation in many applications of visual recognition. However, less research attention has been focused on multi-class active learning. In this paper, we propose a novel Gaussian process classifier model with multiple annotators for multi-class visual recognition. Expectation propagation (EP) is adopted for efficient approximate Bayesian inference of our probabilistic model for classification. Based on the EP approximation inference, a generalized Expectation Maximization (GEM) algorithm is derived to estimate both the parameters for instances and the quality of each individual annotator. Also, we incorporate the idea of reinforcement learning to actively select both the informative samples and the high-quality annotators, which better explores the trade-off between exploitation and exploration. The experiments clearly demonstrate the efficacy of the proposed model. Chengjiang Long, Gang Hua 0001 |
ICCV | 2 |
| 2015 | Learning Discriminative Reconstructions for Unsupervised Outlier RemovalabstractWe study the problem of automatically removing outliers from noisy data, with application for removing outlier images from an image collection. We address this problem by utilizing the reconstruction errors of an autoencoder. We observe that when data are reconstructed from low-dimensional representations, the inliers and the outliers can be well separated according to their reconstruction errors. Based on this basic observation, we gradually inject discriminative information in the learning process of an autoencoder to make the inliers and the outliers more separable. Experiments on a variety of image datasets validate our approach. Xudong Cao, Fang Wen 0001, Gang Hua 0001, Jian Sun 0001 |
ICCV | 4 |
| 2015 | CANNET: Context aware nonlocal convolutional networks for semantic image segmentationabstractSemantic segmentation has long been a hot topic, most methods are the region based method, which lost connection information to their neighbors. In this paper we propose to encode context information into convolutional networks on this semantic labeling task. Firstly, we propose the nonlocal convolution kernel, which extracts feature from larger neighbor regions without introducing more parameters. Then we build up a context aware module, which takes both local patch and nonlocal neighbor information into account. At last we embed the module into convolutional networks and tested the improvement on benchmark datasets. Lingyan Ran, Yanning Zhang 0001, Gang Hua 0001 |
ICIP | 3 |
| 2015 | Modeling Inter- and Intra-Part Deformations for Object Structure Parsing
Ling Cai 0003, Rongrong Ji, Wei Liu 0005, Gang Hua 0001 |
IJCAI | 4 |
| 2015 | Multi-View Visual Recognition of Imperfect Testing DataabstractA practical yet under-explored problem often encountered by multimedia researchers is the recognition of imperfect testing data, where multiple sensing channels are deployed but interference or transmission distortion corrupts some of them. Typical cases of imperfect testing data include missing features and feature misalignments. To address these challenges, we choose the latent space model and introduce a new similarity learning canonical-correlation analysis (SLCCA) method to capture the semantic consensus between views. The consensus information is preserved by projection matrices learned with modified canonical-correlation analysis (CCA) optimization terms with new, explicit class-similarity constraints. To make it computationally tractable, we propose to combine a practical relaxation and an alternating scheme to solve the optimization problem. Experiments on four challenging multi-view visual recognition datasets demonstrate the efficacy of the proposed method. Qilin Zhang 0004, Gang Hua 0001 |
ACM Multimedia | 2 |
| 2015 | Visual Topic Network: Building better image representations for images in social media
Zhenxing Niu, Gang Hua 0001, Qi Tian 0001, Xinbo Gao 0001 |
Comput. Vis. Image Underst. | 2 |
| 2015 | Guest editorial: selected papers from ICIMCS 2013
Meng Wang 0001, Ke Lu 0002, Gang Hua 0001, Cees Snoek |
Multim. Syst. | 3 |
| 2015 | The VLSI implementation of a high-resolution depth-sensing SoC based on active structured light
Huimin Yao, Chenyang Ge, Gang Hua 0001, Nanning Zheng 0001 |
Mach. Vis. Appl. | 3 |
| 2015 | Topical Video Object Discovery From Key Frames by Modeling Word Co-Occurrence PriorabstractA topical video object refers to an object, that is, frequently highlighted in a video. It could be, e.g., the product logo and the leading actor/actress in a TV commercial. We propose a topic model that incorporates a word co-occurrence prior for efficient discovery of topical video objects from a set of key frames. Previous work using topic models, such as latent Dirichelet allocation (LDA), for video object discovery often takes a bag-of-visual-words representation, which ignored important co-occurrence information among the local features. We show that such data driven co-occurrence information from bottom-up can conveniently be incorporated in LDA with a Gaussian Markov prior, which combines top-down probabilistic topic modeling with bottom-up priors in a unified model. Our experiments on challenging videos demonstrate that the proposed approach can discover different types of topical objects despite variations in scale, view-point, color and lighting changes, or even partial occlusions. The efficacy of the co-occurrence prior is clearly demonstrated when compared with topic models without such priors. Gangqiang Zhao, Junsong Yuan 0001, Gang Hua 0001 |
IEEE Trans. Image Process. | 3 |
| 2014 | Eigen-PEP for Video Face Recognition
Gang Hua 0001, Xiaohui Shen, Zhe Lin 0001, Jonathan Brandt |
ACCV (3) | 2 |
| 2014 | Accurate Object Detection with Location Relaxation and Regionlets Re-localization
Chengjiang Long, Xiaoyu Wang 0002, Gang Hua 0001, Ming Yang 0007, Yuanqing Lin |
ACCV (1) | 3 |
| 2014 | Can Visual Recognition Benefit from Auxiliary Information in Training?
Qilin Zhang 0004, Gang Hua 0001, Wei Liu 0005, Zicheng Liu 0001, Zhengyou Zhang |
ACCV (1) | 2 |
| 2014 | Efficient Boosted Exemplar-Based Face DetectionabstractDespite the fact that face detection has been studied intensively over the past several decades, the problem is still not completely solved. Challenging conditions, such as extreme pose, lighting, and occlusion, have historically hampered traditional, model-based methods. In contrast, exemplar-based face detection has been shown to be effective, even under these challenging conditions, primarily because a large exemplar database is leveraged to cover all possible visual variations. However, relying heavily on a large exemplar database to deal with the face appearance variations makes the detector impractical due to the high space and time complexity. We construct an efficient boosted exemplar-based face detector which overcomes the defect of the previous work by being faster, more memory efficient, and more accurate. In our method, exemplars as weak detectors are discriminatively trained and selectively assembled in the boosting framework which largely reduces the number of required exemplars. Notably, we propose to include non-face images as negative exemplars to actively suppress false detections to further improve the detection accuracy. We verify our approach over two public face detection benchmarks and one personal photo album, and achieve significant improvement over the state-of-the-art algorithms in terms of both accuracy and efficiency. Zhe Lin 0001, Jonathan Brandt, Xiaohui Shen, Gang Hua 0001 |
CVPR | 5 |
| 2014 | Unsupervised One-Class Learning for Automatic Outlier RemovalabstractOutliers are pervasive in many computer vision and pattern recognition problems. Automatically eliminating outliers scattering among practical data collections becomes increasingly important, especially for Internet inspired vision applications. In this paper, we propose a novel one-class learning approach which is robust to contamination of input training data and able to discover the outliers that corrupt one class of data source. Our approach works under a fully unsupervised manner, differing from traditional one-class learning supervised by known positive labels. By design, our approach optimizes a kernel-based max-margin objective which jointly learns a large margin one-class classifier and a soft label assignment for inliers and outliers. An alternating optimization algorithm is then designed to iteratively refine the classifier and the labeling, achieving a provably convergent solution in only a few iterations. Extensive experiments conducted on four image datasets in the presence of artificial and real-world outliers demonstrate that the proposed approach is considerably superior to the state-of-the-arts in obliterating outliers from contaminated one class of images, exhibiting strong robustness at a high outlier proportion up to 60%. Wei Liu 0005, Gang Hua 0001, John R. Smith |
CVPR | 2 |
| 2014 | Hash-SVM: Scalable Kernel Machines for Large-Scale Visual ClassificationabstractThis paper presents a novel algorithm which uses compact hash bits to greatly improve the efficiency of non-linear kernel SVM in very large scale visual classification problems. Our key idea is to represent each sample with compact hash bits, over which an inner product is defined to serve as the surrogate of the original nonlinear kernels. Then the problem of solving the nonlinear SVM can be transformed into solving a linear SVM over the hash bits. The proposed Hash-SVM enjoys dramatic storage cost reduction owing to the compact binary representation, as well as a (sub-)linear training complexity via linear SVM. As a critical component of Hash-SVM, we propose a novel hashing scheme for arbitrary non-linear kernels via random subspace projection in reproducing kernel Hilbert space. Our comprehensive analysis reveals a well behaved theoretic bound of the deviation between the proposed hashing-based kernel approximation and the original kernel function. We also derive requirements on the hash bits for achieving a satisfactory accuracy level. Several experiments on large-scale visual classification benchmarks are conducted, including one with over 1 million images. The results show that Hash-SVM greatly reduces the computational complexity (more than ten times faster in many cases) while keeping comparable accuracies. Yadong Mu, Gang Hua 0001, Wei Fan 0001, Shih-Fu Chang |
CVPR | 2 |
| 2014 | Semi-supervised Relational Topic Model for Weakly Annotated Image Recognition in Social MediaabstractIn this paper, we address the problem of recognizing images with weakly annotated text tags. Most previous work either cannot be applied to the scenarios where the tags are loosely related to the images, or simply take a pre-fusion at the feature level or a post-fusion at the decision level to combine the visual and textual content. Instead, we first encode the text tags as the relations among the images, and then propose a semi-supervised relational topic model (ss-RTM) to explicitly model the image content and their relations. In such way, we can efficiently leverage the loosely related tags, and build an intermediate level representation for a collection of weakly annotated images. The intermediate level representation can be regarded as a mid-level fusion of the visual and textual content, which is able to explicitly model their intrinsic relationships. Moreover, image category labels are also modeled in the ss-RTM, and recognition can be conducted without training an additional discriminative classifier. Our extensive experiments on social multimedia datasets (images+tags) demonstrated the advantages of the proposed model. Zhenxing Niu, Gang Hua 0001, Xinbo Gao 0001, Qi Tian 0001 |
CVPR | 2 |
| 2014 | Description-Discrimination Collaborative Tracking
Dapeng Chen, Zejian Yuan, Gang Hua 0001, Yang Wu 0001, Nanning Zheng 0001 |
ECCV (1) | 3 |
| 2014 | Video Object Discovery and Co-segmentation with Extremely Weak Supervision
Le Wang 0003, Gang Hua 0001, Rahul Sukthankar, Jianru Xue, Nanning Zheng 0001 |
ECCV (4) | 2 |
| 2014 | The IJCB 2014 PaSC video face and person recognition competitionabstractThe Point-and-Shoot Face Recognition Challenge (PaSC) is a performance evaluation challenge including 1401 videos of 265 people acquired with handheld cameras and depicting people engaged in activities with non-frontal head pose. This report summarizes the results from a competition using this challenge problem. In the Video-to-video Experiment a person in a query video is recognized by comparing the query video to a set of target videos. Both target and query videos are drawn from the same pool of 1401 videos. In the Still-to-video Experiment the person in a query video is to be recognized by comparing the query video to a larger target set consisting of still images. Algorithm performance is characterized by verification rate at a false accept rate of 0.01 and associated receiver operating characteristic (ROC) curves. Participants were provided eye coordinates for video frames. Results were submitted by 4 institutions: (i) Advanced Digital Science Center, Singapore; (ii) CPqD, Brasil; (iii) Stevens Institute of Technology, USA; and (iv) University of Ljubljana, Slovenia. Most competitors demonstrated video face recognition performance superior to the baseline provided with PaSC. The results represent the best performance to date on the handheld video portion of the PaSC. J. Ross Beveridge, Hao Zhang 0013, Patrick J. Flynn, Yooyoung Lee, Venice Erin Liong, Jiwen Lu, Marcus A. Angeloni, Tiago de Freitas Pereira, Gang Hua 0001, Vitomir Struc, Janez Krizaj, P. Jonathon Phillips |
IJCB | 10 |
| 2014 | ObjectPatchNet: Towards scalable and semantic image annotation and retrieval
Shiliang Zhang, Qi Tian 0001, Gang Hua 0001, Qingming Huang, Wen Gao 0001 |
Comput. Vis. Image Underst. | 3 |
| 2014 | Hyperspectral Image Classification Through Bilayer Graph-Based LearningabstractHyperspectral image classification with limited number of labeled pixels is a challenging task. In this paper, we propose a bilayer graph-based learning framework to address this problem. For graph-based classification, how to establish the neighboring relationship among the pixels from the high dimensional features is the key toward a successful classification. Our graph learning algorithm contains two layers. The first-layer constructs a simple graph, where each vertex denotes one pixel and the edge weight encodes the similarity between two pixels. Unsupervised learning is then conducted to estimate the grouping relations among different pixels. These relations are subsequently fed into the second layer to form a hypergraph structure, on top of which, semisupervised transductive learning is conducted to obtain the final classification results. Our experiments on three data sets demonstrate the merits of our proposed approach, which compares favorably with state of the art. Yue Gao 0002, Rongrong Ji, Peng Cui 0001, Qionghai Dai, Gang Hua 0001 |
IEEE Trans. Image Process. | 5 |
| 2014 | Weakly Supervised Visual Dictionary Learning by Harnessing Image AttributesabstractBag-of-features (BoFs) representation has been extensively applied to deal with various computer vision applications. To extract discriminative and descriptive BoF, one important step is to learn a good dictionary to minimize the quantization loss between local features and codewords. While most existing visual dictionary learning approaches are engaged with unsupervised feature quantization, the latest trend has turned to supervised learning by harnessing the semantic labels of images or regions. However, such labels are typically too expensive to acquire, which restricts the scalability of supervised dictionary learning approaches. In this paper, we propose to leverage image attributes to weakly supervise the dictionary learning procedure without requiring any actual labels. As a key contribution, our approach establishes a generative hidden Markov random field (HMRF), which models the quantized codewords as the observed states and the image attributes as the hidden states, respectively. Dictionary learning is then performed by supervised grouping the observed states, where the supervised information is stemmed from the hidden states of the HMRF. In such a way, the proposed dictionary learning approach incorporates the image attributes to learn a semantic-preserving BoF representation without any genuine supervision. Experiments in large-scale image retrieval and classification tasks corroborate that our approach significantly outperforms the state-of-the-art unsupervised dictionary learning approaches. Yue Gao 0002, Rongrong Ji, Wei Liu 0005, Qionghai Dai, Gang Hua 0001 |
IEEE Trans. Image Process. | 5 |
| 2014 | Joint Segmentation and Recognition of Categorized Objects From Noisy Web Image CollectionabstractThe segmentation of categorized objects addresses the problem of joint segmentation of a single category of object across a collection of images, where categorized objects are referred to objects in the same category. Most existing methods of segmentation of categorized objects made the assumption that all images in the given image collection contain the target object. In other words, the given image collection is noise free. Therefore, they may not work well when there are some noisy images which are not in the same category, such as those image collections gathered by a text query from modern image search engines. To overcome this limitation, we propose a method for automatic segmentation and recognition of categorized objects from noisy Web image collections. This is achieved by cotraining an automatic object segmentation algorithm that operates directly on a collection of images, and an object category recognition algorithm that identifies which images contain the target object. The object segmentation algorithm is trained on a subset of images from the given image collection which are recognized to contain the target object with high confidence, while training the object category recognition model is guided by the intermediate segmentation results obtained from the object segmentation algorithm. This way, our co-training algorithm automatically identifies the set of true positives in the noisy Web image collection, and simultaneously extracts the target objects from all the identified images. Extensive experiments validated the efficacy of our proposed approach on four datasets: 1) the Weizmann horse dataset, 2) the MSRC object category dataset, 3) the iCoseg dataset, and 4) a new 30-categories dataset including 15,634 Web images with both hand-annotated category labels and ground truth segmentation labels. It is shown that our method compares favorably with the state-of-the-art, and has the ability to deal with noisy image collections. Le Wang 0003, Gang Hua 0001, Jianru Xue, Zhanning Gao, Nanning Zheng 0001 |
IEEE Trans. Image Process. | 2 |
| 2013 | Probabilistic Elastic Matching for Pose Variant Face VerificationabstractPose variation remains to be a major challenge for real-world face recognition. We approach this problem through a probabilistic elastic matching method. We take a part based representation by extracting local features (e.g., LBP or SIFT) from densely sampled multi-scale image patches. By augmenting each feature with its location, a Gaussian mixture model (GMM) is trained to capture the spatial-appearance distribution of all face images in the training corpus. Each mixture component of the GMM is confined to be a spherical Gaussian to balance the influence of the appearance and the location terms. Each Gaussian component builds correspondence of a pair of features to be matched between two faces/face tracks. For face verification, we train an SVM on the vector concatenating the difference vectors of all the feature pairs to decide if a pair of faces/face tracks is matched or not. We further propose a joint Bayesian adaptation algorithm to adapt the universally trained GMM to better model the pose variations between the target pair of faces/face tracks, which consistently improves face verification accuracy. Our experiments show that our method outperforms the state-of-the-art in the most restricted protocol on Labeled Face in the Wild (LFW) and the YouTube video face database by a significant margin. Gang Hua 0001, Zhe Lin 0001, Jonathan Brandt, Jianchao Yang |
CVPR | 2 |
| 2013 | Topical Video Object Discovery from Key Frames by Modeling Word Co-occurrence PriorabstractA topical video object refers to an object that is frequently highlighted in a video. It could be, e.g., the product logo and the leading actor/actress in a TV commercial. We propose a topic model that incorporates a word co-occurrence prior for efficient discovery of topical video objects from a set of key frames. Previous work using topic models, such as Latent Dirichelet Allocation (LDA), for video object discovery often takes a bag-of-visual-words representation, which ignored important co-occurrence information among the local features. We show that such data driven co-occurrence information from bottom-up can conveniently be incorporated in LDA with a Gaussian Markov prior, which combines top down probabilistic topic modeling with bottom up priors in a unified model. Our experiments on challenging videos demonstrate that the proposed approach can discover different types of topical objects despite variations in scale, view-point, color and lighting changes, or even partial occlusions. The efficacy of the co-occurrence prior is clearly demonstrated when comparing with topic models without such priors. Gangqiang Zhao, Junsong Yuan 0001, Gang Hua 0001 |
CVPR | 3 |
| 2013 | Collaborative Active Learning of a Kernel Machine Ensemble for RecognitionabstractActive learning is an effective way of engaging users to interactively train models for visual recognition. The vast majority of previous works, if not all of them, focused on active learning with a single human oracle. The problem of active learning with multiple oracles in a collaborative setting has not been well explored. Moreover, most of the previous works assume that the labels provided by the human oracles are noise free, which may often be violated in reality. We present a collaborative computational model for active learning with multiple human oracles. It leads to not only an ensemble kernel machine that is robust to label noises, but also a principled label quality measure to online detect irresponsible labelers. Instead of running independent active learning processes for each individual human oracle, our model captures the inherent correlations among the labelers through shared data among them. Our simulation experiments and experiments with real crowd-sourced noisy labels demonstrated the efficacy of our model. Gang Hua 0001, Chengjiang Long, Ming Yang 0007, Yan Gao 0003 |
ICCV | 1 |
| 2013 | Probabilistic Elastic Part Model for Unsupervised Face Detector AdaptationabstractWe propose an unsupervised detector adaptation algorithm to adapt any offline trained face detector to a specific collection of images, and hence achieve better accuracy. The core of our detector adaptation algorithm is a probabilistic elastic part (PEP) model, which is offline trained with a set of face examples. It produces a statistically aligned part based face representation, namely the PEP representation. To adapt a general face detector to a collection of images, we compute the PEP representations of the candidate detections from the general face detector, and then train a discriminative classifier with the top positives and negatives. Then we re-rank all the candidate detections with this classifier. This way, a face detector tailored to the statistics of the specific image collection is adapted from the original detector. We present extensive results on three datasets with two state-of-the-art face detectors. The significant improvement of detection accuracy over these state of-the-art face detectors strongly demonstrates the efficacy of the proposed face detector adaptation algorithm. Gang Hua 0001, Zhe Lin 0001, Jonathan Brandt, Jianchao Yang |
ICCV | 2 |
| 2013 | Active Visual Recognition with Expertise Estimation in CrowdsourcingabstractWe present a noise resilient probabilistic model for active learning of a Gaussian process classifier from crowds, i.e., a set of noisy labelers. It explicitly models both the overall label noises and the expertise level of each individual labeler in two levels of flip models. Expectation propagation is adopted for efficient approximate Bayesian inference of our probabilistic model for classification, based on which, a generalized EM algorithm is derived to estimate both the global label noise and the expertise of each individual labeler. The probabilistic nature of our model immediately allows the adoption of the prediction entropy and estimated expertise for active selection of data sample to be labeled, and active selection of high quality labelers to label the data, respectively. We apply the proposed model for three visual recognition tasks, i.e., object category recognition, gender recognition, and multi-modal activity recognition, on three datasets with real crowd-sourced labels from Amazon Mechanical Turk. The experiments clearly demonstrated the efficacy of the proposed model. Chengjiang Long, Gang Hua 0001, Ashish Kapoor |
ICCV | 2 |
| 2013 | Large-scale video event classification using dynamic temporal pyramid matching of visual semanticsabstractVideo event classification and retrieval has recently emerged as a challenging research topic. In addition to the variation in appearance of visual content and the large scale of the collections to be analyzed, this domain presents new and unique challenges in the modeling of the explicit temporal structure and implicit temporal trends of content within the video events. In this study, we present a technique for video event classification that captures temporal information over semantics using a scalable and efficient modeling scheme. An architecture for partitioning videos into a linear temporal pyramid, using segments of equal length and segments determined by the patterns of the underlying data, is applied over a rich underlying semantic description at the frame level using a taxonomy of nearly 1000 concepts containing 500,000 training images. Forward model selection with data bagging is used to prune the space of temporal features and data for efficiency. The system is implemented in the Hadoop Map-Reduce environment for arbitrary scalability. Our method is applied to the TRECVID Multimedia Event Detection 2012 task. Results demonstrate a significant boost in performance of over 50%, in terms of mean average precision, compared to common max or average pooling, and 17.7% compared to more complex pooling strategies that ignore temporal content. Noel Codella, Gang Hua 0001, Liangliang Cao, Michele Merler, Leiguang Gong, Matthew L. Hill, John R. Smith |
ICIP | 2 |
| 2013 | Semi-Supervised Learning with Manifold Fitted Graphs
Tongtao Zhang, Rongrong Ji, Wei Liu 0005, Dacheng Tao, Gang Hua 0001 |
IJCAI | 5 |
| 2013 | Automatic salient object extraction with contextual cue and its applications to recognition and alpha matting
Jianru Xue, Le Wang 0003, Nanning Zheng 0001, Gang Hua 0001 |
Pattern Recognit. | 4 |
| 2013 | Introduction to the special section of best papers of ACM multimedia 2012abstractNo abstract available. Ioannis Kompatsiaris, Wenjun Zeng 0001, Gang Hua 0001, Liangliang Cao |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2012 | Context aware topic model for scene recognitionabstractWe present a discriminative latent topic model for scene recognition. The capacity of our model is originated from the modeling of two types of visual contexts, i.e., the category specific global spatial layout of different scene elements, and the reinforcement of the visual coherence in uniform local regions. In contrast, most previous methods for scene recognition either only modeled one of these two visual contexts, or just totally ignored both of them. We cast these two coupled visual contexts in a discriminative Latent Dirichlet Allocation framework, namely context aware topic model. Then scene recognition is achieved by Bayesian inference given a target image. Our experiments on several scene recognition benchmarks clearly demonstrated the advantages of the proposed model. Zhenxing Niu, Gang Hua 0001, Xinbo Gao 0001, Qi Tian 0001 |
CVPR | 2 |
| 2012 | Detection by detections: Non-parametric detector adaptation for a videoabstractWe propose an approach to improving the detection results of a generic offline trained detector on a specific video. Our method does not leverage visual tracking as most detection by tracking methods do. Instead, the proposed detection by detections approach can serve as a more confident initialization for detection by tracking methods. Different from other supervised detector adaptation methods, we constrain the task to videos and no supervised labels for the target video are required for the adaptation; we intend to fill the gap between detection by tracking and pure detection by frames. As a non-parametric detector adaptation method, confident detections are collected to re-rank and to group other detections. We focus on methods with high precision detection results since it is necessitated in real application. Extensive experiments with two state-of-the-art detectors demonstrate the efficacy of our approach. Xiaoyu Wang 0002, Gang Hua 0001, Tony X. Han |
CVPR | 2 |
| 2012 | Scene Aligned Pooling for Complex Video Recognition
Liangliang Cao, Yadong Mu, Apostol Natsev, Shih-Fu Chang, Gang Hua 0001, John R. Smith |
ECCV (2) | 5 |
| 2012 | Multi-scale shared features for cascade object detectionabstractWe introduce an efficient computational framework to extract multi-scale feature descriptors. The framework is based on sharing of descriptor elements across the image and scale space to minimize redundant computation. Any type of local patch or grid-based features can be computed through this framework for capturing coarse-to-fine object appearances. We apply it to human detection by boosting a strong soft cascade classifier. Our experiments demonstrate that the proposed descriptors achieve superior performance both in computational efficiency and detection accuracy. Zhe Lin 0001, Gang Hua 0001, Larry Davis 0001 |
ICIP | 2 |
| 2012 | Video Event Detection Using Temporal Pyramids of Visual Semantics with Kernel Optimization and Model Subspace BoostingabstractIn this study, we present a system for video event classification that generates a temporal pyramid of static visual semantics using minimum-value, maximum-value, and average-value aggregation techniques. Kernel optimization and model subspace boosting are then applied to customize the pyramid for each event. SVM models are independently trained for each level in the pyramid using kernel selection according to 3-fold cross-validation. Kernels that both enforce static temporal order and permit temporal alignment are evaluated. Model subspace boosting is used to select the best combination of pyramid levels and aggregation techniques for each event. The NIST TRECVID Multimedia Event Detection (MED) 2011 dataset was used for evaluation. Results demonstrate that kernel optimizations using both temporally static and dynamic kernels together achieves better performance than any one particular method alone. In addition, model sub-space boosting reduces the size of the model by 80%, while maintaining 96% of the performance gain. Noel Codella, Apostol Natsev, Gang Hua 0001, Matthew L. Hill, Liangliang Cao, Leiguang Gong, John R. Smith |
ICME | 3 |
| 2012 | Concurrent segmentation of categorized objects from an image collection
Le Wang 0003, Jianru Xue, Nanning Zheng 0001, Gang Hua 0001 |
ICPR | 4 |
| 2012 | Introduction to the Special Issue on Mobile Vision
Gang Hua 0001, Yun Fu 0001, Matthew Turk 0001, Marc Pollefeys, Zhengyou Zhang |
Int. J. Comput. Vis. | 1 |
| 2012 | Dynamic hand gesture recognition: An exemplar-based approach from motion divergence fields
Xiaohui Shen, Gang Hua 0001, Lance Williams, Ying Wu 0001 |
Image Vis. Comput. | 2 |
| 2012 | Semantic Model Vectors for Complex Video Event RecognitionabstractWe propose semantic model vectors, an intermediate level semantic representation, as a basis for modeling and detecting complex events in unconstrained real-world videos, such as those from YouTube. The semantic model vectors are extracted using a set of discriminative semantic classifiers, each being an ensemble of SVM models trained from thousands of labeled web images, for a total of 280 generic concepts. Our study reveals that the proposed semantic model vectors representation outperforms-and is complementary to-other low-level visual descriptors for video event modeling. We hence present an end-to-end video event detection system, which combines semantic model vectors with other static or dynamic visual descriptors, extracted at the frame, segment, or full clip level. We perform a comprehensive empirical study on the 2010 TRECVID Multimedia Event Detection task (http://www.nist.gov/itl/iad/mig/med10.cfm), which validates the semantic model vectors representation not only as the best individual descriptor, outperforming state-of-the-art global and local static features as well as spatio-temporal HOG and HOF descriptors, but also as the most compact. We also study early and late feature fusion across the various approaches, leading to a 15% performance boost and an overall system performance of 0.46 mean average precision. In order to promote further research in this direction, we made our semantic model vectors for the TRECVID MED 2010 set publicly available for the community to use (http://www1.cs.columbia.edu/~mmerler/SMV.html). Michele Merler, Bert Huang, Lexing Xie, Gang Hua 0001, Apostol Natsev |
IEEE Trans. Multim. | 4 |
| 2012 | Introduction to the special section of best papers of ACM multimedia 2011abstractNo abstract available. Daniel Gatica-Perez, Gang Hua 0001, Wei Tsang Ooi, Pål Halvorsen |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2011 | Spatial-DiscLDA for visual recognitionabstractTopic models such as pLSA, LDA and their variants have been widely adopted for visual recognition. However, most of the adopted models, if not all, are unsupervised, which neglected the valuable supervised labels during model training. In this paper, we exploit recent advancement in supervised topic modeling, more particularly, the DiscLDA model for object recognition. We extend it to a part based visual representation to automatically identify and model different object parts. We call the proposed model as Spatial-DiscLDA (S-DiscLDA). It models the appearances and locations of the object parts simultaneously, which also takes the supervised labels into consideration. It can be directly used as a classifier to recognize the object. This is performed by an approximate inference algorithm based on Gibbs sampling and bridge sampling methods. We examine the performance of our model by comparing its performance with another supervised topic model on two scene category datasets, i.e., LabelMe and UIUC-sport dataset. We also compare our approach with other approaches which model spatial structures of visual features on the popular Caltech-4 dataset. The experimental results illustrate that it provides competitive performance. Zhenxing Niu, Gang Hua 0001, Xinbo Gao 0001, Qi Tian 0001 |
CVPR | 2 |
| 2011 | Motion divergence fields for dynamic hand gesture recognitionabstractAlthough it is in general difficult to track articulated hand motion, exemplar-based approaches provide a robust solution for hand gesture recognition. Presumably, a rich set of dynamic hand gestures are needed for a meaningful recognition system. How to build the visual representation for the motion patterns is the key for scalable recognition. We propose a novel representation based on the divergence map of the gestural motion field, which transforms motion patterns into spatial patterns. Given the motion divergence maps, we leverage modern image feature detectors to extract salient spatial patterns, such as Maximum Stable Extremal Regions (MSER). A local descriptor is extracted from each region to capture the local motion pattern. The descriptors from gesture exemplars are subsequently indexed using a pre-trained vocabulary tree. New gestures are then matched efficiently with the database gestures with a TF-IDF scheme. Our extensive experiments on a large hand gesture database with 10 categories and 1050 video samples validate the efficacy of the extracted motion patterns for gesture recognition. The proposed approach achieves an overall recognition rate of 97.62%, while the average recognition time is only 34.53 ms. Xiaohui Shen, Gang Hua 0001, Lance Williams, Ying Wu 0001 |
FG | 2 |
| 2011 | What characterizes a shadow boundary under the sun and sky?abstractDespite decades of study, robust shadow detection remains difficult, especially within a single color image. We describe a new approach to detect shadow boundaries in images of outdoor scenes lit only by the sun and sky. The method first extracts visual features of candidate edges that are motivated by physical models of illumination and occluders. We feed these features into a Support Vector Machine (SVM) that was trained to discriminate between most-likely shadow-edge candidates and less-likely ones. Finally, we connect edges to help reject non-shadow edge candidates, and to encourage closed, connected shadow boundaries. On benchmark shadow-edge data sets from Lalonde et al. and Zhu et al., our method showed substantial improvements when compared to other recent shadow-detection methods based on statistical learning. Xiang Huang 0006, Gang Hua 0001, Jack Tumblin, Lance Williams |
ICCV | 2 |
| 2011 | Automatic salient object extraction with contextual cueabstractWe present a method for automatically extracting salient object from a single image, which is cast in an energy minimization framework. Unlike most previous methods that only leverage appearance cues, we employ an auto-context cue as a complementary data term. Benefitting from a generic saliency model for bootstrapping, the segmentation of the salient object and the learning of the auto-context model are iteratively performed without any user intervention. Upon convergence, we obtain not only a clear separation of the salient object, but also an auto-context classifier which can be used to recognize the same type of object in other images. Our experiments on four benchmarks demonstrated the efficacy of the added contextual cue. It is shown that our method compares favorably with the state-of-the-art, some of which even embraced user interactions. Le Wang 0003, Jianru Xue, Nanning Zheng 0001, Gang Hua 0001 |
ICCV | 4 |
| 2011 | Modeling spatial and semantic cues for large-scale near-duplicated image retrieval
Shiliang Zhang, Qi Tian 0001, Gang Hua 0001, Wengang Zhou 0001, Qingming Huang, Houqiang Li, Wen Gao 0001 |
Comput. Vis. Image Underst. | 3 |
| 2011 | Discriminative Learning of Local Image DescriptorsabstractIn this paper, we explore methods for learning local image descriptors from training data. We describe a set of building blocks for constructing descriptors which can be combined together and jointly optimized so as to minimize the error of a nearest-neighbor classifier. We consider both linear and nonlinear transforms with dimensionality reduction, and make use of discriminant learning techniques such as Linear Discriminant Analysis (LDA) and Powell minimization to solve for the parameters. Using these techniques, we obtain descriptors that exceed state-of-the-art performance with low dimensionality. In addition to new experiments and recommendations for descriptor learning, we are also making available a new and realistic ground truth data set based on multiview stereo data. Matthew A. Brown, Gang Hua 0001, Simon A. J. Winder |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2011 | Introduction to the Special Section on Real-World Face RecognitionabstractThe motivations for organizing this special section were to better address the challenges of face recognition in real-world scenarios, to promote systematic research and evaluation of promising methods and systems, to provide a snapshot of where we are in this domain, and to stimulate discussion about future directions. We solicited original contributions of research on all aspects of real-world face recognition, including: the design of robust face similarity features and metrics; robust face clustering and sorting algorithms; novel user interaction models and face recognition algorithms for face tagging; novel applications of web face recognition; novel computational paradigms for face recognition; challenges in large scale face recognition tasks, e.g., on the Internet; face recognition with contextual information; face recognition benchmarks and evaluation methodology for moderately controlled or uncontrolled environments; and video face recognition. We received 42 original submissions, four of which were rejected without review; the other 38 papers entered the normal review process. Each paper was reviewed by three reviewers who are experts in their respective topics. More than 100 expert reviewers have been involved in the review process. The papers were equally distributed among the guest editors. A final decision for each paper was made by at least two guest editors assigned to it. To avoid conflict of interest, no guest editor submitted any papers to this special section. Gang Hua 0001, Ming-Hsuan Yang 0001, Erik G. Learned-Miller, Yi Ma 0001, Matthew Turk 0001, David J. Kriegman, Thomas S. Huang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2011 | Generating Descriptive Visual Words and Visual Phrases for Large-Scale Image ApplicationsabstractBag-of-visual Words (BoWs) representation has been applied for various problems in the fields of multimedia and computer vision. The basic idea is to represent images as visual documents composed of repeatable and distinctive visual elements, which are comparable to the text words. Notwithstanding its great success and wide adoption, visual vocabulary created from single-image local descriptors is often shown to be not as effective as desired. In this paper, descriptive visual words (DVWs) and descriptive visual phrases (DVPs) are proposed as the visual correspondences to text words and phrases, where visual phrases refer to the frequently co-occurring visual word pairs. Since images are the carriers of visual objects and scenes, a descriptive visual element set can be composed by the visual words and their combinations which are effective in representing certain visual objects or scenes. Based on this idea, a general framework is proposed for generating DVWs and DVPs for image applications. In a large-scale image database containing 1506 object and scene categories, the visual words and visual word pairs descriptive to certain objects or scenes are identified and collected as the DVWs and DVPs. Experiments show that the DVWs and DVPs are informative and descriptive and, thus, are more comparable with the text words than the classic visual words. We apply the identified DVWs and DVPs in several applications including large-scale near-duplicated image retrieval, image search re-ranking, and object recognition. The combination of DVW and DVP performs better than the state of the art in large-scale near-duplicated image retrieval in terms of accuracy, efficiency and memory consumption. The proposed image search re-ranking algorithm: DWPRank outperforms the state-of-the-art algorithm by 12.4% in mean average precision and about 11 times faster in efficiency. Shiliang Zhang, Qi Tian 0001, Gang Hua 0001, Qingming Huang, Wen Gao 0001 |
IEEE Trans. Image Process. | 3 |
| 2010 | Interest seam imageabstractWe propose interest seam image, an efficient visual synopsis for video. To extract an interest seam image, a spatiotemporal energy map is constructed for the target video shot. Then an optimal seam which encompasses the highest energy is identified by an efficient dynamic programming algorithm. The optimal seam is used to extract a seam of pixels from each video frame to form one column of an image, based on which an interest seam image is finally composited. The interest seam image is efficient both in terms of computation and memory cost. Therefore it is able to power a wide variety of web-scale video content analysis applications, such as near duplicate video clip search, video genre recognition and classification, as well as video clustering, etc. The representation capacity of the proposed interest seam image is demonstrated in a large scale video retrieval task. Its advantages are clearly exhibited when compared with previous works, as reported in our experiments. Gang Hua 0001, Lei Zhang 0001, Harry Shum |
CVPR | 2 |
| 2010 | Joint People, Event, and Location Recognition in Personal Photo Collections Using Cross-Domain Context
Dahua Lin, Ashish Kapoor, Gang Hua 0001, Simon Baker |
ECCV (1) | 3 |
| 2010 | Discriminative Tracking by Metric Learning
Xiaoyu Wang 0002, Gang Hua 0001, Tony X. Han |
ECCV (3) | 2 |
| 2010 | A nonnegative sparsity induced similarity measure with application to cluster analysis of spam imagesabstractImage spam is an email spam that embeds text content into graphical images to bypass traditional spam filters. The majority of previous approaches focus on filtering image spam from client side. To effectively detect the attack activities of the spammers and fast trace back the spam sources, it is also essential to employ cluster analysis to comprehensively filter the image emails on the server side. In this paper, we present a nonnegative sparsity induced similarity measure for cluster analysis of spam images. This similarity measure is based on an assumption that a spam image should be represented well by the nonnegative linear combination of a small number of spam images in the same cluster. It is due to the observation that spammers generate large number of varieties from a single image source with different image processing and manipulation techniques. Experiments on a spam image dataset collected from our department email server demonstrated the advantages of the proposed approach. Yan Gao 0003, Alok N. Choudhary, Gang Hua 0001 |
ICASSP | 3 |
| 2010 | ACM workshop on mobile cloud media computingabstractSmart mobile devices such as camera phones typically will be carried by people all the time. These devices are true "multimedia" devices that acquire, process, transmit and present text, image, video and audio data. However, due to the limitations in hardware and networking, multimedia applications and systems have not been adequately supported on mobile devices. With the recent developments mobile hardware, wireless network, and cloud computing, it is now the prime time for us to realize intelligent mobile device centered multimedia applications with the support of a cloud computing platform. The focus of this workshop is on exploring challenges and opportunities of intelligent multimedia technologies, applications and systems on mobile devices, especially when a media cloud computing platform can be appropriately leveraged. Xian-Sheng Hua 0001, Gang Hua 0001, Chang Wen Chen |
ACM Multimedia | 2 |
| 2010 | Building contextual visual vocabulary for large-scale image applicationsabstractNot withstanding its great success and wide adoption in Bag-of-visual Words representation, visual vocabulary created from single image local features is often shown to be ineffective largely due to three reasons. First, many detected local features are not stable enough, resulting in many noisy and non-descriptive visual words in images. Second, single visual word discards the rich spatial contextual information among the local features, which has been proven to be valuable for visual matching. Third, the distance metric commonly used for generating visual vocabulary does not take the semantic context into consideration, which renders them to be prone to noise. To address these three confrontations, we propose an effective visual vocabulary generation framework containing three novel contributions: 1) we propose an effective unsupervised local feature refinement strategy; 2) we consider local features in groups to model their spatial contexts; 3) we further learn a discriminant distance metric between local feature groups, which we call discriminant group distance. This group distance is further leveraged to induce visual vocabulary from groups of local features. We name it contextual visual vocabulary, which captures both the spatial and semantic contexts. We evaluate the proposed local feature refinement strategy and the contextual visual vocabulary in two large-scale image applications: large-scale near-duplicate image retrieval on a dataset containing 1.5 million images and image search re-ranking tasks. Our experimental results show that the contextual visual vocabulary shows significant improvement over the classic visual vocabulary. Moreover, it outperforms the state-of-the-art Bundled Feature in the terms of retrieval precision, memory consumption and efficiency. Shiliang Zhang, Qingming Huang, Gang Hua 0001, Shuqiang Jiang, Wen Gao 0001, Qi Tian 0001 |
ACM Multimedia | 3 |
| 2010 | Visual quality assessment for web videos
Tian Xia 0002, Tao Mei 0001, Gang Hua 0001, Yongdong Zhang 0001, Xian-Sheng Hua 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2010 | A Hierarchical Visual Model for Video Object SummarizationabstractWe propose a novel method for removing irrelevant frames from a video given user-provided frame-level labeling for a very small number of frames. We first hypothesize a number of windows which possibly contain the object of interest, and then determine which window(s) truly contain the object of interest. Our method enjoys several favorable properties. First, compared to approaches where a single descriptor is used to describe a whole frame, each window's feature descriptor has the chance of genuinely describing the object of interest; hence it is less affected by background clutter. Second, by considering the temporal continuity of a video instead of treating frames as independent, we can hypothesize the location of the windows more accurately. Third, by infusing prior knowledge into the patch-level model, we can precisely follow the trajectory of the object of interest. This allows us to largely reduce the number of windows and hence reduce the chance of overfitting the data during learning. We demonstrate the effectiveness of the method by comparing it to several other semi-supervised learning approaches on challenging video clips. David Liu 0001, Gang Hua 0001, Tsuhan Chen |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2010 | A Comprehensive Approach to Image Spam Detection: From Server to Client SolutionabstractImage spam is a type of e-mail spam that embeds spam text content into graphical images to bypass traditional text-based e-mail spam filters. To effectively detect image spam, it is desirable to leverage image content analysis technologies. However, most previous works of image spam detection focus on filtering the image spam on the client side. We propose a more desirable comprehensive solution which embraces both server-side filtering and client-side detection to effectively mitigate image spam. On the server side, we present a nonnegative sparsity induced similarity measure for cluster analysis of spam images to filter the attack activities of spammers and fast trace back the spam sources. On the client side, we employ the principle of active learning where the learner guides the users to label as few images as possible while maximizing the classification accuracy. The server-side filtering identifies large image clusters as suspicious spam sources and further analysis can be performed to identify the real sources and block them from the beginning. For those spam images which survived the server-side filter, our active learner on the client side will further guide the users to interactively and efficiently filter them out. Our experiments on an image spam data-set collected from the e-mail server of our department demonstrate the efficacy of the proposed comprehensive solution. Yan Gao 0003, Alok N. Choudhary, Gang Hua 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2009 | Efficient Scale-Space Spatiotemporal Saliency Tracking for Distortion-Free Video Retargeting
Gang Hua 0001, Cha Zhang, Zicheng Liu 0001, Zhengyou Zhang, Ying Shan |
ACCV (2) | 1 |
| 2009 | Multiple instance fFeature for robust part-based object detectionabstractFeature misalignment in object detection refers to the phenomenon that features which fire up in some positive detection windows do not fire up in other positive detection windows. Most often it is caused by pose variation and local part deformation. Previous work either totally ignores this issue, or naively performs a local exhaustive search to better position each feature. We propose a learning framework to mitigate this problem, where a boosting algorithm is performed to seed the position of the object part, and a multiple instance boosting algorithm further pursues an aggregated feature for this part, namely multiple instance feature. Unlike most previous boosting based object detectors, where each feature value produces a single classification result, the value of the proposed multiple instance feature is the Noisy-OR integration of a bag of classification results. Our approach is applied to the task of human detection and is tested on two popular benchmarks. The proposed approach brings significant improvement in performance, i.e., smaller number of features used in the cascade and better detection accuracy. Zhe Lin 0001, Gang Hua 0001, Larry Davis 0001 |
CVPR | 2 |
| 2009 | Picking the best DAISYabstractLocal image descriptors that are highly discriminative, computational efficient, and with low storage footprint have long been a dream goal of computer vision research. In this paper, we focus on learning such descriptors, which make use of the DAISY configuration and are simple to compute both sparsely and densely. We develop a new training set of match/non-match image patches which improves on previous work. We test a wide variety of gradient and steerable filter based configurations and optimize over all parameters to obtain low matching errors for the descriptors. We further explore robust normalization, dimension reduction and dynamic range reduction to increase the discriminative power and yet reduce the storage requirement of the learned descriptors. All these enable us to obtain highly efficient local descriptors: e.g, 13.2% error at 13 bytes storage per descriptor, compared with 26.1% error at 128 bytes for SIFT. Simon A. J. Winder, Gang Hua 0001, Matthew A. Brown |
CVPR | 2 |
| 2009 | Implicit elastic matching with random projections for pose-variant face recognitionabstractWe present a new approach to robust pose-variant face recognition, which exhibits excellent generalization ability even across completely different datasets due to its weak dependence on data. Most face recognition algorithms assume that the face images are very well-aligned. This assumption is often violated in real-life face recognition tasks, in which face detection and rectification have to be performed automatically prior to recognition. Although great improvements have been made in face alignment recently, significant pose variations may still occur in the aligned faces. We propose a multiscale local descriptor-based face representation to mitigate this issue. First, discriminative local image descriptors are extracted from a dense set of multiscale image patches. The descriptors are expanded by their spatial locations. Each expanded descriptor is quantized by a set of random projection trees. The final face representation is a histogram of the quantized descriptors. The location expansion constrains the quantization regions to be localized not just in feature space but also in image space, allowing us to achieve an implicit elastic matching for face images. Our experiments on challenging face recognition benchmarks demonstrate the advantages of the proposed approach for handling large pose variations, as well as its superb generalization ability. John Wright 0001, Gang Hua 0001 |
CVPR | 2 |
| 2009 | A robust elastic and partial matching metric for face recognitionabstractWe present a robust elastic and partial matching metric for face recognition. To handle challenges such as pose, facial expression and partial occlusion, we enable both elastic and partial matching by computing a part based face representation. In which N local image descriptors are extracted from densely sampled overlapping image patches. We then define a distance metric where each descriptor in one face is matched against its spatial neighborhood in the other face and the minimal distance is recorded. For implicit partial matching, the list of all minimal distances are sorted in ascending order and the distance at the αN-th position is picked up as the final distance. The parameter 0 ≤ α ≤ 1 controls how much occlusion, facial expression changes, or pixel degradations we would allow. The optimal parameter values of this new distance metric are extensively studied and identified with real-life photo collections. We also reveal that filtering the face image by a simple difference of Gaussian brings significant robustness to lighting variations and beats the more utilized self-quotient image. Extensive evaluations on face recognition benchmarks show that our method is leading or is competitive in performance when compared to state-of-the-art. Gang Hua 0001, Amir Akbarzadeh |
ICCV | 1 |
| 2009 | Which faces to tag: Adding prior constraints into active learningabstractWe introduce an algorithm that guides the user to tag faces in the best possible order during a face recognition assisted tagging scenario. In particular, we extend the active learning paradigm to take advantage of constraints known a priori. For example, in the context of personal photo collections, if two faces come from the same source photograph, we know that they must be of different people. Similarly, in the context of video, we know that the faces from a single track must be of the same person. Given a set of unlabeled images and constraints, we use a probabilistic discriminative model that models the posterior distributions by propagating label information using a message passing scheme. The uncertainty estimate provided by the model naturally allows for active learning paradigms where the user is consulted after each iteration to tag additional faces. Our experiments show that performing active learning while incorporating a priori constraints provides a significant boost in many real-world face recognition tasks. Ashish Kapoor, Gang Hua 0001, Amir Akbarzadeh, Simon Baker |
ICCV | 2 |
| 2009 | What can visual content analysis do for text based image search?abstractModern image search engines such as Google, Yahoo!, Microsoft Live image search are all text meta word based. To search for images, the users type in a text query and the search engines rank the result images almost sorely based on the text meta-words. The abundant visual information in the images themselves is largely neglected. Recently, we have observed several new features released in the aforementioned image search engines, especially Microsoft Live image search, which are clearly based on the analysis of the visual content. We summarize some of these features, give insights about how they are designed, and motivate new content analysis based features for text based image search engines. Gang Hua 0001, Qi Tian 0001 |
ICME | 1 |
| 2009 | Descriptive visual words and visual phrases for image applicationsabstractThe Bag-of-visual Words (BoW) image representation has been applied for various problems in the fields of multimedia and computer vision. The basic idea is to represent images as visual documents composed of repeatable and distinctive visual elements, which are comparable to the words in texts. However, massive experiments show that the commonly used visual words are not as expressive as the text words, which is not desirable because it hinders their effectiveness in various applications. In this paper, Descriptive Visual Words (DVWs) and Descriptive Visual Phrases (DVPs) are proposed as the visual correspondences to text words and phrases, where visual phrases refer to the frequently co-occurring visual word pairs. Since images are the carriers of visual objects and scenes, novel descriptive visual element set can be composed by the visual words and their combinations which are effective in representing certain visual objects or scenes. Based on this idea, a general framework is proposed for generating DVWs and DVPs from classic visual words for various applications. In a large-scale image database containing 1506 object and scene categories, the visual words and visual word pairs descriptive to certain scenes or objects are identified as the DVWs and DVPs. Experiments show that the DVWs and DVPs are compact and descriptive, thus are more comparable with the text words than the classic visual words. We apply the identified DVWs and DVPs in several applications including image retrieval, image re-ranking, and object recognition. The DVW and DVP combination outperforms the classic visual words by 19.5% and 80% in image retrieval and object recognition tasks, respectively. The DVW and DVP based image re-ranking algorithm: DWPRank outperforms the state-of-the-art VisualRank by 12.4% in accuracy and about 11 times faster in efficiency. Shiliang Zhang, Qi Tian 0001, Gang Hua 0001, Qingming Huang, Shipeng Li 0001 |
ACM Multimedia | 3 |
| 2009 | Face Relighting from a Single Image under Arbitrary Unknown Lighting ConditionsabstractIn this paper, we present a new method to modify the appearance of a face image by manipulating the illumination condition, when the face geometry and albedo information is unknown. This problem is particularly difficult when there is only a single image of the subject available. Recent research demonstrates that the set of images of a convex Lambertian object obtained under a wide variety of lighting conditions can be approximated accurately by a low-dimensional linear subspace using a spherical harmonic representation. Moreover, morphable models are statistical ensembles of facial properties such as shape and texture. In this paper, we integrate spherical harmonics into the morphable model framework by proposing a 3D spherical harmonic basis morphable model (SHBMM). The proposed method can represent a face under arbitrary unknown lighting and pose simply by three low-dimensional vectors, i.e., shape parameters, spherical harmonic basis parameters, and illumination coefficients, which are called the SHBMM parameters. However, when the image was taken under an extreme lighting condition, the approximation error can be large, thus making it difficult to recover albedo information. In order to address this problem, we propose a subregion-based framework that uses a Markov random field to model the statistical distribution and spatial coherence of face texture, which makes our approach not only robust to extreme lighting conditions, but also insensitive to partial occlusions. The performance of our framework is demonstrated through various experimental results, including the improved rates for face recognition under extreme lighting conditions. Yang Wang 0001, Lei Zhang 0002, Zicheng Liu 0001, Gang Hua 0001, Zhengyou Zhang, Dimitris Samaras |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2009 | Context-Aware Visual TrackingabstractEnormous uncertainties in unconstrained environments lead to a fundamental dilemma that many tracking algorithms have to face in practice: Tracking has to be computationally efficient, but verifying whether or not the tracker is following the true target tends to be demanding, especially when the background is cluttered and/or when occlusion occurs. Due to the lack of a good solution to this problem, many existing methods tend to be either effective but computationally intensive by using sophisticated image observation models or efficient but vulnerable to false alarms. This greatly challenges long-duration robust tracking. This paper presents a novel solution to this dilemma by considering the context of the tracking scene. Specifically, we integrate into the tracking process a set of auxiliary objects that are automatically discovered in the video on the fly by data mining. Auxiliary objects have three properties, at least in a short time interval: 1) persistent co-occurrence with the target, 2) consistent motion correlation to the target, and 3) easy to track. Regarding these auxiliary objects as the context of the target, the collaborative tracking of these auxiliary objects leads to efficient computation as well as strong verification. Our extensive experiments have exhibited exciting performance in very challenging real-world testing cases. Ming Yang 0007, Ying Wu 0001, Gang Hua 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2008 | Integrated feature selection and higher-order spatial feature extraction for object categorizationabstractIn computer vision, the bag-of-visual words image representation has been shown to yield good results. Recent work has shown that modeling the spatial relationship between visual words further improves performance. Previous work extracts higher-order spatial features exhaustively. However, these spatial features are expensive to compute. We propose a novel method that simultaneously performs feature selection and feature extraction. Higher-order spatial features are progressively extracted based on selected lower order ones, thereby avoiding exhaustive computation. The method can be based on any additive feature selection algorithm such as boosting. Experimental results show that the method is computationally much more efficient than previous approaches, without sacrificing accuracy. David Liu 0001, Gang Hua 0001, Paul A. Viola, Tsuhan Chen |
CVPR | 2 |
| 2008 | VideoCut: Removing Irrelevant Frames by Discovering the Object of Interest
David Liu 0001, Gang Hua 0001, Tsuhan Chen |
ECCV (1) | 2 |
| 2007 | Face Recognition using Discriminatively Trained Orthogonal Rank One Tensor ProjectionsabstractWe propose a method for face recognition based on a discriminative linear projection. In this formulation images are treated as tensors, rather than the more conventional vector of pixels. Projections are pursued sequentially and take the form of a rank one tensor, i.e., a tensor which is the outer product of a set of vectors. A novel and effective technique is proposed to ensure that the rank one tensor projections are orthogonal to one another. These constraints on the tensor projections provide a strong inductive bias and result in better generalization on small training sets. Our work is related to spectrum methods, which achieve orthogonal rank one projections by pursuing consecutive projections in the complement space of previous projections. Although this may be meaningful for applications such as reconstruction, it is less meaningful for pursuing discriminant projections. Our new scheme iteratively solves an eigenvalue problem with orthogonality constraints on one dimension, and solves unconstrained eigenvalue problems on the other dimensions. Experiments demonstrate that on small and medium sized face recognition datasets, this approach outperforms previous embedding methods. On large face datasets this approach achieves results comparable with the best, often using fewer discriminant projections. Gang Hua 0001, Paul A. Viola, Steven Mark Drucker |
CVPR | 1 |
| 2007 | Face Re-Lighting from a Single Image under Harsh Lighting ConditionsabstractIn this paper, we present a new method to change the illumination condition of a face image, with unknown face geometry and albedo information. This problem is particularly difficult when there is only one single image of the subject available and it was taken under a harsh lighting condition. Recent research demonstrates that the set of images of a convex Lambertian object obtained under a wide variety of lighting conditions can be approximated accurately by a low-dimensional linear subspace using spherical harmonic representation. However, the approximation error can be large under harsh lighting conditions thus making it difficult to recover albedo information. In order to address this problem, we propose a subregion based framework that uses a Markov Random Field to model the statistical distribution and spatial coherence of face texture, which makes our approach not only robust to harsh lighting conditions, but insensitive to partial occlusions as well. The performance of our framework is demonstrated through various experimental results, including the improvement to the face recognition rate under harsh lighting conditions. Yang Wang 0001, Zicheng Liu 0001, Gang Hua 0001, Zhengyou Zhang, Dimitris Samaras |
CVPR | 3 |
| 2007 | Discriminant Embedding for Local Image DescriptorsabstractInvariant feature descriptors such as SIFT and GLOH have been demonstrated to be very robust for image matching and visual recognition. However, such descriptors are generally parameterised in very high dimensional spaces e.g. 128 dimensions in the case of SIFT. This limits the performance of feature matching techniques in terms of speed and scalability. Furthermore, these descriptors have traditionally been carefully hand crafted by manually tuning many parameters. In this paper, we tackle both of these problems by formulating descriptor design as a non- parametric dimensionality reduction problem. In contrast to previous approaches that use only the global statistics of the inputs, we adopt a discriminative approach. Starting from a large training set of labelled match/non-match pairs, we pursue lower dimensional embeddings that are optimised for their discriminative power. Extensive comparative experiments demonstrate that we can exceed the performance of the current state of the art techniques such as SIFT with far fewer dimensions, and with virtually no parameters to be tuned by hand. Gang Hua 0001, Matthew A. Brown, Simon A. J. Winder |
ICCV | 1 |
| 2007 | A decentralized probabilistic approach to articulated body tracking
Gang Hua 0001, Ying Wu 0001 |
Comput. Vis. Image Underst. | 1 |
| 2006 | Efficient Optimal Kernel Placement for Reliable Visual TrackingabstractThis paper describes a novel approach to optimal kernel placement in kernel-based tracking. If kernels are placed at arbitrary places, kernel-based methods are likely to be trapped in ill-conditioned locations, which prevents the reliable recovery of the motion parameters and jeopardizes the tracking performance. The theoretical analysis presented in this paper indicates that the optimal kernel placement can be evaluated based on a closed-form criterion, and achieved efficiently by a novel gradient-based algorithm. Based on that, new methods for temporal-stable multiple kernel placement and scale-invariant kernel placement are proposed. These new theoretical results and new algorithms greatly advance the study of kernel-based tracking in both theory and practice. Extensive real-time experimental results demonstrate the improved tracking reliability. Zhimin Fan 0002, Ming Yang 0007, Ying Wu 0001, Gang Hua 0001, Ting Yu 0003 |
CVPR (1) | 4 |
| 2006 | Measurement integration under inconsistency for robust trackingabstractThe solutions to many vision problems involve integrating measurements from multiple sources. Most existing methods rely on a hidden assumption, i.e., these measurements are consistent. In reality, unfortunately, this may not hold. The fact that naively fusing inconsistent measurements amounts to failing these methods indicates that this is not a trivial problem. This paper presents a novel approach to handling it. A new theorem is proven that gives two algebraic criteria to examine the consistency and inconsistency. In addition, a more general criterion is presented. Based on the theoretical analysis, a new information integration method is proposed and leads to encouraging results when applied to the task of visual tracking. Gang Hua 0001, Ying Wu 0001 |
CVPR (1) | 1 |
| 2006 | Automatic Business Card Scanning with a CameraabstractIn this paper, we present a system to automatically extract, rectify and enhance business card images. First the business card image patch is automatically segmented by minimizing a novel local-global variational energy. Second a quadrangle is fitted to the segmented image patch. With the four corner points of the quadrangle, we then estimate the physical aspect ratio of the business card and obtain a homography to rectify the quadrangle back to rectangular shape. We finally enhance the contrast of the rectified business card image using a S-shaped curve. Extensive experiments demonstrated the efficacy and robustness of our system. Gang Hua 0001, Zicheng Liu 0001, Zhengyou Zhang, Ying Wu 0001 |
ICIP | 1 |
| 2006 | Sequential mean field variational analysis of structured deformable shapes
Gang Hua 0001, Ying Wu 0001 |
Comput. Vis. Image Underst. | 1 |
| 2006 | Iterative Local-Global Energy Minimization for Automatic Extraction of Objects of InterestabstractWe propose a novel global-local variational energy to automatically extract objects of interest from images. Previous formulations only incorporate local region potentials, which are sensitive to incorrectly classified pixels during iteration. We introduce a global likelihood potential to achieve better estimation of the foreground and background models and, thus, better extraction results. Extensive experiments demonstrate its efficacy. Gang Hua 0001, Zicheng Liu 0001, Zhengyou Zhang, Ying Wu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2005 | Learning to Estimate Human Pose with Data Driven Belief PropagationabstractWe propose a statistical formulation for 2D human pose estimation from single images. The human body configuration is modeled by a Markov network and the estimation problem is to infer pose parameters from image cues such as appearance, shape, edge, and color. From a set of hand labeled images, we accumulate prior knowledge of 2D body shapes by learning their low-dimensional representations for inference of pose parameters. A data driven belief propagation Monte Carlo algorithm, utilizing importance sampling functions built from bottom-up visual cues, is proposed for efficient probabilistic inference. Contrasted to the few sequential statistical formulations in the literature, our algorithm integrates both top-down as well as bottom-up reasoning mechanisms, and can carry out the inference tasks in parallel. Experimental results demonstrate the potency and effectiveness of the proposed algorithm in estimating 2D human pose from single images. Gang Hua 0001, Ming-Hsuan Yang 0001, Ying Wu 0001 |
CVPR (2) | 1 |
| 2005 | A Statistical Field Model for Pedestrian DetectionabstractThis paper presents a new statistical model for detecting and tracking deformable objects such as pedestrians, where large shape variations induced by local shape deformation can not be well captured by global methods such as PCA. The proposed model employs a Boltzmann distribution to capture the prior of local deformation, and embeds it into a Markov network which can be learned from data. A mean field variational analysis of this model provides computationally efficient algorithms for computing the likelihood of image observations and facilitate fast model training. Based on that, effective detection and tracking algorithms for deformable objects are proposed and applied to pedestrian detection and tracking. The proposed method has several advantages. Firstly, it captures local deformation well and thus is robust to occlusions and clutter. In addition, it is computationally tractable. Moreover, it divides deformation into local deformation and global deformation, then conquers them by combining bottom-up and top-down methodologies. Extensive experiments demonstrate the effectiveness of the proposed model for deformable objects. Ying Wu 0001, Ting Yu 0003, Gang Hua 0001 |
CVPR (1) | 3 |
| 2005 | Variational Maximum A Posteriori by Annealed Mean Field AnalysisabstractThis paper proposes a novel probabilistic variational method with deterministic annealing for the maximum a posteriori (MAP) estimation of complex stochastic systems. Since the MAP estimation involves global optimization, in general, it is very difficult to achieve. Therefore, most probabilistic inference algorithms are only able to achieve either the exact or the approximate posterior distributions. Our method constrains the mean field variational distribution to be multivariate Gaussian. Then, a deterministic annealing scheme is nicely incorporated into the mean field fix-point iterations to obtain the optimal MAP estimate. This is based on the observation that when the covariance of the variational Gaussian distribution approaches to zero, the infimum point of the Kullback-Leibler (KL) divergence between the variational Gaussian and the real posterior will be the same as the supreme point of the real posterior. Although global optimality may not be guaranteed, our extensive synthetic and real experiments demonstrate the effectiveness and efficiency of the proposed method. Gang Hua 0001, Ying Wu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2004 | Multi-Scale Visual Tracking by Sequential Belief Propagation
Gang Hua 0001, Ying Wu 0001 |
CVPR (1) | 1 |
| 2003 | Switching Observation Models for Contour Tracking in ClutterabstractWe propose a generative model approach to contour tracking against nonstationary clutter and to coping with occlusions by explicit modelling and inferring. The proposed dynamic Bayesian networks consist of multiple hidden processes, which model the target, the clutter and the occlusions. The image observation models, which depict the generation of the image features, are conditioned on all the hidden processes. Based on this framework, the tracker can automatically switch among different observation models according to the hidden states of the clutter and occlusions. In addition, the inference of these hidden states provides self-evaluations for the tracker. The tracking and inference are implemented based on sequence Monte Carlo techniques. The effectiveness of the proposed approach to robust tracking and inferring nonstationary clutter and occlusion is demonstrated for a variety of image sequences. Ying Wu 0001, Gang Hua 0001, Ting Yu 0003 |
CVPR (1) | 2 |
| 2003 | Tracking Appearances with OcclusionsabstractOcclusion is a difficult problem for appearance-based target tracking, especially when we need to track multiple targets simultaneously and maintain the target identities during tracking. To cope with the occlusion problem explicitly, this paper proposes a dynamic Bayesian network, which accommodates an extra hidden process for occlusion and stipulates the conditions on which the image observation likelihood is calculated. The statistical inference of such a hidden process can reveal the occlusion relations among different targets, which makes the tracker more robust against partial even complete occlusions. In addition, considering the fact that target appearances change with views, another generative model for multiple view representation is proposed by adding a switching variable to select from different view templates. The integration of the occlusion model and multiple view model results in a complex dynamic Bayesian network, where extra hidden processes describe the switch of targets' templates, the targets' dynamics, and the occlusions among different targets. The tracking and inferring algorithms are implemented by the sampling-based sequential Monte Carlo strategies. Our experiments show the effectiveness of the proposed probabilistic models and the algorithms. Ying Wu 0001, Ting Yu 0003, Gang Hua 0001 |
CVPR (1) | 3 |
| 2003 | Tracking Articulated Body by Dynamic Markov NetworkabstractA new method for visual tracking of articulated objects is presented. Analyzing articulated motion is challenging because the dimensionality increase potentially demands tremendous increase of computation. To ease this problem, we propose an approach that analyzes subparts locally while reinforcing the structural constraints at the mean time. The computational model of the proposed approach is based on a dynamic Markov network, a generative model which characterizes the dynamics and the image observations of each individual subpart as well as the motion constraints among different subparts. Probabilistic variational analysis of the model reveals a mean field approximation to the posterior densities of each subparts given visual evidence, and provides a computationally efficient way for such a difficult Bayesian inference problem. In addition, we design mean field Monte Carlo (MFMC) algorithms, in which a set of low dimensional particle filters interact with each other and solve the high dimensional problem collaboratively. Extensive experiments on tracking human body parts demonstrate the effectiveness, significance and computational efficiency of the proposed method. Ying Wu 0001, Gang Hua 0001, Ting Yu 0003 |
ICCV | 2 |