EDBT 2026 Demo / reviewers in the wild / expert
Shengping Zhang
dblp:60/1866
· DBLP profile ↗
137ranked-venue papers
27as first author
62since 2021 · last 2026
0000-0001-5200-3420ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 97 · 13 first-author · 41 since 2021Artificial intelligence and machine learning · 65 · 11 first-author · 36 since 2021Databases, data management, data science and information retrieval · 6 · 2 first-authorComputer networks · 4 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DialoGen: Towards Dialog Gesture Generation via Identity-Decoupled Style Guidance in Interactive Diffusion ModelabstractWe propose DialoGen, a novel framework for generating realistic gestures for both interlocutors in dialog scenarios, conditioned on conversational audios. Unlike most existing methods that focus solely on a single speaker, DialoGen simultaneously generates synchronized gestures for both participants while also embedding identity-decoupled style into generated gestures that enhance realism and expressiveness. To ensure precise synchronization between interlocutors, DialoGen adopts an interactive dual-diffusion model with mutual interaction estimation, which integrates interaction correlation into the diffusion process. More importantly, by leveraging supervised contrastive learning, we develop the identity-decoupled style guidance to adaptively decompose the identity-specific style of interlocutors into latent space, enabling multi-style dialog gesture generation. Extensive experimental results demonstrate that our model significantly outperforms existing methods in generating realistic, speech-aligned, identity-specific gestures, offering a high-quality solution for various dialog scenarios. Weiyu Zhao, Chenyang Wang 0002, Liangxiao Hu, Zonglin Li 0004, Wei Yu 0004, Shengping Zhang |
AAAI | 6 |
| 2026 | SS-NeRF: Physically Based Sparse Spectral Rendering With Neural Radiance FieldabstractIn this paper, we propose SS-NeRF, the end-to-end Neural Radiance Field (NeRF)-based architectures for high-quality physically based rendering with sparse inputs. We modify the classical spectral rendering into two main steps, 1) the generation of a series of spectrum maps spanning different wavelengths, 2) the combination of these spectrum maps for the RGB output. The proposed architecture follows these two steps through the proposed multi-layer perceptron (MLP)-based architecture (SpectralMLP) and spectrum attention UNet (SAUNet). Given the ray origin and the ray direction, the SpectralMLP constructs the spectral radiance field to obtain spectrum maps of novel views, which are then sent to the SAUNet to produce RGB images of white-light illumination. Applying NeRF to build up the spectral rendering is a more physically-based way from the perspective of ray-tracing. Further, the spectral radiance fields decompose difficult scenes and improve the performance of NeRF-based methods. Previous baseline, such as SpectralNeRF, outperforms recent methods in synthesizing novel views but requires relatively dense viewpoints for accurate scene reconstruction. To tackle this, we propose SS-NeRF to enhance the detail of scene representation with sparse inputs. In SS-NeRF, we first design the depth-aware continuity to optimize the reconstruction based on single-view depth predictions. Then, the geometric-projected consistency is introduced to optimize the multi-view geometry alignment. Additionally, we introduce a superpixel-aligned consistency to ensure that the average color within each superpixel region remains consistent. Comprehensive experimental results demonstrate that the proposed method is superior to recent state-of-the-art methods when synthesizing new views on both synthetic and real-world datasets. Ru Li 0002, Guanghui Liu 0001, Shengping Zhang, Bing Zeng 0001, Shuaicheng Liu |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | Dual Alignment-Enhanced Fashion Vision-Language Pre-TrainingabstractFashion vision-language pre-training (VLP) models have demonstrated remarkable capabilities in excelling at a wide range of fashion cross-modal tasks. However, current models still face three notable limitations (1) an inability to discern varying levels of consistency between textual descriptions and multi-view images, (2) a deficiency in explicit fine-grained alignment between images and text, and (3) a lack of specific supervision mechanisms for facilitating global joint embedding learning. To address these limitations, we propose a novel dual alignment-enhanced fashion VLP model. This model delves deeply into the rich resources of multi-view images and semantic attributes associated with each fashion item. Notably, we introduce two novel pre-training tasks: Multi-grained Adaptive Image-Text Alignment (MAITA) and Joint Embedding-oriented Alignment (JEA). MAITA focuses on optimizing the text/image encoder by orchestrating adaptive alignment between multi-view images and input text. This encompasses both coarse-grained and fine-grained alignment strategies to enrich semantic understanding, while JEA is devised to supervise the fine-grained semantic learning process of the global multimodal joint embedding. Experimental results spanning four diverse downstream tasks, including cross-modal retrieval, text-guided image retrieval, category recognition, and subcategory recognition, substantiate the significant performance superiority of our model over prior state-of-the-art fashion VLP models. Weili Guan, Kejie Wang, Xuemeng Song, Kaihao Zhang, Xiaojun Chang, Shengping Zhang |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2025 | ProsodyTalker: 3D Visual Speech Animation via Prosody DecompositionabstractMost existing 3D visual speech animation methods synthesize lip movements synchronized with speech, which however neglect head poses and therefore degrade the animation realism. The animation of head poses presents two primary challenges: (1) the intricate mapping between speech and head poses remains poorly understood and (2) the absence of 4D face datasets featuring realistic head poses. Inspired by prosody decomposition in speech processing, we discern that head movements correlate with the fundamental frequency (F0) of speech prosody, while lip movements align with the language content. These observations motivate us to propose a novel framework, dubbed ProsodyTalker, that concurrently synthesizes lip and head movements, grounded in the principles of prosody decomposition. The core idea is first to adopt information perturbation to explicitly decompose the speech prosody into pose-related F0 and lip-related language content. Then, an autoregressive content-oriented fusion decoder is employed to enhance lip synchronization in the synthesized facial sequences. To synthesize head poses, we design a transformer-based variational autoencoder to learn a latent distribution of facial sequences and propose an F0-conditioned latent diffusion model to establish a probabilistic mapping from F0 to pose-related latent codes. Furthermore, we contribute a large-scale 4D face dataset containing bunches of variations in identities, head poses and facial motions. Extensive experiments show that our method achieves more realistic animation than state-of-the-art methods. Zonglin Li 0004, Xiaoqian Lv, Qinglin Liu, Quanling Meng, Xin Sun 0003, Shengping Zhang |
AAAI | 6 |
| 2025 | Path-Adaptive Matting for Efficient Inference Under Various Computational Cost ConstraintsabstractIn this paper, we explore a novel image matting task aimed at achieving efficient inference under various computational cost constraints, specifically FLOP limitations, using a single matting network. Existing matting methods which have not explored scalable architectures or path-learning strategies, fail to tackle this challenge. To overcome these limitations, we introduce Path-Adaptive Matting (PAM), a framework that dynamically adjusts network paths based on image contexts and computational cost constraints. We formulate the training of the computational cost-constrained matting network as a bilevel optimization problem, jointly optimizing the matting network and the path estimator. Building on this formalization, we design a path-adaptive matting architecture by incorporating path selection layers and learnable connect layers to estimate optimal paths and perform efficient inference within a unified network. Furthermore, we propose a performance-aware path-learning strategy to generate path labels online by evaluating a few paths sampled from the prior distribution of optimal paths and network estimations, enabling robust and efficient online path learning. Experiments on five image matting datasets demonstrate that the proposed PAM framework achieves competitive performance across a range of computational cost constraints. Qinglin Liu, Zonglin Li 0004, Xiaoqian Lv, Xin Sun 0003, Ru Li 0002, Shengping Zhang |
AAAI | 6 |
| 2025 | Multi-view Consistent 3D Panoptic Scene Understandingabstract3D panoptic scene understanding seeks to create novel view images with 3D-consistent panoptic segmentation, which is crucial for many vision and robotics applications. Mainstream methods (e.g., Panoptic Lifting) directly use machine-generated 2D panoptic segmentation masks as training labels. However, these generated masks often exhibit multi-view inconsistencies, leading to ambiguities during the optimization process. To address this, we present Multi-view Consistent 3D Panoptic Scene Understanding (MVC-PSU), featuring two key components: 1) Probabilistic Semantic Aligner, which associates semantic information of corresponding pixels across multiple views by probabilistic alignment to ensure that predicted panoptic segmentation masks are consistent across different views. 2) Geometric Consistency Enforcer, which uses multi-view projection and monocular depth consistency to ensure that the geometry of the reconstructed scene is accurate and consistent across different views. Experimental results demonstrate that the proposed MVC-PSU surpasses state-of-the-art methods on the ScanNet, Replica, and HyperSim datasets. Xianzhu Liu, Xin Sun 0003, Haozhe Xie, Zonglin Li 0004, Ru Li 0002, Shengping Zhang |
AAAI | 6 |
| 2025 | LVPNet: A Latent-Variable-Based Prediction-Driven End-to-End Framework for Lossless Compression of Medical Images
Chenyue Song, Wei Zhang 0192, Siqiao Li, Haiqi Zhu, Shengping Zhang, Shaohui Liu, Feng Jiang 0001 |
MICCAI (8) | 7 |
| 2025 | REA-Listener: Real-Time Listening Head Generation with Dynamic Emotion Modeling and Flexible Modality AdaptationabstractListening head generation aims to synthesize realistic and responsive non-verbal listener head motions that respond to speakers in conversational scenarios. Existing methods typically rely on fixed audio-visual input modalities and predefined emotion labels, limiting their adaptability and expressiveness in real-world scenarios. In this paper, we propose a novel real-time framework, REA-Listener, to generate high-fidelity listening head videos with flexible modality adaptation and dynamic emotion modeling. Specifically, we first propose a Modality-Adaptive Mixture of Experts (MA-MoE) module to encode arbitrary combinations of speaker audio and visual signals into a unified embedding space, ensuring robustness under partial modality conditions. To further enhance the temporal consistency of listener emotion, we present a lightweight emotional head dynamics generator with a multi-modal emotion predictor, which infers listener emotions dynamically from speaker context alongside head motion coefficient prediction. Finally, we employ a 3D-aware renderer based on 3D Gaussian Splatting to produce high-quality listener head videos in real time. With these components, our approach achieves efficient head motion generation at 30fps on a single NVIDIA RTX 3090 GPU, supporting real-time interaction. Extensive evaluations and applications demonstrate that our method outperforms state-of-the-art methods in listening head generation. Sizhe Zhao, Chenyang Wang 0002, Weiyu Zhao, Zonglin Li 0004, Ming Li 0042, Shengping Zhang |
ACM Multimedia | 6 |
| 2025 | 2D Semantic-Guided Semantic Scene Completion
Xianzhu Liu, Haozhe Xie, Shengping Zhang, Hongxun Yao, Rongrong Ji, Liqiang Nie, Dacheng Tao |
Int. J. Comput. Vis. | 3 |
| 2025 | Towards Universal Modal Tracking With Online Dense Temporal Token LearningabstractWe propose a universal video-level modality-awareness tracking model with online dense temporal token learning (called UM-ODTrack). It is designed to support various tracking tasks, including RGB, RGB+Thermal, RGB+Depth, and RGB+Event, utilizing the same model architecture and parameters. Specifically, our model is designed with three core goals: Video-level Sampling. We expand the model's inputs to a video sequence level, aiming to see a richer video context from an near-global perspective. Video-level Association. Furthermore, we introduce two simple yet effective online dense temporal token association mechanisms to propagate the appearance and motion trajectory information of target via a video stream manner. Modality Scalable. We propose two novel gated perceivers that adaptively learn cross-modal representations via a gated attention mechanism, and subsequently compress them into the same set of model parameters via a one-shot training manner for multi-task inference. This new solution brings the following benefits: (i) The purified token sequences can serve as temporal prompts for the inference in the next video frames, whereby previous information is leveraged to guide future inference. (ii) Unlike multi-modal trackers that require independent training, our one-shot training scheme not only alleviates the training burden, but also improves model representation. Extensive experiments on visible and multi-modal benchmarks show that our UM-ODTrack achieves a new SOTA performance. Yaozong Zheng, Bineng Zhong 0001, Qihua Liang, Shengping Zhang, Guorong Li, Xianxian Li, Rongrong Ji |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Patch-Based Spatio-Temporal Deformable Attention BiRNN for Video DeblurringabstractSuccessful video deblurring relies on effectively using sharp pixels from other frames to recover the blurry pixels of the current frame. However, mainstream methods only use estimated optical flows to align and fuse features from adjacent frames without considering the pixel-wise blur levels, leading to the introduction of blurry pixels from adjacent frames. Furthermore, these methods fail to effectively exploit information from the entire input video. To address these limitations, we propose STDANet++, which redesigns the state-of-the-art method STDANet by introducing patch-based spatio-temporal deformable attention (PSTDA) module and long-term frame fusion (LTFF) module to the BiRNN-based structure. By effectively utilizing sharp information across the entire video, the proposed method outperforms state-of-the-art methods on the GoPro, DVD and BSD datasets, according to our experimental results. The source code is available at https://github.com/huicongzhang/STDANetPP. Huicong Zhang, Haozhe Xie, Shengping Zhang, Hongxun Yao |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Uncertainty-Guided Diffusion Model for Camouflaged Object DetectionabstractRecently, diffusion models have significantly improved the performance of Camouflaged Object Detection (COD) by adding noise to a mask and iteratively denoising it to match the target distributions. Due to the direct extraction of features from noisy masks and the lack of conditional constraints on a prediction area, the diffusion model may deviate from a correct prediction range and produces mispredictions in regions with high uncertainty. To address this issue, we propose an uncertainty-guided diffusion model (UGDNet) for COD, which explicitly quantifies uncertainty and integrates it as an anchor condition into the diffusion models to provide an initialization of the diffusion regions. The core idea is first to utilize a probability representation and transformer to explicitly model uncertainty, aiming to identify areas where a model may generate overconfident mispredictions. Then, we use the uncertainty as an anchor condition to provide a reference prediction range for the diffusion model, guiding each step of the diffusion process. Furthermore, we use uncertainty to guide feature aggregation, prompting the model to pay extra attention to the semantic features of regions with high uncertainty to refine the segmentation results further. The experimental results indicate that our proposed UGDNet achieves higher accuracy than existing state-of-the-art models on five COD benchmarks, including COD10K, NC4K, CAMO, CHAMELEON, and CDS2K. Bineng Zhong 0001, Qihua Liang, Zhiyi Mo, Shengping Zhang, Shuxiang Song 0001 |
IEEE Trans. Multim. | 5 |
| 2024 | SpectralNeRF: Physically Based Spectral Rendering with Neural Radiance FieldabstractIn this paper, we propose SpectralNeRF, an end-to-end Neural Radiance Field (NeRF)-based architecture for high-quality physically based rendering from a novel spectral perspective. We modify the classical spectral rendering into two main steps, 1) the generation of a series of spectrum maps spanning different wavelengths, 2) the combination of these spectrum maps for the RGB output. Our SpectralNeRF follows these two steps through the proposed multi-layer perceptron (MLP)-based architecture (SpectralMLP) and Spectrum Attention UNet (SAUNet). Given the ray origin and the ray direction, the SpectralMLP constructs the spectral radiance field to obtain spectrum maps of novel views, which are then sent to the SAUNet to produce RGB images of white-light illumination. Applying NeRF to build up the spectral rendering is a more physically-based way from the perspective of ray-tracing. Further, the spectral radiance fields decompose difficult scenes and improve the performance of NeRF-based methods. Comprehensive experimental results demonstrate the proposed SpectralNeRF is superior to recent NeRF-based methods when synthesizing new views on synthetic and real datasets. The codes and datasets are available at https://github.com/liru0126/SpectralNeRF. Ru Li 0002, Guanghui Liu 0001, Shengping Zhang, Bing Zeng 0001, Shuaicheng Liu |
AAAI | 4 |
| 2024 | Explicit Visual Prompts for Visual Object TrackingabstractHow to effectively exploit spatio-temporal information is crucial to capture target appearance changes in visual tracking. However, most deep learning-based trackers mainly focus on designing a complicated appearance model or template updating strategy, while lacking the exploitation of context between consecutive frames and thus entailing the when-and-how-to-update dilemma. To address these issues, we propose a novel explicit visual prompts framework for visual tracking, dubbed EVPTrack. Specifically, we utilize spatio-temporal tokens to propagate information between consecutive frames without focusing on updating templates. As a result, we cannot only alleviate the challenge of when-to-update, but also avoid the hyper-parameters associated with updating strategies. Then, we utilize the spatio-temporal tokens to generate explicit visual prompts that facilitate inference in the current frame. The prompts are fed into a transformer encoder together with the image tokens without additional processing. Consequently, the efficiency of our model is improved by avoiding how-to-update. In addition, we consider multi-scale information as explicit visual prompts, providing multiscale template features to enhance the EVPTrack's ability to handle target scale changes. Extensive experimental results on six benchmarks (i.e., LaSOT, LaSOText, GOT-10k, UAV123, TrackingNet, and TNL2K.) validate that our EVPTrack can achieve competitive performance at a real-time speed by effectively exploiting both spatio-temporal and multi-scale information. Code and models are available at https://github.com/GXNU-ZhongLab/EVPTrack. Liangtao Shi, Bineng Zhong 0001, Qihua Liang, Ning Li 0044, Shengping Zhang, Xianxian Li |
AAAI | 5 |
| 2024 | ODTrack: Online Dense Temporal Token Learning for Visual TrackingabstractOnline contextual reasoning and association across consecutive video frames are critical to perceive instances in visual tracking. However, most current top-performing trackers persistently lean on sparse temporal relationships between reference and search frames via an offline mode. Consequently, they can only interact independently within each image-pair and establish limited temporal correlations. To alleviate the above problem, we propose a simple, flexible and effective video-level tracking pipeline, named ODTrack, which densely associates the contextual relationships of video frames in an online token propagation manner. ODTrack receives video frames of arbitrary length to capture the spatio-temporal trajectory relationships of an instance, and compresses the discrimination features (localization information) of a target into a token sequence to achieve frame-to-frame association. This new solution brings the following benefits: 1) the purified token sequences can serve as prompts for the inference in the next video frame, whereby past information is leveraged to guide future inference; 2) the complex online update strategies are effectively avoided by the iterative propagation of token sequences, and thus we can achieve more efficient model representation and computation. ODTrack achieves a new SOTA performance on seven benchmarks, while running at real-time speed. Code and models are available at https://github.com/GXNU-ZhongLab/ODTrack. Yaozong Zheng, Bineng Zhong 0001, Qihua Liang, Zhiyi Mo, Shengping Zhang, Xianxian Li |
AAAI | 5 |
| 2024 | ProxyCap: Real-Time Monocular Full-Body Capture in World Space via Human-Centric Proxy-to-Motion LearningabstractLearning-based approaches to monocular motion capture have recently shown promising results by learning to regress in a data-driven manner. However, due to the challenges in data collection and network designs, it remains challenging to achieve real-time full-body capture while being accurate in world space. In this work, we introduce ProxyCap, a human-centric proxy-to-motion learning scheme to learn world-space motions from a proxy dataset of 2D skeleton sequences and 3D rotational motions. Such proxy data enables us to build a learning-based network with accurate world-space supervision while also mitigating the generalization issues. For more accurate and physically plausible predictions in world space, our network is designed to learn human motions from a human-centric perspective, which enables the understanding of the same motion captured with different camera trajectories. Moreover, a contact-aware neural motion descent module is proposed to improve foot-ground contact and motion misalignment with the proxy observations. With the proposed learning-based solution, we demonstrate the first real-time monocular full-body capture system with plausible foot-ground contact in world space even using hand-held cameras. Yuxiang Zhang 0006, Hongwen Zhang 0001, Liangxiao Hu, Jiajun Zhang 0012, Hongwei Yi, Shengping Zhang, Yebin Liu |
CVPR | 6 |
| 2024 | GaussianAvatar: Towards Realistic Human Avatar Modeling from a Single Video via Animatable 3D GaussiansabstractWe present GaussianAvatar, an efficient approach to cre-ating realistic human avatars with dynamic 3D appear-ances from a single video. We start by introducing animat-able 3D Gaussians to explicitly represent humans in var-ious poses and clothing styles. Such an explicit and ani-matable representation can fuse 3D appearances more effi-ciently and consistently from 2D observations. Our repre-sentation is further augmented with dynamic properties to support pose-dependent appearance modeling, where a dy-namic appearance network along with an optimizable feature tensor is designed to learn the motion-to-appearance mapping. Moreover, by leveraging the differentiable motion condition, our method enables a joint optimization of motions and appearances during avatar modeling, which helps to tackle the long-standing issue of inaccurate motion esti-mation in monocular settings. The efficacy of GaussianA-vatar is validated on both the public dataset and our col-lected dataset, demonstrating its superior performances in terms of appearance quality and rendering efficiency. The code and dataset are available at https://github.com/aipixel/GaussianAvatar. Liangxiao Hu, Hongwen Zhang 0001, Yuxiang Zhang 0006, Boyao Zhou, Boning Liu 0001, Shengping Zhang, Liqiang Nie |
CVPR | 6 |
| 2024 | Fourier Priors-Guided Diffusion for Zero-Shot Joint Low-Light Enhancement and DeblurringabstractExisting joint low-light enhancement and deblurring methods learn pixel-wise mappings from paired synthetic data, which results in limited generalization in real-world scenes. While some studies explore the rich generative prior of pre-trained diffusion models, they typically rely on the assumed degradation process and cannot handle unknown real-world degradations well. To address these problems, we propose a novel zero-shot framework, FourierDiff, which embeds Fourier priors into a pre-trained diffusion model to harmoniously handle the joint degradation of luminance and structures. FourierDiff is appealing in its relaxed requirements on paired training data and degradation assumptions. The key zero-shot insight is motivated by image characteristics in the Fourier domain: most luminance information concentrates on amplitudes while structure and content information are closely related to phases. Based on this observation, we decompose the sampled results of the reverse diffusion process in the Fourier domain and take advantage of the amplitude of the generative prior to align the enhanced brightness with the distribution of natural images. To yield a sharp and content-consistent enhanced result, we further design a spatial-frequency alternating optimization strategy to progressively refine the phase of the input. Extensive experiments demonstrate the superior effectiveness of the proposed method, especially in real-world scenes. The code is available at https://github.com/aipixel/FourierDiff. Xiaoqian Lv, Shengping Zhang, Chenyang Wang 0002, Yichen Zheng, Bineng Zhong 0001, Chongyi Li, Liqiang Nie |
CVPR | 2 |
| 2024 | DiffPerformer: Iterative Learning of Consistent Latent Guidance for Diffusion-Based Human Video GenerationabstractExisting diffusion models for pose-guided human video generation mostly suffer from temporal inconsistency in the generated appearance and poses due to the inherent randomization nature of the generation process. In this paper, we propose a novel framework, DiffPerformer, to synthesize high-fidelity and temporally consistent human video. Without complex architecture modification or costly training, DiffPerformer finetunes a pre-trained diffusion model on a single video of the target character and introduces an implicit video representation as a proxy to learn temporally consistent guidance for the diffusion model. The guidance is encoded into VAE latent space and an iterative optimization loop is constructed between the implicit video representation and the diffusion model, allowing to harness the smooth property of the implicit video representation and the generative capabilities of the diffusion model in a mutually beneficial way. Moreover, we propose 3D-aware human flow as a temporal constraint during the optimization to explicitly model the correspondence between driving poses and human appearance. This alleviates the mis-alignment between driving poses and target performer and therefore maintains the appearance coherence under various motions. Extensive experiments demonstrate that our method outperforms the state-of-the-art methods. The code is available at https://github.com/aipixel/DiffPerformer. Chenyang Wang 0002, Zerong Zheng, Tao Yu 0007, Xiaoqian Lv, Bineng Zhong 0001, Shengping Zhang, Liqiang Nie |
CVPR | 6 |
| 2024 | Autoregressive Queries for Adaptive Tracking with Spatio-Temporal TransformersabstractThe rich spatio-temporal information is crucial to capture the complicated target appearance variations in visual tracking. However, most top-performing tracking algorithms rely on many hand-crafted components for spatio-temporal information aggregation. Consequently, the spatio-temporal information is far away from being fully explored. To alleviate this issue, we propose an adaptive tracker with spatio-temporal transformers (named AQA-Track), which adopts simple autoregressive queries to effectively learn spatio-temporal information without many hand-designed components. Firstly, we introduce a set of learnable and autoregressive queries to capture the instantaneous target appearance changes in a sliding window fashion. Then, we design a novel attention mechanism for the interaction of existing queries to generate a new query in current frame. Finally, based on the initial target template and learnt autoregressive queries, a spatio-temporal information fusion module (STM) is designed for spatiotemporal formation aggregation to locate a target object. Benefiting from the STM, we can effectively combine the static appearance and instantaneous changes to guide robust tracking. Extensive experiments show that our method significantly improves the tracker's performance on six popular tracking benchmarks: LaSOT, LaSOText, TrackingNet, GOT-10k, TNL2K, and UAV123. Code and models will be https://github.com/orgs/GXNU-ZhongLab. Jinxia Xie, Bineng Zhong 0001, Zhiyi Mo, Shengping Zhang, Liangtao Shi, Shuxiang Song 0001, Rongrong Ji |
CVPR | 4 |
| 2024 | GPS-Gaussian: Generalizable Pixel-Wise 3D Gaussian Splatting for Real-Time Human Novel View SynthesisabstractWe present a new approach, termed GPS-Gaussian, for synthesizing novel views of a character in a real-time manner. The proposed method enables 2K-resolution rendering under a sparse-view camera setting. Unlike the original Gaussian Splatting or neural implicit rendering methods that necessitate per-subject optimizations, we introduce Gaussian parameter maps defined on the source views and regress directly Gaussian Splatting properties for instant novel view synthesis without any fine-tuning or optimization. To this end, we train our Gaussian parameter regression module on a large amount of human scan data, jointly with a depth estimation module to lift 2D parameter maps to 3D space. The proposed framework is fully differentiable and experiments on several datasets demonstrate that our method outperforms state-of-the-art methods while achieving an exceeding rendering speed. The code is available at https://github.com/aipixel/GPS-Gaussian. Shunyuan Zheng, Boyao Zhou, Ruizhi Shao, Boning Liu 0001, Shengping Zhang, Liqiang Nie, Yebin Liu |
CVPR | 5 |
| 2024 | Revisiting Context Aggregation for Image MattingabstractTraditional studies emphasize the significance of context information in improving matting performance. Consequently, deep learning-based matting methods delve into designing pooling or affinity-based context aggregation modules to achieve superior results. However, these modules cannot well handle the context scale shift caused by the difference in image size during training and inference, resulting in matting performance degradation. In this paper, we revisit the context aggregation mechanisms of matting networks and find that a basic encoder-decoder network without any context aggregation modules can actually learn more universal context aggregation, thereby achieving higher matting performance compared to existing methods. Building on this insight, we present AEMatter, a matting network that is straightforward yet very effective. AEMatter adopts a Hybrid-Transformer backbone with appearance-enhanced axis-wise learning (AEAL) blocks to build a basic network with strong context aggregation learning capability. Furthermore, AEMatter leverages a large image training strategy to assist the network in learning context aggregation from data. Extensive experiments on five popular matting datasets demonstrate that the proposed AEMatter outperforms state-of-the-art matting methods by a large margin. The source code is available at https://github.com/aipixel/AEMatter. Qinglin Liu, Xiaoqian Lv, Quanling Meng, Zonglin Li 0004, Xiangyuan Lan, Shuo Yang 0006, Shengping Zhang, Liqiang Nie |
ICML | 7 |
| 2024 | Mind the Boundary: Coreset Selection via Reconstructing the Decision BoundaryabstractExisting paradigms of pushing the state of the art require exponentially more training data in many fields. Coreset selection seeks to mitigate this growing demand by identifying the most efficient subset of training data. In this paper, we delve into geometry-based coreset methods and preliminarily link the geometry of data distribution with models’ generalization capability in theoretics. Leveraging these theoretical insights, we propose a novel coreset construction method by selecting training samples to reconstruct the decision boundary of a deep neural network learned on the full dataset. Extensive experiments across various popular benchmarks demonstrate the superiority of our method over multiple competitors. For the first time, our method achieves a 50% data pruning rate on the ImageNet-1K dataset while sacrificing less than 1% in accuracy. Additionally, we showcase and analyze the remarkable cross-architecture transferability of the coresets derived from our approach. Shuo Yang 0006, Zhe Cao 0001, Ruiheng Zhang 0001, Ping Luo 0002, Shengping Zhang, Liqiang Nie |
ICML | 6 |
| 2024 | Learning Scale-Aware Spatio-temporal Implicit Representation for Event-based Motion DeblurringabstractExisting event-based motion deblurring methods mostly focus on restoring images with the same spatial and temporal scales as events. However, the unknown scales of images and events in the real world pose great challenges and have rarely been explored. To address this gap, we propose a novel Scale-Aware Spatio-temporal Network (SASNet) to flexibly restore blurred images with event streams at arbitrary scales. The core idea is to implicitly aggregate both spatial and temporal correspondence features of images and events to generalize at continuous scales. To restore highly blurred local areas, we develop a Spatial Implicit Representation Module (SIRM) to aggregate spatial correlation at any resolution through event encoding sampling. To tackle global motion blur, a Temporal Implicit Representation Module (TIRM) is presented to learn temporal correlation via temporal shift operations with long-term aggregation. Additionally, we build a High-resolution Hybrid Deblur (H2D) dataset using a new-generation hybrid event-based sensor, which comprises images with naturally spatially aligned and temporally synchronized events at various scales. Experiments demonstrate that our SASNet outperforms state-of-the-art methods on both synthetic GoPro and real H2D datasets, especially in high-speed motion scenarios. Code and dataset are available at https://github.com/aipixel/SASNet. Wei Yu 0004, Jianing Li 0001, Shengping Zhang, Xiangyang Ji |
ICML | 3 |
| 2024 | Rethinking Imbalance in Image Super-Resolution for Efficient InferenceabstractExisting super-resolution (SR) methods optimize all model weights equally using $\mathcal{L}_1$ or $\mathcal{L}_2$ losses by uniformly sampling image patches without considering dataset imbalances or parameter redundancy, which limits their performance. To address this, we formulate the image SR task as an imbalanced distribution transfer learning problem from a statistical probability perspective, proposing a plug-and-play Weight-Balancing framework (WBSR) to achieve balanced model learning without changing the original model structure and training data. Specifically, we develop a Hierarchical Equalization Sampling (HES) strategy to address data distribution imbalances, enabling better feature representation from texture-rich samples. To tackle model optimization imbalances, we propose a Balanced Diversity Loss (BDLoss) function, focusing on learning texture regions while disregarding redundant computations in smooth regions. After joint training of HES and BDLoss to rectify these imbalances, we present a gradient projection dynamic inference strategy to facilitate accurate and efficient inference. Extensive experiments across various models, datasets, and scale factors demonstrate that our method achieves comparable or superior performance to existing approaches with about 34\% reduction in computational cost. Wei Yu 0004, Qinglin Liu, Jianing Li 0001, Shengping Zhang, Xiangyang Ji |
NeurIPS | 5 |
| 2024 | High-Resolution Image Harmonization with Adaptive-Interval Color TransformationabstractExisting high-resolution image harmonization methods typically rely on global color adjustments or the upsampling of parameter maps. However, these methods ignore local variations, leading to inharmonious appearances. To address this problem, we propose an Adaptive-Interval Color Transformation method (AICT), which predicts pixel-wise color transformations and adaptively adjusts the sampling interval to model local non-linearities of the color transformation at high resolution. Specifically, a parameter network is first designed to generate multiple position-dependent 3-dimensional lookup tables (3D LUTs), which use the color and position of each pixel to perform pixel-wise color transformations. Then, to enhance local variations adaptively, we separate a color transform into a cascade of sub-transformations using two 3D LUTs to achieve the non-uniform sampling intervals of the color transform. Finally, a global consistent weight learning method is proposed to predict an image-level weight for each color transform, utilizing global information to enhance the overall harmony. Extensive experiments demonstrate that our AICT achieves state-of-the-art performance with a lightweight architecture. The code is available at https://github.com/aipixel/AICT. Quanling Meng, Qinglin Liu, Zonglin Li 0004, Xiangyuan Lan, Shengping Zhang, Liqiang Nie |
NeurIPS | 5 |
| 2024 | Dual-context aggregation for universal image matting
Qinglin Liu, Xiaoqian Lv, Wei Yu 0002, Changyong Guo, Shengping Zhang |
Multim. Tools Appl. | 5 |
| 2024 | Stereo Image Restoration via Attention-Guided Correspondence LearningabstractAlthough stereo image restoration has been extensively studied, most existing work focuses on restoring stereo images with limited horizontal parallax due to the binocular symmetry constraint. Stereo images with unlimited parallax (e.g., large ranges and asymmetrical types) are more challenging in real-world applications and have rarely been explored so far. To restore high-quality stereo images with unlimited parallax, this paper proposes an attention-guided correspondence learning method, which learns both self- and cross-views feature correspondence guided by parallax and omnidirectional attention. To learn cross-view feature correspondence, a Selective Parallax Attention Module (SPAM) is proposed to interact with cross-view features under the guidance of parallax attention that adaptively selects receptive fields for different parallax ranges. Furthermore, to handle asymmetrical parallax, we propose a Non-local Omnidirectional Attention Module (NOAM) to learn the non-local correlation of both self- and cross-view contexts, which guides the aggregation of global contextual features. Finally, we propose an Attention-guided Correspondence Learning Restoration Network (ACLRNet) upon SPAMs and NOAMs to restore stereo images by associating the features of two views based on the learned correspondence. Extensive experiments on five benchmark datasets demonstrate the effectiveness and generalization of the proposed method on three stereo image restoration tasks including super-resolution, denoising, and compression artifact reduction. Shengping Zhang, Wei Yu 0004, Feng Jiang 0001, Liqiang Nie, Hongxun Yao, Qingming Huang, Dacheng Tao |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | Toward Modalities Correlation for RGB-T TrackingabstractRecently, RGB-T tracking methods have made significant progress, demonstrating remarkable capabilities in addressing the complexities of tracking tasks within demanding environments. However, these methods overlook instability of modal validity in real-world scenarios. This limits the model’s ability to understand the correlation between modalities, thereby hindering the model’s ability to fully leverage the synergistic effects of RGB and TIR. To address this challenge, we propose a novel RGB-T tracking model named MCTrack, from the perspective of leveraging correlation among modalities. First, during the feature extraction stage, we design a novel module based on channel matching modeling to construct bidirectional channel context information flow for two modalities. By leveraging information flow, specific modalities correlation information can be transmitted to two modes, augmenting the correlation between the two modes adaptively. Subsequently, after the feature extraction network, the features of each modality are decoded and transformed to generate more correlated feature representations. During this stage, we extract distinctive and collective features by leveraging the correlation among modalities. Then fusing these features and generated search region features specifically for localization. This aids the model in comprehending the correlation between RGB and TIR under complex scenarios, thereby enhancing its ability to capture and utilize key features. Based on extensive experiments conducted on four popular RGB-T tracking benchmarks, our model demonstrates superior performance, particularly showcasing impressive results on the LasHeR dataset with an achieved Precision of 71.6%. Xiantao Hu, Bineng Zhong 0001, Qihua Liang, Shengping Zhang, Ning Li 0044, Xianxian Li |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Transformer Tracking via Frequency FusionabstractTransformer has achieved impressive progress in visual tracking due to their capability of global modeling, which enables them to learn low-frequency features(i.e., high-level semantic information). However, it seems to overlook the high-frequency features(i.e., low-level texture and edge information) which are crucial to identify different intra-class object instances in the tracking task. To address this issue, we propose a transformer based tracker via frequency fusion perspective that investigated whether high-frequency and low-frequency features can be effectively combined to achieve robust tracking. Specifically, we design a simple yet effective two-stage fusion strategy and use an appropriate frequency fusion strategy in tracking process of each stage so as to make full use of frequency domain information. In the feature extraction stage, we use wavelet decomposition of high-frequency subbands to solve the performance loss caused by the transformer’s catastrophic forgetting of high-frequency information. In the prediction head stage, we use a variety of wavelet decomposition subbands to model the multi-frequency information. The two-stage fusion strategy makes our model extract more balanced and beneficial multi-frequency information, enabling it to effectively capture target texture information and local edge information while also being sensitive to global information. Extensive experiments on six challenging benchmarks (i.e., LaSOT$_{ext}$, UAV123, TNL2K, LaSOT, TrackingNet, and GOT-10k) demonstrates the superior performance of our tracker. Xiantao Hu, Bineng Zhong 0001, Qihua Liang, Shengping Zhang, Ning Li 0044, Xianxian Li, Rongrong Ji |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | PBR-GAN: Imitating Physically-Based Rendering With Generative Adversarial NetworksabstractWe propose a Generative Adversarial Network (GAN)-based architecture for achieving high-quality physically based rendering (PBR). Conventional PBR relies heavily on ray tracing, which is computationally expensive in complicated environments. Some recent deep learning-based methods can improve efficiency but cannot deal with illumination variation well. In this paper, we propose PBR-GAN, an end-to-end GAN-based network that solves these problems while generating natural photo-realistic images. Two encoders (the shading encoder and albedo encoder) and two decoders (the image decoder and light decoder) are introduced to achieve our target. The two encoders and the image decoder constitute the generator that learns the mapping between the generated domain and the real domain. The light decoder produces light maps that pay more attention to the highlight and shadow regions. The discriminator aims to optimize the generator by distinguishing target images from the generated ones. Three novel loss items, concentrating on domain translation, overall shading preservation, and light map estimation, are proposed to optimize the photo-realistic outputs. Furthermore, a real dataset is collected to provide realistic information for training GAN architecture. Extensive experiments indicate that PBR-GAN can preserve the illumination variation and improve the image perceptual quality. Ru Li 0002, Peng Dai 0003, Guanghui Liu 0001, Shengping Zhang, Bing Zeng 0001, Shuaicheng Liu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Hybrid Transformers With Attention-Guided Spatial Embeddings for Makeup Transfer and RemovalabstractExisting makeup transfer methods typically transfer simple makeup colors in a well-conditioned face image and fail to handle makeup style details (e.g., complicated colors and shapes) and facial occlusion. To address these problems, this paper proposes Hybrid Transformers with Attention-guided Spatial Embeddings (named HT-ASE) for makeup transfer and removal. Specifically, a makeup context extractor adopts makeup context global-local interactions to aggregate the high-level context and low-level detail features of the makeup styles, which obtains the context-aware makeup features that encode the complicated colors and shapes of the makeup styles. A face identity extractor adopts a face identity local interaction to aggregate the identity-relevant features of shallow layers into identity semantic features, which refines the identity features. A spatially similarity-aware fusion network introduces a spatially-adaptive layer-instance normalization with attention-guided spatial embeddings to perform semantic alignment and fusion between the makeup and identity features, yielding precise and robust transfer results even with large spatial misalignment and facial occlusion. Extensive experimental results demonstrate that the proposed method outperforms the state-of-the-art methods, especially in the preservation of makeup style details and handling facial occlusion. Mingxiu Li, Wei Yu 0002, Qinglin Liu, Zonglin Li 0004, Ru Li 0002, Bineng Zhong 0001, Shengping Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2024 | Identity-Aware Variational Autoencoder for Face SwappingabstractFace swapping aims to transfer the identity of a source face to a target face image while preserving the target attributes (e.g., facial expression, head pose, illumination, and background). Most existing methods use a face recognition model to extract global features from the source face and directly fuse them with the target to generate a swapping result. However, identity-irrelevant attributes (e.g., hairstyle and facial appearances) contribute a lot to the recognition task, and thus swapping this task-specific feature inevitably interfuses source attributes with target ones. In this paper, we propose an identity-aware variational autoencoder (ID-VAE) based face swapping framework, dubbed VAFSwap, which learns disentangled identity and attribute representations for high-fidelity face swapping. In particular, we overcome the unpaired training barrier of VAE and impose a proxy identity on the latent space by exploiting the weak supervision from an auxiliary image set whose identity is averaged from multiple collected face images. To explicitly guide the identity fusion, we further devise an identity-associated matrix that corresponds different face regions with their identity representations to perform identity-related feature interactions. Finally, we incorporate spatial dimensions into the latent space and exploit the generative priors of a pre-trained face generator, allowing the effective elimination of noticeable swapping artifacts. Extensive experiments on the FaceForensics++ and CelebA-HQ datasets demonstrate that our method outperforms the state-of-the-art significantly. Zonglin Li 0004, Shengfeng He, Quanling Meng, Shengping Zhang, Bineng Zhong 0001, Rongrong Ji |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | End-to-End Human Instance MattingabstractHuman instance matting aims to estimate an alpha matte for each human instance in an image, which is extremely challenging and has rarely been studied so far. Despite some efforts to use instance segmentation to generate a trimap for each instance and apply trimap-based matting methods, the resulting alpha mattes are often inaccurate due to inaccurate segmentation. In addition, this approach is computationally inefficient due to multiple executions of the matting method. To address these problems, this paper proposes a novel End-to-End Human Instance Matting (E2E-HIM) framework for simultaneous multiple instance matting in a more efficient manner. Specifically, a general perception network first extracts image features and decodes instance contexts into latent codes. Then, a united guidance network exploits spatial attention and semantics embedding to generate united semantics guidance, which encodes the locations and semantic correspondences of all instances. Finally, an instance matting network decodes the image features and united semantics guidance to predict all instance-level alpha mattes. In addition, we construct a large-scale human instance matting dataset (HIM-100K) comprising over 100,000 human images with instance alpha matte labels. Experiments on HIM-100K demonstrate the proposed E2E-HIM outperforms the existing methods on human instance matting with 50% lower errors and 5× faster speed (6 instances in a 640 × 640 image). Experiments on the PPM-100, RWP-636, and P3M datasets demonstrate that E2E-HIM also achieves competitive performance on traditional human matting. Qinglin Liu, Shengping Zhang, Quanling Meng, Bineng Zhong 0001, Peiqiang Liu, Hongxun Yao |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Positive-Sample-Free Object Tracking via a Soft ConstraintabstractMost of the existing bounding box-based trackers rely on a classification subnetwork and a regression subnetwork to predict the location and scale of the bounding box. They learn the classification subnetwork by processing each sample individually and applying the suggested classification confidence to produce the final prediction. They typically involve heuristic positive sample configurations, which inevitably introduce mislabelled training samples and therefore deteriorate their tracking performance. Moreover, the parallel prediction of the bounding box position and scale may lead to misalignment of classification and regression. To address these issues,we propose a simple yet effective soft constraint-based tracking framework without positive samples (named SoftCT). SoftCT adaptively senses the target’s pixel position through a soft constraint mechanism, which eliminates potential performance gaps caused by artificially marking the target’s pixel position. In addition, SoftCT computes the state of the bounding box by aggregating such positional information, thereby allowing the tracker to avoid misalignment in classification and regression due to uninformed communication. Specifically, SoftCT directly senses the position of the target pixel and fuses this information into the bounding box prediction, rather than requiring explicit annotation or regression of the target pixel. Extensive experiments on six tracking benchmarks including GOT-10k, TrackingNet, LaSOT, UAV123, LaSOText and TNL2K demonstrate that our tracker achieves state-of-the-art performance, confirming its effectiveness and efficiency. Jiaxin Ye, Bineng Zhong 0001, Qihua Liang, Shengping Zhang, Xianxian Li, Rongrong Ji |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Rate-Adaptive Neural Network for Image Compressive SensingabstractDeep learning-based image compressive sensing (CS) methods have achieved great success in the past few years. However, most of them are content-independent, with a spatially uniform sampling rate allocation for the entire image. Such practises may potentially degrade the performance of image CS with block-based sampling, since the content of different blocks in an image is different. In this article, we propose a novel rate-adaptive image CS neural network (dubbed RACSNet) to achieve adaptive sampling rate allocation based on the content characteristics of the image with a single model. Specifically, a measurement domain-based reconstruction distortion is first used to guide the sampling rate allocation for different blocks in an image without access to the ground truth image. Then, a step-wise training strategy is designed to train a reusable sampling matrix, which is capable of sampling image blocks to generate the compressed measurements under arbitrary sampling rates. Subsequently, a pyramid-shaped initial reconstruction sub-network and a hierarchical deep reconstruction sub-network that fuse the measurement information of different scales are put forward to reconstruct image blocks from the compressed measurements. Finally, a reconstruction distortion map and an improved loss function are developed to eliminate the blocking artifacts and further enhance the CS reconstruction. Experimental results on both objective metrics and subjective visual qualities show that the proposed RACSNet achieves significant improvements over the state-of-the-art methods. Shengping Zhang, Wenxue Cui, Shaohui Liu, Feng Jiang 0001, Debin Zhao |
IEEE Trans. Multim. | 2 |
| 2024 | Limb-Aware Virtual Try-On Network With Progressive Clothing WarpingabstractImage-based virtual try-on aims to transfer an in-shop clothing image to a person image. Most existing methods adopt a single global deformation to perform clothing warping directly, which lacks fine-grained modeling of in-shop clothing and leads to distorted clothing appearance. In addition, existing methods usually fail to generate limb details well because they are limited by the used clothing-agnostic person representation without referring to the limb textures of the person image. To address these problems, we propose Limb-aware Virtual Try-on Network named PL-VTON, which performs fine-grained clothing warping progressively and generates high-quality try-on results with realistic limb details. Specifically, we present Progressive Clothing Warping (PCW) that explicitly models the location and size of in-shop clothing and utilizes a two-stage alignment strategy to progressively align the in-shop clothing with the human body. Moreover, a novel gravity-aware loss that considers the fit of the person wearing clothing is adopted to better handle the clothing edges. Then, we design Person Parsing Estimator (PPE) with a non-limb target parsing map to semantically divide the person into various regions, which provides structural constraints on the human body and therefore alleviates texture bleeding between clothing and body regions. Finally, we introduce Limb-aware Texture Fusion (LTF) that focuses on generating realistic details in limb regions, where a coarse try-on result is first generated by fusing the warped clothing image with the person image, then limb textures are further fused with the coarse result under limb-aware guidance to refine limb details. Extensive experiments demonstrate that our PL-VTON outperforms the state-of-the-art methods both qualitatively and quantitatively. Shengping Zhang, Weigang Zhang, Xiangyuan Lan, Hongxun Yao, Qingming Huang |
IEEE Trans. Multim. | 1 |
| 2024 | Human Selective MattingabstractExisting human matting methods are incapable of accurately estimating the alpha mattes of arbitrarily selected humans from a group photo. An alternative solution is to apply them to the corresponding cropped image patches. However, this option obtains an inaccurate alpha estimation due to the interference of the body parts of the neighboring humans. In addition, these methods are only trained on finely annotated synthetic data, which causes poor performance in real-world scenarios due to the domain shift. To address these problems, we propose human selective matting (HSMatt), which performs matting for arbitrarily selected humans from a group photo given only a simple bounding box as guidance. Specifically, we design a global–local context network to extract both local and global semantic context features. A human-aware trimap network is then proposed to generate human-aware trimaps for the selected humans, which adopts stacked bidirectional inference modules with intermediate supervision to progressively refine the estimated trimap. Finally, a partially supervised matting network is introduced to estimate the alpha matte, which uses a sample-varying loss to train the network on both the finely annotated synthetic data and coarsely annotated real-world data, resulting in high accuracy and good generalization. To evaluate the proposed HSMatt, we construct the first human selective matting dataset, named HSM-200K, which contains over 200,000 human images with instance-level alpha matte annotations. Experimental results demonstrate that the proposed HSMatt outperforms state-of-the-art methods. Qinglin Liu, Quanling Meng, Xiaoqian Lv, Zonglin Li 0004, Wei Yu 0004, Shengping Zhang |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2023 | CaPhy: Capturing Physical Properties for Animatable Human AvatarsabstractWe present CaPhy, a novel method for reconstructing animatable human avatars with realistic dynamic properties for clothing. Specifically, we aim for capturing the geometric and physical properties of the clothing from real observations. This allows us to apply novel poses to the human avatar with physically correct deformations and wrinkles of the clothing. To this end, we combine unsupervised training with physics-based losses and 3D-supervised training using scanned data to reconstruct a dynamic model of clothing that is physically realistic and conforms to the human scans. We also optimize the physical parameters of the underlying physical model from the scans by introducing gradient constraints of the physics-based losses. In contrast to previous work on 3D avatar reconstruction, our method is able to generalize to novel poses with realistic dynamic cloth deformations. Experiments on several subjects demonstrate that our method can estimate the physical properties of the garments, resulting in superior quantitative and qualitative results compared with previous methods. Zhaoqi Su, Liangxiao Hu, Siyou Lin, Hongwen Zhang 0001, Shengping Zhang, Justus Thies, Yebin Liu |
ICCV | 5 |
| 2023 | Rectangular-Output Image StitchingabstractImage stitching aims to combine two images with overlapping fields to expand the field-of-view (FoV). However, the stitched images of existing methods are irregular, and need to be processed by rectangling methods, which is time-consuming and prone to be unnatural. In this paper, we propose the first end-to-end framework, Rectangular-output Deep Image Stitching Network (RDISNet), to directly stitch two images into a standard rectangular image while learning color consistency between image pairs and maintaining the authenticity of the content. To further preserve the structure of large objects in the stitched image, we design a dilated BN-RCU block to expand the receptive field of RDISNet for extracting enriched spatial context. Furthermore, we design a novel data synthesis pipeline and build the first rectangular-output deep image stitching dataset (RDIS-D) for jointing image stitching and rectangling. Experimental results demonstrate that RDISNet performs favorably against the state-of-the-art methods. Hongfei Zhou, Yuhe Zhu, Xiaoqian Lv, Qinglin Liu, Shengping Zhang |
ICIP | 5 |
| 2023 | Interactive Object Placement with Reinforcement LearningabstractObject placement aims to insert a foreground object into a background image with a suitable location and size to create a natural composition. To predict a diverse distribution of placements, existing methods usually establish a one-to-one mapping from random vectors to the placements. However, these random vectors are not interpretable, which prevents users from interacting with the object placement process. To address this problem, we propose an Interactive Object Placement method with Reinforcement Learning, dubbed IOPRE, to make sequential decisions for producing a reasonable placement given an initial location and size of the foreground. We first design a novel action space to flexibly and stably adjust the location and size of the foreground while preserving its aspect ratio. Then, we propose a multi-factor state representation learning method, which integrates composition image features and sinusoidal positional embeddings of the foreground to make decisions for selecting actions. Finally, we design a hybrid reward function that combines placement assessment and the number of steps to ensure that the agent learns to place objects in the most visually pleasing and semantically appropriate location. Experimental results on the OPA dataset demonstrate that the proposed method achieves state-of-the-art performance in terms of plausibility and diversity. Shengping Zhang, Quanling Meng, Qinglin Liu, Liqiang Nie, Bineng Zhong 0001, Xiaopeng Fan 0001, Rongrong Ji |
ICML | 1 |
| 2023 | Attention guided domain alignment for conditional face image generation
Zonglin Li 0004, Shengping Zhang, Quanling Meng, Qinglin Liu, Huiyu Zhou 0001 |
Comput. Vis. Image Underst. | 2 |
| 2023 | Learning Geometric Transformation for Point Cloud Completion
Shengping Zhang, Xianzhu Liu, Haozhe Xie, Liqiang Nie, Huiyu Zhou 0001, Dacheng Tao, Xuelong Li 0001 |
Int. J. Comput. Vis. | 1 |
| 2023 | Learning Enriched Hop-Aware Correlation for Robust 3D Human Pose Estimation
Shengping Zhang, Chenyang Wang 0002, Liqiang Nie, Hongxun Yao, Qingming Huang, Qi Tian 0001 |
Int. J. Comput. Vis. | 1 |
| 2023 | Correction to: Learning Enriched Hop-Aware Correlation for Robust 3D Human Pose Estimation
Shengping Zhang, Chenyang Wang 0002, Liqiang Nie, Hongxun Yao, Qingming Huang, Qi Tian 0001 |
Int. J. Comput. Vis. | 1 |
| 2023 | Scale-Aware Frequency Attention network for super-resolution
Wei Yu 0004, Zonglin Li 0004, Qinglin Liu, Feng Jiang 0001, Changyong Guo, Shengping Zhang |
Neurocomputing | 6 |
| 2023 | SiamBAN: Target-Aware Tracking With Siamese Box Adaptive NetworkabstractVariation of scales or aspect ratios has been one of the main challenges for tracking. To overcome this challenge, most existing methods adopt either multi-scale search or anchor-based schemes, which use a predefined search space in a handcrafted way and therefore limit their performance in complicated scenes. To address this problem, recent anchor-free based trackers have been proposed without using prior scale or anchor information. However, an inconsistency problem between classification and regression degrades the tracking performance. To address the above issues, we propose a simple yet effective tracker (named Siamese Box Adaptive Network, SiamBAN) to learn a target-aware scale handling schema in a data-driven manner. Our basic idea is to predict the target boxes in a per-pixel fashion through a fully convolutional network, which is anchor-free. Specifically, SiamBAN divides the tracking problem into classification and regression tasks, which directly predict objectiveness and regress bounding boxes, respectively. A no-prior box design is proposed to avoid tuning hyper-parameters related to candidate boxes, which makes SiamBAN more flexible. SiamBAN further uses a target-aware branch to address the inconsistency problem. Experiments on benchmarks including VOT2018, VOT2019, OTB100, UAV123, LaSOT and TrackingNet show that SiamBAN achieves promising performance and runs at 35 FPS. Zedu Chen, Bineng Zhong 0001, Guorong Li, Shengping Zhang, Rongrong Ji, Zhenjun Tang, Xianxian Li |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Unsupervised Low-Light Video Enhancement With Spatial-Temporal Co-Attention TransformerabstractExisting low-light video enhancement methods are dominated by Convolution Neural Networks (CNNs) that are trained in a supervised manner. Due to the difficulty of collecting paired dynamic low/normal-light videos in real-world scenes, they are usually trained on synthetic, static, and uniform motion videos, which undermines their generalization to real-world scenes. Additionally, these methods typically suffer from temporal inconsistency (e.g., flickering artifacts and motion blurs) when handling large-scale motions since the local perception property of CNNs limits them to model long-range dependencies in both spatial and temporal domains. To address these problems, we propose the first unsupervised method for low-light video enhancement to our best knowledge, named LightenFormer, which models long-range intra- and inter-frame dependencies with a spatial-temporal co-attention transformer to enhance brightness while maintaining temporal consistency. Specifically, an effective but lightweight S-curve Estimation Network (SCENet) is first proposed to estimate pixel-wise S-shaped non-linear curves (S-curves) to adaptively adjust the dynamic range of an input video. Next, to model the temporal consistency of the video, we present a Spatial-Temporal Refinement Network (STRNet) to refine the enhanced video. The core module of STRNet is a novel Spatial-Temporal Co-attention Transformer (STCAT), which exploits multi-scale self- and cross-attention interactions to capture long-range correlations in both spatial and temporal domains among frames for implicit motion estimation. To achieve unsupervised training, we further propose two non-reference loss functions based on the invertibility of the S-curve and the noise independence among frames. Extensive experiments on the SDSD and LLIV-Phone datasets demonstrate that our LightenFormer outperforms state-of-the-art methods. Xiaoqian Lv, Shengping Zhang, Chenyang Wang 0002, Weigang Zhang, Hongxun Yao, Qingming Huang |
IEEE Trans. Image Process. | 2 |
| 2023 | Automatic Shadow Generation via Exposure FusionabstractShadow generation aims to generate a plausible shadow for the inserted foreground object in a composite image. Besides the composite image and the associated mask of the inserted foreground object, existing methods also require a mask of all background objects as well as their shadows as an auxiliary input, which is laborious in practical applications. Meanwhile, most existing methods use a linear illumination transformation to darken the shadow region, which is prone to produce unrealistic shadows especially when background illumination is complex. To address these problems, this paper proposes an automatic shadow generation method, which avoids the laborious acquisition of the background object masks while harmonizing the shadow region to achieve plausible shadow effects. Specifically, to implicitly exploit background illumination to infer the shadow shape of the inserted foreground object, we first propose a Hierarchy Attention U-Net (HAU-Net) to sequentially build global interactions between the foreground object and background across spatial and channel dimensions. Since the spatial-variant property of the shadow, we formulate shadow harmonization as an exposure fusion problem and propose an Illumination-Aware Fusion Network (IFNet), which uses an improved illumination model with a double linear transformation to produce multiple under-exposure images of the shadow region. IFNet then learns pixel-wise fusion kernels that consider the local smoothness of the shadow to fuse the composite image with these under-exposure images to generate the realistic shadow of the foreground object. Extensive experiments on the DESOBA and Shadow-AR datasets demonstrate that our method achieves state-of-the-art performance for shadow generation on both the BOS and BOS-free test images. Quanling Meng, Shengping Zhang, Zonglin Li 0004, Chenyang Wang 0002, Weigang Zhang, Qingming Huang |
IEEE Trans. Multim. | 2 |
| 2022 | Lightweight Image Matting via Efficient Non-local Guidance
Zhaoxiang Kang, Zonglin Li 0004, Qinglin Liu, Yuhe Zhu, Hongfei Zhou, Shengping Zhang |
ACCV (2) | 6 |
| 2022 | Natural Image Matting with Shifted Window Self-AttentionabstractNatural image matting is a challenging and significant task in computer vision. Recently, image matting achieves fantastic development by introducing deep learning methods. To the best of our knowledge, there is no image matting method using the Transformer. Compared with CNNs, the Transformer pays more attention to the interest points and the relationships of content, which is beneficial to the image matting task. In this paper, we first present a novel Transformer-based image matting method with Shifted Window self-Attention. Specifically, our method contains two encoders, an alpha encoder and a context encoder. The former leverages the Transformer with Shifted Window self-Attention to extract features of details, such as hairs, feathers and porous parts of foreground objects. Shifted Window self-Attention focuses on patches with the size of the window and connections of adjacent patches. With this, the Transformer is capable of dealing with high-resolution images. The context encoder, which takes rescaled images as input, aims to extract the whole structure information of foreground objects. Then, we propose a novel Hierarchical Pyramid Pooling Module (HPPM) which enables the network to have the flexibility to extract features at various resolutions. Experiments show that our method achieves competitive performance on the Composition-1K dataset. Yang Liu 0119, Zonglin Li 0004, Chenyang Wang 0002, Shengping Zhang |
ICIP | 5 |
| 2022 | Progressive Limb-Aware Virtual Try-OnabstractExisting image-based virtual try-on methods directly transfer specific clothing to a human image without utilizing clothing attributes to refine the transferred clothing geometry and textures, which causes incomplete and blurred clothing appearances. In addition, these methods usually mask the limb textures of the input for the clothing-agnostic person representation, which results in inaccurate predictions for human limb regions (i.e., the exposed arm skin), especially when transforming between long-sleeved and short-sleeved garments. To address these problems, we present a progressive virtual try-on framework, named PL-VTON, which performs pixel-level clothing warping based on multiple attributes of clothing and embeds explicit limb-aware features to generate photo-realistic try-on results. Specifically, we design a Multi-attribute Clothing Warping (MCW) module that adopts a two-stage alignment strategy based on multiple attributes to progressively estimate pixel-level clothing displacements. A Human Parsing Estimator (HPE) is then introduced to semantically divide the person into various regions, which provides structural constraints on the human body and therefore alleviates texture bleeding between clothing and limb regions. Finally, we propose a Limb-aware Texture Fusion (LTF) module to estimate high-quality details in limb regions by fusing textures of the clothing and the human body with the guidance of explicit limb-aware features. Extensive experiments demonstrate that our proposed method outperforms the state-of-the-art virtual try-on methods both qualitatively and quantitatively. Shengping Zhang, Qinglin Liu, Zonglin Li 0004, Chenyang Wang 0002 |
ACM Multimedia | 2 |
| 2022 | BacklitNet: A dataset and network for backlit image enhancement
Xiaoqian Lv, Shengping Zhang, Qinglin Liu, Haozhe Xie, Bineng Zhong 0001, Huiyu Zhou 0001 |
Comput. Vis. Image Underst. | 2 |
| 2022 | Continuous Prediction of Lower-Limb Kinematics From Multi-Modal Biomedical SignalsabstractThe fast-growing techniques of measuring and fusing multi-modal biomedical signals enable advanced motor intent decoding schemes of lower-limb exoskeletons, meeting the increasing demand for rehabilitative or assistive applications of take-home healthcare. Challenges of exoskeletons’ motor intent decoding schemes remain in making a continuous prediction to compensate for the hysteretic response caused by mechanical transmission. In this paper, we solve this problem by proposing an ahead-of-time continuous prediction of lower-limb kinematics, with the prediction of knee angles during level walking as a case study. Firstly, an end-to-end kinematics prediction network(KinPreNet),1consisting of a feature extractor and an angle predictor, is proposed and experimentally compared with features and methods traditionally used in ahead-of-time prediction of gait phases. Secondly, inspired by the electromechanical delay(EMD), we further explore our algorithm’s capability of compensating response delay of mechanical transmission by validating the performance of the different sections of prediction time. And we experimentally reveal the time boundary of compensating the hysteretic response. Thirdly, a comparison of employing EMG signals or not is performed to reveal the EMG and kinematic signals’ collaborated contributions to the continuous prediction. During the experiments, EMG signals of nine muscles and knee angles calculated from inertial measurement unit (IMU) signals are recorded from ten healthy subjects. Our algorithm can predict knee angles with the averaged RMSE of 3.98 deg which is better than the 15.95-deg averaged RMSE of utilizing the traditional methods of ahead-of-time prediction. The best prediction time is in the interval of 27ms and 108ms. To the best of our knowledge, this is the first study of continuously predicting lower-limb kinematics in an ahead-of-time manner based on the electromechanical delay (EMD). Chunzhi Yi, Feng Jiang 0001, Shengping Zhang, Hao Guo 0015, Chifu Yang, Zhen Ding, Baichun Wei, Xiangyuan Lan, Huiyu Zhou 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | Efficient Regional Memory Network for Video Object SegmentationabstractRecently, several Space-Time Memory based networks have shown that the object cues (e.g. video frames as well as the segmented object masks) from the past frames are useful for segmenting objects in the current frame. However, these methods exploit the information from the memory by global-to-global matching between the current and past frames, which lead to mismatching to similar objects and high computational complexity. To address these problems, we propose a novel local-to-local matching solution for semi-supervised VOS, namely Regional Memory Network (RMNet). In RMNet, the precise regional memory is constructed by memorizing local regions where the target objects appear in the past frames. For the current query frame, the query regions are tracked and predicted based on the optical flow estimated from the previous frame. The proposed local-to-local matching effectively alleviates the ambiguity of similar objects in both memory and query frames, which allows the information to be passed from the regional memory to the query region efficiently and effectively. Experimental results indicate that the proposed RM-Net performs favorably against state-of-the-art methods on the DAVIS and YouTube-VOS datasets. Haozhe Xie, Hongxun Yao, Shangchen Zhou, Shengping Zhang, Wenxiu Sun |
CVPR | 4 |
| 2021 | Distractor-Aware Fast Tracking via Dynamic Convolutions and MOT PhilosophyabstractA practical long-term tracker typically contains three key properties, i.e. an efficient model design, an effective global re-detection strategy and a robust distractor awareness mechanism. However, most state-of-the-art long-term trackers (e.g., Pseudo and re-detecting based ones) do not take all three key properties into account and therefore may either be time-consuming or drift to distractors. To address the issues, we propose a two-task tracking framework (named DMTrack), which utilizes two core components (i.e., one-shot detection and re-identification (re-id) association) to achieve distractor-aware fast tracking via Dynamic convolutions (d-convs) and Multiple object tracking (MOT) philosophy. To achieve precise and fast global detection, we construct a lightweight one-shot detector using a novel dynamic convolutions generation method, which provides a unified and more flexible way for fusing target information into the search field. To distinguish the target from distractors, we resort to the philosophy of MOT to reason distractors explicitly by maintaining all potential similarities’ tracklets. Benefited from the strength of high recall detection and explicit object association, our tracker achieves state-of-the-art performance on the LaSOT, Ox-UvA, TLP, VOT2018LT and VOT2019LT benchmarks and runs in real-time (3x faster than comparisons)1. Zikai Zhang 0003, Bineng Zhong 0001, Shengping Zhang, Zhenjun Tang, Xin Liu 0011, Zhaoxiang Zhang 0001 |
CVPR | 3 |
| 2021 | Long-Range Feature Propagating for Natural Image MattingabstractNatural image matting estimates the alpha values of unknown regions in the trimap. Recently, deep learning based methods propagate the alpha values from the known regions to unknown regions according to the similarity between them. However, we find that more than 50% pixels in the unknown regions cannot be correlated to pixels in known regions due to the limitation of small effective reception fields of common convolutional neural networks, which leads to inaccurate estimation when the pixels in the unknown regions cannot be inferred only with pixels in the reception fields. To solve this problem, we propose Long-Range Feature Propagating Network (LFPNet), which learns the long-range context features outside the reception fields for alpha matte estimation. Specifically, we first design the propagating module which extracts the context features from the downsampled image. Then, we present Center-Surround Pyramid Pooling (CSPP) that explicitly propagates the context features from the surrounding context image patch to the inner center image patch. Finally, we use the matting module which takes the image, trimap and context features to estimate the alpha matte. Experimental results demonstrate that the proposed method performs favorably against the state-of-the-art methods on the AlphaMatting and Adobe Image Matting datasets. Qinglin Liu, Haozhe Xie, Shengping Zhang, Bineng Zhong 0001, Rongrong Ji |
ACM Multimedia | 3 |
| 2021 | Toward 3D object reconstruction from stereo images
Haozhe Xie, Hongxun Yao, Shangchen Zhou, Shengping Zhang, Xiaojun Tong, Wenxiu Sun |
Neurocomputing | 4 |
| 2021 | Sketch-specific data augmentation for freehand sketch recognition
Ying Zheng 0009, Hongxun Yao, Xiaoshuai Sun, Shengping Zhang, Sicheng Zhao, Fatih Porikli |
Neurocomputing | 4 |
| 2021 | Is It Easy to Recognize Baby's Age and Gender?
Yang Liu 0119, Ruili He, Xiaoqian Lv, Wei Wang 0107, Xin Sun 0003, Shengping Zhang |
J. Comput. Sci. Technol. | 6 |
| 2021 | Editorial: Recent Advances on the Mobile Multimedia Services and Applications
Shengping Zhang, Yanxiao Zhao, Dalei Wu, Qing Yang 0003 |
Mob. Networks Appl. | 1 |
| 2021 | Low-light image enhancement via deep Retinex decomposition and bilateral learning
Xiaoqian Lv, Yujing Sun 0004, Jun Zhang 0017, Feng Jiang 0001, Shengping Zhang |
Signal Process. Image Commun. | 5 |
| 2020 | Overwater Image Dehazing via Cycle-Consistent Generative Adversarial Network
Shunyuan Zheng, Jiamin Sun, Qinglin Liu, Yuankai Qi, Shengping Zhang |
ACCV (2) | 5 |
| 2020 | Siamese Box Adaptive Network for Visual TrackingabstractMost of the existing trackers usually rely on either a multi-scale searching scheme or pre-defined anchor boxes to accurately estimate the scale and aspect ratio of a target. Unfortunately, they typically call for tedious and heuristic configurations. To address this issue, we propose a simple yet effective visual tracking framework (named Siamese Box Adaptive Network, SiamBAN) by exploiting the expressive power of the fully convolutional network (FCN). SiamBAN views the visual tracking problem as a parallel classification and regression problem, and thus directly classifies objects and regresses their bounding boxes in a unified FCN. The no-prior box design avoids hyper-parameters associated with the candidate boxes, making SiamBAN more flexible and general. Extensive experiments on visual tracking benchmarks including VOT2018, VOT2019, OTB100, NFS, UAV123, and LaSOT demonstrate that SiamBAN achieves state-of-the-art performance and runs at 40 FPS, confirming its effectiveness and efficiency. The code will be available at https://github.com/hqucv/siamban. Zedu Chen, Bineng Zhong 0001, Guorong Li, Shengping Zhang, Rongrong Ji |
CVPR | 4 |
| 2020 | Object-and-Action Aware Model for Visual Language Navigation
Yuankai Qi, Zizheng Pan, Shengping Zhang, Anton van den Hengel, Qi Wu 0001 |
ECCV (10) | 3 |
| 2020 | GRNet: Gridding Residual Network for Dense Point Cloud Completion
Haozhe Xie, Hongxun Yao, Shangchen Zhou, Jiageng Mao, Shengping Zhang, Wenxiu Sun |
ECCV (9) | 5 |
| 2020 | An Effective Way to Boost Black-Box Adversarial Attack
Xinjie Feng, Hongxun Yao, Wenbin Che, Shengping Zhang |
MMM (1) | 4 |
| 2020 | Pix2Vox++: Multi-scale Context-aware 3D Object Reconstruction from Single and Multiple Images
Haozhe Xie, Hongxun Yao, Shengping Zhang, Shangchen Zhou, Wenxiu Sun |
Int. J. Comput. Vis. | 3 |
| 2020 | Modality-correlation-aware sparse representation for RGB-infrared object tracking
Xiangyuan Lan, Mang Ye, Shengping Zhang, Huiyu Zhou 0001, Pong C. Yuen |
Pattern Recognit. Lett. | 3 |
| 2020 | Siamese Local and Global Networks for Robust Face TrackingabstractConvolutional neural networks (CNNs) have achieved great success in several face-related tasks, such as face detection, alignment and recognition. As a fundamental problem in computer vision, face tracking plays a crucial role in various applications, such as video surveillance, human emotion detection and human-computer interaction. However, few CNN-based approaches are proposed for face (bounding box) tracking. In this paper, we propose a face tracking method based on Siamese CNNs, which takes advantages of powerful representations of hierarchical CNN features learned from massive face images. The proposed method captures discriminative face information at both local and global levels. At the local level, representations for attribute patches (i.e:, eyes, nose and mouth) are learned to distinguish a face from another one, which are robust to pose changes and occlusions. At the global level, representations for each whole face are learned, which take into account the spatial relationships among local patches and facial characters, such as skin color and nevus. In addition, we build a new largescale challenging face tracking dataset to evaluate face tracking methods and to facilitate the research forward in this field. Extensive experiments on the collected dataset demonstrate the effectiveness of our method in comparison to several state-of-theart visual tracking methods. Yuankai Qi, Shengping Zhang, Feng Jiang 0001, Huiyu Zhou 0001, Dacheng Tao, Xuelong Li 0001 |
IEEE Trans. Image Process. | 2 |
| 2020 | Light Field Saliency Detection With Deep Convolutional NetworksabstractLight field imaging presents an attractive alternative to RGB imaging because of the recording of the direction of the incoming light. The detection of salient regions in a light field image benefits from the additional modeling of angular patterns. For RGB imaging, methods using CNNs have achieved excellent results on a range of tasks, including saliency detection. However, it is not trivial to use CNN-based methods for saliency detection on light field images because these methods are not specifically designed for processing light field inputs. In addition, current light field datasets are not sufficiently large to train CNNs. To overcome these issues, we present a new Lytro Illum dataset, which contains 640 light fields and their corresponding ground-truth saliency maps. Compared to current publicly available light field saliency datasets [1], [2], our new dataset is larger, of higher quality, contains more variation and more types of light field inputs. This makes our dataset suitable for training deeper networks and benchmarking. Furthermore, we propose a novel end-to-end CNN-based framework for light field saliency detection. Specifically, we propose three novel MAC (Model Angular Changes) blocks to process light field micro-lens images. We systematically study the impact of different architecture variants and compare light field saliency with regular 2D saliency. Our extensive comparisons indicate that our novel network significantly outperforms state-of-the-art methods on the proposed dataset and has desired generalization abilities on other existing datasets. Jun Zhang 0017, Yamei Liu, Shengping Zhang, Ronald Poppe, Meng Wang 0001 |
IEEE Trans. Image Process. | 3 |
| 2020 | Introduction to the Special Issue on Multimodal Machine Learning for Human Behavior AnalysisabstractNo abstract available. Shengping Zhang, Huiyu Zhou 0001, Dong Xu 0001, M. Emre Celebi 0001, Thierry Bouwmans |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2019 | Learning Attribute-Specific Representations for Visual TrackingabstractIn recent years, convolutional neural networks (CNNs) have achieved great success in visual tracking. Most of existing methods train or fine-tune a binary classifier to distinguish the target from its background. However, they may suffer from the performance degradation due to insufficient training data. In this paper, we show that attribute information (e.g., illumination changes, occlusion and motion) in the context facilitates training an effective classifier for visual tracking. In particular, we design an attribute-based CNN with multiple branches, where each branch is responsible for classifying the target under a specific attribute. Such a design reduces the appearance diversity of the target under each attribute and thus requires less data to train the model. We combine all attributespecific features via ensemble layers to obtain more discriminative representations for the final target/background classification. The proposed method achieves favorable performance on the OTB100 dataset compared to state-of-the-art tracking methods. After being trained on the VOT datasets, the proposed network also shows a good generalization ability on the UAV-Traffic dataset, which has significantly different attributes and target appearances with the VOT datasets. Yuankai Qi, Shengping Zhang, Weigang Zhang, Li Su 0003, Qingming Huang, Ming-Hsuan Yang 0001 |
AAAI | 2 |
| 2019 | Pix2Vox: Context-Aware 3D Reconstruction From Single and Multi-View ImagesabstractRecovering the 3D representation of an object from single-view or multi-view RGB images by deep neural networks has attracted increasing attention in the past few years. Several mainstream works (e.g., 3D-R2N2) use recurrent neural networks (RNNs) to fuse multiple feature maps extracted from input images sequentially. However, when given the same set of input images with different orders, RNN-based approaches are unable to produce consistent reconstruction results. Moreover, due to long-term memory loss, RNNs cannot fully exploit input images to refine reconstruction results. To solve these problems, we propose a novel framework for single-view and multi-view 3D reconstruction, named Pix2Vox. By using a well-designed encoder-decoder, it generates a coarse 3D volume from each input image. Then, a context-aware fusion module is introduced to adaptively select high-quality reconstructions for each part (e.g., table legs) from different coarse 3D volumes to obtain a fused 3D volume. Finally, a refiner further refines the fused 3D volume to generate the final output. Experimental results on the ShapeNet and Pix3D benchmarks indicate that the proposed Pix2Vox outperforms state-of-the-arts by a large margin. Furthermore, the proposed method is 24 times faster than 3D-R2N2 in terms of backward inference time. The experiments on ShapeNet unseen 3D categories have shown the superior generalization abilities of our method. Haozhe Xie, Hongxun Yao, Xiaoshuai Sun, Shangchen Zhou, Shengping Zhang |
ICCV | 5 |
| 2019 | Proposal-Refined Weakly Supervised Object Detection in Underwater Images
Xiaoqian Lv, Qinglin Liu, Jiamin Sun, Shengping Zhang |
ICIG (1) | 5 |
| 2019 | Learning a Reliable Decision Making Policy for Robust TrackingabstractRecent years deep learning based visual object trackers have achieved state-of-the-art performance on multiple benchmarks. However, most of these trackers lack an effective mechanism to avoid the wrong template update or re-detect the object when unreliable tracking result appears. In this paper, a novel tracking framework consisting of a tracking network for locating the target and a policy network for decision making is proposed. Firstly, during the off-line training phase, a variant of policy gradient algorithm is adopted, which makes the model converge better and faster. Secondly, current response map and history response map are both fed to the policy network to check the reliability of the tracking result, which effectively distinguishes the response diversity. Finally, an efficient redetection module is proposed to filter a large number of searching areas, which greatly improves the speed. Our proposed algorithm is measured on OTB dataset. Assessment results show that our tracking algorithm improves performance by 5%-6% at the expense of only a small amount of speed. Xiaofeng Huang, Kanghao Wang, Haibing Yin, Shengsheng Zheng, Shengping Zhang |
VCIP | 6 |
| 2019 | Robust visual tracking via scale-and-state-awareness
Yuankai Qi, Shengping Zhang, Qingming Huang, Hongxun Yao |
Neurocomputing | 3 |
| 2019 | Action recognition with multi-scale trajectory-pooled 3D convolutional descriptors
Xiusheng Lu, Hongxun Yao, Sicheng Zhao, Xiaoshuai Sun, Shengping Zhang |
Multim. Tools Appl. | 5 |
| 2019 | Hedging Deep Features for Visual TrackingabstractConvolutional Neural Networks (CNNs) have been applied to visual tracking with demonstrated success in recent years. Most CNN-based trackers utilize hierarchical features extracted from a certain layer to represent the target. However, features from a certain layer are not always effective for distinguishing the target object from the backgrounds especially in the presence of complicated interfering factors (e.g., heavy occlusion, background clutter, illumination variation, and shape deformation). In this work, we propose a CNN-based tracking algorithm which hedges deep features from different CNN layers to better distinguish target objects and background clutters. Correlation filters are applied to feature maps of each CNN layer to construct a weak tracker, and all weak trackers are hedged into a strong one. For robust visual tracking, we propose a hedge method to adaptively determine weights of weak classifiers by considering both the difference between the historical as well as instantaneous performance, and the difference among all weak trackers over time. In addition, we design a Siamese network to define the loss of each weak tracker for the proposed hedge method. Extensive experiments on large benchmark datasets demonstrate the effectiveness of the proposed algorithm against the state-of-the-art tracking methods. Yuankai Qi, Shengping Zhang, Qingming Huang, Hongxun Yao, Jongwoo Lim, Ming-Hsuan Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2019 | Hierarchical residual learning for image denoising
Wuzhen Shi, Feng Jiang 0001, Shengping Zhang, Rui Wang 0093, Debin Zhao, Huiyu Zhou 0001 |
Signal Process. Image Commun. | 3 |
| 2019 | Context-Aware Mouse Behavior Recognition Using Hidden Markov ModelsabstractAutomated recognition of mouse behaviors is crucial in studying psychiatric and neurologic diseases. To achieve this objective, it is very important to analyze the temporal dynamics of mouse behaviors. In particular, the change between mouse neighboring actions is swift in a short period. In this paper, we develop and implement a novel hidden Markov model (HMM) algorithm to describe the temporal characteristics of mouse behaviors. In particular, we here propose a hybrid deep learning architecture, where the first unsupervised layer relies on an advanced spatial-temporal segment Fisher vector encoding both visual and contextual features. Subsequent supervised layers based on our segment aggregate network are trained to estimate the state-dependent observation probabilities of the HMM. The proposed architecture shows the ability to discriminate between visually similar behaviors and results in high recognition rates with the strength of processing imbalanced mouse behavior datasets. Finally, we evaluate our approach using JHuang's and our own datasets, and the results show that our method outperforms other state-of-the-art approaches. Zheheng Jiang, Danny Crookes, Brian Desmond Green, Haiping Ma, Ling Li 0010, Shengping Zhang, Dacheng Tao, Huiyu Zhou 0001 |
IEEE Trans. Image Process. | 7 |
| 2018 | Robust Collaborative Discriminative Learning for RGB-Infrared TrackingabstractTracking target of interests is an important step for motion perception in intelligent video surveillance systems. While most recently developed tracking algorithms are grounded in RGB image sequences, it should be noted that information from RGB modality is not always reliable (e.g. in a dark environment with poor lighting condition), which urges the need to integrate information from infrared modality for effective tracking because of the insensitivity to illumination condition of infrared thermal camera. However, several issues encountered during the tracking process limit the fusing performance of these heterogeneous modalities: 1) the cross-modality discrepancy of visual and motion characteristics, 2) the uncertainty of degree of reliability in different modalities, and 3) large target appearance variations and background distractions within each modality. To address these issues, this paper proposes a novel and optimal discriminative learning framework for multi-modality tracking. In particular, the proposed discriminative learning framework is able to: 1) jointly eliminate outlier samples caused by large variations and learn discriminability-consistent features from heterogeneous modalities, and 2) collaboratively perform modality reliability measurement and target-background separation. Extensive experiments on RGB-infrared image sequences demonstrate the effectiveness of the proposed method. Xiangyuan Lan, Mang Ye, Shengping Zhang, Pong C. Yuen |
AAAI | 3 |
| 2018 | An Efficient Deep Convolutional Laplacian Pyramid Architecture for Cs Reconstruction At Low Sampling RatiosabstractThe compressed sensing (CS) has been successfully applied to image compression in the past few years as most image signals are sparse in a certain domain. Several CS reconstruction models have been proposed and obtained superior performance. However, these methods suffer from blocking artifacts or ringing effects at low sampling ratios in most cases. To address this problem, we propose a deep convolutional Laplacian Pyramid Compressed Sensing Network (LapC-SNet) for CS, which consists of a sampling sub-network and a reconstruction sub-network. In the sampling sub-network, we utilize a convolutional layer to mimic the sampling operator. In contrast to the fixed sampling matrices used in traditional CS methods, the filters used in our convolutional layer are jointly optimized with the reconstruction sub-network. In the reconstruction sub-network, two branches are designed to reconstruct multi -scale residual images and muti -scale target images progressively using a Laplacian pyramid architecture. The proposed LapCSNet not only integrates multi-scale information to achieve better performance but also reduces computational cost dramatically. Experimental results on benchmark datasets demonstrate that the proposed method is capable of reconstructing more details and sharper edges against the state-of-the-arts methods. Wenxue Cui, Heyao Xu, Xinwei Gao, Shengping Zhang, Feng Jiang 0001, Debin Zhao |
ICASSP | 4 |
| 2018 | Classification Guided Deep Convolutional Network for Compressed SensingabstractCompressed Sensing (CS) has been successfully applied to image compression in the past few years. However, there are still several challenges that restrict its applications in practice including large memory requirement and unsatisfactory reconstruction performance. To address these challenges, in this paper, we propose a classification guided deep convolutional network for image compressed sensing (CCSNet), which includes a sampling sub-network and a reconstruction sub-network. In the sampling sub-network, multiple convolutional layers are used to sample the original image, which significantly reduces the parameters of the sampling matrix while causes performance degradation moderately compared against existing convolution based sampling methods. In the reconstruction sub-network, a novel two-branch architecture is proposed to improve the adaptability of the model to various textures in natural images. The first branch, named the classification branch, is to classify the sampled measurements of the original image to one of the predefined textural classes. The second branch, named the reconstruction branch, consists of multiple sub-branches, which are responsible for reconstructing the original images belonging to the corresponding textural classes. By jointly utilizing two sub-networks, the entire network can be trained in the form of end-to-end metric with a joint loss function. Experimental results demonstrate that the proposed method provides a significant quality improvement in terms of PSNR compared against state-of-the-art methods. Wenxue Cui, Shaohui Liu, Shengping Zhang, Yashu Liu 0003, Heyao Xu, Xinwei Gao, Feng Jiang 0001, Debin Zhao |
ICPR | 3 |
| 2018 | An Efficient Deep Quantized Compressed Sensing Coding Framework of Natural ImagesabstractTraditional image compressed sensing (CS) coding frameworks solve an inverse problem that is based on the measurement coding tools (prediction, quantization, entropy coding, etc.) and the optimization based image reconstruction method. These CS coding frameworks face the challenges of improving the coding efficiency at the encoder, while also suffering from high computational complexity at the decoder. In this paper, we move forward a step and propose a novel deep network based CS coding framework of natural images, which consists of three sub-networks: sampling sub-network, offset sub-network and reconstruction sub-network that responsible for sampling, quantization and reconstruction, respectively. By cooperatively utilizing these sub-networks, it can be trained in the form of an end-to-end metric with a proposed rate-distortion optimization loss function. The proposed framework not only improves the coding performance, but also reduces the computational cost of the image reconstruction dramatically. Experimental results on benchmark datasets demonstrate that the proposed method is capable of achieving superior rate-distortion performance against state-of-the-art methods. Wenxue Cui, Feng Jiang 0001, Xinwei Gao, Shengping Zhang, Debin Zhao |
ACM Multimedia | 4 |
| 2018 | Plant identification based on very deep convolutional neural networks
Heyan Zhu, Qinglin Liu, Yuankai Qi, Feng Jiang 0001, Shengping Zhang |
Multim. Tools Appl. | 6 |
| 2018 | BoMW: Bag of Manifold Words for One-Shot Learning Gesture Recognition From KinectabstractIn this paper, we study one-shot learning gesture recognition on RGB-D data recorded from Microsoft's Kinect. To this end, we propose a novel bag of manifold words (BoMW)-based feature representation on symmetric positive definite (SPD) manifolds. In particular, we use covariance matrices to extract local features from RGB-D data due to its compact representation ability as well as the convenience of fusing both RGB and depth information. Since covariance matrices are SPD matrices and the space spanned by them is the SPD manifold, traditional learning methods in the Euclidean space, such as sparse coding, cannot be directly applied to them. To overcome this problem, we propose a unified framework to transfer the sparse coding on SPD manifolds to the one on the Euclidean space, which enables any existing learning method to be used. After building BoMW representation on a video from each gesture class, a nearest neighbor classifier is adopted to perform the one-shot learning gesture recognition. Experimental results on the ChaLearn gesture data set demonstrate the outstanding performance of the proposed one-shot learning gesture recognition method compared against the state-of-the-art methods. The effectiveness of the proposed feature extraction method is also validated on a new RGB-D action recognition data set. Lei Zhang 0036, Shengping Zhang, Feng Jiang 0001, Yuankai Qi, Jun Zhang 0017, Yuliang Guo, Huiyu Zhou 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2018 | Learning Common and Feature-Specific Patterns: A Novel Multiple-Sparse-Representation-Based TrackerabstractThe use of multiple features has been shown to be an effective strategy for visual tracking because of their complementary contributions to appearance modeling. The key problem is how to learn a fused representation from multiple features for appearance modeling. Different features extracted from the same object should share some commonalities in their representations while each feature should also have some feature-specific representation patterns which reflect its complementarity in appearance modeling. Different from existing multi-feature sparse trackers which only consider the commonalities among the sparsity patterns of multiple features, this paper proposes a novel multiple sparse representation framework for visual tracking which jointly exploits the shared and feature-specific properties of different features by decomposing multiple sparsity patterns. Moreover, we introduce a novel online multiple metric learning to efficiently and adaptively incorporate the appearance proximity constraint, which ensures that the learned commonalities of multiple features are more representative. Experimental results on tracking benchmark videos and other challenging videos demonstrate the effectiveness of the proposed tracker. Xiangyuan Lan, Shengping Zhang, Pong C. Yuen, Rama Chellappa |
IEEE Trans. Image Process. | 2 |
| 2018 | Structure-Aware Local Sparse Coding for Visual TrackingabstractSparse coding has been applied to visual tracking and related vision problems with demonstrated success in recent years. Existing tracking methods based on local sparse coding sample patches from a target candidate and sparsely encode these using a dictionary consisting of patches sampled from target template images. The discriminative strength of existing methods based on local sparse coding is limited as spatial structure constraints among the template patches are not exploited. To address this problem, we propose a structure-aware local sparse coding algorithm, which encodes a target candidate using templates with both global and local sparsity constraints. For robust tracking, we show the local regions of a candidate region should be encoded only with the corresponding local regions of the target templates that are the most similar from the global view. Thus, a more precise and discriminative sparse representation is obtained to account for appearance changes. To alleviate the issues with tracking drifts, we design an effective template update scheme. Extensive experiments on challenging image sequences demonstrate the effectiveness of the proposed algorithm against numerous state-of-the-art methods. Yuankai Qi, Jian Zhang 0018, Shengping Zhang, Qingming Huang, Ming-Hsuan Yang 0001 |
IEEE Trans. Image Process. | 4 |
| 2018 | Point-to-Set Distance Metric Learning on Deep Representations for Visual TrackingabstractFor autonomous driving application, a car shall be able to track objects in the scene in order to estimate where and how they will move such that the tracker embedded in the car can efficiently alert the car for effective collision-avoidance. Traditional discriminative object tracking methods usually train a binary classifier via a support vector machine (SVM) scheme to distinguish the target from its background. Despite demonstrated success, the performance of the SVM-based trackers is limited because the classification is carried out only depending on support vectors (SVs) but the target's dynamic appearance may look similar to the training samples that have not been selected as SVs, especially when the training samples are not linearly classifiable. In such cases, the tracker may drift to the background and fail to track the target eventually. To address this problem, in this paper, we propose to integrate the point-to-set/image-to-imageSet distance metric learning (DML) into visual tracking tasks and take full advantage of all the training samples when determining the best target candidate. The point-to-set DML is conducted on convolutional neural network features of the training data extracted from the starting frames. When a new frame comes, target candidates are first projected to the common subspace using the learned mapping functions, and then the candidate having the minimal distance to the target template sets is selected as the tracking result. Extensive experimental results show that even without model update the proposed method is able to achieve favorable performance on challenging image sequences compared with several state-of-the-art trackers. Shengping Zhang, Yuankai Qi, Feng Jiang 0001, Xiangyuan Lan, Pong C. Yuen, Huiyu Zhou 0001 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2017 | Convolutional Neural Networks Based Intra Prediction for HEVCabstractSummary form only given. Traditional intra prediction methods for HEVC rely on using the nearest reference lines for predicting a block, which ignore much richer context between the current block and its neighboring blocks and therefore cause inaccurate prediction especially when weak spatial correlation exists between the current block and the reference lines. To overcome this problem, in this paper, an intra-prediction convolutional neural network (IPCNN) is proposed for intra prediction, which exploits the rich context of the current block and therefore is capable of improving the accuracy of predicting the current block. Meanwhile, the reconstruction of the three nearest blocks can also be refined. To the best of our knowledge, this is the first paper that directly applies CNNs to intra prediction for HEVC. Experimental results validate the effectiveness of applying CNNs to intra prediction and the proposed method can achieve 0.70% bitrate reduction compared to HEVC reference software HM-14.0. Wenxue Cui, Tao Zhang 0013, Shengping Zhang, Feng Jiang 0001, Wangmeng Zuo, Zhaolin Wan, Debin Zhao |
DCC | 3 |
| 2017 | An End-to-End Compression Framework Based on Convolutional Neural NetworksabstractSummary form only given. Traditional image coding standards (such as JPEG and JPEG2000) make the decoded image suffer from many blocking artifacts or noises since the use of big quantization steps. To overcome this problem, we proposed an end-to-end compression framework based on two CNNs, as shown in Figure 1, which produce a compact representation for encoding using a third party coding standard and reconstruct the decoded image, respectively. To make two CNNs effectively collaborate, we develop a unified end-to-end learning framework to simultaneously learn CrCNN and ReCNN such that the compact representation obtained by CrCNN preserves the structural information of the image, which facilitates to accurately reconstruct the decoded image using ReCNN and also makes the proposed compression framework compatible with existing image coding standards. Wen Tao, Feng Jiang 0001, Shengping Zhang, Jie Ren 0016, Wuzhen Shi, Wangmeng Zuo, Xun Guo 0002, Debin Zhao |
DCC | 3 |
| 2017 | Deep networks for compressed image sensingabstractThe compressed sensing (CS) theory has been successfully applied to image compression in the past few years as most image signals are sparse in a certain domain. Several CS reconstruction models have been recently proposed and obtained superior performance. However, there still exist two important challenges within the CS theory. The first one is how to design a sampling mechanism to achieve an optimal sampling efficiency, and the second one is how to perform the reconstruction to get the highest quality to achieve an optimal signal recovery. In this paper, we try to deal with these two problems with a deep network. First of all, we train a sampling matrix via the network training instead of using a traditional manually designed one, which is much appropriate for our deep network based reconstruct process. Then, we propose a deep network to recover the image, which imitates traditional compressed sensing reconstruction processes. Experimental results demonstrate that our deep networks based CS reconstruction method offers a very significant quality improvement compared against state-of-the-art ones. Wuzhen Shi, Feng Jiang 0001, Shengping Zhang, Debin Zhao |
ICME | 3 |
| 2017 | Behavior Recognition in Mouse Videos using Contextual Features Encoded by Spatial-temporal Stacked Fisher VectorsabstractManual measurement of mouse behavior is highly labor intensive and prone to error. This investigation aims to efficiently and accurately recognize individual mouse behaviors in action videos and continuous videos. In our system each mouse action video is expressed as the collection of a set of interest points. We extract both appearance and contextual features from the interest points collected from the training datasets, and then obtain two Gaussian Mixture Model (GMM) dictionaries for the visual and contextual features. The two GMM dictionaries are leveraged by our spatial-temporal stacked Fisher Vector (FV) to represent each mouse action video. A neural network is used to classify mouse action and finally applied to annotate continuous video. The novelty of our proposed approach is: (i) our method exploits contextual features from spatiotemporal interest points, leading to enhanced performance, (ii) we encode contextual features and then fuse them with appearance features, and (iii) location information of a mouse is extracted from spatio-temporal interest points to support mouse behavior recognition. We evaluate our method against the database of Jhuang et al. (Jhuang et al., 2010) and the results show that our method outperforms several state-of-the-art approaches. Zheheng Jiang, Danny Crookes, Brian Desmond Green, Shengping Zhang, Huiyu Zhou 0001 |
ICPRAM | 4 |
| 2017 | Positive influence maximization in signed social networks based on simulated annealing
Cuihua Wang, Shengping Zhang, Guanglu Zhou, Chong Wu 0001 |
Neurocomputing | 3 |
| 2017 | Actor identification via mining representative actions
Wenlong Xie, Hongxun Yao, Xiaoshuai Sun, Sicheng Zhao, Wei Yu 0004, Shengping Zhang |
Neurocomputing | 6 |
| 2017 | Plant identification via multipath sparse coding
Heyan Zhu, Shengping Zhang, Pong C. Yuen |
Multim. Tools Appl. | 3 |
| 2017 | Robust Visual Tracking via Basis MatchingabstractMost existing tracking approaches are based on either the tracking by detection framework or the tracking by matching framework. The former needs to learn a discriminative classifier using positive and negative samples, which will cause tracking drift due to unreliable samples. The latter usually performs tracking by matching local interest points between a target candidate and the tracked target, which is not robust to target appearance changes over time. In this paper, we propose a novel tracking by matching framework for robust tracking based on basis matching rather than point matching. In particular, we learn the target model from target images using a set of Gabor basis functions, which have large responses on the corresponding spatial positions after a max pooling. During tracking, a target candidate is evaluated by computing the responses of the Gabor basis functions on their corresponding spatial positions. The experimental results on a set of challenging sequences validate that the performance of the proposed tracking method outperforms those of several state-of-the-art methods. Shengping Zhang, Xiangyuan Lan, Yuankai Qi, Pong C. Yuen |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2017 | Modeling Information Diffusion over Social Networks for Temporal Dynamic PredictionabstractModeling the process of information diffusion is a challenging problem. Although numerous attempts have been made in order to solve this problem, very few studies are actually able to simulate and predict temporal dynamics of the diffusion process. In this paper, we propose a novel information diffusion model, namely GT model, which treats the nodes of a network as intelligent and rational agents and then calculates their corresponding payoffs, given different choices to make strategic decisions. By introducing time-related payoffs based on the diffusion data, the proposed GT model can be used to predict whether or not the user's behaviors will occur in a specific time interval. The user's payoff can be divided into two parts: social payoff from the user's social contacts and preference payoff from the user's idiosyncratic preference. We here exploit the global influence of the user and the social influence between any two users to accurately calculate the social payoff. In addition, we develop a new method of presenting social influence that can fully capture the temporal dynamics of social influence. Experimental results from two different datasets, Sina Weibo and Flickr demonstrate the rationality and effectiveness of the proposed prediction method with different evaluation metrics. Shengping Zhang, Xin Sun 0003, Huiyu Zhou 0001, Sheng Li 0003, Xuelong Li 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2017 | A Biologically Inspired Appearance Model for Robust Visual TrackingabstractIn this paper, we propose a biologically inspired appearance model for robust visual tracking. Motivated in part by the success of the hierarchical organization of the primary visual cortex (area V1), we establish an architecture consisting of five layers: whitening, rectification, normalization, coding, and pooling. The first three layers stem from the models developed for object recognition. In this paper, our attention focuses on the coding and pooling layers. In particular, we use a discriminative sparse coding method in the coding layer along with spatial pyramid representation in the pooling layer, which makes it easier to distinguish the target to be tracked from its background in the presence of appearance variations. An extensive experimental study shows that the proposed method has higher tracking accuracy than several state-of-the-art trackers. Shengping Zhang, Xiangyuan Lan, Hongxun Yao, Huiyu Zhou 0001, Dacheng Tao, Xuelong Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2016 | Hedged Deep TrackingabstractIn recent years, several methods have been developed to utilize hierarchical features learned from a deep convolutional neural network (CNN) for visual tracking. However, as features from a certain CNN layer characterize an object of interest from only one aspect or one level, the performance of such trackers trained with features from one layer (usually the second to last layer) can be further improved. In this paper, we propose a novel CNN based tracking framework, which takes full advantage of features from different CNN layers and uses an adaptive Hedge method to hedge several CNN based trackers into a single stronger one. Extensive experiments on a benchmark dataset of 100 challenging image sequences demonstrate the effectiveness of the proposed algorithm compared to several state-of-theart trackers. Yuankai Qi, Shengping Zhang, Hongxun Yao, Qingming Huang, Jongwoo Lim, Ming-Hsuan Yang 0001 |
CVPR | 2 |
| 2016 | 3D Mask Face Anti-spoofing with Remote Photoplethysmography
Si-Qi Liu 0003, Pong C. Yuen, Shengping Zhang, Guoying Zhao 0001 |
ECCV (7) | 3 |
| 2016 | Robust Joint Discriminative Feature Learning for Visual Tracking
Xiangyuan Lan, Shengping Zhang, Pong C. Yuen |
IJCAI | 2 |
| 2016 | Sparsity analysis versus sparse representation classifier
Baochang Zhang 0001, Suli Ji, Shengping Zhang, Wankou Yang |
Neurocomputing | 4 |
| 2016 | Spatiochromatic Context Modeling for Color Saliency AnalysisabstractVisual saliency is one of the most noteworthy perceptual abilities of human vision. Recent progress in cognitive psychology suggests that: 1) visual saliency analysis is mainly completed by the bottom-up mechanism consisting of feedforward low-level processing in primary visual cortex (area V1) and 2) color interacts with spatial cues and is influenced by the neighborhood context, and thus it plays an important role in a visual saliency analysis. From a computational perspective, the most existing saliency modeling approaches exploit multiple independent visual cues, irrespective of their interactions (or are not computed explicitly), and ignore contextual influences induced by neighboring colors. In addition, the use of color is often underestimated in the visual saliency analysis. In this paper, we propose a simple yet effective color saliency model that considers color as the only visual cue and mimics the color processing in V1. Our approach uses region-/boundary-defined color features with spatiochromatic filtering by considering local color-orientation interactions, therefore captures homogeneous color elements, subtle textures within the object and the overall salient object from the color image. To account for color contextual influences, we present a divisive normalization method for chromatic stimuli through the pooling of contrary/complementary color units. We further define a color perceptual metric over the entire scene to produce saliency maps for color regions and color boundaries individually. These maps are finally globally integrated into a one single saliency map. The final saliency map is produced by Gaussian blurring for robustness. We evaluate the proposed method on both synthetic stimuli and several benchmark saliency data sets from the visual saliency analysis to salient object detection. The experimental results demonstrate that the use of color as a unique visual cue achieves competitive results on par with or better than 12 state-of-the-art approaches. Jun Zhang 0017, Meng Wang 0001, Shengping Zhang, Xuelong Li 0001, Xindong Wu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2015 | Online Dictionary Learning on Symmetric Positive Definite Manifolds with Vision ApplicationsabstractSymmetric Positive Definite (SPD) matrices in the form of region covariances are considered rich descriptors for images and videos. Recent studies suggest that exploiting the Riemannian geometry of the SPD manifolds could lead to improved performances for vision applications. For tasks involving processing large-scale and dynamic data in computer vision, the underlying model is required to progressively and efficiently adapt itself to the new and unseen observations. Motivated by these requirements, this paper studies the problem of online dictionary learning on the SPD manifolds. We make use of the Stein divergence to recast the problem of online dictionary learning on the manifolds to a problem in Reproducing Kernel Hilbert Spaces, for which, we develop efficient algorithms by taking into account the geometric structure of the SPD manifolds. To our best knowledge, our work is the first study that provides a solution for online dictionary learning on the SPD manifolds. Empirical results on both large-scale image classification task and dynamic video processing tasks validate the superior performance of our approach as compared to several state-of-the-art algorithms. Shengping Zhang, Shiva Prasad Kasiviswanathan, Pong C. Yuen, Mehrtash Harandi |
AAAI | 1 |
| 2015 | Histograms of locally aggregated oriented gradientsabstractMotivated by the Vector of Locally Aggregated Descriptors (VLAD), we propose a new Histograms of Locally Aggregated Oriented Gradients descriptor (called HLAOG). In the Histograms of Oriented Gradients descriptor (HOG), the zero-order information of the gradients is captured. By contrast, in the HLAOG descriptor we accumulate the differences between gradient orientations and their nearest bin centers, which characterizes the distribution of the gradient orientations in regard to the bin centers. The HLAOG descriptor is demonstrated to be complementary to HOG in the experiments. Then, for setting the weights of the votes on different bins in a better way, we choose Gaussian function as the weighting method and present another new Gaussian Weighted Histograms of Oriented Gradients descriptor (called GWHOG) based on HOG. Evaluations on two public object recognition datasets (Caltech-101 and VOC2007) show that the combination of HOG and HLAOG outperforms HOG and the combination of HLAOG and GWHOG gets the best result. Xiusheng Lu, Shengping Zhang, Hongxun Yao, Xin Sun 0003, Yanhao Zhang 0001 |
ICIP | 2 |
| 2015 | Formation Period Matters: Towards Socially Consistent Group Detection via Dense Subgraph SeekingabstractGroup detection becomes an important task in crowd behavior surveillance. However, most existing methods ignore the formation persistency characteristics, which predict unreliable interactions when the crowd is realistic and complex. To address this issue, we propose a novel graph-based method to declare that the formation period really matters for detecting social groups in crowd. First, we develop a socially motivated representation by modeling the formation period probability in a Bayesian manner, which results in social and temporal consistency for group member interactions. A graph is then established using individuals as nodes and formation periods as edge weights to reflect pedestrian relationships. In this way, seeking of socially consistent groups is converted into an optimization problem which seeks dense subgraphs with maximum formation likelihood within the graph structure. We employ graph shift optimization to detect groups by finding all the dense subgraphs due to its robust performance. In the experimental results on public datasets, our proposed method clearly outperforms other related state-of-the-art methods. Yanhao Zhang 0001, Shengping Zhang, Hongxun Yao, Qingming Huang |
ICMR | 3 |
| 2015 | Multi-layered gesture recognition with Kinect
Feng Jiang 0001, Shengping Zhang, Debin Zhao |
J. Mach. Learn. Res. | 2 |
| 2015 | Adaptive NormalHedge for robust visual tracking
Shengping Zhang, Huiyu Zhou 0001, Hongxun Yao, Yanhao Zhang 0001, Kuanquan Wang, Jun Zhang 0017 |
Signal Process. | 1 |
| 2015 | Robust Visual Tracking Using Structurally Random Projection and Weighted Least SquaresabstractSparse representation-based visual tracking approaches have attracted increasing interests in the community in recent years. The main idea is to linearly represent each target candidate using a set of target and trivial templates, while imposing a sparsity constraint onto the representation coefficients. After we obtain the coefficients using ℓ1-norm minimization methods, the candidate with the lowest error, when it is reconstructed using only the target templates and the associated coefficients, is considered as the tracking result. In spite of promising system performance widely reported, it is unclear if the performance of these trackers can be maximized. In addition, computational complexity caused by the dimensionality of the feature space limits these algorithms in real-time applications. In this paper, we propose a real-time visual tracking method based on structurally random projection (RP) and weighted least squares (WLS) techniques. In particular, to enhance the discriminative capability of the tracker, we introduce background templates to the linear representation framework. To handle appearance variations over time, we relax the sparsity constraint using a WLS method to obtain the representation coefficients. To further reduce the computational complexity, structurally RP is used to reduce the dimensionality of the feature space, while preserving the pairwise distances between the data points in the feature space. Experimental results show that the proposed approach outperforms several state-of-the-art tracking methods. Shengping Zhang, Huiyu Zhou 0001, Feng Jiang 0001, Xuelong Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2015 | Non-Rigid Object Contour Tracking via a Novel Supervised Level Set ModelabstractWe present a novel approach to non-rigid objects contour tracking in this paper based on a supervised level set model (SLSM). In contrast to most existing trackers that use bounding box to specify the tracked target, the proposed method extracts the accurate contours of the target as tracking output, which achieves better description of the non-rigid objects while reduces background pollution to the target model. Moreover, conventional level set models only emphasize the regional intensity consistency and consider no priors. Differently, the curve evolution of the proposed SLSM is object-oriented and supervised by the specific knowledge of the targets we want to track. Therefore, the SLSM can ensure a more accurate convergence to the exact targets in tracking applications. In particular, we firstly construct the appearance model for the target in an online boosting manner due to its strong discriminative power between the object and the background. Then, the learnt target model is incorporated to model the probabilities of the level set contour by a Bayesian manner, leading the curve converge to the candidate region with maximum likelihood of being the target. Finally, the accurate target region qualifies the samples fed to the boosting procedure as well as the target model prepared for the next time step. We firstly describe the proposed mechanism of two-phase SLSM for single target tracking, then give its generalized multi-phase version for dealing with multi-target tracking cases. Positive decrease rate is used to adjust the learning pace over time, enabling tracking to continue under partial and total occlusion. Experimental results on a number of challenging sequences validate the effectiveness of the proposed method. Xin Sun 0003, Hongxun Yao, Shengping Zhang |
IEEE Trans. Image Process. | 3 |
| 2014 | Image Restoration via Multi-prior Collaboration
Feng Jiang 0001, Shengping Zhang, Debin Zhao, Sun-Yuan Kung |
ACCV (3) | 2 |
| 2014 | Action recognition based on overcomplete independent components analysis
Shengping Zhang, Hongxun Yao, Xin Sun 0003, Kuanquan Wang, Jun Zhang 0017, Xiusheng Lu, Yanhao Zhang 0001 |
Inf. Sci. | 1 |
| 2014 | Dynamic image mosaic via SIFT and dynamic programming
Shengping Zhang, Jun Zhang 0017, Yunlu Zhang |
Mach. Vis. Appl. | 2 |
| 2013 | Night video enhancement using improved dark channel priorabstractVideos taken under low lighting condition usually have serious loss of visibility and contrast and are inconvenient for observation and analysis. To solve this problem, this paper presents a real-time night video enhancement approach. As observed that a pixel-wise inversion of a night video has quite similar appearance with the video acquired at foggy days, we use the similar idea of haze removal method to enhance the perceptual quality of the night videos. We present an improved dark channel prior model and integrate it with local smoothing and image Gaussian Pyramid operators. The experimental results demonstrate that the proposed approach can improve the perceptual quality of night videos in real-time in terms of not only enhancing details, but also effectively avoiding excessive enhancement phenomenon. Xuesong Jiang, Hongxun Yao, Shengping Zhang, Xiusheng Lu, Wei Zeng 0006 |
ICIP | 3 |
| 2013 | Non-rigid object tracking by adaptive data-driven kernelabstractWe derive an adaptive data-driven kernel in this paper to simultaneously address the kernel scale/orientation selection problem as well as the constant kernel shape in deformable object tracking applications. Level set technique is novelly introduced into the mean shift sample space to implement kernel evolution and update. Since the active contour model is designed to drive the kernel constantly to the direction that maximizes target likelihood, the kernel can adapt to target shape variation simultaneously with the mean shift iterations. Thus, it can give a better estimation bias to produce accurate shift of the mean and successfully avoid performance loss stemmed from pollution of the non-object regions hiding inside the kernel. Experimental results on a number of challenging sequences validate the effectiveness of the technique. Xin Sun 0003, Hongxun Yao, Shengping Zhang, Mingui Sun |
ICIP | 3 |
| 2013 | Sparse coding based motion attention for abnormal event detectionabstractIn this paper, we present a novel method based on sparsely coded motion attention for detecting abnormal events in crowded scenes. Unlike existing sparse coding based approaches, our model does not need to learn a dictionary and directly sparsely codes the motion features of the center patches with features of its surrounding patches. The sparse coding error is used to measure the motion attention intensity of the center patch. To reflect the crowd abnormal intensity, an online updated weighting scheme is designed to obtain the global activity intensity map. Two publicly available datasets-UMN dataset and UCSD Ped1 dataset are utilized to evaluate our approach in detecting global abnormal event and local abnormal event, respectively. The experiments show our method achieves the promising performance and is competitive with the state-of-the-art approaches. Shengping Zhang, Hongxun Yao |
ICIP | 2 |
| 2013 | Robust visual tracking based on online learning sparse representation
Shengping Zhang, Hongxun Yao, Huiyu Zhou 0001, Xin Sun 0003, Shaohui Liu |
Neurocomputing | 1 |
| 2013 | Sparse coding based visual tracking: Review and experimental comparison
Shengping Zhang, Hongxun Yao, Xin Sun 0003, Xiusheng Lu |
Pattern Recognit. | 1 |
| 2012 | Robust Visual Tracking Using an Effective Appearance Model Based on Sparse CodingabstractIntelligent video surveillance is currently one of the most active research topics in computer vision, especially when facing the explosion of video data captured by a large number of surveillance cameras. As a key step of an intelligent surveillance system, robust visual tracking is very challenging for computer vision. However, it is a basic functionality of the human visual system (HVS). Psychophysical findings have shown that the receptive fields of simple cells in the visual cortex can be characterized as being spatially localized, oriented, and bandpass, and it forms a sparse, distributed representation of natural images. In this article, motivated by these findings, we propose an effective appearance model based on sparse coding and apply it in visual tracking. Specifically, we consider the responses of general basis functions extracted by independent component analysis on a large set of natural image patches as features and model the appearance of the tracked target as the probability distribution of these features. In order to make the tracker more robust to partial occlusion, camouflage environments, pose changes, and illumination changes, we further select features that are related to the target based on an entropy-gain criterion and ignore those that are not. The target is finally represented by the probability distribution of those related features. The target search is performed by minimizing the Matusita distance between the distributions of the target model and a candidate using Newton-style iterations. The experimental results validate that the proposed method is more robust and effective than three state-of-the-art methods. Shengping Zhang, Hongxun Yao, Xin Sun 0003, Shaohui Liu |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2011 | A novel supervised level set method for non-rigid object trackingabstractWe present a novel approach to non-rigid object tracking based on a supervised level set model (SLSM). In contrast with conventional level set models, which emphasize the intensity consistency only and consider no priors, the curve evolution of the proposed SLSM is object-oriented and supervised by the specific knowledge of the target we want to track. Therefore, the SLSM can ensure a more accurate convergence to the target in tracking applications. In particular, we firstly construct the appearance model for the target in an on-line boosting manner due to its strong discriminative power between objects and background. Then the probability of the contour is modeled by considering both the region and edge cues in a Bayesian manner, leading the curve converge to the candidate region with maximum likelihood of being the target. Finally, accurate target region qualifies the samples fed the boosting procedure as well as the target model prepared for the next time step. Positive decrease rate is used to adjust the learning pace over time, enabling tracking to continue under partial and total occlusion. Experimental results on a number of challenging sequences validate the effectiveness of the technique. Xin Sun 0003, Hongxun Yao, Shengping Zhang |
CVPR | 3 |
| 2011 | Contour tracking via on-line discriminative appearance modeling based level setsabstractA novel level set method based on on-line discriminative appearance modeling (DAMLSM) is presented for contour tracking. In contrast with traditional level set models which emphasize the intensity consistent segmentation and consider no priors, the proposed DAMLSM takes the context of tracking into account and use a discriminative patch based target model to guide the curve evolution. By modeling both the region and edge cues in a Bayesian manner, the proposed level set method can lead an accurate convergence to the candidate region with maximum likelihood of being the target. Finally, we update the target model to adapt to the appearance variation, enabling tracking to continue under occlusion. Experiments confirm the robustness and reliability of our method. Xin Sun 0003, Hongxun Yao, Shengping Zhang |
ICIP | 3 |
| 2011 | Robust visual tracking via context objects computingabstractOcclusions are challenging issue for robust visual tracking. In this paper, motivated by the fact that a tracked object is usual- ly embedded into context that provides useful information for estimating the target, we propose a novel tracking algorithm named Tracking with Context Prediction (TCP). The context here includes the neighboring objects and specific parts of tar- get. The proposed method simultaneously track the target and context objects using the existing tracking methods. The positions of the context objects are used to predict the position of the target. Thus, the target can be stably tracked even when it is partially or fully occluded. By computing the probability of each prediction being target, our algorithm allows the drifting of context objects during tracking and do not require predictions from all context objects are correct. Experiments on challenging sequences show significant improvements especially in the case of occlusions and appearance changes. Zhongqian Sun, Hongxun Yao, Shengping Zhang, Xin Sun 0003 |
ICIP | 3 |
| 2011 | Sparse regression analysis for object recognitionabstractThis paper proposes a new method named Sparse Regression Analysis (SRA) for object representation and recognition. In SRA, ℓ1-norm minimization is combined with regression analysis to represent the input signal. The discriminative ability of SRA derives from the fact that the subset which most compactly expresses the input signal is activated in the regression analysis. To achieve a further improvement, Kernelized SRA (KSRA) is developed to make a nonlinear extension of SRA. The experiments are conducted on both palmprint and face recognition, which show that the proposed methods achieve a much better performance than sparse representation classifier, principal component analysis, and linear discriminant analysis. Baochang Zhang 0001, Shengping Zhang, Jianzhuang Liu |
ICIP | 2 |
| 2010 | Robust visual tracking using feature-based visual attentionabstractPsychophysical findings have shown that human vision system has an ability to improve target search by enhancing the representation of image components that are related to the searched target, which is the so-called feature-based visual attention. In this paper, motivated by these psychophysical findings, we propose a robust visual tracking algorithm by simulating such feature-based visual attention. Specially, we consider the general sparse basis functions extracted on a large set of natural image patches as features. We define that a feature is related to the target when succeeding activations of that feature cannot increase system's entropy. The target is finally represented by the probability distribution of those related features. The target search is performed by minimizing the Matusita distance measure between the distributions of the target model and candidate using Newton-style iterations. The experimental results verify that the proposed method is more robust and effective than widely used mean shift based methods. Shengping Zhang, Hongxun Yao, Shaohui Liu |
ICASSP | 1 |
| 2010 | A steganography strategy based on equivalence partitions of hiding unitsabstractThis paper designs a novel hiding strategy based on an equivalence relation, which can remarkably enhance the quality of stego image without sacrificing the security and capacity of original steganography schemes. According to a constructed equivalence relation based on the capacity of hiding units, all hiding units can be partitioned into equivalence classes. Following that, the hiding procedure is performed in predefined order in equivalence classes as the traditional steganography scheme. Because of considering the relation between the length of message and capacity, the performance of the hiding method using proposed hiding strategy outperforms the original approaches when embedding the same length message. Experimental results indicate that the gain from the proposed strategy over existing hiding schemes can reach up to 4.0 dB. Shaohui Liu, Hongxun Yao, Shengping Zhang, Wen Gao 0001 |
ICME | 3 |
| 2010 | A refined particle filter method for contour trackingabstractTraditional particle filter which uses simple geometric shapes for representation cannot track objects with complex shape accurately. In this paper, we propose a refined particle filter method for contour tracking based on a binary level set model. In contrast with other previous work, the computational efficiency is greatly improved due to the simple form of the level set function. In addition, we perform curve evolution in the update step to make good use of the observation at current time. Finally, we consider some appearance information as well as the energy function to measure the weight for particles, which can identify the target more accurately. Experiment results on several challenging video sequences have verified the proposed algorithm is efficient and effective in many complicated scenes. Xin Sun 0003, Hongxun Yao, Shengping Zhang |
VCIP | 3 |
| 2010 | Robust object tracking based on sparse representationabstractIn this paper, we propose a novel and robust object tracking algorithm based on sparse representation. Object tracking is formulated as a object recognition problem rather than a traditional search problem. All target candidates are considered as training samples and the target template is represented as a linear combination of all training samples. The combination coefficients are obtained by solving for the minimum l1-norm solution. The final tracking result is the target candidate associated with the non-zero coefficient. Experimental results on two challenging test sequences show that the proposed method is more effective than the widely used mean shift tracker. Shengping Zhang, Hongxun Yao, Xin Sun 0003, Shaohui Liu |
VCIP | 1 |
| 2010 | Robust object tracking combining color and scale invariant featuresabstractObject tracking plays a very important role in many computer vision applications. However its performance will significantly deteriorate due to some challenges in complex scene, such as pose and illumination changes, clustering background and so on. In this paper, we propose a robust object tracking algorithm which exploits both global color and local scale invariant (SIFT) features in a particle filter framework. Due to the expensive computation cost of SIFT features, the proposed tracker adopts a speed-up variation of SIFT, SURF, to extract local features. Specially, the proposed method first finds matching points between the target model and target candidate, than the weight of the corresponding particle based on scale invariant features is computed as the the proportion of matching points of that particle to matching points of all particles, finally the weight of the particle is obtained by combining weights of color and SURF features with a probabilistic way. The experimental results on a variety of challenging videos verify that the proposed method is robust to pose and illumination changes and is significantly superior to the standard particle filter tracker and the mean shift tracker. Shengping Zhang, Hongxun Yao, Peipei Gao |
VCIP | 1 |
| 2010 | Partial occlusion robust object tracking using an effective appearance modelabstractPartial occlusion is one of the most challenging difficulties for object tracking. In this paper, we present an approach to address this problem by using an effective appearance model which has two innovations. First, in contrast to widely used color histogram that models the appearance of an object using only color information, we assert that both color and texture are important cues for tracking, especially in the presence of complex background. We thus propose a novel local descriptor, named local color texture pattern (LCTP), to model the appearance of the object with color and texture information simultaneously. Second, global color histogram completely ignores the spatial layout information of an object and are sensitive to partial occlusion. In this work, we overcome this limitation based on a block-dividing way: 1) divide target into multiple blocks and then represent each block with LCTP histogram, 2) with a selectivity strategy, we select blocks that are not occluded and then combine similarities of those selected blocks to obtain final similarity measure. Experimental results demonstrate that the proposed method is more robust to partial occlusion than two state-of-the-art algorithms. Shengping Zhang, Hongxun Yao, Shaohui Liu |
VCIP | 1 |
| 2009 | Transfer Function Design Using Acting Force ModelabstractTransfer function plays an important role in volume rendering. Multi-dimension transfer function can achieve high rendering quality, but suffer from troublesome user specifications. In order to reduce the complexity of manipulations, a transfer function design method based on dimension reduction model is proposed in this paper. The idea of dimension reduction model is to integrate several data features into one index for transfer function. This paper also presents one instance of dimension reduction model called acting force model (AFM) to verity the effectiveness. In AFM, one integrated index is calculated by applying transformative universal gravitation formula and kinetic energy formula to three data features: scalar value, gradient and coherence distance. The experiments demonstrate the proposed method can produce flexible results without overloading the users. Yi Zhang 0070, Jiawan Zhang, Shengping Zhang |
ICIG | 4 |
| 2009 | Spatial-temporal nonparametric background subtraction in dynamic scenesabstractTraditional background subtraction methods model only temporal variation of each pixel. However, there is also spatial variation in real word due to dynamic background such as waving trees, spouting fountain and camera jitters, which causes the significant performance degradation of traditional methods. In this paper, a novel spatial-temporal nonparametric background subtraction approach (STNBS) is proposed to effectively handle dynamic background by modeling the spatial and temporal variations simultaneously. Specially, for each pixel in an image, we adaptively maintain a sample consisting of pixels observed in previous frames. At current frame, for a particular pixel, the proposed method estimates the probabilities of observing this pixel based on samples of its neighboring pixels. The pixel is labeled as background if one of these estimated probabilities is larger than a fixed threshold. All samples are adaptively updated over time. Experimental results on several challenging sequences show that the proposed method achieves the best performance than two state-of-the-art algorithms. Shengping Zhang, Hongxun Yao, Shaohui Liu |
ICME | 1 |
| 2009 | Dynamic Background Subtraction Based on Local Dependency HistogramabstractTraditional background subtraction methods perform poorly when scenes contain dynamic backgrounds such as waving tree branches, spouting fountain, illumination changes, camera jitters, etc. In this paper, from the view of spatial context, we present a novel and effective dynamic background method with three contributions. First, we present a novel local dependency descriptor, called local dependency histogram (LDH), to effectively model the spatial dependencies between a pixel and its neighboring pixels. The spatial dependencies contain substantial evidence for differentiating dynamic background regions from moving objects of interest. Second, based on the proposed LDH, an effective approach to dynamic background subtraction is proposed, in which each pixel is modeled as a group of weighted LDHs. Labeling a pixel as foreground or background is done by comparing the LDH computed in current frame against its model LDHs. The model LDHs are adaptively updated by the current LDH. Finally, unlike traditional approaches using a fixed threshold to judge whether a pixel matches to its model, an adaptive thresholding technique is also proposed. Experimental results on a diverse set of dynamic scenes validate that the proposed method significantly outperforms traditional methods for dynamic background subtraction. Shengping Zhang, Hongxun Yao, Shaohui Liu |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2008 | Dynamic background modeling and subtraction using spatio-temporal local binary patternsabstractTraditional background modeling and subtraction methods have a strong assumption that the scenes are of static structures with limited perturbation. These methods will perform poorly in dynamic scenes. In this paper, we present a solution to this problem. We first extend the local binary patterns from spatial domain to spatio-temporal domain, and present a new online dynamic texture extraction operator, named spatio-temporal local binary patterns (STLBP). Then we present a novel and effective method for dynamic background modeling and subtraction using STLBP. In the proposed method, each pixel is modeled as a group of STLBP dynamic texture histograms which combine spatial texture and temporal motion information together. Compared with traditional methods, experimental results show that the proposed method adapts quickly to the changes of the dynamic background. It achieves accurate detection of moving objects and suppresses most of the false detections for dynamic changes of nature scenes. Shengping Zhang, Hongxun Yao, Shaohui Liu |
ICIP | 1 |
| 2008 | A covariance-based method for dynamic background subtractionabstractBackground subtraction in dynamic scenes is an important and challenging task. In this paper, we present a novel and effective method for dynamic background subtraction based on covariance matrix descriptor. The algorithm integrates two distinct levels: pixel level and region level. At the pixel level, spatial properties that are obtained from pixel coordinate values, and appearance properties, i.e., intensity, texture, gradient, etc, are used as features of each pixel. In the region level, the correlation of features extracted at the pixel level is represented by a covariance matrix that is calculated over a rectangle region around the pixel. Each pixel is modeled as a group of weighted adaptive covariance matrices. Experimental results on a diverse set of dynamic scenes show that the proposed method dramatically out-performs traditional methods for dynamic background subtraction. Shengping Zhang, Hongxun Yao, Shaohui Liu, Xilin Chen 0001, Wen Gao 0001 |
ICPR | 1 |
| 2007 | Combining Global and Local Classifiers for Lipreading
Shengping Zhang, Hongxun Yao, Yuqi Wan |
ACII | 1 |