EDBT 2026 Demo / reviewers in the wild / expert
Zhigang Tu 0001
dblp:142/4070-1
· DBLP profile ↗
63ranked-venue papers
17as first author
44since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 44 · 10 first-author · 31 since 2021Artificial intelligence and machine learning · 31 · 8 first-author · 19 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Logic-Driven Network for long-term action anticipation
Na Zhou, Renjie Yang, Zhigang Tu 0001 |
Eng. Appl. Artif. Intell. | 4 |
| 2026 | AstroHSP: A hybrid supervision framework for robust monocular astronaut pose estimation
Haohang Jian, Xiongwu Xiao, Jianguo Yan, Zhigang Tu 0001 |
Neural Networks | 5 |
| 2026 | OwlSight: A Robust Illumination Adaptation Framework for Dark Video Human Action RecognitionabstractHuman action recognition in low-light environments is crucial for various real-world applications. However, the existing methods overlook the full utilization of brightness information throughout the training phase, leading to suboptimal performance. To address this issue, we propose OwlSight, a biomimetic framework with whole-stage illumination enhancement to interact with action classification for accurate dark video human action recognition. Specifically, OwlSight incorporates a Time-Consistency Module (TCM) to capture shallow spatiotemporal features meanwhile maintaining temporal coherence, which are then processed by a Luminance Adaptation Module (LAM) to dynamically adjust the brightness based on the input luminance distribution. Furthermore, a Reflect Augmentation Module (RAM) is presented to maximize illumination utilization and simultaneously enhance action recognition via two interactive paths. Additionally, we build a large-scale datasetDark-101, which comprises 21,030 dark videos across 101 action categories, significantly surpassing the existing datasets (e.g., ARID1.5 and Dark-48) in scale and diversity. Our method establishes new state-of-the-art (SOTA) performance across all benchmarks, achieving Top-1 accuracies of 99.27% on ARID (a 2.0% improvement), 94.85% on ARID1.5 (a 5.36% improvement), 48.24% on Dark-48 (a 1.56% improvement), and 53.85% on Dark-101 (a 1.72% improvement), demonstrating its superior effectiveness in challenging dark video environments. Shihao Cheng, Jinlu Zhang 0001, Aoran Xiao, Zhigang Tu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | EBPersons: A Dataset for Person Detection at the Edges of BuildingsabstractWith the increasing prevalence of buildings, incidents of falling from heights have become more and more frequent. Accurately detecting individuals at the edges of buildings through surveillance videos is crucial for timely intervention and accident prevention. However, this task, termed Person Detection at the Edges of Buildings (PDEB), presents significant challenges including variations in lighting conditions, occlusions, and small size of person instances. Existing person detection datasets are inadequate for PDEB due to domain gaps. To address this issue, we construct EBPersons, a completely new dataset specifically designed for PDEB. Comprising 1,314 videos captured across over 300 diverse building scenes with diverse lighting conditions, EBPersons provides a rich and challenging benchmark for PDEB research. Furthermore, we propose a baseline method specifically designed for PDEB, named STASH, which includes three key components: a Scale Match strategy to improve small object detection, a Temporal ROI Align Operator to leverage temporal context, and a Sequential-level Semantics Aggregation head to enhance feature representation. Extensive experiments are conducted on EBPersons to compare our method with other detectors, including generic object detectors, pedestrian detectors, and video object detectors. The results demonstrate the superior performance of the proposed STASH, providing a strong baseline for future research on PDEB. Our EBPer sons dataset and the baseline code are publicly available at https://ebpersons.github.io/. Zitao Gao, Bing Qu, Chunluan Zhou, Junsong Yuan 0001, Zhigang Tu 0001 |
IEEE Trans. Multim. | 5 |
| 2026 | Frequency-enhanced diffusion models: curriculum-guided semantic alignment for zero-shot skeleton action recognition
Zhengbo Zhang, Jingyu Pan, Zhigang Tu 0001 |
Vis. Comput. | 5 |
| 2025 | Visual Prompting for One-shot Controllable Video Editing without InversionabstractOne-shot controllable video editing (OCVE) is an important yet challenging task, aiming to propagate user edits that are made – using any image editing tool – on the first frame of a video to all subsequent frames, while ensuring content consistency between edited frames and source frames. To achieve this, prior methods employ DDIM inversion to transform source frames into latent noise, which is then fed into a pre-trained diffusion model, conditioned on the user-edited first frame, to generate the edited video. However, the DDIM inversion process accumulates errors, which hinder the latent noise from accurately reconstructing the source frames, ultimately compromising content consistency in the generated edited frames. To overcome it, our method eliminates the need for DDIM inversion by performing OCVE through a novel perspective based on visual prompting. Furthermore, inspired by consistency models that can perform multi-step consistency sampling to generate a sequence of content-consistent images, we propose a content consistency sampling (CCS) to ensure content consistency between the generated edited frames and the source frames. Moreover, we introduce a temporal-content consistency sampling (TCS) based on Stein Variational Gradient Descent to ensure temporal consistency across the edited frames. Extensive experiments validate the effectiveness of our approach. Zhengbo Zhang, Duo Peng, Joo-Hwee Lim, Zhigang Tu 0001, De Wen Soh, Lin Geng Foo |
CVPR | 5 |
| 2025 | SemTalk: Holistic Co-Speech Motion Generation with Frame-Level Semantic EmphasisabstractCo-speech gesture generation must carefully integrate common rhythmic motion with rare yet essential semantic gestures. In this work, we propose SemTalk for holistic co-speech gesture generation with frame-level semantic emphasis. Our key insight is to separately learn base motions and sparse motions, and then adaptively fuse them. In particular, coarse2fine cross-attention module and rhythmic consistency learning are explored to establish rhythm-related base motion, ensuring a coherent foundation that synchronizes gestures with the speech rhythm. Subsequently, semantic emphasis learning is designed to generate semantic-aware sparse motion, focusing on frame-level semantic cues. Finally, to integrate sparse motion into the base motion and generate semantic-emphasized co-speech gestures, we further leverage a learned semantic score for adaptive synthesis. Qualitative and quantitative comparisons on two public datasets demonstrate that our method outperforms the state-of-the-art, delivering high-quality co-speech motion with enhanced semantic richness over a stable base motion. Xiangyue Zhang, Jianfang Li 0001, Ziqiang Dang, Jianqiang Ren, Liefeng Bo, Zhigang Tu 0001 |
ICCV | 7 |
| 2025 | MikuDance: Animating Character Art With Mixed Motion DynamicsabstractWe propose MikuDance, a diffusion-based pipeline incorporating mixed motion dynamics to animate stylized character art. MikuDance consists of two key techniques: Mixed Motion Modeling and Mixed-Control Diffusion, to address the challenges of high-dynamic motion and reference-guidance misalignment in character art animation. Specifically, a Scene Motion Tracking strategy is presented to explicitly model the dynamic camera in pixel-wise space, enabling unified character-scene motion modeling. Building on this, the Mixed-Control Diffusion implicitly aligns the scale and body shape of diverse characters with motion guidance, allowing flexible control of local character motion. Subsequently, a Motion-Adaptive Normalization module is incorporated to effectively inject global scene motion, paving the way for comprehensive character art animation. Through extensive experiments, we demonstrate the effectiveness and generalizability of MikuDance across various character art and motion guidance, consistently producing high-quality animations with remarkable motion dynamics. Xianfang Zeng, Xin Chen 0040, Wei Zuo, Gang Yu 0002, Zhigang Tu 0001 |
ICCV | 6 |
| 2025 | EchoMask: Speech-Queried Attention-based Mask Modeling for Holistic Co-Speech Motion Generation
Xiangyue Zhang, Jianfang Li 0001, Jianqiang Ren, Liefeng Bo, Zhigang Tu 0001 |
ACM Multimedia | 6 |
| 2025 | Bidirectional-Modulation Frequency-Heterogeneous Network for Remote Sensing Image DehazingabstractRecently, deep neural networks have been extensively explored in remote sensing image haze removal and achieved remarkable performance. However, existing methods fail to effectively fuse the features extracted from Convolutional Neural Networks (CNNs) and Transformer networks, leading to performance degradation. Moreover, most dehazing methods lack further exploration of the distinct properties of high- and low-frequency features, which are crucial for texture restoration and haze removal. To address these issues, we propose a Bidirectional-Modulation Frequency-Heterogeneous Network (BMFH-Net). Specifically, we propose a Differential-Expert Guided Bidirectional Modulation (DGBM) module that incorporates Differential experts and physical inversion models to exploit the complementarity of CNN-Transformer features and extract their latent haze-related physical characteristics, thereby enabling more effective bidirectional alignment. Furthermore, a Wavelet Frequency Heterogeneous Enhancement (WFHE) Module is designed to capture the most representative high-frequency features to refine image texture details, while enhancing the global perception of haze and reconstructing structural information during low-frequency processing. Experiments on challenging remote sensing image datasets demonstrate that our BMFH-Net outperforms several state-of-the-art haze removal methods. The code is released publicly at https://github.com/zqf2024/BMFH-Net. Qingfei Zhong, Bo Du 0001, Zhigang Tu 0001, Jun Wan 0005, Wenbin Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Robust 2D Skeleton Action Recognition via Decoupling and Distilling 3D Latent FeaturesabstractHuman skeletons provide a compact representation for action recognition. Compared to 3D skeletons, 2D skeletons lack view-independence and depth, making them less robust for motion analysis. However, 3D skeleton data requires specialized hardware, limiting its practicality, especially in outdoor or dynamic settings. In contrast, 2D skeletons can be extracted from standard RGB videos, making them more accessible. To address this, we propose 2D³-SkelAct, a 2D skeleton-based action recognition model. It maps 2D inputs to a 3D latent space, where pose and view features are decoupled. Additionally, 2D³-SkelAct distills motion cues from 3D models, enhancing motion detail capture while keeping the benefits of 2D data. Specifically, the pipeline of our 2D3-SkelAct consists of two steps:pose-view decouplingandpose-view distilling. First, we use a spatio-temporal transformer to decouple 2D skeleton sequences into latent pose and view features, enhancing the model’s ability to learn motion dynamics. Next, these decoupled features are separately integrated into the 2D skeleton model through two cross-attention modules, allowing it to extract discriminative motion features while mitigating uncertainties in 3D viewpoint and depth. Additionally, we distill motion cues from 3D models to compensate for the limitations of 2D skeletons. Remarkably, our model can seamless integrate with various skeleton feature extractors. We validate the proposed 2D3-SkelAct through extensive experiments, demonstrating its adaptability across different model architectures as where consistent improvement achieving. When combined with advanced skeleton feature extractors, 2D3-SkelAct achieves state-of-the-art performance in 2D skeleton-based action recognition. Xiangyue Zhang, Yifan Jia 0007, Zhigang Tu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | FADE: A Dataset for Detecting Falling Objects Around Buildings in VideoabstractObjects falling from buildings, a frequently occurring event in daily life, can cause severe injuries to pedestrians due to the high impact force they exert. Surveillance cameras are often installed around buildings to detect falling objects, but such detection remains challenging due to the small size and fast motion of the objects. Moreover, the field of falling object detection around buildings (FODB) lacks a large-scale dataset for training learning-based detection methods and for standardized evaluation. To address these challenges, we propose a large and diverse video benchmark dataset named FADE. Specifically, FADE contains 2,611 videos from 25 scenes, featuring 8 falling object categories, 4 weather conditions, and 4 video resolutions. Additionally, we develop a novel detection method for FODB that effectively leverages motion information and generates small-sized yet high-quality detection proposals. The efficacy of our method is evaluated on the proposed FADE dataset by comparing it with state-of-the-art approaches in generic object detection, video object detection, and moving object detection. The dataset and code are publicly available at https://fadedataset.github.io/FADE.github.io/. Zhigang Tu 0001, Zhengbo Zhang, Zitao Gao, Chunluan Zhou, Junsong Yuan 0001, Bo Du 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2025 | Informative Sample Selection Model for Skeleton-Based Action Recognition With Limited Training SamplesabstractSkeleton-based human action recognition aims to classify human skeletal sequences, which are spatiotemporal representations of actions, into predefined categories. To reduce the reliance on costly annotations of skeletal sequences while maintaining competitive recognition accuracy, the task of 3D Action Recognition with Limited Training Samples, also known as semi-supervised 3D Action Recognition, has been proposed. In addition, active learning, which aims to proactively select the most informative unlabeled samples for annotation, has been explored in semi-supervised 3D Action Recognition for training sample selection. Specifically, researchers adopt an encoder-decoder framework to embed skeleton sequences into a latent space, where clustering information, combined with a margin-based selection strategy using a multi-head mechanism, is utilized to identify the most informative sequences in the unlabeled set for annotation. However, the most representative skeleton sequences may not necessarily be the most informative for the action recognizer, as the model may have already acquired similar knowledge from previously seen skeleton samples. To solve it, we reformulate Semi-supervised 3D action recognition via active learning from a novel perspective by casting it as a Markov Decision Process (MDP). Built upon the MDP framework and its training paradigm, we train an informative sample selection model to intelligently guide the selection of skeleton sequences for annotation. To enhance the representational capacity of the factors in the state-action pairs within our method, we project them from Euclidean space to hyperbolic space. Furthermore, we introduce a meta tuning strategy to accelerate the deployment of our method in real-world scenarios. Extensive experiments on three 3D action recognition benchmarks demonstrate the effectiveness of our method. Zhigang Tu 0001, Zhengbo Zhang, Jia Gong, Junsong Yuan 0001, Bo Du 0001 |
IEEE Trans. Image Process. | 1 |
| 2025 | Expressive Keypoints for Skeleton-Based Action Recognition via Progressive Skeleton EvolutionabstractIn the realm of skeleton-based human action recognition, the traditional methods which rely on coarse body keypoints fall short of capturing subtle human actions. In this work, we propose Expressive Keypoints that incorporates hand and foot details to form a fine-grained skeletal representation, to improve the discriminative ability for existing models in discerning intricate human actions. However, the increased computational cost from processing nearly three times more joints becomes a new challenge. To address this, we present the Progressive Skeleton Evolution strategy, which significantly improves efficiency while preserving the benefits of fine-grained keypoints. The core idea involves utilizing learnable mapping matrices, semantically initialized to progressively downsample keypoints and prioritize prominent joints by allocating importance weights. Additionally, a plug-and-play Instance Pooling module is exploited to extend our approach to multi-person scenarios without surging computation cost. Extensive experimental results over seven datasets demonstrate the superiority of our method compared to the state-of-the-arts for skeleton-based human action recognition. Code has been made available at https://github.com/YijieYang23/PSE-GCN. Jinlu Zhang 0001, Bo Du 0001, Zhigang Tu 0001 |
IEEE Trans. Image Process. | 5 |
| 2025 | DED-net: multi-scale fusion and illumination-guided enhancement for low-light image restoration
Haosen Dong, Shihao Cheng, Tianyou Fang, Zhigang Tu 0001 |
Vis. Comput. | 5 |
| 2025 | End-to-end pose-action recognition via implicit pose encoding and multi-scale skeleton modeling
Jianlin Zhou, Zhigang Tu 0001 |
Vis. Comput. | 4 |
| 2024 | TapMo: Shape-aware Motion Generation of Skeleton-free CharactersabstractPrevious motion generation methods are limited to the pre-rigged 3D human model, hindering their applications in the animation of various non-rigged characters. In this work, we present TapMo, a Text-driven Animation PIpeline for synthesizing Motion in a broad spectrum of skeleton-free 3D characters. The pivotal innovation in TapMo is its use of shape deformation-aware features as a condition to guide the diffusion model, thereby enabling the generation of mesh-specific motions for various characters. Specifically, TapMo comprises two main components - Mesh Handle Predictor and Shape-aware Diffusion Module. Mesh Handle Predictor predicts the skinning weights and clusters mesh vertices into adaptive handles for deformation control, which eliminates the need for traditional skeletal rigging. Shape-aware Motion Diffusion synthesizes motion with mesh-specific adaptations. This module employs text-guided motions and mesh features extracted during the first stage, preserving the geometric integrity of the animations by accounting for the character's shape and deformation. Trained in a weakly-supervised manner, TapMo can accommodate a multitude of non-human meshes, both with and without associated text motions. We demonstrate the effectiveness and generalizability of TapMo through rigorous qualitative and quantitative experiments. Our results reveal that TapMo consistently outperforms existing auto-animation methods, delivering superior-quality animations for both seen or unseen heterogeneous 3D characters. Shaoli Huang, Zhigang Tu 0001, Xin Chen 0040, Xiaohang Zhan, Gang Yu 0002, Ying Shan |
ICLR | 3 |
| 2024 | Generative Motion Stylization of Cross-structure Characters within Canonical Motion Space
Xin Chen 0040, Gang Yu 0002, Zhigang Tu 0001 |
ACM Multimedia | 4 |
| 2024 | Dark-DSAR: Lightweight one-step pipeline for action recognition in dark videos
Yuwei Yin, Renjie Yang, Yuanzhong Liu, Zhigang Tu 0001 |
Neural Networks | 5 |
| 2024 | A Modular Neural Motion Retargeting System Decoupling Skeleton and Shape PerceptionabstractMotion mapping between characters with different structures but corresponding to homeomorphic graphs, meanwhile preserving motion semantics and perceiving shape geometries, poses significant challenges in skinned motion retargeting. We propose M-R2ET, a modular neural motion retargeting system to comprehensively address these challenges. The key insight driving M-R2ET is its capacity to learn residual motion modifications within a canonical skeleton space. Specifically, a cross-structure alignment module is designed to learn joint correspondences among diverse skeletons, enabling motion copy and forming a reliable initial motion for semantics and geometry perception. Besides, two residual modification modules, i.e., the skeleton-aware module and shape-aware module, preserving source motion semantics and perceiving target character geometries, effectively reduce interpenetration and contact-missing. Driven by our distance-based losses that explicitly model the semantics and geometry, these two modules learn residual motion modifications to the initial motion in a single inference without post-processing. To balance these two motion modifications, we further present a balancing gate to conduct linear interpolation between them. Extensive experiments on the public dataset Mixamo demonstrate that our M-R2ET achieves the state-of-the-art performance, enabling cross-structure motion retargeting, and providing a good balance among the preservation of motion semantics as well as the attenuation of interpenetration and contact-missing Zhigang Tu 0001, Junwu Weng, Junsong Yuan 0001, Bo Du 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Towards robust image matching in low-luminance environments: Self-supervised keypoint detection and descriptor-free cross-fusion matching
Sikang Liu 0001, Yida Wei, Zhichao Wen, Xueli Guo 0001, Zhigang Tu 0001, You Li 0001 |
Pattern Recognit. | 5 |
| 2024 | Patch Similarity Self-Knowledge Distillation for Cross-View Geo-LocalizationabstractCross-view geo-localization is an extremely challenging task due to drastic discrepancies in scene context and object scale between different views. Existing works normally concentrate on aligning the global appearance between two views but underestimate these two discrepancies. In practice, only a small region in the retrieved aerial image can be matched to the whole query ground image (i.e. scene context change). On the other hand, the retrieved aerial images are only able to describe the coarse-grained information but the query ground images can capture the fine-grained details (i.e. object scale change). In this paper, we propose a novel self-distillation framework called Patch Similarity Self-Knowledge Distillation (PaSS-KD), which provides the local and multi-scale knowledge as fine-grained location-related supervision to guide cross-view image feature extraction and representation in a self-enhanced manner. Specifically, we develop an auxiliary image-to-patch retrieval task to explore the scene context change and devise a multi-scale patch partition strategy to sense the object scale change across views. Additionally, our self-distilling framework can be removed to avoid additional computation cost at the inference stage. Extensive experiments show that our method not only achieves the state-of-the-art image retrieval performance on the CVUSA and CVACT benchmarks, but also significantly boosts the fine-grained localization accuracy on the VIGOR dataset. Remarkably, for 10 meter-level localization, we improve the relative accuracy by a factor of 0.8× and 1.6× on the VIGOR dataset under same-area and cross-area evaluation, respectively. Songlian Li, Xiongwu Xiao, Zhigang Tu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Zero-Shot Parameter Learning Network for Low-Light Image Enhancement in Permanently Shadowed RegionsabstractObtaining high-visibility images of the lunar polar permanently shadowed region (PSR) is quite important for internal landforms and material existence exploration. However, PSR images usually have poor quality due to a lack of sufficient illumination. Existing researches, that attempt to address this problem, face challenges caused by relying on virtual assumptions, manual processing, and paired data. To solve these problems, we aim to avoid using paired datasets and directly optimize PSR images, and accordingly propose a zero-shot parameter learning model (ZSPL-PSR) for PSR image enhancement. Our ZSPL-PSR, which enhances PSR images by estimating parameters to adjust image properties, consists of a parameter learning network and a parameter weight learning structure. Particularly, first, a parameter learning network that integrates robust information is constructed to separately estimate the midtone brightness parameters, shadow brightness parameters, and contrast parameters. Where these parameters are beneficial for iteratively improve the overall brightness, shadow brightness, and contrast of the image. Second, a parameter weight learning structure is exploited to coordinate the priority of different parameter maps. In addition, to highlight the terrain details in the enhanced PSR image, we use USM sharpening for postprocessing. The experimental results display the fully interpretable enhanced PSR maps of the lunar north and south poles and their sharpened versions, showcasing rich landforms in PSR. To validate the model performance, a benchmark PSR testing set has been constructed, and extensive comparisons conducted on it demonstrated that ZSPL-PSR exceeds other zero-shot learning methods significantly in image quality. Our code is available athttps://github.com/dl-zfq/ZSPL-PSR. Fengqi Zhang, Zhigang Tu 0001, Weifeng Hao, Fei Li 0023, Mao Ye 0009 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | SGSR-Net: Structure Semantics Guided LiDAR Super-Resolution Network for Indoor LiDAR SLAMabstractMulti-Beam LiDAR (MBL) sensors sample the real-world with discrete 3D point clouds (PC) and have become a major and essential 3D sensing capability for autonomous robots. To ensure an accurate point sampling on surfaces, high-resolution MBL sensors (e.g., Ouster OS0-128) are commonly used to collect dense point clouds for robot tasks, including object detection and tracking, simultaneous localization and mapping (SLAM), in applications such as autonomous driving vehicles (ADVs). However, the high cost and large volume/weight/energy consumption of such sensors limit their usage in broader applications such as UAV/UGV swarms with small-scale agents with limited payload. Existing studies on Super-Resolution (SR) upsampling of the PC from low-resolution MBL have not considered the geometry semantics of the scenes, thus resulting in less optimal SR points for downstream subtasks (e.g., SLAM). Thus, this article proposes SGSR-Net, a structure semantics-guided MBL Super-Resolution network. SGSR-Net takes the low-resolution range images of the MBL sensors as input and produces dense and structure-aware Super-Resolution point cloud from those sparse measurements through a vertical spatial and channel attention-enhanced CNN model coupling with guided Monte Carlo filtering, for indoor LiDAR-SLAM applications. The SGSR-Net is validated using datasets collected by a UGV equipped with multiple MBL sensors. The results demonstrate that the proposed CG-LSR (CASE Attention Guided Encoder-Decoder LiDAR Super-Resolution Network) reduces the MAE of the SR points by 12.4% down to 0.177 m when compared with the state-of-the-art (SOTA) method Shan et al. (2020), Ren et al. (2021), Kwon et al. (2022), Long and Wang (2022). The indoor SLAM results with SR-points produced by SGSR-Net show that the mean and RMSE of the absolute pose error (APE) are decreased by 27% and 30%, down to 0.849 m and 0.902 m, respectively, which significantly improve the indoor-SLAM performance and stability of SOTA LiDAR-SLAM systems (i.e. LeGO-LOAM Shan and Englot (2018), Dellenbach et al. (2022), Vizzo et al. (2023), Zhang and Singh (2014)). Chi Chen 0002, Ang Jin, Zhiye Wang, Yongwei Zheng, Bisheng Yang, Jian Zhou 0011, Zhigang Tu 0001 |
IEEE Trans. Multim. | 8 |
| 2023 | Skinned Motion Retargeting with Residual Perception of Motion Semantics & GeometryabstractA good motion retargeting cannot be reached without reasonable consideration of source-target differences on both the skeleton and shape geometry levels. In this work, we propose a novel Residual RETargeting network (R2ET) structure, which relies on two neural modification modules, to adjust the source motions to fit the target skeletons and shapes progressively. In particular, a skeleton-aware module is introduced to preserve the source motion semantics. A shape-aware module is designed to perceive the geometries of target characters to reduce interpenetration and contact-missing. Driven by our explored distance-based losses that explicitly model the motion semantics and geometry, these two modules can learn residual motion modifications on the source motion to generate plausible retargeted motion in a single inference without postprocessing. To balance these two modifications, we further present a balancing gate to conduct linear interpolation between them. Extensive experiments on the public dataset Mixamo demonstrate that our R2ET achieves the state-of-the-art performance, and provides a good balance between the preservation of motion semantics as well as the attenuation of interpenetration and contact-missing. Code is available at https://github.com/Kebii/R2ET. Junwu Weng, Fang Zhao 0006, Shaoli Huang, Xuefei Zhe, Linchao Bao, Ying Shan, Jue Wang 0001, Zhigang Tu 0001 |
CVPR | 10 |
| 2023 | PHRIT: Parametric Hand Representation with Implicit TemplateabstractWe propose PHRIT, a novel approach for parametric hand mesh modeling with an implicit template that combines the advantages of both parametric meshes and implicit representations. Our method represents deformable hand shapes using signed distance fields (SDFs) with part-based shape priors, utilizing a deformation field to execute the deformation. The model offers efficient high-fidelity hand reconstruction by deforming the canonical template at infinite resolution. Additionally, it is fully differentiable and can be easily used in hand modeling since it can be driven by the skeleton and shape latent codes. We evaluate PHRIT on multiple downstream tasks, including skeleton-driven hand reconstruction, shapes from point clouds, and singleview 3D reconstruction, demonstrating that our approach achieves realistic and immersive hand modeling with state- of-the-art performance. Zhisheng Huang, Yujin Chen, Jinlu Zhang 0001, Zhigang Tu 0001 |
ICCV | 5 |
| 2023 | Consistent 3D Hand Reconstruction in Video via Self-Supervised LearningabstractWe present a method for reconstructing accurate and consistent 3D hands from a monocular video. We observe that the detected 2D hand keypoints and the image texture provide important cues about the geometry and texture of the 3D hand, which can reduce or even eliminate the requirement on 3D hand annotation. Accordingly, in this work, we propose$\mathrm{{S}^{2}HAND}$, a self-supervised 3D hand reconstruction model, that can jointly estimate pose, shape, texture, and the camera viewpoint from a single RGB input through the supervision of easily accessible 2D detected keypoints. We leverage the continuous hand motion information contained in the unlabeled video data and explore$\mathrm{{S}^{2}HAND(V)}$, which uses a set of weights shared$\mathrm{{S}^{2}HAND}$to process each frame and exploits additional motion, texture, and shape consistency constrains to obtain more accurate hand poses, and more consistent shapes and textures. Experiments on benchmark datasets demonstrate that our self-supervised method produces comparable hand reconstruction performance compared with the recent full-supervised methods in single-frame as input setup, and notably improves the reconstruction accuracy and consistency when using the video training data. Zhigang Tu 0001, Zhisheng Huang, Yujin Chen, Linchao Bao, Bisheng Yang, Junsong Yuan 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | DTCM: Joint Optimization of Dark Enhancement and Action Recognition in VideosabstractRecognizing human actions in dark videos is a useful yet challenging visual task in reality. Existing augmentation-based methods separate action recognition and dark enhancement in a two-stage pipeline, which leads to inconsistently learning of temporal representation for action recognition. To address this issue, we propose a novel end-to-end framework termed Dark Temporal Consistency Model (DTCM), which is able to jointly optimize dark enhancement and action recognition, and force the temporal consistency to guide downstream dark feature learning. Specifically, DTCM cascades the action classification head with the dark augmentation network to perform dark video action recognition in a one-stage pipeline. Our explored spatio-temporal consistency loss, which utilizes the RGB-Difference of dark video frames to encourage temporal coherence of the enhanced video frames, is effective for boosting spatio-temporal representation learning. Extensive experiments demonstrated that our DTCM has remarkable performance: 1) Competitive accuracy, which outperforms the state-of-the-arts on the ARID dataset by 2.32% and the UAVHuman-Fisheye dataset by 4.19% in accuracy, respectively; 2) High efficiency, which surpasses the current most advanced method (Chen et al., 2021) with only 6.4% GFLOPs and 71.3% number of parameters; 3) Strong generalization, which can be used in various action recognition methods (e.g., TSM, I3D, 3D-ResNext-101, Video-Swin) to promote their performance significantly. Zhigang Tu 0001, Yuanzhong Liu, Qizi Mu, Junsong Yuan 0001 |
IEEE Trans. Image Process. | 1 |
| 2023 | Joint-Bone Fusion Graph Convolutional Network for Semi-Supervised Skeleton Action RecognitionabstractIn recent years, graph convolutional networks (GCNs) play an increasingly critical role in skeleton-based human action recognition. However, most GCN-based methods still have two main limitations: 1) They only consider the motion information of the joints or process the joints and bones separately, which are unable to fully explore the latent functional correlation between joints and bones for action recognition. 2) Most of these works are performed in the supervised learning way, which heavily relies on massive labeled training data. To address these issues, we propose a semi-supervised skeleton-based action recognition method which has been rarely exploited before. We design a novel correlation-driven joint-bone fusion graph convolutional network (CD-JBF-GCN) as an encoder and use a pose prediction head as a decoder to achieve semi-supervised learning. Specifically, the correlation-driven joint-bone fusion graph convolution (CD-JBF-GC) can explore the motion transmission between the joint stream and the bone stream, so as to promote both streams to learn more discriminative feature representations. The pose prediction based auto-encoder in the self-supervised training fashion allows the network to learn motion representation from the unlabeled data, which is essential for action recognition. Extensive experiments on two popular datasets, i.e. NTU-RGB+D and Kinetics-Skeleton, demonstrate that our model achieves the state-of-the-art performance for semi-supervised skeleton-based action recognition and is also useful for fully-supervised methods. Zhigang Tu 0001, Hongyan Li 0003, Yujin Chen, Junsong Yuan 0001 |
IEEE Trans. Multim. | 1 |
| 2023 | Flow-pose Net: an effective two-stream network for fall detection
Kexin Fei, Chao Wang 0010, Yuanzhong Liu, Zhigang Tu 0001 |
Vis. Comput. | 6 |
| 2023 | Graph-aware transformer for skeleton-based action recognition
Wei Xie 0008, Chao Wang 0010, Ruide Tu, Zhigang Tu 0001 |
Vis. Comput. | 5 |
| 2022 | MixSTE: Seq2seq Mixed Spatio-Temporal Encoder for 3D Human Pose Estimation in VideoabstractRecent transformer-based solutions have been introduced to estimate 3D human pose from 2D keypoint sequence by considering body joints among all frames globally to learn spatio-temporal correlation. We observe that the motions of different joints differ significantly. However, the previous methods cannot efficiently model the solid inter-frame correspondence of each joint, leading to insufficient learning of spatial-temporal correlation. We propose MixSTE (Mixed Spatio-Temporal Encoder), which has a temporal transformer block to separately model the temporal motion of each joint and a spatial transformer block to learn inter-joint spatial correlation. These two blocks are utilized alternately to obtain better spatio-temporal feature encoding. In addition, the network output is extended from the central frame to entire frames of the input video, thereby improving the coherence between the input and output sequences. Extensive experiments are conducted on three benchmarks (i.e. Human3.6M, MPI-INF-3DHP, and HumanEva). The results show that our model outperforms the state-of-the-art approach by 10.9% P-MPJPE and 7.6% MPJPE. The code is available at https://github.com/JinluZhang1126/MixSTE. Jinlu Zhang 0001, Zhigang Tu 0001, Jianyu Yang 0002, Yujin Chen, Junsong Yuan 0001 |
CVPR | 2 |
| 2022 | Distilling Inter-Class Distance for Semantic SegmentationabstractKnowledge distillation is widely adopted in semantic segmentation to reduce the computation cost. The previous knowledge distillation methods for semantic segmentation focus on pixel-wise feature alignment and intra-class feature variation distillation, neglecting to transfer the knowledge of the inter-class distance in the feature space, which is important for semantic segmentation such a pixel-wise classification task. To address this issue, we propose an Inter-class Distance Distillation (IDD) method to transfer the inter-class distance in the feature space from the teacher network to the student network. Furthermore, semantic segmentation is a position-dependent task, thus we exploit a position information distillation module to help the student network encode more position information. Extensive experiments on three popular datasets: Cityscapes, Pascal VOC and ADE20K show that our method is helpful to improve the accuracy of semantic segmentation models and achieves the state-of-the-art performance. E.g. it boosts the benchmark model (``PSPNet+ResNet18") by 7.50% in accuracy on the Cityscapes dataset. Zhengbo Zhang, Chunluan Zhou, Zhigang Tu 0001 |
IJCAI | 3 |
| 2022 | Multi-Hyperedge Hypergraph for Group Activity RecognitionabstractGroup activity recognition aims to identify group activities from the videos. Most of the previous methods focus on modeling between individuals (one-to-one), which ignores the fact that a single individual's behavior may be jointly determined by multiple individual behaviors (many-to-one). For this reason, we propose a Multi-Hyperedge Hypergraph (MHH) to capture high-order relationships between multiple people. Specifically, we build three different types of hyperedges on the hypergraph structure. Each hyperedge can accommodate the characteristics of multiple nodes to capture different types of high-order relationships between nodes. Then, we use the late fusion method to fuse the three features to further enhance the overall behavioral representation. Finally, we perform a series of experiments on two of the most widely used benchmarks in group activity recognition, which have proved the effectiveness of MHH. More importantly, as far as we know, this is the first case of using a hypergraph structure for group activity recognition. Wanxin Li, Wei Xie 0008, Zhigang Tu 0001, Lianghao Jin |
IJCNN | 3 |
| 2022 | Multi-Part Adaptive Graph Convolutional Network for Skeleton-Based Action RecognitionabstractIn skeleton-based action recognition task, graph convolutional network has attracted widespread attention and achieved remarkable results. However, most of the current methods are performing graph convolution on the entire skeleton graph, ignoring the fact that people are composed of different body parts. In addition, previous work ignores the temporal and spatial independence and relevance of different parts. Thus, to solve these issues, we optimize the representation of the skeleton graph, graph convolution and temporal convolution respectively. In this work, we propose multi-part adaptive graph convolution (MPA-GC) to adaptively learn the topology of each part of the body and dynamically aggregate the relevance between them. Meanwhile, we add a multi-scale temporal convolution module to better obtain temporal dimension features. Ultimately, we develop a powerful graph convolutional network named MPA-GCN, and extensive experiments on two public large-scale datasets NTU-RGB+D and NTU-RGB+D120 demonstrate the effectiveness of our module, which outperforms state-of-the-art methods. Wei Xie 0008, Zhigang Tu 0001, Wanxin Li, Lianghao Jin |
IJCNN | 3 |
| 2022 | Uncertainty-Aware 3D Human Pose Estimation from Monocular VideoabstractEstimating the 3D human pose from the monocular video is challenging mainly due to the depth ambiguity and inaccurate 2D detected keypoints. To quantify the depth uncertainty of 3D human pose via the neural network, we imbue the uncertainty modeling to depth prediction by using evidential deep learning (EDL). Meanwhile, to calibrate the distribution uncertainty of the 2D detection, we explore a probabilistic representation to model the realistic distribution. Specifically, we exploit the EDL to measure the depth prediction uncertainty of the network, and decompose the x-y coordinates into individual distributions to model the deviation uncertainty of the inaccurate 2D keypoints. Then we optimize the depth uncertainty parameters and calibrate the 2D deviations to obtain accurate 3D human poses. Besides, to provide effective latent features for uncertainty learning, we design an encoder which combines graph convolutional network (GCN) and transformer to learn discriminative spatio-temporal representations. Extensive experiments are conducted on three benchmarks (Human3.6M, MPI-INF-3DHP, and HumanEva-I) and the comprehensive results show that our model surpasses the state-of-the-arts by a large margin. Jinlu Zhang 0001, Yujin Chen, Zhigang Tu 0001 |
ACM Multimedia | 3 |
| 2022 | Video anomaly detection with spatio-temporal dissociation
Yunpeng Chang, Zhigang Tu 0001, Wei Xie 0008, Bin Luo 0005, Shifu Zhang, Haigang Sui, Junsong Yuan 0001 |
Pattern Recognit. | 2 |
| 2022 | Zoom Transformer for Skeleton-Based Group Activity RecognitionabstractSkeleton-based human action recognition has attracted increasing attention and many methods have been proposed to boost the performance. However, these methods still confront three main limitations: 1) Focusing on single-person action recognition while neglecting the group activity of multiple people (more than 5 people). In practice, multi-person group activity recognition via skeleton data is also a meaningful problem. 2) Unable to mine high-level semantic information from the skeleton data, such as interactions among multiple people and their positional relationships. 3) Existing datasets used for multi-person group activity recognition are all RGB videos involved, which cannot be directly applied to skeleton-based group activity analysis. To address these issues, we propose a novel Zoom Transformer to exploit both the low-level single-person motion information and the high-level multi-person interaction information in a uniform model structure with carefully designed Relation-aware Maps. Besides, we estimate the multi-person skeletons from the existing real-world video datasets i.e. Kinetics and Volleyball-Activity, and release two new benchmarks to verify the effectiveness of our Zoom Transfromer. Extensive experiments demonstrate that our model can effectively cope with the skeleton-based multi-person group activity. Additionally, experiments on the large-scale NTU-RGB+D dataset validate that our model also achieves remarkable performance for single-person action recognition. The code and the skeleton data are publicly available athttps://github.com/Kebii/Zoom-Transformer Yifan Jia 0007, Wei Xie 0008, Zhigang Tu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | A General Dynamic Knowledge Distillation Method for Visual AnalyticsabstractExisting knowledge distillation (KD) method normally fixes the weight of the teacher network, and uses the knowledge from the teacher network to guide the training of the student network no-ninteractively, thus it is called static knowledge distillation (SKD). SKD is widely used in model compression on the homologous data and knowledge transfer on the heterogeneous data. However, the teacher network that with fixed-weight constrains the student network to learn knowledge from it. It is worth expecting that the teacher network itself can be continuously optimized to promote the learning ability of the student network dynamically. To overcome this limitation, we propose a novel dynamic knowledge distillation (DKD) method, in which the teacher network and the student network can learn from each other interactively. Importantly, we analyzed the effectiveness of DKD mathematically (see Eq. 4), and addressed one crucial issue caused by the continuous change of the teacher network in the dynamic distillation process via designing a valid loss function. We verified the practicality of our DKD by extensive experiments on various visual tasks,e.g. for model compression, we conducted experiments on image classification and object detection. For knowledge transfer, video-based human action recognition is chosen for analysis. The experimental results on benchmark datasets (i.e. ILSVRC2012, COCO2017, HMDB51, UCF101) demonstrated that the proposed DKD is valid to improve the performance of these visual tasks for a large margin. The source code is publicly available online at1. Zhigang Tu 0001, Xiangjian Liu |
IEEE Trans. Image Process. | 1 |
| 2022 | Motion-Driven Visual Tempo Learning for Video-Based Action RecognitionabstractAction visual tempo characterizes the dynamics and the temporal scale of an action, which is helpful to distinguish human actions that share high similarities in visual dynamics and appearance. Previous methods capture the visual tempo either by sampling raw videos with multiple rates, which require a costly multi-layer network to handle each rate, or by hierarchically sampling backbone features, which rely heavily on high-level features that miss fine-grained temporal dynamics. In this work, we propose a Temporal Correlation Module (TCM), which can be easily embedded into the current action recognition backbones in a plug-in-and-play manner, to extract action visual tempo from low-level backbone features at single-layer remarkably. Specifically, our TCM contains two main components: a Multi-scale Temporal Dynamics Module (MTDM) and a Temporal Attention Module (TAM). MTDM applies a correlation operation to learn pixel-wise fine-grained temporal dynamics for both fast-tempo and slow-tempo. TAM adaptively emphasizes expressive features and suppresses inessential ones via analyzing the global information across various tempos. Extensive experiments conducted on several action recognition benchmarks, e.g. Something-Something V1&V2, Kinetics-400, UCF-101, and HMDB-51, have demonstrated that the proposed TCM is effective to promote the performance of the existing video-based action recognition models for a large margin. The source code is publicly released at https://github.com/zphyix/TCM. Yuanzhong Liu, Junsong Yuan 0001, Zhigang Tu 0001 |
IEEE Trans. Image Process. | 3 |
| 2021 | Model-Based 3D Hand Reconstruction via Self-Supervised LearningabstractReconstructing a 3D hand from a single-view RGB image is challenging due to various hand configurations and depth ambiguity. To reliably reconstruct a 3D hand from a monocular image, most state-of-the-art methods heavily rely on 3D annotations at the training stage, but obtaining 3D annotations is expensive. To alleviate reliance on labeled training data, we propose S2HAND, a self-supervised 3D hand reconstruction network that can jointly estimate pose, shape, texture, and the camera viewpoint. Specifically, we obtain geometric cues from the input image through easily accessible 2D detected keypoints. To learn an accurate hand reconstruction model from these noisy geometric cues, we utilize the consistency between 2D and 3D representations and propose a set of novel losses to rationalize outputs of the neural network. For the first time, we demonstrate the feasibility of training an accurate 3D hand reconstruction network without relying on manual annotations. Our experiments show that the proposed self-supervised method achieves comparable performance with recent fully-supervised methods. The code is available at https://github.com/TerenceCYJ/S2HAND. Yujin Chen, Zhigang Tu 0001, Linchao Bao, Ying Zhang 0021, Xuefei Zhe, Ruizhi Chen, Junsong Yuan 0001 |
CVPR | 2 |
| 2021 | Multi-Attribute Enhancement Network for Person SearchabstractPerson Search is designed to jointly solve the problems of Person Detection and Person Re-identification (Re-ID), in which the target person will be located in a large number of uncut images. Over the past few years, Person Search based on deep learning has made great progress. Visual character attributes play a key role in retrieving the query person, which has been explored in Re-ID but has been ignored in Person Search. So, we introduce attribute learning into the model, allowing the use of attribute features for retrieval task. Specifically, we propose a simple and effective model called Multi-Attribute Enhancement (MAE) which introduces attribute tags to learn local features. In addition to learning the global representation of pedestrians, it also learns the local representation, and combines the two aspects to learn robust features to promote the search performance. Additionally, we verify the effectiveness of our module on the existing benchmark dataset, CUHK-SYSU and PRW. Ultimately, our model achieves state-of-the-art among end-to-end methods, especially reaching 91.8% of mAP and 93.0% of rank-1 on CUHK-SYSU. Codes and models are available at https:// github. com/chenlq123/ MAE. Lequan Chen, Wei Xie 0008, Zhigang Tu 0001, Jinglei Guo, Yaping Tao |
IJCNN | 3 |
| 2021 | Automatic segmentation of left and right ventricles in cardiac MRI using 3D-ASM and deep learning
Huaifei Hu, Ning Pan, Haihua Liu, Liman Liu, Tailang Yin, Zhigang Tu 0001, Alejandro F. Frangi |
Signal Process. Image Commun. | 6 |
| 2021 | Joint Hand-Object 3D Reconstruction From a Single Image With Cross-Branch Feature FusionabstractAccurate 3D reconstruction of the hand and object shape from a hand-object image is important for understanding human-object interaction as well as human daily activities. Different from bare hand pose estimation, hand-object interaction poses a strong constraint on both the hand and its manipulated object, which suggests that hand configuration may be crucial contextual information for the object, and vice versa. However, current approaches address this task by training a two-branch network to reconstruct the hand and object separately with little communication between the two branches. In this work, we propose to consider hand and object jointly in feature space and explore the reciprocity of the two branches. We extensively investigate cross-branch feature fusion architectures with MLP or LSTM units. Among the investigated architectures, a variant with LSTM units that enhances object feature with hand feature shows the best performance gain. Moreover, we employ an auxiliary depth estimation module to augment the input RGB image with the estimated depth map, which further improves the reconstruction accuracy. Experiments conducted on public datasets demonstrate that our approach significantly outperforms existing approaches in terms of the reconstruction accuracy of objects. Yujin Chen, Zhigang Tu 0001, Ruizhi Chen, Linchao Bao, Zhengyou Zhang, Junsong Yuan 0001 |
IEEE Trans. Image Process. | 2 |
| 2020 | Clustering Driven Deep Autoencoder for Video Anomaly Detection
Yunpeng Chang, Zhigang Tu 0001, Wei Xie 0008, Junsong Yuan 0001 |
ECCV (15) | 2 |
| 2020 | Detecting spatiotemporal irregularities in videos via a 3D convolutional autoencoder
Mengjia Yan 0003, Jingjing Meng, Chunluan Zhou, Zhigang Tu 0001, Yap-Peng Tan, Junsong Yuan 0001 |
J. Vis. Commun. Image Represent. | 4 |
| 2020 | Learning motion representation for real-time spatio-temporal action localization
Dejun Zhang, Linchao He, Zhigang Tu 0001, Shifu Zhang, Boxiong Yang |
Pattern Recognit. | 3 |
| 2020 | Unsupervised Learning of Optical Flow With CNN-Based Non-Local FilteringabstractEstimating optical flow from successive video frames is one of the fundamental problems in computer vision and image processing. In the era of deep learning, many methods have been proposed to use convolutional neural networks (CNNs) for optical flow estimation in an unsupervised manner. However, the performance of unsupervised optical flow approaches is still unsatisfactory and often lagging far behind their supervised counterparts, primarily due to over-smoothing across motion boundaries and occlusion. To address these issues, in this paper, we propose a novel method with a new post-processing term and an effective loss function to estimate optical flow in an unsupervised, end-to-end learning manner. Specifically, we first exploit a CNN-based non-local term to refine the estimated optical flow by removing noise and decreasing blur around motion boundaries. This is implemented via automatically learning weights of dependencies over a large spatial neighborhood. Because of its learning ability, the method is effective for various complicated image sequences. Secondly, to reduce the influence of occlusion, a symmetrical energy formulation is introduced to detect the occlusion map from refined bi-directional optical flows. Then the occlusion map is integrated to the loss function. Extensive experiments are conducted on challenging datasets, i.e. FlyingChairs, MPI-Sintel and KITTI to evaluate the performance of the proposed method. The state-of-the-art results demonstrate the effectiveness of our proposed method. Zhigang Tu 0001, Dejun Zhang, Jun Liu 0036, Baoxin Li, Junsong Yuan 0001 |
IEEE Trans. Image Process. | 2 |
| 2020 | Part-based visual tracking with spatially regularized correlation filters
Dejun Zhang, Lu Zou, Zhuyang Xie, Fazhi He, Yiqi Wu, Zhigang Tu 0001 |
Vis. Comput. | 7 |
| 2019 | SO-HandNet: Self-Organizing Network for 3D Hand Pose Estimation With Semi-Supervised Learningabstract3D hand pose estimation has made significant progress recently, where Convolutional Neural Networks (CNNs) play a critical role. However, most of the existing CNN-based hand pose estimation methods depend much on the training set, while labeling 3D hand pose on training data is laborious and time-consuming. Inspired by the point cloud autoencoder presented in self-organizing network (SO-Net), our proposed SO-HandNet aims at making use of the unannotated data to obtain accurate 3D hand pose estimation in a semi-supervised manner. We exploit hand feature encoder (HFE) to extract multi-level features from hand point cloud and then fuse them to regress 3D hand pose by a hand pose estimator (HPE). We design a hand feature decoder (HFD) to recover the input point cloud from the encoded feature. Since the HFE and the HFD can be trained without 3D hand pose annotation, the proposed method is able to make the best of unannotated data during the training phase. Experiments on four challenging benchmark datasets validate that our proposed SO-HandNet can achieve superior performance for 3D hand pose estimation via semi-supervised learning. Yujin Chen, Zhigang Tu 0001, Liuhao Ge, Dejun Zhang, Ruizhi Chen, Junsong Yuan 0001 |
ICCV | 2 |
| 2019 | A survey of variational and CNN-based optical flow techniques
Zhigang Tu 0001, Wei Xie 0008, Dejun Zhang, Ronald Poppe, Remco C. Veltkamp, Baoxin Li, Junsong Yuan 0001 |
Signal Process. Image Commun. | 1 |
| 2019 | Semantic Cues Enhanced Multimodality Multistream CNN for Action RecognitionabstractThis paper addresses the issue of video-based action recognition by exploiting an advanced multistream convolutional neural network (CNN) to fully use semantics-derived multiple modalities in both spatial (appearance) and temporal (motion) domains, since the performance of the CNN-based action recognition methods heavily relates to two factors: semantic visual cues and the network architecture. Our work consists of two major parts. First, to extract useful human-related semantics accurately, we propose a novel spatiotemporal saliency-based video object segmentation (STS) model. By fusing different distinctive saliency maps, which are computed according to object signatures of complementary object detection approaches, a refined STS maps can be obtained. In this way, various challenges in the realistic video can be handled jointly. Based on the estimated saliency maps, an energy function is constructed to segment two semantic cues: the actor and one distinctive acting part of the actor. Second, we modify the architecture of the two-stream network (TS-Net) to design a multistream network that consists of three TS-Nets with respect to the extracted semantics, which is able to use deeper abstract visual features of multimodalities in multi-scale spatiotemporally. Importantly, the performance of action recognition is significantly boosted when integrating the captured human-related semantics into our framework. Experiments on four public benchmarks-JHMDB, HMDB51, UCF-Sports, and UCF101-demonstrate that the proposed method outperforms the state-of-the-art algorithms. Zhigang Tu 0001, Wei Xie 0008, Justin Dauwels, Baoxin Li, Junsong Yuan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2019 | Action-Stage Emphasized Spatiotemporal VLAD for Video Action RecognitionabstractDespite outstanding performance in image recognition, convolutional neural networks (CNNs) do not yet achieve the same impressive results on action recognition in videos. This is partially due to the inability of CNN for modeling long-range temporal structures especially those involving individual action stages that are critical to human action recognition. In this paper, we propose a novel action-stage (ActionS) emphasized spatiotemporal Vector of Locally Aggregated Descriptors (ActionS-STVLAD) method to aggregate informative deep features across the entire video according to adaptive video feature segmentation and adaptive segment feature sampling (AVFS-ASFS). In our ActionSST- VLAD encoding approach, by using AVFS-ASFS, the key frame features are chosen and the corresponding deep features are automatically split into segments with the features in each segment belonging to a temporally coherent ActionS. Then, based on the extracted key frame feature in each segment, a flow-guided warping technique is introduced to detect and discard redundant feature maps, while the informative ones are aggregated by using our exploited similarity weight. Furthermore, we exploit an RGBF modality to capture motion salient regions in the RGB images corresponding to action activity. Extensive experiments are conducted on four public benchmarks - HMDB51, UCF101, Kinetics and ActivityNet for evaluation. Results show that our method is able to effectively pool useful deep features spatiotemporally, leading to state-of-the-art performance for videobased action recognition. Zhigang Tu 0001, Hongyan Li 0003, Dejun Zhang, Justin Dauwels, Baoxin Li, Junsong Yuan 0001 |
IEEE Trans. Image Process. | 1 |
| 2018 | Actor-Action Semantic Segmentation with Region Masks
Kang Dang, Chunluan Zhou, Zhigang Tu 0001, Michael Hoy, Justin Dauwels, Junsong Yuan 0001 |
BMVC | 3 |
| 2018 | Salience Guided Depth Calibration for Perceptually Optimized Compressive Light Field 3D DisplayabstractMulti-layer light field displays are a type of computational three-dimensional (3D) display which has recently gained increasing interest for its holographic-like effect and natural compatibility with 2D displays. However, the major shortcoming, depth limitation, still cannot be overcome in the traditional light field modeling and reconstruction based on multi-layer liquid crystal displays (LCDs). Considering this disadvantage, our paper incorporates a salience guided depth optimization over a limited display range to calibrate the displayed depth and present the maximum area of salience region for multi-layer light field display. Different from previously reported cascaded light field displays that use the fixed initialization plane as the depth center of display content, our method automatically calibrates the depth initialization based on the salience results derived from the proposed contrast enhanced salience detection method. Experiments demonstrate that the proposed method provides a promising advantage in visual perception for the compressive light field displays from both software simulation and prototype demonstration. Shizheng Wang, Wenjuan Liao, Philip Surman, Zhigang Tu 0001, Yuanjin Zheng, Junsong Yuan 0001 |
CVPR | 4 |
| 2018 | Multi-stream CNN: Learning representations based on human-related regions for action recognition
Zhigang Tu 0001, Wei Xie 0008, Qianqing Qin, Ronald Poppe, Remco C. Veltkamp, Baoxin Li, Junsong Yuan 0001 |
Pattern Recognit. | 1 |
| 2017 | Variational method for joint optical flow estimation and edge-aware image restoration
Zhigang Tu 0001, Wei Xie 0008, Coert Van Gemeren, Ronald Poppe, Remco C. Veltkamp |
Pattern Recognit. | 1 |
| 2017 | Fusing disparate object signatures for salient object detection in video
Zhigang Tu 0001, Zuwei Guo, Wei Xie 0008, Mengjia Yan 0003, Remco C. Veltkamp, Baoxin Li, Junsong Yuan 0001 |
Pattern Recognit. | 1 |
| 2016 | MSR-CNN: Applying motion salient region based descriptors for action recognitionabstractIn recent years the most popular video-based human action recognition methods rely on extracting feature representations using Convolutional Neural Networks (CNN) and then using these representations to classify actions. In this work, we propose a fast and accurate video representation that is derived from the motion-salient region (MSR), which represents features most useful for action labeling. By improving a well-performed foreground detection technique, the region of interest (ROI) corresponding to actors in the foreground in both the appearance and the motion field can be detected under various realistic challenges. Furthermore, we propose a complementary motion salient measure to select a secondary ROI - the major moving part of the human. Accordingly, a MSR-based CNN descriptor (MSR-CNN) is formulated to recognize human action, where the descriptor incorporates appearance and motion features along with tracks of MSR. The computation can be efficiently implemented due to two characteristics: 1) only part of the RGB image and the motion field need to be processed; 2) less data is used as input for the CNN feature extraction. Comparative evaluation on JHMDB and UCF Sports datasets shows that our method outperforms the state-of-the-art in both efficiency and accuracy. Zhigang Tu 0001, Yikang Li 0001, Baoxin Li |
ICPR | 1 |
| 2016 | Weighted local intensity fusion method for variational optical flow estimation
Zhigang Tu 0001, Ronald Poppe, Remco C. Veltkamp |
Pattern Recognit. | 1 |
| 2016 | Adaptive guided image filter for warping in variational optical flow computation
Zhigang Tu 0001, Ronald Poppe, Remco C. Veltkamp |
Signal Process. | 1 |
| 2014 | Improved Color Patch Similarity Measure Based Weighted Median Filter
Zhigang Tu 0001, Coert Van Gemeren, Remco C. Veltkamp |
ACCV (5) | 1 |
| 2014 | A combined post-filtering method to improve accuracy of variational optical flow estimation
Zhigang Tu 0001, Nico Van der Aa, Coert Van Gemeren, Remco C. Veltkamp |
Pattern Recognit. | 1 |