VLDB 2026 Research / reviewers in the wild / expert
Yang Xiao 0007
dblp:181/1848-7
· DBLP profile ↗
86ranked-venue papers
6as first author
44since 2021 · last 2026
0000-0002-7739-4146ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 47 · 3 first-author · 21 since 2021Artificial intelligence and machine learning · 36 · 2 first-author · 22 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 6 since 2021Security and privacy · 4 · 3 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 2 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DEFANet: Dual-Path Edge-Target Collaboration with Frequency-Aware Enhancement for Infrared Small Target DetectionabstractInfrared small target detection is challenging due to limited target size and low signal-to-noise ratio. Unlike common targets, infrared small targets contain a higher proportion of edge pixels and exhibit blurred boundaries due to diffraction and quantization artifacts, making boundaries uniquely valuable cues for target perception. However, existing methods often emphasize holistic modeling while underutilizing such informative boundary cues. Motivated by this observation, we propose a Dual-Path Edge-Guided Frequency-Aware Network (DEFANet), which enables edge-target collaborative modeling for enhanced feature representation. DEFANet features a dual-path design, consisting of a main branch for holistic target modeling and an edge branch for boundary transition perception. To facilitate interaction and enhance representation in both branches, we introduce two core modules: Frequency-Aware Dual Enhancement Module (FADE) and Edge-Guided Integration Module (EGI). FADE employs a Frequency-Decoupled Attention Enhancement Mechanism to enhance both branches in the frequency domain, strengthening holistic modeling in the main branch and boundary representation in the edge branch. EGI leverages a Dual-Path Group-Wise Guidance Mechanism to integrate enhanced edge features into the main branch, improving boundary perception. Extensive experiments on four public infrared small target datasets, MDvsFA, LAFT, SIRST, and SIATD, demonstrate that DEFANet achieves SOTA performance. Ablation studies further validate the effectiveness of DEFANet and the soundness of its design motivation. Shuaiyuan Du, Yang Xiao 0007, Zhiguo Cao 0001 |
AAAI | 2 |
| 2026 | DeFB: Decomposed Feature Learning for Real-Time Multi-Person Eyeblink Detection in Untrimmed In-the-Wild VideosabstractMulti-person eyeblink detection in untrimmed in-the-wild videos is a recently emerged and challenging task. Due to its significant spatio-temporal fine-grained characteristics compared to general actions, we empirically find that general action detectors, though effective in general domains, struggle with this task (i.e., Blink-AP < 2%). Specialized eyeblink detection methods alleviate it through fine-grained spatio-temporal operations. SOTA method proposes a unified model combining instance-aware face localization and eyeblink detection through joint multi-task learning and feature sharing. While effective, it exhibits two critical limitations that may contribute to its unsatisfactory performance (i.e., Blink-AP=10.11%): (1) Face localization and eyeblink detection require distinct spatio-temporal feature granularities, making joint modeling in a unified feature space suboptimal. (2) Eyeblink task training could be largely affected by unstable face-eye feature learning under the joint training paradigm. To address this, we propose DeFB, a decomposed feature learning paradigm with favorable effectiveness and efficiency: (1) We model faces and eyes in granularity-specific feature spaces, which enhances fine-grained perception while reducing computational costs compared to a unified feature space. (2) To mitigate face-eye feature learning instability, we adopt an asynchronous learning mechanism where eye feature learning refines well-trained coarse face features, with shared queries acting as a bridge between stages to retain the efficient feature sharing of existing unified models. Compared with SOTA method, DeFB doubles the performance (Blink-AP: 24.65% v.s. 10.11%) while boosting efficiency by nearly 35%. DeFB can also be integrated as a plug-in to substantially augment the eyeblink detection capabilities of general action detectors. Jinfang Gan, Wenzheng Zeng, Yang Xiao 0007, Xintao Zhang, Chaoyang Zheng, Ran Wang 0005, Zhiguo Cao 0001 |
AAAI | 3 |
| 2026 | Phys-Liquid: A Physics-Informed Dataset for Estimating 3D Geometry and Volume of Transparent Deformable LiquidsabstractEstimating the geometric and volumetric properties of transparent deformable liquids is challenging due to optical complexities and dynamic surface deformations induced by container movements. Autonomous robots performing precise liquid manipulation tasks—such as dispensing, aspiration, and mixing—must handle containers in ways that inevitably induce these deformations, complicating accurate liquid state assessment. Current datasets lack comprehensive physics-informed simulation data representing realistic liquid behaviors under diverse dynamic scenarios. To bridge this gap, we introduce Phys-Liquid, a physics-informed dataset comprising 97,200 simulation images and corresponding 3D meshes, capturing liquid dynamics across multiple laboratory scenes, lighting conditions, liquid colors, and container rotations. To validate the realism and effectiveness of Phys-Liquid, we propose a four-stage reconstruction and estimation pipeline involving liquid segmentation, multi-view mask generation, 3D mesh reconstruction, and real-world scaling. Experimental results demonstrate improved accuracy and consistency in reconstructing liquid geometry and volume, outperforming existing benchmarks. The dataset and associated validation methods facilitate future advancements in transparent liquid perception tasks. Ke Ma 0012, Yizhou Fang, Jean-Baptiste Weibel, Xinggang Wang, Yang Xiao 0007, Yi Fang 0006 |
AAAI | 6 |
| 2026 | MoEG-HOI: Mixture of Expert Groups for One-Stage Hand-Object Interaction Motion Generation with Hand-Finger-Joint Semantic GuidanceabstractIn this paper, MoEG-HOI is proposed as a novel method for the challenging 3D hand-object interaction (HOI) motion generation task, by introducing Mixture-of-Experts (MoE) to this field for the first time. Almost all the mainstream approaches in HOI motion generation leverage diffusion model as its strong generative ability. Nevertheless, due to HOI’s fine-grained property, well training diffusion in one-stage way is actually not trivial. Existing state-of-the-art (SOTA) methods (e.g.,Text2HOI and MF-MDM) alleviate this mainly via a coarse-to-fine, multi-stage paradigm. Although effective and practical, this paradigm prevents end-to-end training for optimal performance. In contrast, MoEG-HOI applies MoE to address this in one-stage way, with end-to-end training ability. This allows each expert to specialize in certain distinct HOI patterns, which alleviates individual expert’s training difficulty. However, intuitively applying MoE is not optimal due to the issues of: (1) towards expert design, original MoE cannot well characterize hand’s articulated structure at the levels of hand, finger, and joint explicitly, and (2) for expert routing mechanism, the characteristics of variational HOI action classes and diffusion noise levels have not been concerned. Towards the first problem, MoE’s experts are designed into groups that correspond to motion generation for hand, finger, and joint respectively, under the semantic guidance from global to local. To facilitate this, HOI’s text description will be correspondingly refined at Hand-Finger-Joint levels using LLM. Secondly, during MoE routing, the information of HOI’s action label and diffusion noise level is concerned to select experts jointly, to better reveal actions’ inter-class variation and dynamics of diffusion generation. SOTA performance on ARCTIC, GRAB and H2O datasets demonstrates the effectiveness of our method. Yang Xiao 0007, Changlong Jiang, Haohong Kuang, Ran Wang 0005 |
AAAI | 2 |
| 2026 | 3D Hand Pose Estimation via Articulated Anchor-to-Joint 3D Local RegressorsabstractIn this paper, we propose to address monocular 3D hand pose estimation from a single RGB or depth image via articulated anchor-to-joint 3D local regressors, in form of A2J-Transformer+. The key idea is to make the local regressors (i.e., anchor points) in 3D space be aware of hand's local fine details and global articulated context jointly, to facilitate predicting their 3D offsets toward hand joints with linear weighted aggregation for joint localization. Our intuition is that, local fine details help to estimate accurate offset but may suffer from the issues including serious occlusion, confusing similar patterns, and overfitting risk. On the other hand, hand's global articulated context can essentially provide additional descriptive clues and constraints to alleviate these issues. To set anchor points adaptively in 3D space, A2J-Transformer+ runs in a 2-stage manner. At the first stage, since the input modality property anchor points distribute more densely on X-Y plane, it leads to lower prediction accuracy along Z direction compared with those in the X and Y directions. To alleviate this, at the second stage anchor points are set near the joints yielded by the first stage evenly along X, Y, and Z directions. This treatment brings two main advantages: (1) balancing the prediction accuracy along X, Y, and Z directions, and (2) ensuring the anchor-joint offsets are of small values relatively easy to estimate. Wide-range experiments on three RGB hand datasets (InterHand2.6 M, HO-3D V2 and RHP) and three depth hand datasets (NYU, ICVL and HANDS 2017) verify A2J-Transformer+'s superiority and generalization ability for different modalities (i.e., RGB and depth) and hand cases (i.e., single hand, interacting hands, and hand-object interaction), even outperforming model-based manners. The test on ITOP dataset reveals that, A2J-Transformer+ can also be applied to 3D human pose estimation task. Changlong Jiang, Yang Xiao 0007, Jinghong Zheng 0002, Haohong Kuang, Cunlin Wu, Zhiguo Cao 0001, Joey Tianyi Zhou, Junsong Yuan 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | PandaPose: 3D Human Pose Lifting from a Single Image via Propagating 2D Pose Prior to 3D Anchor Spaceabstract3D human pose lifting from a single RGB image is a challenging task in 3D vision. Existing methods typically establish a direct joint-to-joint mapping from 2D to 3D poses based on 2D features. This formulation suffers from two fundamental limitations: inevitable error propagation from input predicted 2D pose to 3D predictions and inherent difficulties in handling self-occlusion cases.
In this paper, we propose PandaPose, a 3D human pose lifting approach via propagating 2D pose prior to 3D anchor space as the unified intermediate representation. Specifically, our 3D anchor space comprises:
(1) Joint-wise 3D anchors in the canonical coordinate system, providing accurate and robust priors to mitigate 2D pose estimation inaccuracies.
(2) Depth-aware joint-wise feature lifting that hierarchically integrates depth information to resolve self-occlusion ambiguities.
(3) The anchor-feature interaction decoder that incorporates 3D anchors with lifted features to generate unified anchor queries encapsulating joint-wise 3D anchor set, visual cues and geometric depth information.
The anchor queries are further employed to facilitate anchor-to-joint ensemble prediction.
Experiments on three well-established benchmarks (i.e., Human3.6M, MPI-INF-3DHP and 3DPW) demonstrate the superiority of our proposition.
The substantial reduction in error by 14.7% compared to SOTA methods
on the challenging conditions of Human3.6M and qualitative comparisons further showcase the effectiveness and robustness of our approach. Jinghong Zheng 0002, Changlong Jiang, Yang Xiao 0007, Jiaqi Li 0007, Haohong Kuang, Ran Wang 0005, Zhiguo Cao 0001, Joey Tianyi Zhou |
NeurIPS | 3 |
| 2025 | Multilingual-Prompt-Guided Directional Feature Learning for Weakly Supervised Video Anomaly DetectionabstractWeakly supervised video anomaly detection has gained attention for its effective performance and cost-efficient annotation, using video-level labels to distinguish between normal and abnormal patterns. However, challenges arise from the diversity and incompleteness of anomalous events, complicating feature learning. Vision-language models offer promising approaches, but designing precise prompts remains difficult. This is because accommodating the diverse range of normal and anomalous scenarios in real-world settings is challenging, and the workload is significant. To tackle these issues, we propose integrating multilingualism and multiple prompts to improve feature learning. By utilizing prompts in various languages to define "anomaly" and "normalcy," we tackle these concepts across different linguistic domains. In each domain, multiple prompts are employed for adaptive top-K prompt selection of snippets. To enhance visual feature learning, a multi-granularity attention module combining Transformer and Mamba is designed. Mamba's long-range adaptation selection builds fine-grained temporal correlations among coarse-grained snippets, while Transformer enhances fine-grained information guided by coarse-grained information. Alongside a multilingual prompt guidance loss, we introduce a gradual directional loss to jointly optimize visual feature distribution and the top-K prompt selection. Our method demonstrates effectiveness on four video datasets and provides generalizability analyses on two medical datasets, including EMG and ECG temporal data. Chizhuo Xiao, Yang Xiao 0007, Joey Tianyi Zhou, Zhiwen Fang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | One-Shot Cross-Domain Instance Detection With Universal Representation
Chen Feng 0002, Jian Cheng 0001, Yang Xiao 0007, Zhiguo Cao 0001 |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2025 | Toward Micro Lip Reading in the Wild: Dataset, Theory, and PracticesabstractMicro lip reading, is characterized by tiny lip movements when someone is speaking. It is highly preferred in robotics and automation, such as a social robot for geriatric care, a patrol robot system for hospitals, and speech system for hearing impaired individuals, etc. However, it has not been well studied before. In this paper, we shed the light to this research topic. A labelled micro lip reading dataset (i.e. HUST-LMLR) of 399 video samples is first established. The samples are captured from the unconstrained movies. One key challenge for micro lip reading in the wild lies on extracting fine features of lip movements effectively and robustly. In this paper, we pay the research efforts to address this issue from two aspects: facial context attention and feature extraction. First, we propose a multi-task learning model for micro lip reading in the wild for the first time. It concerns micro lip reading and facial landmark detection jointly for capturing global face context with soft attention, to facilitate micro lip reading. Second, we propose that motion feature should be concerned and combined with appearance feature jointly to characterize tiny lip movements effectively. Finally, the experiments on HUST-LMLR demonstrate the challenges of our dataset and our proposed approach results in a remarkable improvement of the state-of-the-art by over 26% WER in HUST-LMLR. Furthermore, results in a well-known public dataset LRS2 also show the generalization and superiority of our approach. If accepted, we will publish our HUST-LMLR dataset and relative supporting matierials athttps://hust-dpkw.github.io/Micro-lip-reading/.Note to Practitioners—This paper aims to construct a challenging public dataset and propose a novel vision-based approach for automatically micro lip reading under the unconstrained conditions. A labelled dataset HUST-LMLR was first established to reveal the “micro” characteristics, with 399 video samples captured from 40 unconstrained movies or documentaries. In addition, two key manners in perspective of algorithm are proposed to address the problem of micro lip reading at sentence level in the wild: multi-task learning with facial landmark detection, feature extraction with appearance and motion combined. The extensive experiments demonstrate the challenges of our dataset HUST-LMLR as well as the superiority of our approach in micro lip reading tasks. Nevertheless, it is still somewhat sensitive to the extremely tiny variation of lip movements and the dramatic variation on human posture. It is worth noting that our proposition can be applied not only for medical support, smart health care but also for human-computer interaction, special education, information security, assistant driving, and virtual reality etc. Ran Wang 0005, Yingsong Cheng, Shengtao Zhong, Yanjiu He, Chutian Sun, Qinlao Zhao, Yang Xiao 0007 |
IEEE Trans Autom. Sci. Eng. | 11 |
| 2025 | Rhythmer: Ranking-Based Skill Assessment With Rhythm-Aware TransformerabstractRanking-based skill assessment is an essential component of video understanding. In this task lacking precise procedure annotations, existing methods place greater emphasis on evaluating the procedure quality via manually normalizing the execution duration. However, the inherent duration-related procedural patterns will undergo alteration. Experimentally, we discover that distinct duration biases are prevalent in duration-sensitive skills, such as those in medical and everyday life. Hence, duration information is crucial for ranking-based skill assessment when dealing with varying durations. Additionally, similar execution processes tend to have closer execution durations. Thus, another critical factor lies in extracting duration-related procedural information alongside similar durations. It is defined as mining rhythm patterns, which are inspired by music rhythms including various duration and duration-related procedures. In our work, a rhythm-aware transformer is proposed to mine the rhythm patterns adaptively. Given pairwise inputs, a co-attention module is designed to mutually highlight duration-related procedure information when comparing pairwise input videos with similar durations, and adaptively attenuate the efficacy when confronted with pairwise inputs featuring significantly different durations. A rhythm-encoding module further embeds duration information into the concatenation of raw features and co-attention features. Following these features, the transformer decoder is designed to learn duration-related queries supervised by a novel duration grouping loss among various duration groups. The experimental results demonstrate that the rhythm-aware transformer is effective for ranking-based skill assessment. Zhuang Luo, Yang Xiao 0007, Feng Yang 0012, Joey Tianyi Zhou, Zhiwen Fang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | CODNet: Infrared Small-Target Detection by Mitigating the Curse of DimensionalityabstractInfrared small-target detection (ISTD) presents significant challenges due to the blurred edges and low signal-to-noise ratio (SNR) of the targets, which often results in the loss of critical target features during convolution and downsampling. While researchers have explored complex architectures to preserve these features, such designs often come at the cost of increased computational cost. Our analysis reveals that the curse of dimensionality (COD) is a fundamental issue in ISTD. Due to the sparse information content of infrared small targets, they struggle to effectively utilize high-dimensional feature spaces, leading to a significant feature loss phenomenon. To address this, we propose the CODNet, which integrates spatial–channel shuffle (SCS) attention, dynamic gated channel attention (DGCA), and DCT-based frequency attention (DFA) in the UNet framework. SCS enhances feature representation through spatial attention and channel shuffling, making the feature space more compact and informative. DGCA performs selective channel compression to improve the feature SNR while reducing redundancy and sparsity. DFA incorporates frequency-domain features to supplement spatial representations, alleviating feature sparsity and improving information utilization. Experimental results demonstrate that the CODNet achieves state-of-the-art performance on four public datasets while maintaining relatively high operation efficiency. Shuaiyuan Du, Chen Feng 0002, Yang Xiao 0007, Zhiguo Cao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Language-Guided 3-D Action Feature Learning Without Ground-Truth Sample Class LabelabstractThis work pays the first research effort to leverage point cloud sequence-based Self-supervised 3-D Action Feature Learning (S3AFL), under text's cross-modality weak supervision. We intend to fill the huge performance gap between point cloud sequence and 3-D skeleton-based manners. The key intuition derives from the observation that skeleton-based manners actually hold the human pose's high-level knowledge that leads to attention on the body's joint-aware local parts. Inspired by this, we propose to introduce the text's weak supervision of high-level semantics into a point cloud sequence-based paradigm. With RGB-point cloud pair sequence acquired via RGB-D camera, text sequence is first generated from RGB component using pretrained image captioning model, as auxiliary weak supervision. Then, S3AFL runs in a cross and intra-modality contrastive learning (CL) way. To resist text's missing and redundant semantics, feature learning is conducted in a multistage way with semantic refinement. Essentially, text is only required for training. To facilitate the feature's representation power on fine-grained actions, a multirank max-pooling (MR-MP) way is also proposed for the point set network to better maintain discriminative clues. Experiments verify that the text's weak supervision can facilitate performance by 10.8%, 10.4%, and 8.0% on NTU RGB+D 60, 120, and N-UCLA at most. The performance gap between point cloud sequence and skeleton-based manners has been remarkably narrowed down. The idea of transferring text's weak supervision to S3AFL can also be applied to a skeleton manner, with strong generality. The source code is available at https://github.com/tangent-T/W3AMT. Yang Xiao 0007, Xingyu Tong 0002, Tingbing Yan, Zhiguo Cao 0001, Joey Tianyi Zhou |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Towards Dual Transparent Liquid Level Estimation in Biomedical Lab: Dataset, Methods and Practices
Xiayu Wang, Ke Ma 0012, Ruiyun Zhong, Xinggang Wang, Yi Fang 0006, Yang Xiao 0007 |
ECCV (65) | 6 |
| 2024 | CrossGLG: LLM Guides One-Shot Skeleton-Based 3D Action Recognition in a Cross-Level Manner
Tingbing Yan, Wenzheng Zeng, Yang Xiao 0007, Xingyu Tong 0002, Zhiwen Fang, Zhiguo Cao 0001, Joey Tianyi Zhou |
ECCV (20) | 3 |
| 2024 | C2Net: content-dependent and -independent cross-attention network for anomaly detection in videos
Jiafei Liang, Yang Xiao 0007, Joey Tianyi Zhou, Feng Yang 0012, Zhiwen Fang |
Appl. Intell. | 2 |
| 2024 | Late better than early: A decision-level information fusion approach for RGB-Thermal crowd counting with illumination awareness
Jian Cheng 0001, Chen Feng 0002, Yang Xiao 0007, Zhiguo Cao 0001 |
Neurocomputing | 3 |
| 2024 | Advance One-Shot Multispectral Instance Detection With Text's SupervisionabstractOne key issue within one-shot multispectral instance detection (OMID) is to extract features of strong instance discriminative power, domain adaptation capability, and instance-wise generality. Existing methods generally only rely on visual clues. Comparatively, text is advantageous due to its structured information, high semantics, and low noise. Inspired by recent emergence of large image-text datasets and breakthrough visual-language models, we propose to advance OMID with text's supervision for the first time. To this end, our key idea is to establish the relationship between one-shot multispectral instance with ImageNet class labels via the CLIP model. Particularly, we retrieve, rank, and ensemble the text features of ImageNet labels via instance image feature as query. Then the resulting instance image and text features are realigned and fused to obtain a multimodal feature. Meanwhile, a multispectral contrastive learning approach is proposed to drive multimodal feature learning for OMID. Note that all the procedures are end-to-end trained in a unified network. In this way, the instance discriminative power and domain adaptation capability are facilitated simultaneously. Experiments on two tailored multispectral instance detection datasets verify the effectiveness of our method. Chen Feng 0002, Jian Cheng 0001, Yang Xiao 0007, Zhiguo Cao 0001 |
IEEE Signal Process. Lett. | 3 |
| 2024 | You Will Never Walk Alone: One-Shot 3D Action Recognition With Point Cloud SequenceabstractIn this work, we pay the first effort to address one-shot 3D action recognition in point cloud sequence, without skeleton information. The main contribution lies in two folders. First, a novel one-shot classification approach that considers the feature distribution of 3D action is proposed. We find that, for different 3D actions their dimensional-wise feature distributions are generally in Gaussian form and similar action categories hold approximate feature distributions. Accordingly, K-nearest base classes’ mean value and covariance matrix information help to form one-shot novel class’s pseudo feature distribution. To alleviate the potential ambiguous problem within nearest neighbor search, we divide the base classes into subsets via C-means clustering to facilitate the similarity measure to novel class. Meanwhile, the feature distribution of base class’s whole set and subsets will be jointly considered for generating novel class’s pseudo feature distribution. Multi-dimensional Gaussian sampling is conducted on the acquired pseudo feature distribution for feature-level data augmentation, to make one-shot novel class “never walk alone” for leveraging classifier training. Secondly to better characterize fine-grained 3D action, a temporal attention method is proposed, via introducing vision Transformer (ViT) to capture action’s discriminative short-term motion pattern with densely sampled short-term 3DV (3D dynamic voxel) features along temporal dimension. Experiments on NTU RGB+D 120 and 60 verify superiority of our approach. It outperforms state-of-the-art skeleton-based methods by 13.9% at most. The source code is available athttps://github.com/Tong-XY/YNWA. Xingyu Tong 0002, Yang Xiao 0007, Jianyu Yang 0002, Zhiguo Cao 0001, Joey Tianyi Zhou, Junsong Yuan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | EGST: An Efficient Solution for Human Gaits Recognition Using Neuromorphic Vision SensorabstractTraditional cameras struggle to perform in challenging scenarios such as low latency, high speed and high dynamic range. In contrast, neuromorphic vision sensors (event cameras) have great potential for robotics and computer vision due to the advantages of high temporal resolution, high dynamic range, and ultra-low resource consumption. Event cameras are novel bio-inspired sensors that monitor the brightness change of each pixel asynchronously and provide a stream of events encoding the time, position and sign of the brightness changes. Hence, traditional computer vision methods cannot be directly applied to the event-stream. Finding event representations that completely maintain event attributes, as well as efficient and accurate learning approaches, is the key to unlocking the potential of event cameras. In this study, we reveal the rigid transfer from event-stream to graph that has been overlooked in previous work and introduce a novel event representation, namely event graph sequence (EGS) considering the local and global temporal clues. Coupled with EGS, we propose a spatio-temporal pattern extracting (STPE) module to capture the spatio-temporal correlation and evolution of EGS. Our novel framework, Event Graph Sequence Transformer (EGST), exploits event properties to provide efficient and accurate recognition. This study focuses on the event-based human gaits recognition task, and EGST is evaluated on three different event-based gait datasets. The evaluation results show better or comparable accuracy than the state-of-the-art, while requiring extremely low computation resources. The code will be available athttps://github.com/C19h/EGST. Liaogehao Chen, Zhenjun Zhang, Yang Xiao 0007, Yaonan Wang 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2024 | TaiChiNet: Negative-Positive Cross-Attention Network for Breast Lesion Segmentation in Ultrasound ImagesabstractBreast lesion segmentation in ultrasound images is essential for computer-aided breast-cancer diagnosis. To improve the segmentation performance, most approaches design sophisticated deep-learning models by mining the patterns of foreground lesions and normal backgrounds simultaneously or by unilaterally enhancing foreground lesions via various focal losses. However, the potential of normal backgrounds is underutilized, which could reduce false positives by compacting the feature representation of all normal backgrounds. From a novel viewpoint of bilateral enhancement, we propose a negative-positive cross-attention network to concentrate on normal backgrounds and foreground lesions, respectively. Derived from the complementing opposites of bipolarity in TaiChi, the network is denoted as TaiChiNet, which consists of the negative normal-background and positive foreground-lesion paths. To transmit the information across the two paths, a cross-attention module, a complementary MLP-head, and a complementary loss are built for deep-layer features, shallow-layer features, and mutual-learning supervision, separately. To the best of our knowledge, this is the first work to formulate breast lesion segmentation as a mutual supervision task from the foreground-lesion and normal-background views. Experimental results have demonstrated the effectiveness of TaiChiNet on two breast lesion segmentation datasets with a lightweight architecture. Furthermore, extensive experiments on the thyroid nodule segmentation and retinal optic cup/disc segmentation datasets indicate the application potential of TaiChiNet. Jinting Wang, Jiafei Liang, Yang Xiao 0007, Joey Tianyi Zhou, Zhiwen Fang, Feng Yang 0012 |
IEEE J. Biomed. Health Informatics | 3 |
| 2024 | Lightweight LiDAR-Camera Alignment With Homogeneous Local-Global Aware RepresentationabstractIn this paper, a novel LiDAR-Camera Alignment (LCA) method using homogeneous local-global spatial aware representation is proposed. Compared with the state-of-the-art methods (e.g., LCCNet), our proposition holds 2 main superiorities. First, homogeneous multi-modality representation learned with a uniform CNN model is applied along the iterative prediction stages, instead of the state-of-the-art heterogeneous counterparts extracted from the separated modality-wise CNN models within each stage. In this way, the model size can be significantly decreased (e.g., 12.39M (ours) vs. 333.75M (LCCNet)). Meanwhile, within our proposition the interaction between LiDAR and camera data is built during feature learning to better exploit the descriptive clues, which has not been well concerned by the existing approaches. Secondly, we propose to equip the learned LCA representation with local-global spatial aware capacity via encoding CNN’s local convolutional features with Transformer’s non-local self-attention manner. Accordingly, the local fine details and global spatial context can be jointly captured by the encoded local features. And, they will be jointly used for LCA. On the other hand, the existing methods generally choose to reveal the global spatial property via intuitively concatenating the local features. Additionally at the initial LCA stage, LiDAR is roughly aligned with camera by our pre-alignment method, according to the point distribution characteristics of its 2D projection version with the initial extrinsic parameters. Although its structure is simple, it can essentially alleviate LCA’s difficulty for the consequent stages. To better optimize LCA, a novel loss function that builds the correlation between translation and rotation loss items is also proposed. The experiments on KITTI data verifies the superiority of our proposition both on effectiveness and efficiency. The source code will be released athttps://github.com/Zaf233/Light-weight-LCAupon acceptance. Angfan Zhu, Yang Xiao 0007, Mingkui Tan, Zhiguo Cao 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2024 | Beyond Pattern Variance: Unsupervised 3-D Action Representation Learning With Point Cloud SequenceabstractThis work pays the first research effort to address unsupervised 3-D action representation learning with point cloud sequence, which is different from existing unsupervised methods that rely on 3-D skeleton information. Our proposition is built on the state-of-the-art 3-D action descriptor 3-D dynamic voxel (3DV) with contrastive learning (CL). The 3DV can compress the point cloud sequence into a compact point cloud of 3-D motion information. Spatiotemporal data augmentations are conducted on it to drive CL. However, we find that existing CL methods (e.g., SimCLR or MoCo v2) often suffer from high pattern variance toward the augmented 3DV samples from the same action instance, that is, the augmented 3DV samples are still of high feature complementarity after CL, while the complementary discriminative clues within them have not been well exploited yet. To address this, a feature augmentation adapted CL (FACL) approach is proposed, which facilitates 3-D action representation via concerning the features from all augmented 3DV samples jointly, in spirit of feature augmentation. FACL runs in a global-local way: one branch learns global feature that involves the discriminative clues from the raw and augmented 3DV samples, and the other focuses on enhancing the discriminative power of local feature learned from each augmented 3DV sample. The global and local features are fused to characterize 3-D action jointly via concatenation. To fit FACL, a series of spatiotemporal data augmentation approaches is also studied on 3DV. Wide-range experiments verify the superiority of our unsupervised learning method for 3-D action feature learning. It outperforms the state-of-the-art skeleton-based counterparts by 6.4% and 3.6% with the cross-setup and cross-subject test settings on NTU RGB+D 120, respectively. The source code is available at https://github.com/tangent-T/FACL. Yang Xiao 0007, Yancheng Wang 0002, Jianyu Yang 0002, Zhiguo Cao 0001, Joey Tianyi Zhou, Junsong Yuan 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | GREnet: Gradually REcurrent Network With Curriculum Learning for 2-D Medical Image SegmentationabstractMedical image segmentation is a vital stage in medical image analysis. Numerous deep-learning methods are booming to improve the performance of 2-D medical image segmentation, owing to the fast growth of the convolutional neural network. Generally, the manually defined ground truth is utilized directly to supervise models in the training phase. However, direct supervision of the ground truth often results in ambiguity and distractors as complex challenges appear simultaneously. To alleviate this issue, we propose a gradually recurrent network with curriculum learning, which is supervised by gradual information of the ground truth. The whole model is composed of two independent networks. One is the segmentation network denoted as GREnet, which formulates 2-D medical image segmentation as a temporal task supervised by pixel-level gradual curricula in the training phase. The other is a curriculum-mining network. To a certain degree, the curriculum-mining network provides curricula with an increasing difficulty in the ground truth of the training set by progressively uncovering hard-to-segmentation pixels via a data-driven manner. Given that segmentation is a pixel-level dense-prediction challenge, to the best of our knowledge, this is the first work to function 2-D medical image segmentation as a temporal task with pixel-level curriculum learning. In GREnet, the naive UNet is adopted as the backbone, while ConvLSTM is used to establish the temporal link between gradual curricula. In the curriculum-mining network, UNet++ supplemented by transformer is designed to deliver curricula through the outputs of the modified UNet++ at different layers. Experimental results have demonstrated the effectiveness of GREnet on seven datasets, i.e., three lesion segmentation datasets in dermoscopic images, an optic disc and cup segmentation dataset and a blood vessel segmentation dataset in retinal images, a breast lesion segmentation dataset in ultrasound images, and a lung segmentation dataset in computed tomography (CT). Jinting Wang, Yujiao Tang, Yang Xiao 0007, Joey Tianyi Zhou, Zhiwen Fang, Feng Yang 0012 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | A2J-Transformer: Anchor-to-Joint Transformer Network for 3D Interacting Hand Pose Estimation from a Single RGB Imageabstract3D interacting hand pose estimation from a single RGB image is a challenging task, due to serious self-occlusion and inter-occlusion towards hands, confusing similar appearance patterns between 2 hands, ill-posed joint position mapping from 2D to 3D, etc.. To address these, we propose to extend A2J-the state-of-the-art depth-based 3D single hand pose estimation method-to RGB domain under interacting hand condition. Our key idea is to equip A2J with strong local-global aware ability to well capture interacting hands' local fine details and global articulated clues among joints jointly. To this end, A2J is evolved under Transformer's non-local encoding-decoding framework to build A2J- Transformer. It holds 3 main advantages over A2J. First, self-attention across local anchor points is built to make them global spatial context aware to better capture joints' articulation clues for resisting occlusion. Secondly, each anchor point is regarded as learnable query with adaptive feature learning for facilitating pattern fitting capacity, instead of having the same local representation with the others. Last but not least, anchor point locates in 3D space instead of 2D as in A2J, to leverage 3D pose prediction. Experiments on challenging InterHand 2.6M demonstrate that, A2J-Transformer can achieve state-of-the-art model-free performance (3.38mm MPJPE advancement in 2-hand case) and can also be applied to depth domain with strong generalization. The code is avaliable at https://github.com/ChanglongJiangGit/A2J-Transformer. Changlong Jiang, Yang Xiao 0007, Cunlin Wu, Jinghong Zheng 0002, Zhiguo Cao 0001, Joey Tianyi Zhou |
CVPR | 2 |
| 2023 | Real-time Multi-person Eyeblink Detection in the Wild for Untrimmed VideoabstractReal-time eyeblink detection in the wild can widely serve for fatigue detection, face anti-spoofing, emotion analysis, etc. The existing research efforts generally focus on single-person cases towards trimmed video. However, multi-person scenario within untrimmed videos is also important for practical applications, which has not been well concerned yet. To address this, we shed light on this research field for the first time with essential contributions on dataset, theory, and practices. In particular, a large-scale dataset termed MPEblink that involves 686 untrimmed videos with 8748 eyeblink events is proposed under multi-person conditions. The samples are captured from uncon-strainedfilms to reveal “in the wild“ characteristics. Meanwhile, a real-time multi-person eyeblink detection method is also proposed. Being different from the existing counter-parts, our proposition runs in a one-stage spatio-temporal way with end-to-end learning capacity. Specifically, it simultaneously addresses the sub-tasks of face detection, face tracking, and human instance-level eyeblink detection. This paradigm holds 2 main advantages: (1) eyeblink features can be facilitated via the face's global context (e.g., head pose and illumination condition) with joint optimization and interaction, and (2) addressing these sub-tasks in parallel instead of sequential manner can save time remarkably to meet the real-time running requirement. Experiments on MPEblink verify the essential challenges of real-time multi-person eyeblink detection in the wild for untrimmed video. Our method also outperforms existing approaches by large margins and with a high inference speed. Wenzheng Zeng, Yang Xiao 0007, Sicheng Wei, Jinfang Gan, Xintao Zhang, Zhiguo Cao 0001, Zhiwen Fang, Joey Tianyi Zhou |
CVPR | 2 |
| 2023 | Multi-spectral template matching based object detection in a few-shot learning manner
Chen Feng 0002, Zhiguo Cao 0001, Yang Xiao 0007, Zhiwen Fang, Joey Tianyi Zhou |
Inf. Sci. | 3 |
| 2023 | Learning dynamic relationship between joints for 3D hand pose estimation from single depth map
Huiqin Xing, Jianyu Yang 0002, Yang Xiao 0007 |
J. Vis. Commun. Image Represent. | 3 |
| 2023 | Learning full context feature for human motion prediction
Huiqin Xing, Yicong Zhou, Jianyu Yang 0002, Yang Xiao 0007 |
J. Vis. Commun. Image Represent. | 4 |
| 2023 | End-to-End Video Gaze Estimation via Capturing Head-Face-Eye Spatial-Temporal Interaction ContextabstractIn this letter, we propose a new method, Multi-Clue Gaze (MCGaze), to facilitate video gaze estimation via capturing spatial-temporal interaction context among head, face, and eye in an end-to-end learning way, which has not been well concerned yet. The main advantage of MCGaze is that the tasks of clue localization of head, face, and eye can be solved jointly for gaze estimation in a one-step way, with joint optimization to seek optimal performance. During this, spatial-temporal context exchange happens among the clues on the head, face, and eye. Accordingly, the final gazes obtained by fusing features from various queries can be aware of global clues from heads and faces, and local clues from eyes simultaneously, which essentially leverages performance. Meanwhile, the one-step running way also ensures high running efficiency. Experiments on the challenging Gaze360 dataset verify the superiority of our proposition. The source code will be released athttps://github.com/zgchen33/MCGaze. Yiran Guan, Zhuoguang Chen, Wenzheng Zeng, Zhiguo Cao 0001, Yang Xiao 0007 |
IEEE Signal Process. Lett. | 5 |
| 2023 | Robust LiDAR-Camera Alignment With Modality Adapted Local-to-Global RepresentationabstractLiDAR-Camera alignment (LCA) is an important preprocessing procedure for fusing LiDAR and camera data. For it, one key issue is to extract unified cross-modality representation for characterizing the heterogeneous LiDAR and camera data effectively and robustly. The main challenge is to resist the modality gap and visual data degradation during feature learning, while still maintaining strong representative power. To address this, a novel modality adapted local-to-global representation learning method is proposed. The research efforts are paid in 2 main folders via modality adaptation and capturing global spatial context. First for modality gap resistance, LiDAR and camera data is projected into the same depth map domain for unified representation learning. Particularly, LiDAR data is converted to depth map according to pre-acquired extrinsic parameters. Thanks to the recent advantage of deep learning based monocular depth estimation, camera data is transformed into depth map in data driven manner, which is jointly optimized with LCA. Secondly to capture global spatial context, ViT (vision transformer) is introduced to LCA. The concept of LCA token is proposed for aggregating the local spatial patterns to form global spatial representation with transformer encoding. And, it is shared by all the samples. In this way, it can involve global sample-level information to leverage generalization ability. The experiments on KITTI dataset verify superiority of our proposition. Furthermore, the proposed approach is more robust to camera data degeneration (e.g., imaging blurring and noise) often faced by the practical applications. Under some challenging test cases, the performance advancement of our method is over$1.9~cm$/4.1° on translation / rotation error. While our model size (8.77M) is much smaller than existing methods (e.g., LCCNet of 66.75M). The source code will be released athttps://github.com/Zaf233/RLCAupon acceptance. Angfan Zhu, Yang Xiao 0007, Zhiguo Cao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Eyelid's Intrinsic Motion-Aware Feature Learning for Real-Time Eyeblink Detection in the WildabstractReal-time eyeblink detection in the wild is a recently emerged challenging task that suffers from dramatic variations in face attribute, pose, illumination, camera view and distance, etc. One key issue is to well characterize eyelid’s intrinsic motion (i.e., approaching and departure between upper and lower eyelid) robustly, under unconstrained conditions. Towards this, a novel eyelid’s intrinsic motion-aware feature learning approach is proposed. Our proposition lies in 3 folds. First, the feature extractor is led to focus on informative eye region adaptively via introducing visual attention in a coarse-to-fine way, to guarantee robustness and fine-grained descriptive ability jointly. Then, 2 constraints are proposed to make feature learning be aware of eyelid’s intrinsic motion. Particularly, one concerns the fact that the inter-frame feature divergence within eyeblink processes should be greater than non-eyeblink ones to better reveal eyelid’s intrinsic motion. The other constraint minimizes the inter-frame feature divergence of non-eyeblink samples, to suppress motion clues due to head or camera movement, illumination change, etc. Meanwhile, concerning the high ambiguity between eyeblink and non-eyeblink samples, soft sample labels are acquired via self-knowledge distillation to conduct feature learning with finer supervision than the hard ones. The experiments verify that, our proposition is significantly superior to the state-of-the-art ones (i.e., advantage on F1-score over 7%) and with real-time running efficiency. It is also of strong generalization capacity towards constrained conditions. The source code is available athttps://github.com/wenzhengzeng/blink_eyelid. Wenzheng Zeng, Yang Xiao 0007, Guilei Hu, Zhiguo Cao 0001, Sicheng Wei, Zhiwen Fang, Joey Tianyi Zhou, Junsong Yuan 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2022 | PIZZA: A Powerful Image-only Zero-Shot Zero-CAD Approach to 6 DoF TrackingabstractEstimating the relative pose of a new object without prior knowledge is a hard problem, while it is an ability very much needed in robotics and Augmented Reality. We present a method for tracking the 6D motion of objects in RGB video sequences when neither the training images nor the 3D geometry of the objects are available. In contrast to previous works, our method can therefore consider unknown objects in open world instantly, without requiring any prior information or a specific training phase. We consider two architectures, one based on two frames, and the other relying on a Transformer Encoder, which can exploit an arbitrary number of past frames. We train our architectures using only synthetic renderings with domain randomization. Our results on challenging datasets are on par with previous works that require much more information (training images of the target objects, 3D models, and/or depth data). Our source code is available at https://github.com/nv-nguyen/pizza. Van Nguyen Nguyen, Yuming Du, Yang Xiao 0007, Michaël Ramamonjisoa, Vincent Lepetit |
3DV | 3 |
| 2022 | C3P: Cross-Domain Pose Prior Propagation for Weakly Supervised 3D Human Pose Estimation
Cunlin Wu, Yang Xiao 0007, Boshen Zhang, Zhiguo Cao 0001, Joey Tianyi Zhou |
ECCV (5) | 2 |
| 2022 | MAT: Multianchor Visual Tracking With Selective Search RegionabstractThe core prerequisite of most modern trackers is a motion assumption, defined as predicting the current location in a limited search region centering at the previous prediction. For clarity, the central subregion of a search region is denoted as the tracking anchor (e.g., the location of the previous prediction in the current frame). However, providing accurate predictions in all frames is very challenging in the complex nature scenes. In addition, the target locations in consecutive frames often change violently under the attribute of fast motion. Both facts are likely to lead the previous prediction to an unbelievable tracking anchor, which will make the aforementioned prerequisite invalid and cause tracking drift. To enhance the reliability of tracking anchors, we propose a real-time multianchor visual tracking mechanism, called multianchor tracking (MAT). Instead of directly relying on the tracking anchor inherited from the previous prediction, MAT selects the best anchor from an anchor ensemble, which includes several objectness-based anchor proposals and the anchor inherited from the previous prediction. The objectness-based anchors provide several complementary selective search regions, and an entropy-minimization-based selection method is introduced to find the best anchor. Our approach offers two benefits: 1) selective search regions can increase the chance of tracking success with affordable computational load and 2) anchor selection introduces the best anchor for each frame, which breaks the limitation of solo depending on the previous prediction. The extensive experiments of nine base trackers upgraded by MAT on four challenging datasets demonstrate the effectiveness of MAT. Zhiwen Fang, Zhiguo Cao 0001, Yang Xiao 0007, Kaicheng Gong, Junsong Yuan 0001 |
IEEE Trans. Cybern. | 3 |
| 2022 | ECML: An Ensemble Cascade Metric-Learning Mechanism Toward Face VerificationabstractFace verification can be regarded as a two-class fine-grained visual-recognition problem. Enhancing the feature's discriminative power is one of the key problems to improve its performance. Metric-learning technology is often applied to address this need while achieving a good tradeoff between underfitting, and overfitting plays a vital role in metric learning. Hence, we propose a novel ensemble cascade metric-learning (ECML) mechanism. In particular, hierarchical metric learning is executed in a cascade way to alleviate underfitting. Meanwhile, at each learning level, the features are split into nonoverlapping groups. Then, metric learning is executed among the feature groups in the ensemble manner to resist overfitting. Considering the feature distribution characteristics of faces, a robust Mahalanobis metric-learning method (RMML) with a closed-form solution is additionally proposed. It can avoid the computation failure issue on an inverse matrix faced by some well-known metric-learning approaches (e.g., KISSME). Embedding RMML into the proposed ECML mechanism, our metric-learning paradigm (EC-RMML) can run in the one-pass learning manner. The experimental results demonstrate that EC-RMML is superior to state-of-the-art metric-learning methods for face verification. The proposed ECML mechanism is also applicable to other metric-learning approaches. Fu Xiong, Yang Xiao 0007, Zhiguo Cao 0001, Yancheng Wang 0002, Joey Tianyi Zhou, Jianxin Wu 0001 |
IEEE Trans. Cybern. | 2 |
| 2022 | Person Re-Identification With Hierarchical Discriminative Spatial AggregationabstractPractically, person re-identification (re-ID) may suffer from the critical spatial misalignment problem due to inaccurate human detection, variation on human pose and camera viewpoint, etc. To address this, a hierarchical discriminative spatial aggregation method is proposed. The key idea is to conduct spatial aggregation on local human parts via global average-pooling to acquire the strong spatial misalignment tolerance, with VALD encoding on the local parts for facilitating discriminative power jointly. This proposition is built on NetVLAD to ensure end-to-end deep learning capacity. Due to the fine-grained property of person re-ID task that has not been well concerned by the original NetVLAD model for scene recognition, a feature refinement layer that consists of 1 fully-connected (FC) layer and 2 batch normalization (BN) layers is added on top of the raw NetVLAD layer to enhance the discriminative power and training convergence. And, a human body occlusion and background component dropout manner is also proposed to resist the effect of serious occlusion. Technically, a refined codeword initialization manner is proposed to alleviate the potential codeword imbalance problem caused by naive random initialization. The proposed discriminative spatial aggregation approach is then conducted on multi-resolution convolutional feature map layers hierarchically via early feature fusion, to involve richer semantic and fine-grained visual clues jointly. Wide-range experiments on 6 datasets (i.e., CUHK03, DukeMTMC-reID, Occluded-DukeMTMC, Market-1501, MSMT17 and Occluded-REID) verifies the effectiveness of our proposition. The source code and supporting material is available athttps://github.com/zmyme/HDSA-reID. Yang Xiao 0007, Fu Xiong, Zhiguo Cao 0001, Zhiwen Fang, Joey Tianyi Zhou |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2022 | Vision-Based Finger Tapping Test in Patients With Parkinson's Disease via Spatial-Temporal 3D Hand Pose EstimationabstractFinger tapping test is crucial for diagnosing Parkinson's Disease (PD), but manual visual evaluations can result in score discrepancy due to clinicians' subjectivity. Moreover, applying wearable sensors requires making physical contact and may hinder PD patient's raw movement patterns. Accordingly, a novel computer-vision approach is proposed using depth camera and spatial-temporal 3D hand pose estimation to capture and evaluate PD patients' 3D hand movement. Within this approach, a temporal encoding module is leveraged to extend A2J's deep learning framework to counter the pose jittering problem, and a pose refinement process is utilized to alleviate dependency on massive data. Additionally, the first vision-based 3D PD hand dataset of 112 hand samples from 48 PD patients and 11 control subjects is constructed, fully annotated by qualified physicians under clinical settings. Testing on this real-world data, this new model achieves 81.2% classification accuracy, even surpassing that of individual clinicians in comparison, fully demonstrating this proposition's effectiveness. The demo video can be accessed at https://github.com/ZhilinGuo/ST-A2J. Zhilin Guo 0001, Weiqi Zeng, Taidong Yu, Yang Xiao 0007, Xuebing Cao, Zhiguo Cao 0001 |
IEEE J. Biomed. Health Informatics | 5 |
| 2022 | Anomaly Detection With Bidirectional Consistency in VideosabstractThe core component of most anomaly detectors is a self-supervised model, tasked with modeling patterns included in training samples and detecting unexpected patterns as the anomalies in testing samples. To cope with normal patterns, this model is typically trained with reconstruction constraints. However, the model has the risk of overfitting to training samples and being sensitive to hard normal patterns in the inference phase, which results in irregular responses at normal frames. To address this problem, we formulate anomaly detection as a mutual supervision problem. Due to collaborative training, the complementary information of mutual learning can alleviate the aforementioned problem. Based on this motivation, a SIamese generative network (SIGnet), including two subnetworks with the same architecture, is proposed to simultaneously model the patterns of the forward and backward frames. During training, in addition to traditional constraints on improving the reconstruction performance, a bidirectional consistency loss based on the forward and backward views is designed as the regularization term to improve the generalization ability of the model. Moreover, we introduce a consistency-based evaluation criterion to achieve stable scores at the normal frames, which will benefit detecting anomalies with fluctuant scores in the inference phase. The results on several challenging benchmark data sets demonstrate the effectiveness of our proposed method. Zhiwen Fang, Jiafei Liang, Joey Tianyi Zhou, Yang Xiao 0007, Feng Yang 0012 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2022 | Discriminative Multi-View Dynamic Image Fusion for Cross-View 3-D Action RecognitionabstractDramatic imaging viewpoint variation is the critical challenge toward action recognition for depth video. To address this, one feasible way is to enhance view-tolerance of visual feature, while still maintaining strong discriminative capacity. Multi-view dynamic image (MVDI) is the most recently proposed 3-D action representation manner that is able to compactly encode human motion information and 3-D visual clue well. However, it is still view-sensitive. To leverage its performance, a discriminative MVDI fusion method is proposed by us via multi-instance learning (MIL). Specifically, the dynamic images (DIs) from different observation viewpoints are regarded as the instances for 3-D action characterization. After being encoded using Fisher vector (FV), they are then aggregated by sum-pooling to yield the representative 3-D action signature. Our insight is that viewpoint aggregation helps to enhance view-tolerance. And, FV can map the raw DI feature to the higher dimensional feature space to promote the discriminative power. Meanwhile, a discriminative viewpoint instance discovery method is also proposed to discard the viewpoint instances unfavorable for action characterization. The wide-range experiments on five data sets demonstrate that our proposition can significantly enhance the performance of cross-view 3-D action recognition. And, it is also applicable to cross-view 3-D object recognition. The source code is available at https://github.com/3huo/ActionView. Yancheng Wang 0002, Yang Xiao 0007, Zhiguo Cao 0001, Zhenjun Zhang, Joey Tianyi Zhou |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2021 | Learning to Better Segment Objects from Unseen Classes with Unlabeled VideosabstractThe ability to localize and segment objects from unseen classes would open the door to new applications, such as autonomous object learning in active vision. Nonetheless, improving the performance on unseen classes requires additional training data, while manually annotating the objects of the unseen classes can be labor-extensive and expensive. In this paper, we explore the use of unlabeled video sequences to automatically generate training data for objects of unseen classes. It is in principle possible to apply existing video segmentation methods to unlabeled videos and automatically obtain object masks, which can then be used as a training set even for classes with no manual labels available. However, our experiments show that these methods do not perform well enough for this purpose. We therefore introduce a Bayesian method that is specifically designed to automatically create such a training set: Our method starts from a set of object proposals and relies on (non-realistic) analysis-by-synthesis to select the correct ones by performing an efficient optimization over all the frames simultaneously. Through extensive experiments, we show that our method can generate a high-quality training set which significantly boosts the performance of segmenting objects of unseen classes. We thus believe that our method could open the door for open-world instance segmentation by exploiting abundant Internet videos. Yuming Du, Yang Xiao 0007, Vincent Lepetit |
ICCV | 2 |
| 2021 | LPQ++: A discriminative blur-insensitive textural descriptor with spatial-channel interaction
Yang Xiao 0007, Zhiguo Cao 0001, Zhiwen Fang, Joey Tianyi Zhou |
Inf. Sci. | 2 |
| 2021 | Multi-Encoder Towards Effective Anomaly Detection in VideosabstractGiven normal training samples, anomaly detection in videos can be regarded as a challenging problem of identifying unexpected events. The state-of-the-art approaches generally resort to the autoencoder model by using a single encoder to capture the motion and content patterns jointly. Nevertheless, due to the lack of accurate labels of normal and abnormal samples, how to detect anomalies is decided by the subjective understanding of models. It infers that different models will prefer to mine different patterns according to the characteristics of models. We call this problem as a pattern bias problem. To alleviate this problem, a novel Multi-Encoder Single-Decoder network, termed as MESDnet, is proposed in the spirit of encoding motion and content cues individually with multiple encoders. MESDnet is of end-to-end learning ability and real-time running speed. Particularly, the differences between adjacent frames and the raw frames are used as the motion and content sources, respectively. Then, a decoder takes charge of detecting anomalies in the way of observing reconstructing error towards the video frames by using the multi-stream encoded motion and content features simultaneously. The experiments on the CUHK Avenue dataset, the UCSD Pedestrian dataset, and the ShanghaiTech Campus dataset verify the effectiveness of MESDnet. Zhiwen Fang, Joey Tianyi Zhou, Yang Xiao 0007, Yanan Li 0006, Feng Yang 0012 |
IEEE Trans. Multim. | 3 |
| 2021 | Abrupt-motion-aware lightweight visual tracking for unmanned aerial vehicles
Kaicheng Gong, Zhiguo Cao 0001, Yang Xiao 0007, Zhiwen Fang |
Vis. Comput. | 3 |
| 2021 | Survey on depth and RGB image-based 3D hand shape and pose estimationabstractThe field of vision-based human hand three-dimensional (3D) shape and pose estimation has attracted significant attention recently owing to its key role in various applications, such as natural humancomputer interactions. With the availability of large-scale annotated hand datasets and the rapid developments of deep neural networks (DNNs), numerous DNN-based data-driven methods have been proposed for accurate and rapid hand shape and pose estimation. Nonetheless, the existence of complicated hand articulation, depth and scale ambiguities, occlusions, and finger similarity remain challenging. In this study, we present a comprehensive survey of state-of-the-art 3D hand shape and pose estimation approaches using RGB-D cameras. Related RGB-D cameras, hand datasets, and a performance analysis are also discussed to provide a holistic view of recent achievements. We also discuss the research potential of this rapidly growing field. Lin Huang 0004, Boshen Zhang, Zhilin Guo 0001, Yang Xiao 0007, Zhiguo Cao 0001, Junsong Yuan 0001 |
Virtual Real. Intell. Hardw. | 4 |
| 2020 | P2B: Point-to-Box Network for 3D Object Tracking in Point CloudsabstractTowards 3D object tracking in point clouds, a novel point-to-box network termed P2B is proposed in an end-to-end learning manner. Our main idea is to first localize potential target centers in 3D search area embedded with target information. Then point-driven 3D target proposal and verification are executed jointly. In this way, the time-consuming 3D exhaustive search can be avoided. Specifically, we first sample seeds from the point clouds in template and search area respectively. Then, we execute permutation-invariant feature augmentation to embed target clues from template into search area seeds and represent them with target-specific features. Consequently, the augmented search area seeds regress the potential target centers via Hough voting. The centers are further strengthened with seed-wise targetness scores. Finally, each center clusters its neighbors to leverage the ensemble power for joint 3D target proposal and verification. We apply PointNet++ as our backbone and experiments on KITTI tracking dataset demonstrate P2B's superiority (~10%'s improvement over state-of-the-art). Note that P2B can run with 40FPS on a single NVIDIA 1080Ti GPU. Our code and model are available at https://github.com/HaozheQi/P2B. Haozhe Qi, Chen Feng 0002, Zhiguo Cao 0001, Yang Xiao 0007 |
CVPR | 5 |
| 2020 | 3DV: 3D Dynamic Voxel for Action Recognition in Depth VideoabstractFor depth-based 3D action recognition, one essential issue is to represent 3D motion pattern effectively and efficiently. To this end, 3D dynamic voxel (3DV) is proposed as a novel 3D motion representation manner. With 3D space voxelization, the key idea of 3DV is to encode the 3D motion information within depth video into a regular voxel set (i.e., 3DV) compactly, via temporal rank pooling. Each available 3DV voxel intrinsically involves 3D spatial and motion feature for 3D action description. 3DV is then abstracted as a point set and input into PointNet++ for 3D action recognition, in the end-to-end learning way. The intuition for transferring 3DV into the point set form is that, PointNet++ is lightweight and effective for deep feature learning towards point set. Since 3DV may loose appearance clue, a multi-stream 3D action recognition manner is also proposed to learn motion and appearance feature jointly. To extract richer temporal order information of actions, we also split the depth video into temporal segments and encode this procedure in 3DV integrally. The extensive experiments on the well-established benchmark datasets (e.g., NTU RGB+D 120 and NTU RGB+D 60) demonstrate the superiority of our proposition. Impressively, we acquire the accuracy of 82.4% and 93.5% on NTU RGB+D 120 with the cross-subject and cross-setup test setting respectively. 3DV's code is available at https://github.com/3huo/3DV-Action. Yancheng Wang 0002, Yang Xiao 0007, Fu Xiong, Wenxiang Jiang 0001, Zhiguo Cao 0001, Joey Tianyi Zhou, Junsong Yuan 0001 |
CVPR | 2 |
| 2020 | Measuring Generalisation to Unseen Viewpoints, Articulations, Shapes and Objects for 3D Hand Pose Estimation Under Hand-Object Interaction
Anil Armagan, Guillermo Garcia-Hernando, Seungryul Baek, Shreyas Hampali, Mahdi Rad, Shipeng Xie, Mingxiu Chen, Boshen Zhang, Fu Xiong, Yang Xiao 0007, Zhiguo Cao 0001, Junsong Yuan 0001, Pengfei Ren 0001, Weiting Huang, Haifeng Sun 0001, Marek Hrúz, Jakub Kanis, Zdenek Krnoul, Qingfu Wan, Shile Li, Linlin Yang 0001, Dongheui Lee, Angela Yao, Weiguo Zhou, Sijia Mei, Adrian Spurr, Umar Iqbal 0001, Pavlo Molchanov 0001, Philippe Weinzaepfel, Romain Brégier, Grégory Rogez, Vincent Lepetit, Tae-Kyun Kim 0001 |
ECCV (23) | 11 |
| 2020 | Exploiting Distilled Learning for Deep Siamese TrackingabstractExisting deep siamese trackers are typically built on off-the-shelf CNN models for feature learning, with the demand for huge power consumption and memory storage. This limits current deep siamese trackers to be carried on resource-constrained devices like mobile phones, given factor that such a deployment normally requires cost-effective considerations. In this work, we address this issue by presenting a novel Distilled Learning Framework(DLF) for siamese tracking, which aims at learning tracking model with efficiency and high accuracy. Specifically, we propose two simple yet effective knowledge distillation strategies, denote as point-wise distillation and pairwise distillation, which are designed for transferring knowledge from a more discriminative teacher tracker into a compact student tracker. In this way, cost-effective and high performance tracking could be achieved. Extensive experiments on several tracking benchmarks demonstrate the effectiveness of our proposed method. Zhiguo Cao 0001, Wei Li 0132, Yang Xiao 0007, Shuaiyuan Du, Angfan Zhu |
ICPR | 4 |
| 2020 | Attention-Driven Loss for Anomaly Detection in Video SurveillanceabstractRecent video anomaly detection methods focus on reconstructing or predicting frames. Under this umbrella, the long-standing inter-class data-imbalance problem resorts to the imbalance between foreground and stationary background objects in video anomaly detection and this has been less investigated by existing solutions. Naively optimizing the reconstructing loss yields a biased optimization towards background reconstruction rather than the objects of interest in the foreground. To solve this, we proposed a simple yet effective solution, termed attention-driven loss to alleviate the foreground-background imbalance problem in anomaly detection. Specifically, we compute a single mask map that summarizes the frame evolution of moving foreground regions and suppresses the background in the training video clips. After that, we construct an attention map through the combination of the mask map and background to give different weights to the foreground and background region respectively. The proposed attention-driven loss is independent of backbone networks and can be easily augmented in most existing anomaly detection models. Augmented with attention-driven loss, the model is able to achieve AUC 86.0% on Avenue, 83.9% on Ped1, 96% on Ped2 datasets. Extensive experimental results and ablation studies further validate the effectiveness of our model. Joey Tianyi Zhou, Le Zhang 0001, Zhiwen Fang, Jiawei Du 0002, Xi Peng 0001, Yang Xiao 0007 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2020 | Towards Real-Time Eyeblink Detection in the Wild: Dataset, Theory and PracticesabstractEffective and real-time eyeblink detection is of wide-range applications, such as deception detection, drive fatigue detection, face anti-spoofing. Despite previous efforts, most of existing focus on addressing the eyeblink detection problem under constrained indoor conditions with relative consistent subject and environment setup. Nevertheless, towards practical applications, eyeblink detection in the wild is highly preferred, and of greater challenges. In this paper, we shed the light to this research topic. A labelled eyeblink in the wild dataset (i.e., HUST-LEBW) of 673 eyeblink video samples (i.e., 381 positives, and 292 negatives) is first established. These samples are captured from the unconstrained movies, with the dramatic variation on face attribute, head pose, illumination condition, imaging configuration, etc. Then, we formulate eyeblink detection task as a binary spatial-temporal pattern recognition problem. After locating and tracking human eyes using SeetaFace engine and KCF (Kernelized Correlation Filters) tracker respectively, a modified LSTM model able to capture the multi-scale temporal information is proposed to verify eyeblink. A feature extraction approach that reveals the appearance and motion characteristics simultaneously is also proposed. The experiments on HUST-LEBW reveal the superiority and efficiency of our approach. The comparisons with the existing state-of-the-art methods validate the advantages of our manner for eyeblink detection in the wild. Guilei Hu, Yang Xiao 0007, Zhiguo Cao 0001, Lubin Meng, Zhiwen Fang, Joey Tianyi Zhou, Junsong Yuan 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2020 | RoSeq: Robust Sequence LabelingabstractIn this paper, we mainly investigate two issues for sequence labeling, namely, label imbalance and noisy data that are commonly seen in the scenario of named entity recognition (NER) and are largely ignored in the existing works. To address these two issues, a new method termed robust sequence labeling (RoSeq) is proposed. Specifically, to handle the label imbalance issue, we first incorporate label statistics in a novel conditional random field (CRF) loss. In addition, we design an additional loss to reduce the weights of overwhelming easy tokens for augmenting the CRF loss. To address the noisy training data, we adopt an adversarial training strategy to improve model generalization. In experiments, the proposed RoSeq achieves the state-of-the-art performances on CoNLL and English Twitter NER-88.07% on CoNLL-2002 Dutch, 87.33% on CoNLL-2002 Spanish, 52.94% on WNUT-2016 Twitter, and 43.03% on WNUT-2017 Twitter without using the additional data. Joey Tianyi Zhou, Hao Zhang 0048, Di Jin 0005, Xi Peng 0001, Yang Xiao 0007, Zhiguo Cao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2019 | A2J: Anchor-to-Joint Regression Network for 3D Articulated Pose Estimation From a Single Depth ImageabstractFor 3D hand and body pose estimation task in depth image, a novel anchor-based approach termed Anchor-to-Joint regression network (A2J) with the end-to-end learning ability is proposed. Within A2J, anchor points able to capture global-local spatial context information are densely set on depth image as local regressors for the joints. They contribute to predict the positions of the joints in ensemble way to enhance generalization ability. The proposed 3D articulated pose estimation paradigm is different from the state-of-the-art encoder-decoder based FCN, 3D CNN and point-set based manners. To discover informative anchor points towards certain joint, anchor proposal procedure is also proposed for A2J. Meanwhile 2D CNN (i.e., ResNet- 50) is used as backbone network to drive A2J, without using time-consuming 3D convolutional or deconvolutional layers. The experiments on 3 hand datasets and 2 body datasets verify A2J's superiority. Meanwhile, A2J is of high running speed around 100 FPS on single NVIDIA 1080Ti GPU. Fu Xiong, Boshen Zhang, Yang Xiao 0007, Zhiguo Cao 0001, Taidong Yu, Joey Tianyi Zhou, Junsong Yuan 0001 |
ICCV | 3 |
| 2019 | Action recognition for depth video using multi-view dynamic images
Yang Xiao 0007, Jun Chen 0001, Yancheng Wang 0002, Zhiguo Cao 0001, Joey Tianyi Zhou, Xiang Bai |
Inf. Sci. | 1 |
| 2019 | Ranking 3D feature correspondences via consistency voting
Jiaqi Yang 0002, Yang Xiao 0007, Zhiguo Cao 0001, Weidong Yang 0006 |
Pattern Recognit. Lett. | 2 |
| 2019 | Real-Time Detection of Fall From Bed Using a Single Depth CameraabstractToward the medical and living healthcare for the elderly and patients, fall from bed is a critical accident that may lead to serious injuries. To alleviate this, an essential problem is to detect this event in time for earning the rescue time. Although some efforts that resort to the wearable devices and smart healthcare room have already been paid to address this problem, the performance is still not satisfactory enough for the practical applications. In this paper, a novel fall from a bed detection method is proposed. In particular, the depth camera is used as the visual sensor due to its insensitivity to illumination variation and capacity of privacy protection. To characterize the human activity well, an effective human upper body detection approach able to extract human head and upper body center is proposed using random forest. Compared with the existing widely used human body parsing methods (e.g., Microsoft Kinect SDK or OpenNI SDK), our proposition can still work reliably when human-bed interaction happens. According to the motion information of human upper body, the fall from bed detection task is formulated as a two-class classification problem. Then, it is solved using the large margin nearest neighbor classification approach. Our method can meet the real-time running requirement with the normal computer. In experiments, we construct a fall from bed detection data set that contains the samples from 42 volunteers (26 males and 16 females) for test. The experimental results demonstrate the effectiveness and efficiency of our proposition. Zhiguo Cao 0001, Yang Xiao 0007, Jing Mao, Junsong Yuan 0001 |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2019 | Aligning 2.5D Scene Fragments With Distinctive Local Geometric Features and Voting-Based CorrespondencesabstractAligning 2.5D views has been extensively explored in the past decades, where most prior works have concentrated on object data with complex structures. This paper presents a method to align real-word scene scans with challenging features such as noise, poor geometric information, and highly repeatable patterns. Our method consists of two modules: pairwise and multiview alignments. Key to the proposed pairwise alignment method is the rotational contour signature geometric feature and voting-based correspondence selection algorithm. The former promises strong discriminative power for 2.5D scene data, while the latter affords high-quality correspondences via a voting process for all raw feature matches using L2distance and point pair affinity constraints. For the multiview alignment method, we first use a connected graph algorithm to establish the connections of all 2.5D views for coarse merging; then, we propose a shape-growing iterative closest point algorithm for further refinement. Experiments are conducted on scene point cloud datasets addressing both the indoor and outdoor scenarios, whereby we demonstrate that the proposed pairwise alignment method clearly outperforms the state of the art. Moreover, the proposed multiview alignment method manages to put multiple unordered 2.5D scene fragments into a unified coordinate system automatically, accurately, and efficiently. Jiaqi Yang 0002, Yang Xiao 0007, Zhiguo Cao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2019 | Supervised Fine-Grained Cloud Detection and Recognition in Whole-Sky ImagesabstractThe whole-sky imager has been increasingly used for ground-based cloud automatic observation. Many approaches based on image processing have been applied to detect or classify clouds in whole-sky images (WSIs). However, most of the studies only focus on image segmentation for cloud detection or image classification for cloud recognition separately. The cloud detection only does the binary segmentation (sky and cloud) without cloud types, while the cloud recognition only gives the single image-level label without cloud coverage. In this paper, a fine-grained cloud detection and recognition task with a solution is proposed to fill the gap, which can simultaneously detect and classify clouds in a WSI. It can be regarded as a pixel-level fine-grained dense prediction for images. First, a new data set is built with pixel-level annotation of nine different types. Then, a solution based on supervised learning is proposed, in which the pixel-level prediction problem is converted to a superpixel classification problem. Multiview features are extracted, including color, inside texture, neighbor texture, and global relation, to represent the superpixels. Moreover, a class-specific feature space transformation method based on metric learning and subspace alignment is proposed to overcome the challenge brought by the high similarity among cloud types and the feature shifting. Finally, several experiments have verified that our approach is effective to the challenging new task and also outperforms some other methods in the normal tasks of cloud detection and cloud classification, respectively. Zhiguo Cao 0001, Yang Xiao 0007, Zhibiao Yang |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2019 | Detecting Alzheimer's Disease on Small Dataset: A Knowledge Transfer PerspectiveabstractComputer-aided diagnosis (CAD) is an attractive topic in Alzheimer's disease (AD) research. Many algorithms are based on a relatively large training dataset. However, small hospitals are usually unable to collect sufficient training samples for robust classification. Although data sharing is expanding in scientific research, it is unclear whether a model based on one dataset is well suited for other data sources. Using a small dataset from a local hospital and a large shared dataset from the AD neuroimaging initiative, we conducted a heterogeneity analysis and found that different functional magnetic resonance imaging data sources show different sample distributions in feature space. In addition, we proposed an effective knowledge transfer method to diminish the disparity among different datasets and improve the classification accuracy on datasets with insufficient training samples. The accuracy increased by approximately 20% compared with that of a model based only on the original small dataset. The results demonstrated that the proposed approach is a novel and effective method for CAD in hospitals with only small training datasets. It solved the challenge of limited sample size in detection of AD, which is a common issue but lack of adequate attention. Furthermore, this paper sheds new light on effective use of multi-source data for neurological disease diagnosis. Wei Li 0086, Xi Chen 0002, Yang Xiao 0007, Yuanyuan Qin |
IEEE J. Biomed. Health Informatics | 4 |
| 2018 | Monocular Relative Depth Perception With Web Stereo Data SupervisionabstractIn this paper we study the problem of monocular relative depth perception in the wild. We introduce a simple yet effective method to automatically generate dense relative depth annotations from web stereo images, and propose a new dataset that consists of diverse images as well as corresponding dense relative depth maps. Further, an improved ranking loss is introduced to deal with imbalanced ordinal relations, enforcing the network to focus on a set of hard pairs. Experimental results demonstrate that our proposed approach not only achieves state-of-the-art accuracy of relative depth perception in the wild, but also benefits other dense per-pixel prediction tasks, e.g., metric depth estimation and semantic segmentation. Ke Xian, Chunhua Shen, Zhiguo Cao 0001, Hao Lu 0003, Yang Xiao 0007, Ruibo Li, Zhenbo Luo |
CVPR | 5 |
| 2018 | Counting Fish in Sonar ImagesabstractThe goal of this paper is to estimate the population of fishes in sonar images. Compared to natural images, sonar images present substantially different visual characteristics. Fishes in sonar images exhibit unreliable appearance cues, expose under imaging noise, vary significantly in shape and size. These pose great challenges for counting even for a human expert. In Computer Vision, a possible solution to this task is object counting with deep networks. This paradigm is typically formulated as a regression problem. The regression, however, greatly suffers from the issue of sample imbalance caused by fish variations in size and density, leading to underestimates in high-density regions and over-estimates in low-density regions. To address this, we build upon a recent local counts regression network and propose two novel losses to regularize a modified l1 loss with slack constraints. In particular, a challenging sonar fish counting dataset with 537 images and manually labeled dotted annotations is constructed. Experimental results on the dataset justify the effectiveness of our proposition and show improved performance of our method over other state-of-the-art approaches. Liang Liu 0001, Hao Lu 0003, Zhiguo Cao 0001, Yang Xiao 0007 |
ICIP | 4 |
| 2018 | Scalable Multi-Consistency Feature Matching with Non-Cooperative GamesabstractCorrespondence selection aiming at seeking correct relationships between two images is a fundamental and critical task in computer vision. This paper attempts to select consistent correspondences in the context of dynamic scenarios where multiple matching consistencies are normally incorporated. To this end, we present a grid-based game-theoretic matching (Grid-GTM) method which is divided into three processes, i.e., grid matching, local games and enrichment. Specifically, grid matching translates the multi-consistency problem into several independent single-consistency problems to decrease difficulties of selection and boost the efficiency. Local games extended under the guidance of a novel payoff function guarantee that mismatches are effectively removed. Enrichment is added to recover correct matches neglected by local games. Crucially, our approach achieves the state-of-the-art performance compared with seven algorithms in comprehensive evaluations. In addition, we construct a dataset that involves multiple consistencies under three different scenes in this paper. Chen Zhao 0025, Jiaqi Yang 0002, Yang Xiao 0007, Zhiguo Cao 0001 |
ICIP | 3 |
| 2018 | RGB-D Co-Segmentation on Indoor Scene with Geometric Prior and Hypothesis Filtering
Lingxiao Hang, Zhiguo Cao 0001, Yang Xiao 0007, Hao Lu 0003 |
PRCV (1) | 3 |
| 2018 | The Accurate Guidance for Image Caption Generation
Xinyuan Qi, Zhiguo Cao 0001, Yang Xiao 0007 |
PRCV (3) | 3 |
| 2018 | Toward Good Practices for Fine-Grained Maize Cultivar Identification With Filter-Specific Convolutional ActivationsabstractCrop cultivar identification is an important aspect in agricultural systems. Traditional solutions involve excessive human interventions, which is labor-intensive and timeconsuming. In addition, cultivar identification is a typical task of fine-grained visual categorization (FGVC). Compared with other common topics in FGVC, studies of this problem are somewhat lagging and limited. In this paper, targeting four Chinese maize cultivars of Jundan No.20, Wuyue No.3, Nongda No.108, and Zhengdan No.958, we first consider the problem of identifying the maize cultivar based on its tassel characteristics by computer vision. In particular, a novel fine-grained maize cultivar identification data set termed HUST-FG-MCI that contains 5000 images is first constructed. To better capture the textual differences in a weakly supervised manner, we proposed an effective deep convolutional neural network and Fisher vector (FV)based feature encoding mechanism. The mechanism tends to highlight subtle object patterns via filter-specific convolutional representations and thus provides strong discrimination for cultivar identification. Experimental results demonstrate that our method outperforms other state-of-the-art approaches. We show also that FV encoding can weaken the linear dependency between convolutional activations, redundant filters exist in the convolutional layer, and high accuracy can be maintained with relatively low-dimensional convolutional features and one or two Gaussian components in FV. Hao Lu 0003, Zhiguo Cao 0001, Yang Xiao 0007, Zhiwen Fang, Yanjun Zhu |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2018 | An Embarrassingly Simple Approach to Visual Domain AdaptationabstractWe show that it is possible to achieve high-quality domain adaptation without explicit adaptation. The nature of the classification problem means that when samples from the same class in different domains are sufficiently close, and samples from differing classes are separated by large enough margins, there is a high probability that each will be classified correctly. Inspired by this, we propose an embarrassingly simple yet effective approach to domain adaptation-only the class mean is used to learn class-specific linear projections. Learning these projections is naturally cast into a linear-discriminant-analysis-like framework, which gives an efficient, closed form solution. Furthermore, to enable to application of this approach to unsupervised learning, an iterative validation strategy is developed to infer target labels. Extensive experiments on cross-domain visual recognition demonstrate that, even with the simplest formulation, our approach outperforms existing non-deep adaptation methods and exhibits classification performance comparable with that of modern deep adaptation methods. An analysis of potential issues effecting the practical application of the method is also described, including robustness, convergence, and the impact of small sample sizes. Hao Lu 0003, Chunhua Shen, Zhiguo Cao 0001, Yang Xiao 0007, Anton van den Hengel |
IEEE Trans. Image Process. | 4 |
| 2018 | Toward the Repeatability and Robustness of the Local Reference Frame for 3D Shape Matching: An EvaluationabstractThe local reference frame (LRF), as an independent coordinate system constructed on the local 3D surface, is broadly employed in 3D local feature descriptors. The benefits of the LRF include rotational invariance and full 3D spatial information, thereby greatly boosting the distinctiveness of a 3D feature descriptor. There are numerous LRF methods in the literature; however, no comprehensive study comparing their repeatability and robustness performance under different application scenarios and nuisances has been conducted. This paper evaluates eight state-of-the-art LRF proposals on six benchmarks with different data modalities (e.g., LiDAR, Kinect, and Space Time) and application contexts (e.g., shape retrieval, 3D registration, and 3D object recognition). In addition, the robustness of each LRF to a variety of nuisances, including varying support radii, Gaussian noise, outliers (shot noise), mesh resolution variation, distance to boundary, keypoint localization error, clutter, occlusion, and partial overlap, is assessed. The experimental study also measures the performance under different keypoint detectors, descriptor matching performance when using different LRFs and feature representation combinations, as well as computational efficiency. Considering the evaluation outcomes, we summarize the traits, advantages, and current limitations of the tested LRF methods. Jiaqi Yang 0002, Yang Xiao 0007, Zhiguo Cao 0001 |
IEEE Trans. Image Process. | 2 |
| 2017 | Performance Evaluation of 3D Correspondence Grouping AlgorithmsabstractThis paper presents a thorough evaluation of several widely-used 3D correspondence grouping algorithms, motived by their significance in vision tasks relying on correct feature correspondences. A good correspondence grouping algorithm is desired to retrieve as many as inliers from initial feature matches, giving a rise in both precision and recall. Towards this rule, we deploy the experiments on three benchmarks respectively addressing shape retrieval, 3D object recognition and point cloud registration scenarios. The variety in application context brings a rich category of nuisances including noise, varying point densities, clutter, occlusion and partial overlaps. It also results to different ratios of inliers and correspondence distributions for comprehensive evaluation. Based on the quantitative outcomes, we give a summarization of the merits/demerits of the evaluated algorithms from both performance and efficiency perspectives. Jiaqi Yang 0002, Ke Xian, Yang Xiao 0007, Zhiguo Cao 0001 |
3DV | 3 |
| 2017 | Rotational contour signatures for both real-valued and binary feature representations of 3D local shape
Jiaqi Yang 0002, Qian Zhang 0046, Ke Xian, Yang Xiao 0007, Zhiguo Cao 0001 |
Comput. Vis. Image Underst. | 4 |
| 2017 | Local phase quantization plus: A principled method for embedding local phase quantization into Fisher vector for blurred image recognition
Yang Xiao 0007, Zhiguo Cao 0001 |
Inf. Sci. | 1 |
| 2017 | Two-dimensional subspace alignment for convolutional activations adaptation
Hao Lu 0003, Zhiguo Cao 0001, Yang Xiao 0007, Yanjun Zhu |
Pattern Recognit. | 3 |
| 2017 | TOLDI: An effective and robust approach for 3D local shape description
Jiaqi Yang 0002, Qian Zhang 0046, Yang Xiao 0007, Zhiguo Cao 0001 |
Pattern Recognit. | 3 |
| 2017 | DeepCloud: Ground-Based Cloud Image Categorization Using Deep Convolutional FeaturesabstractAccurate ground-based cloud image categorization is a critical but challenging task that has not been well addressed. One of the essential issues that affect the performance is to extract the representative visual features. Nearly all of the existing methods rely on the hand-crafted descriptors (e.g., local binary patterns, CENsus TRsansform hISTogram, and scale-invariant feature transform). Their limited discriminative power indeed leads to the unsatisfactory performance. To alleviate this, we propose “DeepCloud” as a novel cloud image feature extraction approach by resorting to the deep convolutional visual features. In the recent years, the deep convolutional neural network (CNN) has achieved the promising results in lots of computer vision and image understanding fields. Nevertheless, it has not been applied to cloud image classification yet. Thus, we actually pay the first effort to fill this blank. Since cloud image classification can be attributed to a multi-instance learning problem, simply employing the convolutional features within CNN cannot achieve the promising result. To address this, Fisher vector encoding is applied to executing the spatial feature aggregation and high-dimensional feature mapping on the raw deep convolutional features. Moreover, the hierarchical convolutional layers are used simultaneously to capture the fine textural characteristics and high-level semantic information in the unified manner. To further leverage the performance, a cloud pattern mining and selection method are also proposed. It targets at finding the discriminative local patterns to better distinguish the different kinds of clouds. The experiments on a challenging ground-based cloud image data set demonstrate the superiority of the proposition over the state-of-the-art methods. Zhiguo Cao 0001, Yang Xiao 0007 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2016 | Fine-grained maize cultivar identification using filter-specific convolutional activationsabstractCultivar identification is an important aspect in agriculture and also a typical task of fine-grained visual categorization (FGVC). In comparison with other common topics in FGVC, studies on this problem are somewhat lagged and limited. In this paper, targeting four Chinese maize cultivars of Jundan No.20, Wuyue No.3, Nongda No.108, and Zhengdan No.958, we first consider the problem of identifying the maize cultivar based on its tassel characteristics. Technically, an effective convolutional neural network (CNN) based feature encoding pipeline that allows integration of deep CNN based column feature extraction, filter-specific Fisher vector (FV) encoding and mutual information (MI) based filter selection is proposed to better address this problem. In particular, a novel fine-grained maize cultivar identification dataset termed MCI-4000 that contains 4000 images is first constructed by our team. Experimental results demonstrate that our method outperforms other stat-of-the-art approaches by at least 5% in accuracy. We also show that, there exists redundant filters in the last convolutional layer, and high accuracy can be achieved with only relatively low-dimensional column features and a small number of Gaussian components in FV. Hao Lu 0003, Zhiguo Cao 0001, Yang Xiao 0007, Zhiwen Fang, Yanjun Zhu |
ICIP | 3 |
| 2016 | Rotational contour signatures for robust local surface descriptionabstractThis paper presents a novel local surface descriptor called rotational contour signatures (RCS) for 3D rigid objects. RCS comprises several signatures that characterize the 2D contour information derived from 3D-to-2D projection of the local surface. The inspiration of our encoding technique comes from that, viewing towards an object, its contour is an effective and robust cue for representing its shape. In order to achieve a comprehensive geometry encoding, the local surface is continually rotated in a predefined local reference frame (LRF) so that multi-view information is obtained. Experiments on two publicly available datasets demonstrate the effectiveness and robustness of the proposed descriptor. Further, comparisons with five state-of-the-art descriptors show the superiority of our RCS descriptor. Jiaqi Yang 0002, Qian Zhang 0046, Ke Xian, Yang Xiao 0007, Zhiguo Cao 0001 |
ICIP | 4 |
| 2016 | Efficient Airport Detection Using Line Segment Detector and Fisher Vector RepresentationabstractIn this letter, a two-stage method for airport detection on remote sensing images is proposed. In the first stage, a new algorithm composed of several line-based processing steps is used for extraction of candidate airport regions. In the second stage, the scale-invariant feature transformation and Fisher vector coding are used for efficient representation of the airport and nonairport regions and support vector machines employed for classification. In order to evaluate the performance of the proposed method, extensive experiments are conducted on airports around the world with different layouts. The measures used in the evaluation are accuracy, sensitivity, and specificity. The proposed method achieved an accuracy of 94.6%, which was benchmarked with two previous methods to prove its superiority. Ümit Budak, Ugur Halici, Abdulkadir Sengür, Murat Karabatak, Yang Xiao 0007 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2016 | Exploiting Attribute Dependency for Attribute Assignment in Crowded ScenesabstractAttributes now play a vital role for characterizing a crowded scene. Compared to low-level visual features, processing informed by attributes can capture rich semantic information. However, to effectively assign attributes to a crowded scene still remains a challenging task. In this letter, inspired by a recently proposed zero-shot learning framework, a novel attribute assignment method that maps low-level features to predefined attributes is proposed. In particular, we propose to exploit the attribute dependency during the phase of attribute assignment, which can be regarded as our main contribution. In addition, to further enhance the performance, an effective low-level feature extraction mechanism is also proposed. More precisely, appearance and motion features are first simultaneously extracted from several sampled video frames and corresponding optical flow fields via deep convolutional neural network and then, respectively, aggregated by using Fisher vector encoding to form the low-level representation of crowded scenes. Experimental results on the challenging WWW dataset demonstrate that both the proposed attribute assignment method and the low-level feature extraction mechanism outperform the state of the art. Chunhua Deng, Zhiguo Cao 0001, Yang Xiao 0007, Hao Lu 0003, Ke Xian |
IEEE Signal Process. Lett. | 3 |
| 2016 | Adobe Boxes: Locating Object Proposals Using Object AdobesabstractDespite the previous efforts of object proposals, the detection rates of the existing approaches are still not satisfactory enough. To address this, we propose Adobe Boxes to efficiently locate the potential objects with fewer proposals, in terms of searching the object adobes that are the salient object parts easy to be perceived. Because of the visual difference between the object and its surroundings, an object adobe obtained from the local region has a high probability to be a part of an object, which is capable of depicting the locative information of the proto-object. Our approach comprises of three main procedures. First, the coarse object proposals are acquired by employing randomly sampled windows. Then, based on local-contrast analysis, the object adobes are identified within the enlarged bounding boxes that correspond to the coarse proposals. The final object proposals are obtained by converging the bounding boxes to tightly surround the object adobes. Meanwhile, our object adobes can also refine the detection rate of most state-of-the-art methods as a refinement approach. The extensive experiments on four challenging datasets (PASCAL VOC2007, VOC2010, VOC2012, and ILSVRC2014) demonstrate that the detection rate of our approach generally outperforms the state-of-the-art methods, especially with relatively small number of proposals. The average time consumed on one image is about 48 ms, which nearly meets the real-time requirement. Zhiwen Fang, Zhiguo Cao 0001, Yang Xiao 0007, Lei Zhu 0010, Junsong Yuan 0001 |
IEEE Trans. Image Process. | 3 |
| 2015 | Blurred image recognition using domain adaptationabstractImage blurring significantly degrades the image recognition performance. In this paper, we novelly address the blurred image recognition task from the perspective of domain adaptation (DA). The scenario is that, the training set (source domain) only comprises of the labelled clear images, and the test set (target domain) is composed of the unlabelled blurred images. DA is executed to eliminate the domain shift by subspace alignment. In this way, the clear and blurred image domains are pushed closer in the feature space. The supervised LMDR metric learning method is employed by us to construct the source domain subspace for further performance enhancement, compared to the unsupervised one (i.e., PCA). The experimental results on two datasets demonstrate that, the proposed DA-based blurred image recognition mechanism can significantly enhance the performance of different kinds of visual descriptors, especially when the blurring degree is strong. Xiaokang Xie, Zhiguo Cao 0001, Yang Xiao 0007, Mengyu Zhu, Hao Lu 0003 |
ICIP | 3 |
| 2015 | Ground-based cloud image categorization using deep convolutional visual featuresabstractGround-based cloud image categorization is an essential and challenging task in automatic sky and cloud observation field. Till now, it still has not been well addressed in both meteorology and image processing communities, due to the large variation of cloud appearance. One feasible way to solve this is to find more discriminative visual representation to characterize the different kinds of clouds. Many efforts have been paid in this way. However, to our knowledge, most of the existing methods only resort to the hand-craft visual descriptors (e.g., LBP, CENTRIST and color histogram). The resulting performance is unfortunately not satisfied enough. Inspired by the great success of deep convolutional neural networks (CNN) in large-scale image classification task (e.g., ImageNet challenge), we first propose to transfer CNN to solve our relative small-scale cloud classification issue. The experiments on two challenging cloud datasets demonstrate that, using the deep convolutional visual features generated by CNN can significantly outperform all the state-of-the-art methods in most cases. Another important contribution of our work is that, we find that applying Fisher Vector (FV) to encoding the off-the-shelf CNN features can further leverage the performance. Zhiguo Cao 0001, Yang Xiao 0007, Wei Li 0086 |
ICIP | 3 |
| 2015 | Beyond local phase quantization: Mid-level blurred image representation using fisher vectorabstractBlurred image recognition is still remaining as a challenging task, while with the wide applications. One principal way for solving this problem is to extract the blur-invariant visual descriptor. To this end, local phase quantization (LPQ) was ever proposed, and achieved promising results. In this paper, to further enhance LPQ's performance, we propose to apply Fisher Vector (FV) encoding approach to acquire the mid-level blurred image representation. To our knowledge, it is the first time that the descriptive power of FV for blurred image recognition has been investigated. Instead of being extracted holistically from the whole image as previously, LPQ is acquired in a densely sampled way. That is, a sliding sub-window will screen the image with certain vertical and horizontal strides. LPQs are then extracted from all the resulting sub-windows respectively. In addition, to maintain local spatial structure information, each sub-window will be divided into finer cells. After being FV encoded, the local LPQs are aggregated using sum-pooling to generate the image signature. The experimental results on three datasets demonstrate that FV can enhance LPQ's performance significantly, and our proposition also outperforms the other blur-invariant descriptors by large margins in most cases. Mengyu Zhu, Zhiguo Cao 0001, Yang Xiao 0007, Xiaokang Xie |
ICIP | 3 |
| 2014 | Entropic image thresholding based on GLGM histogram
Yang Xiao 0007, Zhiguo Cao 0001, Junsong Yuan 0001 |
Pattern Recognit. Lett. | 1 |
| 2014 | mCENTRIST: A Multi-Channel Feature Generation Mechanism for Scene CategorizationabstractmCENTRIST, a new multichannel feature generation mechanism for recognizing scene categories, is proposed in this paper. mCENTRIST explicitly captures the image properties that are encoded jointly by two image channels, which is different from popular multichannel descriptors. In order to avoid the curse of dimensionality, tradeoffs at both feature and channel levels have been executed to make mCENTRIST computationally practical. As a result, mCENTRIST is both efficient and easy to implement. In addition, a hyperopponent color space is proposed by embedding Sobel information into the opponent color space for further performance improvements. Experiments show that mCENTRIST outperforms established multichannel descriptors on four RGB and RGB-near infrared data sets, including aerial orthoimagery, indoor, and outdoor scene category recognition tasks. Experiments also verify that the hyper opponent color space enhances descriptors' performance effectively. Yang Xiao 0007, Jianxin Wu 0001, Junsong Yuan 0001 |
IEEE Trans. Image Process. | 1 |
| 2013 | Human-virtual human interaction by upper body gesture understandingabstractIn this paper, a novel human-virtual human interaction system is proposed. This system supports a real human to communicate with a virtual human using natural body language. Meanwhile, the virtual human is capable of understanding the meaning of human upper body gestures and reacting with its own personality by the means of body action, facial expression and verbal language simultaneously. In total, 11 human upper body gestures with and without human-object interaction are currently involved in the system. They can be characterized by human head, hand and arm posture. In our system implementation, the wearable Immersion CyberGlove II is used to capture the hand posture and the vision-based Microsoft Kinect takes charge of capturing the head and arm posture. This is a new sensor solution for human-gesture capture, and can be regarded as the most important contribution of this paper. Based on the posture data from the CyberGlove II and the Kinect, an effective and real-time human gesture recognition algorithm is also proposed. To verify the effectiveness of the gesture recognition method, we build a human gesture sample dataset. Additionally, the experiments demonstrate that our algorithm can recognize human gestures with high accuracy in real time. Yang Xiao 0007, Junsong Yuan 0001, Daniel Thalmann |
VRST | 1 |
| 2012 | Image classification using HTM cortical learning algorithms
Wen Zhuo, Zhiguo Cao 0001, Yueming Qin, Zhenghong Yu, Yang Xiao 0007 |
ICPR | 5 |
| 2008 | User Behavior Modeling and Traffic Analysis of IMS Presence ServersabstractPresence is a service that allows a user to be informed about the reachability, availability, and willingness of communication of another user. Presence service has become a key enabler for many popular applications such as instant messaging and push-to-talk. Traffic to a presence server is resulted by user behaviors such as login/logout, online status modification, and automatic status refresh by client application. Mathematical models are proposed in this paper to study the user behaviors associated with a presence server during the time of a day such that traffic characteristics to a presence server can be inferred. The correctness of our estimation and analysis are verified by extensive simulations using Matlab. This study on the relationship between user behaviors and traffic to a presence server can help network operators to plan network capacity, optimize server performance and detect traffic anomaly. Zhiguo Cao 0001, Caixia Chi, Ruibing Hao, Yang Xiao 0007 |
GLOBECOM | 4 |
| 2008 | Entropic thresholding based on gray-level spatial correlation histogramabstractIn this paper, an entropic thresholding method based on the gray-level spatial correlation (GLSC) histogram defined by ourselves is presented. Compared with traditional two-dimensional histogram, we take into account the image local property in a different way by GLSC histogram. In experiment, we make comparison of the proposed method with two-dimensional entropic thresholding method proposed by Abutaleb and one-dimensional entropic thresholding method proposed by Kapur. The experiment demonstrates that generally our method could yield equivalent or even better result than Abutaleb’s method while saving time remarkably and perform much better than Kapur’s method without too more time consumption. Yang Xiao 0007, Zhiguo Cao 0001, Tianxu Zhang |
ICPR | 1 |