Li Cheng 0001

dblp:13/4938-1 · DBLP profile ↗
← Back
131ranked-venue papers
17as first author
59since 2021 · last 2026
0000-0003-3261-3533ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 81 · 12 first-author · 35 since 2021Graphics, computer vision, multimedia, augmented reality and games · 70 · 9 first-author · 37 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 2 since 2021Security and privacy · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 VPHO: Joint Visual-Physical Cue Learning and Aggregation for Hand-Object Pose Estimation
abstract
Estimating the 3D poses of hands and objects from a single RGB image is a fundamental yet challenging problem, with broad applications in augmented reality and human-computer interaction. Existing methods largely rely on visual cues alone, often producing results that violate physical constraints such as interpenetration or non-contact. Recent efforts to incorporate physics reasoning typically depend on post-optimization or non-differentiable physics engines, which compromise visual consistency and end-to-end trainability. To overcome these limitations, we propose a novel framework that jointly integrates visual and physical cues for hand-object pose estimation. This integration is achieved through two key ideas: 1) joint visual-physical cue learning: The model is trained to extract 2D visual cues and 3D physical cues, thereby enabling more comprehensive representation learning for hand-object interactions; 2) candidate pose aggregation: A novel refinement process that aggregates multiple diffusion-generated candidate poses by leveraging both visual and physical predictions, yielding a final estimate that is visually consistent and physically plausible. Extensive experiments demonstrate that our method significantly outperforms existing state-of-the-art approaches in both pose accuracy and physical plausibility.
Jun Zhou 0026, Chi Xu 0002, Kaifeng Tang, Yuting Ge, Tingrui Guo, Li Cheng 0001
AAAI6
2026 SAM3-I: Segment Anything with Instructions
abstract
Jingjing Li, Yue Feng, Yuchen Guo, Jincai Huang, Wei Ji, Qi Bi, Yongri Piao, Miao Zhang, Xiaoqi Zhao, Qiang Chen, Shihao Zou, Huchuan Lu, Li Cheng. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Jincai Huang 0003, Wei Ji 0011, Qi Bi, Yongri Piao, Miao Zhang 0004, Xiaoqi Zhao 0003, Qiang Chen 0007, Shihao Zou, Huchuan Lu, Li Cheng 0001
ACL (1)13
2026 SDAE: Similarity-guided dual-masking feature level multi-class anomaly explanation for cluster-based models on data streams
Bin Li 0030, Li Cheng 0001, Jingtao Qi, Haoxiang Jin
Inf. Sci.2
2026 ParkinsonNet: A unified end-to-end framework for estimating Parkinson's disease motor symptom severity
abstract
• ParkinsonNet: We propose a novel end-to-end network, ParkinsonNet, for automatically qualifying the severity of PD motor symptoms. This network is capable of processing multiple motor tests without any manual feature design. • Temporal Self-Attention Enhancement Module (TAEM): To accurately perceive the gradual progression of motor symptoms, we introduce TAEM, which combines temporal compression with long-term dependency modeling, enabling robust and comprehensive temporal feature extraction. • Similarity Matching Module (SMM): To address the challenges of class imbalance and limited datasets, we design the SMM, that transforms the conventional classification or regression task into a similarity matching problem. This module aligns skeleton features with their most similar text-based features, leveraging semantic relationships for improved performance. • Empirical evaluations: Extensive experiments are conducted on two recent PD datasets, demonstrating the superiority of ParkinsonNet, and providing essential benchmark performance for future algorithm development and evaluation. Parkinson’s Disease (PD) is a progressive neurodegenerative disorder characterized by worsening motor symptoms such as bradykinesia, imbalance, tremor, rigidity, and gait disturbances. Clinician assessments are often time-consuming and costly, and the limited availability of specialists, along with patient mobility issues, complicates frequent evaluations. In this paper, we propose a novel end-to-end network to automatically qualify the severity of motor symptoms in PD, referred to as ParkinsonNet. Unlike most existing methods that focus on isolated tests, ParkinsonNet provides a unified learning framework that is evaluated across multiple PD motor symptoms, as demonstrated on finger tapping and gait. Specifically, to accurately perceive the gradual progression of motor symptoms throughout an entire test cycle (e.g., decrementing amplitude), a temporal self-attention enhancement module is designed by combining temporal compression with long-term temporal dependency modeling. To ease the issues of class imbalance and limited datasets, a similarity matching module is proposed that transforms the conventional classification or regression task into a similarity matching problem, matching the skeleton feature with its most similar texture feature. Additionally, a vector quantization module is incorporated to encode spatiotemporal features into a discrete-valued space, compressing and abstracting motion representations while retaining critical information for more accurate classification. Extensive experiments on two newly identified benchmark datasets demonstrate the superiority of our ParkinsonNet and set new benchmark performance for future algorithm development and evaluation. Our code will be released at this https URL .
Yande Li, Fang Ba, Minglun Gong, Li Cheng 0001
Pattern Recognit.4
2025 HiPoser: 3D Human Pose Estimation with Hierarchical Shared Learning at Parts-Level Using Inertial Measurement Units
abstract
This paper considers the challenging problem of 3D Human Pose Estimation (HPE) from a sparse set of Inertial Measurement Units (IMUs). Existing efforts typically reconstruct a pose sequence by either directly tackling whole-body motions or focusing on distinctive spatio-temporal features of local body parts. Unfortunately, these methods ignore existing interdependent motor synergies amongst body parts, which may lead to pose estimation with ambiguous local parts. This observation motivates us to propose a hierarchical learning-based approach, HiPoser, which utilizes a hierarchical shared structure using Mamba blocks as the backbone to focus on the following estimation tasks, involving: 1) torso pose, 2) lower limbs pose, 3) upper limbs pose, and finally 4) global translation. These tasks selectively incorporate body motion states and are to be carried out sequentially in reconstructing part-based poses, which are amalgamated to estimate the final full-body pose with the global translation that satisfies inter-part consistencies. Our hierarchical structure allows HiPoser the flexibility in prioritizing different aspects of pose estimation, to emphasize more on detail or stability. Empirical evaluations over three benchmark datasets demonstrate the superiority of HiPoser over existing state-of-the-art models, suggesting that analyzing the synergistic movement of body parts is indeed important for advancing IMU-based 3D HPE.
Guorui Liao, Chunyuan Zheng 0001, Li Cheng 0001, Shanshan Huang 0004, Jun Liao 0001, Haoxuan Li 0001, Li Liu 0001
AAAI3
2025 CASAGPT: Cuboid Arrangement and Scene Assembly for Interior Design
abstract
We present a novel approach for indoor scene synthesis, which learns to arrange decomposed cuboid primitives to represent 3D objects within a scene. Unlike conventional methods that use bounding boxes to determine the placement and scale of 3D objects, our approach leverages cuboids as a straightforward yet highly effective alternative for modeling objects. This allows for compact scene generation while minimizing object intersections. Our approach, coined CasaGPT for Cuboid Arrangement and Scene Assembly, employs an autoregressive model to sequentially arrange cuboids, producing physically plausible scenes. By applying rejection sampling during the fine-tuning stage to filter out scenes with object collisions, our model further reduces intersections and enhances scene quality. Additionally, we introduce a refined dataset, 3DFRONT-NC, which eliminates significant noise presented in the original dataset, 3D-FRONT. Extensive experiments on the 3D-FRONT dataset as well as our dataset demonstrate that our approach consistently outperforms the state-of-the-art methods, enhancing the realism of generated scenes, and providing a promising direction for 3D scene synthesis. Code is available at https://github.com/CASAGPT/CASA-GPT
Weitao Feng 0001, Hang Zhou 0007, Jing Liao 0001, Li Cheng 0001, Wenbo Zhou 0004
CVPR4
2025 BOOTPLACE: Bootstrapped Object Placement with Detection Transformers
abstract
In this paper, we tackle the copy-paste image-to-image composition problem with a focus on object placement learning. Prior methods have leveraged generative models to reduce the reliance for dense supervision. However, this often limits their capacity to model complex data distributions. Alternatively, transformer networks with a sparse contrastive loss have been explored, but their over-relaxed regularization often leads to imprecise object placement. We introduce BootPlace, a novel paradigm that formulates object placement as a placement-by-detection problem. Our approach begins by identifying suitable regions of interest for object placement. This is achieved by training a specialized detection transformer on object-subtracted backgrounds, enhanced with multi-object supervisions. It then semantically associates each target compositing object with detected regions based on their complementary characteristics. Through a boostrapped training approach applied to randomly object-subtracted images, our model enforces meaningful placements through extensive paired data augmentation. Experimental results on established benchmarks demonstrate BootPlace’s superior performance in object repositioning, markedly surpassing state-of-the-art baselines on Cityscapes and OPA datasets with notable improvements in IOU scores. Additional ablation studies further showcase the compositionality and generalizability of our approach, supported by user study evaluations. Code is available at https://github.com/RyanHangZhou/BootPlace
Hang Zhou 0007, Xinxin Zuo, Li Cheng 0001
CVPR4
2025 InterMask: 3D Human Interaction Generation via Collaborative Masked Modeling
abstract
Generating realistic 3D human-human interactions from textual descriptions remains a challenging task. Existing approaches, typically based on diffusion models, often produce results lacking realism and fidelity. In this work, we introduce *InterMask*, a novel framework for generating human interactions using collaborative masked modeling in discrete space. InterMask first employs a VQ-VAE to transform each motion sequence into a 2D discrete motion token map. Unlike traditional 1D VQ token maps, it better preserves fine-grained spatio-temporal details and promotes *spatial awareness* within each token. Building on this representation, InterMask utilizes a generative masked modeling framework to collaboratively model the tokens of two interacting individuals. This is achieved by employing a transformer architecture specifically designed to capture complex spatio-temporal inter-dependencies. During training, it randomly masks the motion tokens of both individuals and learns to predict them. For inference, starting from fully masked sequences, it progressively fills in the tokens for both individuals. With its enhanced motion representation, dedicated architecture, and effective learning strategy, InterMask achieves state-of-the-art results, producing high-fidelity and diverse human interactions. It outperforms previous methods, achieving an FID of $5.154$ (vs $5.535$ of in2IN) on the InterHuman dataset and $0.399$ (vs $5.207$ of InterGen) on the InterX dataset. Additionally, InterMask seamlessly supports reaction generation without the need for model redesign or fine-tuning.
Muhammad Gohar Javed, Chuan Guo 0002, Li Cheng 0001
ICLR3
2025 MotionDreamer: One-to-Many Motion Synthesis with Localized Generative Masked Transformer
abstract
Generative masked transformer have demonstrated remarkable success across various content generation tasks, primarily due to their ability to effectively model large-scale dataset distributions with high consistency. However, in the animation domain, large datasets are not always available. Applying generative masked modeling to generate diverse instances from a single MoCap reference may lead to overfitting, a challenge that remains unexplored. In this work, we present MotionDreamer, a localized masked modeling paradigm designed to learn motion internal patterns from a given motion with arbitrary topology and duration. By embedding the given motion into quantized tokens with a novel distribution regularization method, MotionDreamer constructs a robust and informative codebook for local motion patterns. Moreover, a sliding window local attention is introduced in our masked transformer, enabling the generation of natural yet diverse animations that closely resemble the reference motion patterns. As demonstrated through comprehensive experiments, MotionDreamer outperforms the state-of-the-art methods that are typically GAN or Diffusion-based in both faithfulness and diversity. Thanks to the consistency and robustness of quantization-based approach, MotionDreamer can also effectively perform downstream tasks such as temporal motion editing, crowd motion synthesis, and beat-aligned dance generation, all using a single reference motion. Our implementation, learned models and results are to be made publicly available upon paper acceptance.
Chuan Guo 0002, Yuxuan Mu, Muhammad Gohar Javed, Xinxin Zuo, Juwei Lu, Hai Jiang 0001, Li Cheng 0001
ICLR8
2025 tHPM-LDM: Integrating Individual Historical Record with Population Memory in Latent Diffusion-Based Glaucoma Forecasting
Jianyang Xie, Yimin Luo, Yanda Meng, Savita Madhusudhan, Gregory Yoke Hong Lip, Li Cheng 0001, Yalin Zheng, He Zhao 0002
MICCAI (1)7
2025 Density-aware and Cluster-based Federated Anomaly Detection on Data Streams
abstract
Federated active anomaly detection on data streams becomes a crucial research problem, since it attempts to discover anomalous data with protecting data privacy and avoiding extensive data labeling. Although extensive work has been conducted on anomaly detection, distinguishing similar anomalies of different categories still remains quite a challenging issue. The requirement of privacy protection in federated settings aggravates the difficulties for instance query and scoring in active anomaly detection when solving this issue. To the best of our knowledge, limited work has focused on this research area. Therefore, we propose Density-aware and cluster-based Federated Active anomaly detection on data Streams, called DFAS. We design a novel lightweight federated anomaly detection clusters with density-aware hash cells, which successfully capture evolving data distribution. The federated anomaly detection clusters are incrementally updated with an acceptable theoretical reconstruction error guarantee. In addition, we propose a straightforward but effective metric divergences accompanied by a greedy search algorithm, which takes both global aggregation bias mitigation and efficiency into account. At last, DFAS detects anomalies and queries the instances for manual labels by measuring the density in hash cells of each cluster, effectively distinguishing closely distributed anomaly classes while maintaining data privacy in the federated setting. Comprehensive experiments on several real-world data sets show that DFAS outperforms previous methods, improving F1 scores by up to 26.7%.
Bin Li 0030, Li Cheng 0001, Zheng Qin 0002, Yunlong Wu 0002
WSDM2
2025 Highly Efficient 3D Human Pose Tracking From Events With Spiking Spatiotemporal Transformer
abstract
Event camera, as an asynchronous vision sensor capturing scene dynamics, presents new opportunities for highly efficient 3D human pose tracking. Existing approaches typically adopt modern-day Artificial Neural Networks (ANNs), such as CNNs or Transformer, where sparse events are converted into dense images or paired with additional gray-scale images as input. Such practices, however, ignore the inherent sparsity of events, resulting in redundant computations, increased energy consumption, and potentially degraded performance. Motivated by these observations, we introduce the first sparse Spiking Neural Networks (SNNs) framework for 3D human pose tracking based solely on events. Our approach eliminates the need to convert sparse data to dense formats or incorporate additional images, thereby fully exploiting the innate sparsity of input events. Central to our framework is a novel Spiking Spatio-temporal Transformer, which enables bi-directional spatio-temporal fusion of spike pose features and provides a guaranteed similarity measurement between binary spike features in spiking attention. Moreover, we have constructed a largescale synthetic dataset, SynEventHPD, that features a broad and diverse set of 3D human motions, as well as much longer hours of event streams. Empirical experiments demonstrate the superiority of our approach over existing state-of-the-art (SOTA) ANN-based methods, requiring only 19.1% FLOPs and 3.6% energy cost. Furthermore, our approach outperforms existing SNN-based benchmarks in this task, highlighting the effectiveness of our proposed SNN framework. The dataset will be released upon acceptance, and code can be found at https://github.com/JimmyZou/HumanPoseTracking_SNN.
Shihao Zou, Yuxuan Mu, Wei Ji 0011, Zi-An Wang, Xinxin Zuo, Sen Wang 0003, Weixin Si, Li Cheng 0001
IEEE Trans. Circuits Syst. Video Technol.8
2025 A Coarse-to-Fine Multi-Hypothesis Method for Ambiguous Hand Pose Estimation
abstract
In hand pose estimation, challenges such as occlusion often result in partial observation of a human hand, making it difficult to uniquely determine the hand pose, thus leading to ambiguity in certain hand regions. Heatmap-based methods may struggle with locating ambiguous joints and end up violating physiological constraints in their predictions. Parametric model based single-solution methods often fail to adequately address this ambiguity issue due to the inherent one-to-many mappings between input and output, resulting in unstable regression. While some existing multi-hypothesis methods have improved diversity by directly modeling the distribution of ambiguous hypotheses, their localization accuracy still falls short compared to the recent single-solution methods. To achieve quality results in both diversity and accuracy, we propose a novel multi-hypothesis approach for hand pose estimation, by progressively integrating heatmap information into the distribution of ambiguous poses using a RANSAC-like strategy. It starts with a conditional-flow model to provide an initial estimate of a coarse distribution over ambiguous joint poses. This is followed by randomly sampling multiple hypotheses, projecting each of them onto 2D heatmap plane, and employing consensus checks to identify unambiguous joints that adhere to skeletal constraints. Joint features are then resampled, with mismatches due to incorrect estimations being eliminated. Finally, we refine the distribution of ambiguous poses using graph neural networks and attention mechanisms. Extensive empirical experiments are carried out, where our approach are carefully examined both qualitatively and quantitatively. It is shown to not only produce more diverse & feasible pose hypotheses than existing multi-hypothesis methods, but also achieves accurate localization results comparable to the state-of-the-art single-solution methods.
Yuting Ge, Chi Xu 0002, Li Cheng 0001
IEEE Trans. Image Process.3
2025 A Hough Voting-Based 2-Point RANSAC Solution to the Perspective-n-Point Problem
abstract
Perspective- $n$ -point is a fundamental problem in multi-view geometry, yet two critical challenges persist: 1) The issues of high outlier rate and near degenerate cases exert a substantial impact on the robustness of existing P $n$ P methods. In the worst-case where both issues are in presence, existing methods tend to either produce erroneous results or become computationally prohibitive. 2) Conventionally, the hypothetical pose with the maximum inlier-set is assumed to be correct. However, it remains unclear whether this assumption holds when the outlier rate approaches ultra-high levels, and along this line what is the maximum amount of outliers that can be robustly handled. To address these challenges, this paper proposes a novel Hough voting based 2-point RANSAC solution. To our knowledge, it is the first P $n$ P solution capable of accurately and efficiently handling high outlier rates in near-degenerate cases. Extensive empirical evaluations have been conducted using the proposed approach, with a particular focus on a systematic examination under ultra-high outlier rates. The results show that, on random synthetic data, our approach works robustly even when dealing with up to 99% outliers. Meanwhile on real-world datasets, the maximum inlier-set assumption oftentimes fails when the outlier rate exceeds 97%, as the incorrect hypothetical poses may yield more inliers than the ground-truths. Our dataset and source code are to be made available at https://github.com/xuchi7/RPnP_plusplus.
Chi Xu 0002, Tingrui Guo, Li Cheng 0001
IEEE Trans. Image Process.4
2025 Hand Gesture Recognition From an Open-Set Perspective
abstract
Existing hand gesture recognition methods predominantly rely on a close-set assumption, which in essence limits the viewpoints, gesture categories, and hand shapes at test time to closely resemble those seen during training. This requirement is however rarely met in practice, as images are often captured from unconstrained viewpoints, with novel gestures and unseen hand shapes that can differ significantly from the training data. This motivates us to investigate an open-set hand gesture recognition problem, where hand gestures are still recognizable from unconstrained viewpoints, and novel gesture classes and hand shapes can be incrementally learned with just a few examples. To address this, we propose a viewpoint influence elimination network that extracts view-independent features, significantly improving performance in scenarios with unconstrained viewpoints. Moreover, a joint-weighted classification scheme is introduced to augment the cosine similarity metric for evaluating few-shot incremental learning of novel gestures and shapes. Finally, as existing hand gesture recognition datasets primarily adhere to the close-set assumption, a new hand gesture recognition dataset, OHG, is introduced in this paper, that includes a wide range of viewpoints, diverse gesture classes, and distinct hand shapes. Experimental hand gesture recognition results demonstrate the superior performance of our approach in both unconstrained viewpoint and few-shot incremental learning scenarios.
Jun Zhou 0026, Chi Xu 0002, Li Cheng 0001
IEEE Trans. Multim.3
2025 Generating High-Fidelity Clothed Human Dynamics with Temporal Diffusion
abstract
Clothed human modeling plays a crucial role in multimedia research, with applications spanning virtual reality, gaming, and fashion design. The goal is to learn clothed human dynamics from observations and then generate humans with high-fidelity clothing details for motion animation. Despite tremendous advancements in clothing shape analysis by existing approaches, the community still faces challenges in generating convincing visual effects of cloth dynamics, maintaining temporally smooth clothing details, and handling diverse clothing patterns. To address these challenges, we introduce ClothDiffuse, a temporal diffusion model that seamlessly integrates three key components into this task—temporal dynamics modeling, iterative refinement, and diversified generation. Our approach begins by using an encoder to extract high-level temporal features from input human body motions. These features are combined with a learnable pixel-aligned garment feature, serving as prior conditions for the shape decoder. The decoder then iteratively denoise Gaussian noise to produce clothing deformations over time on the input unclothed human bodies. To ensure that the results align with observations and adhere to physical plausibility for clothing shape inference, we propose two physics-inspired loss functions that preserve the intra-frame distances and inter-frame forces of clothing points. Additionally, the stochastic nature of the denoising process allows for the generation of diverse and plausible clothing shapes. Experiments show that our approach outperforms state-of-the-art methods in chamfer distance and visual effects, particularly for loose clothing such as dresses and skirts. Furthermore, our approach effectively adapts to out-of-domain clothing types and generates realistic clothes dynamics.
Shihao Zou, Yuanlu Xu, Nikolaos Sarafianos, Federica Bogo, Tony Tung, Weixin Si, Li Cheng 0001
ACM Trans. Multim. Comput. Commun. Appl.7
2024 MoMask: Generative Masked Modeling of 3D Human Motions
abstract
We introduce MoMask, a novel masked modeling framework for text-driven 3D human motion generation. In Mo-Mask, a hierarchical quantization scheme is employed to represent human motion as multi-layer discrete motion tokens with high-fidelity details. Starting at the base layer, with a sequence of motion tokens obtained by vector quan-tization, the residual tokens of increasing orders are de-rived and stored at the subsequent layers of the hierar-chy. This is consequently followed by two distinct bidirectional transformers. For the base-layer motion tokens, a Masked Transformer is designated to predict randomly masked motion tokens conditioned on text input at training stage. During generation (i. e. inference) stage, starting from an empty sequence, our Masked Transformer iteratively fills up the missing tokens; Subsequently, a Residual Transformer learns to progressively predict the next-layer tokens based on the results from current layer. Extensive experiments demonstrate that MoMask outperforms the state-of-art methods on the text-to-motion generation task, with an FID of 0.045 (vs e.g. 0.141 of T2M-GPT) on the HumanML3D dataset, and 0.228 (vs 0.514) on KIT-ML, respectively. MoMask can also be seamlessly applied in related tasks without further model fine-tuning, such as text-guided temporal inpainting.
Chuan Guo 0002, Yuxuan Mu, Muhammad Gohar Javed, Sen Wang 0003, Li Cheng 0001
CVPR5
2024 TexGen: Text-Guided 3D Texture Generation with Multi-view Sampling and Resampling
Dong Huo, Zixin Guo, Xinxin Zuo, Zhihao Shi, Juwei Lu, Peng Dai 0002, Songcen Xu, Li Cheng 0001, Yee-Hong Yang
ECCV (38)8
2024 GSD: View-Guided Gaussian Splatting Diffusion for 3D Reconstruction
Yuxuan Mu, Xinxin Zuo, Chuan Guo 0002, Juwei Lu, Songcen Xu, Peng Dai 0002, Youliang Yan, Li Cheng 0001
ECCV (79)10
2024 Generative Human Motion Stylization in Latent Space
abstract
Human motion stylization aims to revise the style of an input motion while keeping its content unaltered. Unlike existing works that operate directly in pose space, we leverage the \textit{latent space} of pretrained autoencoders as a more expressive and robust representation for motion extraction and infusion. Building upon this, we present a novel \textit{generative} model that produces diverse stylization results of a single motion (latent) code. During training, a motion code is decomposed into two coding components: a deterministic content code, and a probabilistic style code adhering to a prior distribution; then a generator massages the random combination of content and style codes to reconstruct the corresponding motion codes. Our approach is versatile, allowing the learning of probabilistic style space from either style labeled or unlabeled motions, providing notable flexibility in stylization as well. In inference, users can opt to stylize a motion using style cues from a reference motion or a label. Even in the absence of explicit style input, our model facilitates novel re-stylization by sampling from the unconditional style prior distribution. Experimental results show that our proposed stylization models, despite their lightweight design, outperform the state-of-the-arts in style reeanactment, content preservation, and generalization across various applications and settings.
Chuan Guo 0002, Yuxuan Mu, Xinxin Zuo, Peng Dai 0002, Youliang Yan, Juwei Lu, Li Cheng 0001
ICLR7
2024 RACon: Retrieval-Augmented Simulated Character Locomotion Control
abstract
In computer animation, driving a simulated character with lifelike motion is challenging. Current generative models, though able to generalize to diverse motions, often pose challenges to the responsiveness of end-user control. To address these issues, we introduce RACon: Retrieval-Augmented Simulated Character Locomotion Control. Our end-to-end hierarchical reinforcement learning method utilizes a retriever and a motion controller. The retriever searches motion experts from a user-specified database in a task-oriented fashion, which boosts the responsiveness to the user’s control. The selected motion experts and the manipulation signal are then transferred to the controller to drive the simulated character. In addition, a retrieval-augmented discriminator is designed to stabilize the training process. Our method surpasses existing techniques in both quality and quantity in locomotion control, as demonstrated in our empirical study. Moreover, by switching extensive databases for retrieval, it can adapt to distinctive motion types at run time. We will release our code upon acceptance.
Yuxuan Mu, Shihao Zou, Kangning Yin, Zheng Tian 0002, Li Cheng 0001, Weinan Zhang 0001, Jun Wang 0012
ICME5
2024 Unleashing Multispectral Video's Potential in Semantic Segmentation: A Semi-supervised Viewpoint and New UAV-View Benchmark
abstract
Thanks to the rapid progress in RGB & thermal imaging, also known as multispectral imaging, the task of multispectral video semantic segmentation, or MVSS in short, has recently drawn significant attentions. Noticeably, it offers new opportunities in improving segmentation performance under unfavorable visual conditions such as poor light or overexposure. Unfortunately, there are currently very few datasets available, including for example MVSeg dataset that focuses purely toward eye-level view; and it features the sparse annotation nature due to the intensive demands of labeling process. To address these key challenges of the MVSS task, this paper presents two major contributions: the introduction of MVUAV, a new MVSS benchmark dataset, and the development of a dedicated semi-supervised MVSS baseline - SemiMV. Our MVUAV dataset is captured via Unmanned Aerial Vehicles (UAV), which offers a unique oblique bird’s-eye view complementary to the existing MVSS datasets; it also encompasses a broad range of day/night lighting conditions and over 30 semantic categories. In the meantime, to better leverage the sparse annotations and extra unlabeled RGB-Thermal videos, a semi-supervised learning baseline, SemiMV, is proposed to enforce consistency regularization through a dedicated Cross-collaborative Consistency Learning (C3L) module and a denoised temporal aggregation strategy. Comprehensive empirical evaluations on both MVSeg and MVUAV benchmark datasets have showcased the efficacy of our SemiMV baseline.
Wei Ji 0011, Wenbo Li 0001, Yilin Shen, Li Cheng 0001, Hongxia Jin
NeurIPS5
2024 LAPAID: A Lightweight, Adaptive and Perspicacious Active Intrusion Detection Method on Network Traffic Streams
abstract
Active intrusion detection on network traffic streams becomes a crucial research problem, since it tries to maximize the effectiveness with limited labeled instances. The closely distributed anomalies of different categories and dynamic network traffic stream make it still quite a challenging issue. Therefore, we propose a Lightweight, Adaptive and Perspicacious Active Intrusion Detection method on network traffic streams, called LAPAID. We design novel lightweight anomaly detection clusters with density-aware hash cells, which successfully capture evolving data distribution. The anomaly detection clusters are incrementally updated in an approximate manner, and aggregated with a straightforward but effective metric divergences. At last, LAPAID detects anomalies and queries the instances for manual labels by measuring the density in hash cells of each cluster, effectively distinguishing closely distributed anomaly classes. Comprehensive experiments on several real-world data sets show that LAPAID outperforms previous methods, improving F1scores by 16.17% on average.
Bin Li 0030, Li Cheng 0001, Zhongshan Zhang
TrustCom2
2024 Adaptive and augmented active anomaly detection on dynamic network traffic streams
abstract
Active anomaly detection queries labels of sampled instances and uses them to incrementally update the detection model, and has been widely adopted in detecting network attacks. However, existing methods cannot achieve desirable performance on dynamic network traffic streams because (1) their query strategies cannot sample informative instances to make the detection model adapt to the evolving stream and (2) their model updating relies on limited query instances only and fails to leverage the enormous unlabeled instances on streams. To address these issues, we propose an active tree based model, adaptive and augmented active prior-knowledge forest (A 3 PF), for anomaly detection on network traffic streams. A prior-knowledge forest is constructed using prior knowledge of network attacks to find feature subspaces that better distinguish network anomalies from normal traffic. On one hand, to make the model adapt to the evolving stream, a novel adaptive query strategy is designed to sample informative instances from two aspects: the changes in dynamic data distribution and the uncertainty of anomalies. On the other hand, based on the similarity of instances in the neighborhood, we devise an augmented update method to generate pseudo labels for the unlabeled neighbors of query instances, which enables usage of the enormous unlabeled instances during model updating. Extensive experiments on two benchmarks, CIC-IDS2017 and UNSW-NB15, demonstrate that A 3 PF achieves significant improvements over previous active methods in terms of the area under the receiver operating characteristic curve (AUC-ROC) (20.9% and 21.5%) and the area under the precision-recall curve (AUC-PR) (44.6% and 64.1%).
Bin Li 0030, Yijie Wang 0001, Li Cheng 0001
Frontiers Inf. Technol. Electron. Eng.3
2024 Domain-Invariant Prototypes for Semantic Segmentation
abstract
Deep learning has greatly advanced the performance of semantic segmentation, however, its success relies on the availability of large amounts of annotated data for training. Hence, many efforts have been devoted to domain adaptive semantic segmentation that focuses on transferring semantic knowledge from a labeled source domain to an unlabeled target domain. Existing self-training methods typically require multiple rounds of training, while another popular framework based on adversarial training is known to be sensitive to hyper-parameters. We propose an easy-to-train framework that learns domain-invariant prototypes for domain adaptive semantic segmentation. In particular, we show that domain adaptation shares a common character with few-shot learning in that both aim to recognize some types of unseen data with knowledge learned from large amounts of seen data. Thus, we propose a unified framework for domain adaptation and few-shot learning. The core idea is to use the class prototypes extracted from few-shot annotated target images to classify pixels of both source images and target images. Our method involves only one-stage training and does not need to be trained on large-scale un-annotated target images. Moreover, our method can be extended to variants of both domain adaptation and few-shot learning. Competitive performances achieved on GTA5-to-Cityscapes and SYNTHIA-to-Cityscapes adaptation tasks show the effectiveness of the proposed novel while simple domain adaptation framework. The source code used in this paper is available at https://github.com/zgyang-hnu/DIP-hunnu.
Zhengeng Yang, Hongshan Yu, Wei Sun 0028, Li Cheng 0001, Ajmal Mian
IEEE Trans. Circuits Syst. Video Technol.4
2024 Realistic Depth Image Synthesis for 3D Hand Pose Estimation
abstract
The training of depth image-based hand pose estimation model typically relies on real-life datasets which are expected to be 1) largescale and cover a diverse range of hand poses and hand shapes, and 2) always come with high-precision annotations. However, existing datasets in reality are rather limited in the above regards due to multitude practical constraints, with time and cost being the major concerns. This observation motivates us to propose an alternative approach, where hand pose model is primarily trained with synthesized hand depth images that closely mimicking the characteristic noise patterns of a specific depth camera make under consideration. It is achieved by firstly mapping a Gaussian distributed variable to certain specific non-i.i.d. (independent and identically distributed) depth noise pattern, and then transforming a “vanilla” noise-free synthetic depth image to a realistic-looking image. Extensive empirical experiments demonstrate that our approach is capable of generating camera-specific realistic-looking hand depth images with precise annotations; comparing to entirely relying on annotated real images, a hand pose model with better performance is obtained by using only a small fraction (10%) of annotated real images as well as our synthesized images.
Jun Zhou 0026, Chi Xu 0002, Yuting Ge, Li Cheng 0001
IEEE Trans. Multim.4
2023 Multispectral Video Semantic Segmentation: A Benchmark Dataset and Baseline
abstract
Robust and reliable semantic segmentation in complex scenes is crucial for many real-life applications such as autonomous safe driving and nighttime rescue. In most approaches, it is typical to make use of RGB images as input. They however work well only in preferred weather conditions; when facing adverse conditions such as rainy, overexposure, or low-light, they often fail to deliver satisfactory results. This has led to the recent investigation into multispectral semantic segmentation, where RGB and thermal infrared (RGBT) images are both utilized as input. This gives rise to significantly more robust segmentation of image objects in complex scenes and under adverse conditions. Nevertheless, the present focus in single RGBT image input restricts existing methods from well addressing dynamic real-world scenes. Motivated by the above observations, in this paper, we set out to address a relatively new task of semantic segmentation of multispectral video input, which we refer to as Multispectral Video Semantic Segmentation, or MVSS in short. An in-house MVSeg dataset is thus curated, consisting of 738 calibrated RGB and thermal videos, accompanied by 3,545 fine-grained pixel-level semantic annotations of 26 categories. Our dataset contains a wide range of challenging urban scenes in both daytime and nighttime. Moreover, we propose an effective MVSS baseline, dubbed MVNet, which is to our knowledge the first model to jointly learn semantic representations from multispectral and temporal contexts. Comprehensive experiments are conducted using various semantic segmentation models on the MVSeg dataset. Empirically, the engagement of multispectral video input is shown to lead to significant improvement in semantic segmentation; the effectiveness of our MVNet baseline has also been verified.
Wei Ji 0011, Cheng Bian, Zongwei Zhou, Jiaying Zhao, Alan L. Yuille, Li Cheng 0001
CVPR7
2023 SemanticRT: A Large-Scale Dataset and Method for Robust Semantic Segmentation in Multispectral Images
abstract
Growing interests in multispectral semantic segmentation (MSS) have been witnessed in recent years, thanks to the unique advantages of combining RGB and thermal infrared images to tackle challenging scenarios with adverse conditions. However, unlike traditional RGB-only semantic segmentation, the lack of a large-scale MSS dataset has become a hindrance to the progress of this field. To address this issue, we introduce a SemanticRT dataset - the largest MSS dataset to date, comprising 11,371 high-quality, pixel-level annotated RGB-thermal image pairs. It is 7 times larger than the existing MFNet dataset, and covers a wide variety of challenging scenarios in adverse lighting conditions such as low-light and pitch black. Further, a novel Explicit Complement Modeling (ECM) framework is developed to extract modality-specific information, which is propagated through a robust cross-modal feature encoding and fusion process. Extensive experiments demonstrate the advantages of our approach and dataset over the existing counterparts. Our new dataset may also facilitate further development and evaluation of existing and new MSS algorithms.
Wei Ji 0011, Cheng Bian, Zhicheng Zhang 0005, Li Cheng 0001
ACM Multimedia5
2023 Decorate3D: Text-Driven High-Quality Texture Generation for Mesh Decoration in the Wild
abstract
This paper presents Decorate3D, a versatile and user-friendly method for the creation and editing of 3D objects using images. Decorate3D models a real-world object of interest by neural radiance field (NeRF) and decomposes the NeRF representation into an explicit mesh representation, a view-dependent texture, and a diffuse UV texture. Subsequently, users can either manually edit the UV or provide a prompt for the automatic generation of a new 3D-consistent texture. To achieve high-quality 3D texture generation, we propose a structure-aware score distillation sampling method to optimize a neural UV texture based on user-defined text and empower an image diffusion model with 3D-consistent generation capability. Furthermore, we introduce a few-view resampling training method and utilize a super-resolution model to obtain refined high-resolution UV textures (2048$\times$2048) for 3D texturing. Extensive experiments collectively validate the superior performance of Decorate3D in retexturing real-world 3D objects. Project page: https://decorate3d.github.io/Decorate3D/.
Xinxin Zuo, Peng Dai 0002, Juwei Lu, Li Cheng 0001, Youliang Yan, Songcen Xu
NeurIPS6
2023 DVSOD: RGB-D Video Salient Object Detection
abstract
Salient object detection (SOD) aims to identify standout elements in a scene, with recent advancements primarily focused on integrating depth data (RGB-D) or temporal data from videos to enhance SOD in complex scenes. However, the unison of two types of crucial information remains largely underexplored due to data constraints. To bridge this gap, we in this work introduce the DViSal dataset, fueling further research in the emerging field of RGB-D video salient object detection (DVSOD). Our dataset features 237 diverse RGB-D videos alongside comprehensive annotations, including object and instance-level markings, as well as bounding boxes and scribbles. These resources enable a broad scope for potential research directions. We also conduct benchmarking experiments using various SOD models, affirming the efficacy of multimodal video input for salient object detection. Lastly, we highlight some intriguing findings and promising future research avenues. To foster growth in this field, our dataset and benchmark results are publicly accessible at: https://dvsod.github.io/.
Wei Ji 0011, Size Wang, Wenbo Li 0001, Li Cheng 0001
NeurIPS5
2023 Delving into Calibrated Depth for Accurate RGB-D Salient Object Detection
Wei Ji 0011, Miao Zhang 0004, Yongri Piao, Huchuan Lu, Li Cheng 0001
Int. J. Comput. Vis.6
2023 Special issue on human-centric intelligent multimedia understanding
Zhenguang Liu, Roger Zimmermann, Li Cheng 0001
Multim. Syst.3
2023 Investigating Pose Representations and Motion Contexts Modeling for 3D Motion Prediction
abstract
Predicting human motion from historical pose sequence is crucial for a machine to succeed in intelligent interactions with humans. One aspect that has been obviated so far, is the fact that how we represent the skeletal pose has a critical impact on the prediction results. Yet there is no effort that investigates across different pose representation schemes. We conduct an indepth study on various pose representations with a focus on their effects on the motion prediction task. Moreover, recent approaches build upon off-the-shelf RNN units for motion prediction. These approaches process input pose sequence sequentially and inherently have difficulties in capturing long-term dependencies. In this paper, we propose a novel RNN architecture termed AHMR (Attentive Hierarchical Motion Recurrent network) for motion prediction which simultaneously models local motion contexts and a global context. We further explore a geodesic loss and a forward kinematics loss for the motion prediction task, which have more geometric significance than the widely employed L2 loss. Interestingly, we applied our method to a range of articulate objects including human, fish, and mouse. Empirical results show that our approach outperforms the state-of-the-art methods in short-term prediction and achieves much enhanced long-term prediction proficiency, such as retaining natural human-like motions over 50 seconds predictions. Our codes are released.
Zhenguang Liu, Shuang Wu 0002, Shuyuan Jin, Shouling Ji, Qi Liu 0049, Shijian Lu, Li Cheng 0001
IEEE Trans. Pattern Anal. Mach. Intell.7
2023 Snipper: A Spatiotemporal Transformer for Simultaneous Multi-Person 3D Pose Estimation Tracking and Forecasting on a Video Snippet
abstract
Multi-person pose understanding from RGB videos involves three complex tasks: pose estimation, tracking and motion forecasting. Intuitively, accurate multi-person pose estimation facilitates robust tracking, and robust tracking builds crucial history for correct motion forecasting. Most existing works either focus on a single task or employ multi-stage approaches to solving multiple tasks separately, which tends to make sub-optimal decision at each stage and also fail to exploit correlations among the three tasks. In this paper, we propose Snipper, a unified framework to perform multi-person 3D pose estimation, tracking, and motion forecasting simultaneously in a single stage. We propose an efficient yet powerful deformable attention mechanism to aggregate spatiotemporal information from the video snippet. Building upon this deformable attention, a video transformer is learned to encode the spatiotemporal features from the multi-frame snippet and to decode informative pose features for multi-person pose queries. Finally, these pose queries are regressed to predict multi-person pose trajectories and future motions in a single shot. In the experiments, we show the effectiveness of Snipper on three challenging public datasets where our generic model rivals specialized state-of-art baselines for pose estimation, tracking, and forecasting. Code is available athttps://github.com/JimmyZou/Snipper.
Shihao Zou, Yuanlu Xu, Chao Li 0021, Lingni Ma, Li Cheng 0001, Minh Vo
IEEE Trans. Circuits Syst. Video Technol.5
2023 Human Pose and Shape Estimation From Single Polarization Images
abstract
This paper focuses on a new problem of estimating human pose and shape from single polarization images. Polarization camera is known to be able to capture the polarization of reflected lights that preserves rich geometric cues of an object surface. Inspired by the recent applications in surface normal reconstruction from polarization images, in this paper, we attempt to estimate human pose and shape from single polarization images by leveraging the polarization-induced geometric cues. A dedicated two-stage pipeline is proposed: given a single polarization image, stage one (Polar2Normal) focuses on the fine detailed human body surface normal estimation; stage two (Polar2Shape) then reconstructs clothed human shape from the polarization image and the estimated surface normal. To empirically validate our approach, a dedicated dataset (PHSPD) is constructed, consisting of over 500 K frames with accurate pose and parametric shape annotations. Empirical evaluations on this real-world dataset as well as a synthetic dataset, SURREAL, demonstrate the effectiveness of our approach. It suggests polarization camera as a promising alternative to the more conventional RGB camera for human pose and shape estimation.
Shihao Zou, Xinxin Zuo, Sen Wang 0003, Yiming Qian, Chuan Guo 0002, Li Cheng 0001
IEEE Trans. Multim.6
2022 Generating Diverse and Natural 3D Human Motions from Text
abstract
Automated generation of 3D human motions from text is a challenging problem. The generated motions are expected to be sufficiently diverse to explore the text-grounded motion space, and more importantly, accurately depicting the content in prescribed text descriptions. Here we tackle this problem with a two-stage approach: text2length sampling and text2motion generation. Text2length involves sampling from the learned distribution function of motion lengths conditioned on the input text. This is followed by our text2motion module using temporal variational autoen-coder to synthesize a diverse set of human motions of the sampled lengths. Instead of directly engaging with pose sequences, we propose motion snippet code as our internal motion representation, which captures local semantic motion contexts and is empirically shown to facilitate the generation of plausible motions faithful to the input text. Moreover, a large-scale dataset of scripted 3D Human motions, HumanML3D, is constructed, consisting of 14,616 motion clips and 44,970 text descriptions.
Chuan Guo 0002, Shihao Zou, Xinxin Zuo, Sen Wang 0003, Wei Ji 0011, Li Cheng 0001
CVPR7
2022 Exploring Denoised Cross-video Contrast for Weakly-supervised Temporal Action Localization
abstract
Weakly-supervised temporal action localization aims to localize actions in untrimmed videos with only video-level labels. Most existing methods address this problem with a “localization-by-classification” pipeline that localizes action regions based on snippet-wise classification sequences. Snippet-wise classifications are unfortunately error prone due to the sparsity of video-level labels. Inspired by recent success in unsupervised contrastive representation learning, we propose a novel denoised cross-video contrastive algorithm, aiming to enhance the feature discrimination ability of video snippets for accurate temporal action localization in the weakly-supervised setting. This is enabled by three key designs: 1) an effective pseudo-label denoising module to alleviate the side effects caused by noisy contrastive features, 2) an efficient region-level feature contrast strategy with a region-level memory bank to capture “global” contrast across the entire dataset, and 3) a diverse contrastive learning strategy to enable action-background separation as well as intra-class compactness & inter-class separability. Extensive experiments on THUMOS14 and ActivityNet v1.3 demonstrate the superior performance of our approach.
Tianyu Yang 0003, Wei Ji 0011, Jue Wang 0001, Li Cheng 0001
CVPR5
2022 TM2T: Stochastic and Tokenized Modeling for the Reciprocal Generation of 3D Human Motions and Texts
Chuan Guo 0002, Xinxin Zuo, Sen Wang 0003, Li Cheng 0001
ECCV (35)4
2022 Object Wake-Up: 3D Object Rigging from a Single Image
Xinxin Zuo, Sen Wang 0003, Zhenbo Yu, Bingbing Ni, Minglun Gong, Li Cheng 0001
ECCV (2)8
2022 Promoting Saliency From Depth: Deep Unsupervised RGB-D Saliency Detection
Wei Ji 0011, Qi Bi, Chuan Guo 0002, Jie Liu 0044, Li Cheng 0001
ICLR6
2022 Music-to-Dance Generation with Optimal Transport
abstract
Dance choreography for a piece of music is a challenging task, having to be creative in presenting distinctive stylistic dance elements while taking into account the musical theme and rhythm. It has been tackled by different approaches such as similarity retrieval, sequence-to-sequence modeling and generative adversarial networks, but their generated dance sequences are often short of motion realism, diversity and music consistency. In this paper, we propose a Music-to-Dance with Optimal Transport Network (MDOT-Net) for learning to generate 3D dance choreographies from music. We introduce an optimal transport distance for evaluating the authenticity of the generated dance distribution and a Gromov-Wasserstein distance to measure the correspondence between the dance distribution and the input music. This gives a well defined and non-divergent training objective that mitigates the limitation of standard GAN training which is frequently plagued with instability and divergent generator loss issues. Extensive experiments demonstrate that our MDOT-Net can synthesize realistic and diverse dances which achieve an organic unity with the input music, reflecting the shared intentionality and matching the rhythmic articulation. Sample results are found at https://www.youtube.com/watch?v=dErfBkrlUO8.
Shuang Wu 0002, Shijian Lu, Li Cheng 0001
IJCAI3
2022 DFAID: Density-aware and feature-deviated active intrusion detection over network traffic streams
Bin Li 0030, Yijie Wang 0001, Kele Xu, Li Cheng 0001, Zhiquan Qin
Comput. Secur.4
2022 Action2video: Generating Videos of Human 3D Actions
Chuan Guo 0002, Xinxin Zuo, Sen Wang 0003, Xinshuang Liu, Shihao Zou, Minglun Gong, Li Cheng 0001
Int. J. Comput. Vis.7
2022 3D pose estimation and future motion prediction from 2D images
Youdong Ma, Xinxin Zuo, Sen Wang 0003, Minglun Gong, Li Cheng 0001
Pattern Recognit.6
2022 Least Squares Approximation via Sparse Subsampled Randomized Hadamard Transform
abstract
Solving least squares (LS) problems is a major topic in many applications. With recent data explosion, traditional approach is no longer suitable while working with large datasets, instead, randomized algorithms become popular in addressing this issue. In this article we propose a new randomized algorithm - sparse subsampled randomized Hadamard transform (SpSRHT) for solving overdetermined least squares problems. Its unique block structure connects two most commonly used randomized algorithms subsampled randomized Hadamard transform (SRHT) and sparse subspace embedding (SpEmb) and creates a general framework which contains them as special cases. We have shown theoretically that SpSRHT with different parameters reaches the relative-error bound with sketch size ranging from the sketch size required by SpEmb to SRHT. The new algorithm closes the gap between SRHT and SpEmb which provides the possibility of balancing accuracy and efficiency demonstrated in them. This advantage of SpSRHT is also well illustrated in our numerical experiments.
Dan Teng, Xiaowei Zhang 0002, Li Cheng 0001, Delin Chu
IEEE Trans. Big Data3
2022 DMRA: Depth-Induced Multi-Scale Recurrent Attention Network for RGB-D Saliency Detection
abstract
In this work, we propose a novel depth-induced multi-scale recurrent attention network for RGB-D saliency detection, named as DMRA. It achieves dramatic performance especially in complex scenarios. There are four main contributions of our network that are experimentally demonstrated to have significant practical merits. First, we design an effective depth refinement block using residual connections to fully extract and fuse cross-modal complementary cues from RGB and depth streams. Second, depth cues with abundant spatial information are innovatively combined with multi-scale contextual features for accurately locating salient objects. Third, a novel recurrent attention module inspired by Internal Generative Mechanism of human brain is designed to generate more accurate saliency results via comprehensively learning the internal semantic relation of the fused feature and progressively optimizing local details with memory-oriented scene understanding. Finally, a cascaded hierarchical feature fusion strategy is designed to promote efficient information interaction of multi-level contextual features and further improve the contextual representability of model. In addition, we introduce a new real-life RGB-D saliency dataset containing a variety of complex scenarios that has been widely used as a benchmark dataset in recent RGB-D saliency detection research. Extensive empirical experiments demonstrate that our method can accurately identify salient objects and achieve appealing performance against 18 state-of-the-art RGB-D saliency models on nine benchmark datasets.
Wei Ji 0011, Ge Yan 0006, Yongri Piao, Shunyu Yao 0004, Miao Zhang 0004, Li Cheng 0001, Huchuan Lu
IEEE Trans. Image Process.7
2022 Deep Learning for Visual Tracking: A Comprehensive Survey
abstract
Visual target tracking is one of the most sought-after yet challenging research topics in computer vision. Given the ill-posed nature of the problem and its popularity in a broad range of real-world scenarios, a number of large-scale benchmark datasets have been established, on which considerable methods have been developed and demonstrated with significant progress in recent years – predominantly by recentdeep learning(DL)-based methods. This survey aims to systematically investigate the current DL-based visual tracking methods, benchmark datasets, and evaluation metrics. It also extensively evaluates and analyzes the leading visual tracking methods. First, the fundamental characteristics, primary motivations, and contributions of DL-based methods are summarized from nine key aspects of: network architecture, network exploitation, network training for visual tracking, network objective, network output, exploitation of correlation filter advantages, aerial-view tracking, long-term tracking, and online tracking. Second, popular visual tracking benchmarks and their respective properties are compared, and their evaluation metrics are summarized. Third, the state-of-the-art DL-based methods are comprehensively examined on a set of well-established benchmarks of OTB2013, OTB2015, VOT2018, LaSOT, UAV123, UAVDT, and VisDrone2019. Finally, by conducting critical analyses of these state-of-the-art trackers quantitatively and qualitatively, their pros and cons under various common scenarios are investigated. It may serve as a gentle use guide for practitioners to weigh when and under what conditions to choose which method(s). It also facilitates a discussion on ongoing issues and sheds light on promising research directions.
Seyed Mojtaba Marvasti-Zadeh, Li Cheng 0001, Hossein Ghanei-Yakhdan, Shohreh Kasaei
IEEE Trans. Intell. Transp. Syst.2
2022 Stabilizing Training of Generative Adversarial Nets via Langevin Stein Variational Gradient Descent
abstract
Generative adversarial networks (GANs), which are famous for the capability of learning complex underlying data distribution, are, however, known to be tricky in the training process, which would probably result in mode collapse or performance deterioration. Current approaches of dealing with GANs' issues almost utilize some practical training techniques for the purpose of regularization, which, on the other hand, undermines the convergence and theoretical soundness of GAN. In this article, we propose to stabilize GAN training via a novel particle-based variational inference-Langevin Stein variational gradient descent (LSVGD), which not only inherits the flexibility and efficiency of original SVGD but also aims to address its instability issues by incorporating an extra disturbance into the update dynamics. We further demonstrate that, by properly adjusting the noise variance, LSVGD simulates a Langevin process whose stationary distribution is exactly the target distribution. We also show that LSVGD dynamics has an implicit regularization, which is able to enhance particles' spread-out and diversity. Finally, we present an efficient way of applying particle-based variational inference on a general GAN training procedure no matter what loss function is adopted. Experimental results on one synthetic data set and three popular benchmark data sets-Cifar-10, Tiny-ImageNet, and CelebA-validate that LSVGD can remarkably improve the performance and stability of various GAN models.
Dong Wang 0015, Xiaoqian Qin, Fengyi Song, Li Cheng 0001
IEEE Trans. Neural Networks Learn. Syst.4
2021 Neighborhood Consensus Networks for Unsupervised Multi-view Outlier Detection
abstract
Multi-view outlier detection recently attracted rapidly growing attention with the development of multi-view learning. Although promising performance demonstrated, we observe that identifying outliers in multi-view data is still a challenging task due to the complicated characteristics of multi-view data. Specifically, an effective multi-view outlier detection method should be able to handle (1) different types of outliers; (2) two or more views; (3) samples without clusters; (4) high dimensional data. Unfortunately, little is known about how these four issues can be handled simultaneously. In this paper, we propose an unsupervised multi-view outlier detection method to address these issues. Our method is based on the proposed novel neighborhood consensus networks termed NC-Nets, which automatically encodes intrinsic information into a comprehensive latent space for each view (for issue (4)) and uniforms the neighborhood structures among different views (for issue (2)). Accordingly, we propose an outlier score measurement which consists of two parts: the within-view reconstruction score and the cross-view neighborhood consensus score. The measurement is designed based on the characteristics of the different outlier types (for issue (1)) and no cluster assumption is needed (for issue (3)). Experimental results show that our method significantly outperforms state-of-the-art methods. On average, our method achieves 11.2% ~ 96.2% improvement in term of AUC and 33.5% ~ 352.7% improvement in term of F1-Score.
Li Cheng 0001, Yijie Wang 0001
AAAI1
2021 CHASE: Robust Visual Tracking via Cell-Level Differentiable Neural Architecture Search
Seyed Mojtaba Marvasti-Zadeh, Javad Khaghani, Li Cheng 0001, Hossein Ghanei-Yakhdan, Shohreh Kasaei
BMVC3
2021 Enhancing Human Motion Assessment by Self-supervised Representation Learning
Mahdiar Nekoui, Li Cheng 0001
BMVC2
2021 Calibrated RGB-D Salient Object Detection
abstract
Complex backgrounds and similar appearances between objects and their surroundings are generally recognized as challenging scenarios in Salient Object Detection (SOD). This naturally leads to the incorporation of depth information in addition to the conventional RGB image as input, known as RGB-D SOD or depth-aware SOD. Meanwhile, this emerging line of research has been considerably hindered by the noise and ambiguity that prevail in raw depth images. To address the aforementioned issues, we propose a Depth Calibration and Fusion (DCF) framework that contains two novel components: 1) a learning strategy to calibrate the latent bias in the original depth maps towards boosting the SOD performance; 2) a simple yet effective cross reference module to fuse features from both RGB and depth modalities. Extensive empirical experiments demonstrate that the proposed approach achieves superior performance against 27 state-of-the-art methods. Moreover, our depth calibration strategy alone can work as a preprocessing step; empirically it results in noticeable improvements when being applied to existing cutting-edge RGB-D SOD models. Source code is available at https://github.com/jiwei0921/DCF.
Wei Ji 0011, Miao Zhang 0004, Yongri Piao, Shunyu Yao 0004, Qi Bi, Kai Ma 0002, Yefeng Zheng 0001, Huchuan Lu, Li Cheng 0001
CVPR11
2021 Learning Calibrated Medical Image Segmentation via Multi-Rater Agreement Modeling
abstract
In medical image analysis, it is typical to collect multiple annotations, each from a different clinical expert or rater, in the expectation that possible diagnostic errors could be mitigated. Meanwhile, from the computer vision practitioner viewpoint, it has been a common practice to adopt the ground-truth labels obtained via either the majority-vote or simply one annotation from a preferred rater. This process, however, tends to overlook the rich information of agreement or disagreement ingrained in the raw multi-rater annotations. To address this issue, we propose to explicitly model the multi-rater (dis-)agreement, dubbed MRNet, which has two main contributions. First, an expertise-aware inferring module or EIM is devised to embed the expertise level of individual raters as prior knowledge, to form high-level semantic features. Second, our approach is capable of reconstructing multi-rater gradings from coarse predictions, with the multi-rater (dis-)agreement cues being further exploited to improve the segmentation performance. To our knowledge, our work is the first in producing calibrated predictions under different expertise levels for medical image segmentation. Extensive empirical experiments are conducted across five medical segmentation tasks of diverse imaging modalities. In these experiments, superior performance of our MRNet is observed comparing to the state-of-the-arts, indicating the effectiveness and applicability of our MRNet toward a wide range of medical segmentation tasks. Source code is publicly available.
Wei Ji 0011, Kai Ma 0002, Cheng Bian, Qi Bi, Hanruo Liu, Li Cheng 0001, Yefeng Zheng 0001
CVPR9
2021 Fden: Mining Effective Information of Features in Detecting Network Anomalies
abstract
Network anomaly detection is important for detecting and reacting to the presence of network attacks. In this paper, we propose a novel method to effectively leverage the features in detecting network anomalies, named FDEn, consisting of flow-based Feature Derivation (FD) and prior knowledge incorporated Ensemble models (Enpk). To mine the effective information in features, 149 features are derived to enrich the feature set of the original data with covering more characteristics of network traffic. To leverage these features effectively, an ensemble model Enpk, including CatBoost and XGBoost, based on the bagging strategy is proposed to first detect anomalies by combining numerical features and categorical features. And then, Enpkadjusts the predicted label of specific data by incorporating the prior knowledge of network security. We conduct empirically experiments on the data set provided by the Network Anomaly Detection Challenge (NADC), in which we obtain average improvement up to 61.6%, 31.7%, 50.2%, and 45.0%, in terms of the cost score, precision, recall and F1-score, respectively.
Bin Li 0030, Yijie Wang 0001, Kele Xu, Li Cheng 0001
ICASSP6
2021 EventHPE: Event-based 3D Human Pose and Shape Estimation
abstract
Event camera is an emerging imaging sensor for capturing dynamics of moving objects as events, which motivates our work in estimating 3D human pose and shape from the event signals. Events, on the other hand, have their unique challenges: rather than capturing static body postures, the event signals are best at capturing local motions. This leads us to propose a two-stage deep learning approach, called EventHPE. The first-stage, FlowNet, is trained by unsupervised learning to infer optical flow from events. Both events and optical flow are closely related to human body dynamics, which are fed as input to the ShapeNet in the second stage, to estimate 3D human shapes. To mitigate the discrepancy between image-based flow (optical flow) and shape-based flow (vertices movement of human body shape), a novel flow coherence loss is introduced by exploiting the fact that both flows are originated from the identical human motion. An in-house event-based 3D human dataset is curated that comes with 3D pose and shape annotations, which is by far the largest one to our knowledge. Empirical evaluations on DHP19 dataset and our in-house dataset demonstrate the effectiveness of our approach.
Shihao Zou, Chuan Guo 0002, Xinxin Zuo, Sen Wang 0003, Pengyu Wang 0007, Xiaoqin Hu, Shoushun Chen, Minglun Gong, Li Cheng 0001
ICCV9
2021 Dual Learning Music Composition and Dance Choreography
abstract
Music and dance have always co-existed as pillars of human activities, contributing immensely to the cultural, social, and entertainment functions in virtually all societies. Notwithstanding the gradual systematization of music and dance into two independent disciplines, their intimate connection is undeniable and one art-form often appears incomplete without the other. Recent research works have studied generative models for dance sequences conditioned on music. The dual task of composing music for given dances, however, has been largely overlooked. In this paper, we propose a novel extension, where we jointly model both tasks in a dual learning approach. To leverage the duality of the two modalities, we introduce an optimal transport objective to align feature embeddings, as well as a cycle consistency loss to foster overall consistency. Experimental results demonstrate that our dual learning framework improves individual task performance, delivering generated music compositions and dance choreographs that are realistic and faithful to the conditioned inputs.
Shuang Wu 0002, Zhenguang Liu, Shijian Lu, Li Cheng 0001
ACM Multimedia4
2021 Joint Semantic Mining for Weakly Supervised RGB-D Salient Object Detection
abstract
Training saliency detection models with weak supervisions, e.g., image-level tags or captions, is appealing as it removes the costly demand of per-pixel annotations. Despite the rapid progress of RGB-D saliency detection in fully-supervised setting, it however remains an unexplored territory when only weak supervision signals are available. This paper is set to tackle the problem of weakly-supervised RGB-D salient object detection. The key insight in this effort is the idea of maintaining per-pixel pseudo-labels with iterative refinements by reconciling the multimodal input signals in our joint semantic mining (JSM). Considering the large variations in the raw depth map and the lack of explicit pixel-level supervisions, we propose spatial semantic modeling (SSM) to capture saliency-specific depth cues from the raw depth and produce depth-refined pseudo-labels. Moreover, tags and captions are incorporated via a fill-in-the-blank training in our textual semantic modeling (TSM) to estimate the confidences of competing pseudo-labels. At test time, our model involves only a light-weight sub-network of the training pipeline, i.e., it requires only an RGB image as input, thus allowing efficient inference. Extensive evaluations demonstrate the effectiveness of our approach under the weakly-supervised setting. Importantly, our method could also be adapted to work in both fully-supervised and unsupervised paradigms. In each of these scenarios, superior performance has been attained by our approach with comparing to the state-of-the-art dedicated methods. As a by-product, a CapS dataset is constructed by augmenting existing benchmark training set with additional image tags and captions.
Wei Ji 0011, Qi Bi, Miao Zhang 0004, Yongri Piao, Huchuan Lu, Li Cheng 0001
NeurIPS8
2021 EAGLE-Eye: Extreme-pose Action Grader using detaiL bird's-Eye view
abstract
Measuring the quality of a sports action entails attending to the execution of the short-term components as well as overall impression of the whole program. In this assessment, both appearance clues and pose dynamics features should be involved. Current approaches often treat a sports routine as a simple fine-grained action, while taking little heed of its complex temporal structure. Besides, they rely solely on either appearance or pose features to score the performance. In this paper, we present JCA and ADA blocks that are responsible for reasoning about the coordination among the joints and appearance dynamics throughout the performance. We build our two-stream network upon the separate stack of these blocks. The early blocks capture the fine-grained temporal dependencies while the last ones reason about the long-term coarse-grained relations. We further introduce an annotated dataset of sports images with unusual pose configurations to boost the performance of pose estimation in such scenarios. Our experiments show that the proposed method not only outperforms the previous works in short-term action assessment but also is the first to generalize well to minute-long figure-skating scoring.
Mahdiar Nekoui, Fidel Omar Tito Cruz, Li Cheng 0001
WACV3
2021 SparseFusion: Dynamic Human Avatar Modeling From Sparse RGBD Images
abstract
In this paper, we propose a novel approach to reconstruct 3D human body shapes based on a sparse set of RGBD frames using a single RGBD camera. We specifically focus on the realistic settings where human subjects move freely during the capture. The main challenge is how to robustly fuse these sparse frames into a canonical 3D model, under pose changes and surface occlusions. This is addressed by our new framework consisting of the following steps. First, based on a generative human template, for every two frames having sufficient overlap, an initial pairwise alignment is performed; It is followed by a global non-rigid registration procedure, in which partial results from RGBD frames are collected into a unified 3D shape, under the guidance of correspondences from the pairwise alignment; Finally, the texture map of the reconstructed human model is optimized to deliver a clear and spatially consistent texture. Empirical evaluations on synthetic and real datasets demonstrate both quantitatively and qualitatively the superior performance of our framework in reconstructing complete 3D human models with high fidelity. It is worth noting that our framework is flexible, with potential applications going beyond shape reconstruction. As an example, we showcase its use in reshaping and reposing to a new avatar.
Xinxin Zuo, Sen Wang 0003, Jiangbin Zheng 0001, Minglun Gong, Ruigang Yang, Li Cheng 0001
IEEE Trans. Multim.7
2020 Outlier Detection Ensemble with Embedded Feature Selection
abstract
Feature selection places an important role in improving the performance of outlier detection, especially for noisy data. Existing methods usually perform feature selection and outlier scoring separately, which would select feature subsets that may not optimally serve for outlier detection, leading to unsatisfying performance. In this paper, we propose an outlier detection ensemble framework with embedded feature selection (ODEFS), to address this issue. Specifically, for each random sub-sampling based learning component, ODEFS unifies feature selection and outlier detection into a pairwise ranking formulation to learn feature subsets that are tailored for the outlier detection method. Moreover, we adopt the thresholded self-paced learning to simultaneously optimize feature selection and example selection, which is helpful to improve the reliability of the training set. After that, we design an alternate algorithm with proved convergence to solve the resultant optimization problem. In addition, we analyze the generalization error bound of the proposed framework, which provides theoretical guarantee on the method and insightful practical guidance. Comprehensive experimental results on 12 real-world datasets from diverse domains validate the superiority of the proposed ODEFS.
Li Cheng 0001, Yijie Wang 0001, Bin Li 0030
AAAI1
2020 COMET: Context-Aware IoU-Guided Network for Small Object Tracking
Seyed Mojtaba Marvasti-Zadeh, Javad Khaghani, Hossein Ghanei-Yakhdan, Shohreh Kasaei, Li Cheng 0001
ACCV (2)5
2020 Feature Selection on Data Stream via Multi-Cluster Structure Preservation
abstract
The modern data arrive continuously in a rapid and time-varying stream, which appears to generate unstable associations on the data structure. However, most of the existing methods focus on dealing with the static data, and they cannot fully take them into the structure construction. To address this issue, we propose an online unsupervised Feature Selection method via Multi-Cluster structure Preservation (FSMCP for short). FSMCP weighs all features by minimizing the differences between the Multi-Cluster structures in the original and the selected feature space. The structure integrates the three-level associations, i.e., the individual-level associations, the aggregation-level associations, and the streaming-level associations. To provide informative features in time, FSMCP check and update the associations as soon as new instances arrive. In comparison with the baseline methods, FSMCP holds better efficiency than offline methods, while still providing almost similar or even better quantitative feature subsets. It outperforms the existing online methods with average NMI improvement of 10.33%.
Yijie Wang 0001, Li Cheng 0001
CIKM3
2020 3D Human Shape Reconstruction from a Polarization Image
Shihao Zou, Xinxin Zuo, Yiming Qian, Sen Wang 0003, Chi Xu 0002, Minglun Gong, Li Cheng 0001
ECCV (14)7
2020 Action2Motion: Conditioned Generation of 3D Human Motions
abstract
Action recognition is a relatively established task, where given an input sequence of human motion, the goal is to predict its action category. This paper, on the other hand, considers a relatively new problem, which could be thought of as an inverse of action recognition: given a prescribed action type, we aim to generate plausible human motion sequences in 3D. Importantly, the set of generated motions are expected to maintain its diversity to be able to explore the entire action-conditioned motion space; meanwhile, each sampled sequence faithfully resembles a natural human body articulation dynamics. Motivated by these objectives, we follow the physics law of human kinematics by adopting the Lie Algebra theory to represent the natural human motions; we also propose a temporal Variational Auto-Encoder (VAE) that encourages a diverse sampling of the motion space. A new 3D human motion dataset, HumanAct12, is also constructed. Empirical experiments over three distinct human motion datasets (including ours) demonstrate the effectiveness of our approach.
Chuan Guo 0002, Xinxin Zuo, Sen Wang 0003, Shihao Zou, Qingyao Sun, Annan Deng, Minglun Gong, Li Cheng 0001
ACM Multimedia8
2020 An end-to-end distance measuring for mixed data based on deep relevance learning
abstract
Distance Measuring between two mixed data objects is the basis of many learning algorithms. The complex relevance between heterogeneous – various types/scales – attributes has a significant influence on the measured results. In this paper, we propose an End-to-End Distance Measuring method for mixe d data based on deep relevance learning, called E2DM. Existing methods confuse the attributes space by mapping the discrete attribute values to new continuous values, or discretize continuous attributes values without considering the relevance. In contrast, E2DM directly manipulates on the original data with data conversion and relevance learning simultaneously to avoid information loss and attribute space confusion. E2DM firstly estimates internal relevance (i.e., relevance within the attribute) influenced distance by considering the categorical attribute value frequency and mapping numerical attribute values into multiple bins. Then it takes a wrapper approach to iteratively optimize relevance influenced distance and bin boundaries using a Frobenius-norm deviation as its objective function. Co-occurrence Mover’s Distance is proposed to explicitly explore relevance between attributes in each iteration. Finally, the distance for numerical attribute values is refined based on the original values and the fallen bin centers. Experimental results on a number of real-world datasets demonstrate that E2DM outperforms the state-of-the-art methods.
Li Cheng 0001, Yijie Wang 0001, Xingkong Ma
Intell. Data Anal.1
2020 Improving retinal vessel segmentation with joint local loss by matting
He Zhao 0002, Huiqi Li, Li Cheng 0001
Pattern Recognit.3
2019 Towards Natural and Accurate Future Motion Prediction of Humans and Animals
abstract
Anticipating the future motions of 3D articulate objects is challenging due to its non-linear and highly stochastic nature. Current approaches typically represent the skeleton of an articulate object as a set of 3D joints, which unfortunately ignores the relationship between joints, and fails to encode fine-grained anatomical constraints. Moreover, conventional recurrent neural networks, such as LSTM and GRU, are employed to model motion contexts, which inherently have difficulties in capturing long-term dependencies. To address these problems, we propose to explicitly encode anatomical constraints by modeling their skeletons with a Lie algebra representation. Importantly, a hierarchical recurrent network structure is developed to simultaneously encodes local contexts of individual frames and global contexts of the sequence. We proceed to explore the applications of our approach to several distinct quantities including human, fish, and mouse. Extensive experiments show that our approach achieves more natural and accurate predictions over state-of-the-art methods.
Zhenguang Liu, Shuang Wu 0002, Shuyuan Jin, Qi Liu 0049, Shijian Lu, Roger Zimmermann, Li Cheng 0001
CVPR7
2019 Unsupervised Feature Selection via Local Total-Order Preservation
Yijie Wang 0001, Li Cheng 0001
ICANN (2)3
2019 Multi-Hierarchy Attribute Relationship Mining Based Outlier Detection for Categorical Data
abstract
Outlier detection for categorical data is very important in many practical scenarios, such as intrusion detection, fraud detection, early detection of diseases, etc. However, there is no inherent difference measure for categorical data. The differences are hidden in complex attribute value relationships. Existing methods do not properly handle the internal relationship and external relationship of attributes, resulting in low accuracy of outlier detection.This paper proposes a novel unsupervised outlier detection method for categorical data based on Multi-Hierarchy Attribute Relationship Mining (MHARM). It detects outliers by mining the hierarchical and complex relationships between attribute values. MHARM first calculates the internal relationship. It processes each attribute independently via an information-theoretic difference to get an internal distance matrix. Then it handles different subhierarchy of external relationship. It divides attributes into two clusters, using mutual information as the correlation measure. For the external relationship of intra-cluster attributes, it iteratively updates an external distance matrix by using an entropy weighted Earth Mover's Distance (EMD) and the internal distance until convergence; for the external relationship of inter-cluster attributes, the joint entropy weighted sum is obtained to be the whole difference between objects. Finally, MHARM uses the sum of whole difference between objects as the outlier score, sorting it for outlier detection. Experimental results show that MHARM has an average AUC value of 13.84% higher than the state-of-the-art methods and significantly reduced the detection volume multiples (dvM) for unearthing 90% outliers on the given nine data sets.
Yijie Wang 0001, Li Cheng 0001
IJCNN3
2019 A Neural Probabilistic outlier detection method for categorical data
Li Cheng 0001, Yijie Wang 0001, Xingkong Ma
Neurocomputing1
2019 Multivariate Regression with Gross Errors on Manifold-Valued Data
abstract
We consider the topic of multivariate regression on manifold-valued output, that is, for a multivariate observation, its output response lies on a manifold. Moreover, we propose a new regression model to deal with the presence of grossly corrupted manifold-valued responses, a bottleneck issue commonly encountered in practical scenarios. Our model first takes a correction step on the grossly corrupted responses via geodesic curves on the manifold, then performs multivariate linear regression on the corrected data. This results in a nonconvex and nonsmooth optimization problem on Riemannian manifolds. To this end, we propose a dedicated approach named PALMR, by utilizing and extending the proximal alternating linearized minimization techniques for optimization problems on euclidean spaces. Theoretically, we investigate its convergence property, where it is shown to converge to a critical point under mild conditions. Empirically, we test our model on both synthetic and real diffusion tensor imaging data, and show that our model outperforms other multivariate regression models when manifold-valued responses contain gross errors, and is effective in identifying gross errors.
Xiaowei Zhang 0002, Xudong Shi 0005, Yu Sun 0014, Li Cheng 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2019 Supervised Segmentation of Un-Annotated Retinal Fundus Images by Synthesis
abstract
We focus on the practical challenge of segmenting new retinal fundus images that are dissimilar to existing well-annotated data sets. It is addressed in this paper by a supervised learning pipeline, with its core being the construction of a synthetic fundus image data set using the proposed R-sGAN technique. The resulting synthetic images are realistic-looking in terms of the query images while maintaining the annotated vessel structures from the existing data set. This helps to bridge the mismatch between the query images and the existing well-annotated data set. As a consequence, any known supervised fundus segmentation technique can be directly utilized on the query images, after training on this synthetic data set. Extensive experiments on different fundus image data sets demonstrate the competitiveness of the proposed approach in dealing with a diverse range of mismatch settings.
He Zhao 0002, Huiqi Li, Sebastian Maurer-Stroh, Yuhong Guo, Qiuju Deng, Li Cheng 0001
IEEE Trans. Medical Imaging6
2018 Multi-Modal Multi-Task Learning for Automatic Dietary Assessment
abstract
We investigate the task of automatic dietary assessment: given meal images and descriptions uploaded by real users, our task is to automatically rate the meals and deliver advisory comments for improving users' diets. To address this practical yet challenging problem, which is multi-modal and multi-task in nature, an end-to-end neural model is proposed. In particular, comprehensive meal representations are obtained from images, descriptions and user information. We further introduce a novel memory network architecture to store meal representations and reason over the meal representations to support predictions. Results on a real-world dataset show that our method outperforms two strong image captioning baselines significantly.
Qi Liu 0049, Yue Zhang 0004, Zhenguang Liu, Ye Yuan 0001, Li Cheng 0001, Roger Zimmermann
AAAI5
2018 Exploring a High-quality Outlying Feature Value Set for Noise-Resilient Outlier Detection in Categorical Data
abstract
Unavoidable noise in real-world categorical data presents significant challenges to existing outlier detection methods because they normally fail to separate noisy values from outlying values. Feature subspace-based methods inevitably mix noisy values when retaining an entire feature because a feature may contain both outlying values and noisy values. Pattern-based methods are normally based on frequency and are easily misled by noisy values, resulting in many faulty patterns. This paper introduces a novel unsupervised framework termed OUVAS, and its parameter-free instantiation RHAC to explore a high-quality outlying value set for detecting outliers in noisy categorical data. Based on the observation that the relations between values reflect their essence, OUVAS investigates value similarities to cluster values into different groups and combines cluster-level analysis and value-level refinement to identify an outlying value set. RHAC instantiates OUVAS by three successive modules (i.e., the combination of Ochiai coefficient and LOUVAIN algorithm to cluster values, hierarchical value coupling learning to perform cluster-level analysis, and a threshold to divide fake and real outlying values in value-level refinement). We show that (i) RHAC-based outlier detector significantly outperforms five state-of-the-art outlier detection methods; (ii) Extended RHAC-based feature selection method successfully improves the performance of existing outlier detectors and performs better than two latest outlying feature selection methods.
Hongzuo Xu, Li Cheng 0001, Yijie Wang 0001, Xingkong Ma
CIKM3
2018 FROD: Fast and Robust Distance-Based Outlier Detection with Active-Inliers-Patterns in Data Streams
Zongren Li, Yijie Wang 0001, Guohong Zhao, Li Cheng 0001, Xingkong Ma
ICANN (1)4
2018 Correction to: Lie-X: Depth Image Based Articulated Object Pose Estimation, Tracking, and Action Recognition on Lie Groups
Chi Xu 0002, Lakshmi Narasimhan Govindarajan, Yu Zhang 0004, Zoe Bichler, Suresh Jesuthasan, Adam Claridge-Chang, Ajay Sriram Mathuru, Wenlong Tang, Peixin Zhu, Li Cheng 0001
Int. J. Comput. Vis.11
2018 Synthesizing retinal and neuronal images with generative adversarial nets
abstract
This paper aims at synthesizing multiple realistic-looking retinal (or neuronal) images from an unseen tubular structured annotation that contains the binary vessel (or neuronal) morphology. The generated phantoms are expected to preserve the same tubular structure, and resemble the visual appearance of the training images. Inspired by the recent progresses in generative adversarial nets (GANs) as well as image style transfer, our approach enjoys several advantages. It works well with a small training set with as few as 10 training examples, which is a common scenario in medical image analysis. Besides, it is capable of synthesizing diverse images from the same tubular structured annotation. Extensive experimental evaluations on various retinal fundus and neuronal imaging applications demonstrate the merits of the proposed approach.
He Zhao 0002, Huiqi Li, Sebastian Maurer-Stroh, Li Cheng 0001
Medical Image Anal.4
2018 Transduction on Directed Graphs via Absorbing Random Walks
abstract
In this paper we consider the problem of graph-based transductive classification, and we are particularly interested in the directed graph scenario which is a natural form for many real world applications. Different from existing research efforts that either only deal with undirected graphs or circumvent directionality by means of symmetrization, we propose a novel random walk approach on directed graphs using absorbing Markov chains, which can be regarded as maximizing the accumulated expected number of visits from the unlabeled transient states. Our algorithm is simple, easy to implement, and works with large-scale graphs on binary, multiclass, and multi-label prediction problems. Moreover, it is capable of preserving the graph structure even when the input graph is sparse and changes over time, as well as retaining weak signals presented in the directed edges. We present its intimate connections to a number of existing methods, including graph kernels, graph Laplacian based methods, and spanning forest of graphs. Its computational complexity and the generalization error are also studied. Empirically, our algorithm is evaluated on a wide range of applications, where it has shown to perform competitively comparing to a suite of state-of-the-art methods. In particular, our algorithm is shown to work exceptionally well with large sparse directed graphs with e.g., millions of nodes and tens of millions of edges, where it significantly outperforms other state-of-the-art methods. In the dynamic graph setting involving insertion or deletion of nodes and edge-weight changes over time, it also allows efficient online updates that produce the same results as of the batch update counterparts.
Jaydeep De, Xiaowei Zhang 0002, Feng Lin 0002, Li Cheng 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2018 Too Far to See? Not Really! - Pedestrian Detection With Scale-Aware Localization Policy
abstract
A major bottleneck of pedestrian detection lies on the sharp performance deterioration in the presence of small-size pedestrians that are relatively far from the camera. Motivated by the observation that pedestrians of disparate spatial scales exhibit distinct visual appearances, we propose in this paper an active pedestrian detector that explicitly operates over multiple-layer neuronal representations of the input still image. More specifically, convolutional neural nets, such as ResNet and faster R-CNNs, are exploited to provide a rich and discriminative hierarchy of feature representations, as well as initial pedestrian proposals. Here each pedestrian observation of distinct size could be best characterized in terms of the ResNet feature representation at a certain layer of the hierarchy. Meanwhile, initial pedestrian proposals are attained by the faster R-CNNs techniques, i.e., region proposal network and follow-up region of interesting pooling layer employed right after the specific ResNet convolutional layer of interest, to produce joint predictions on the bounding-box proposals' locations and categories (i.e., pedestrian or not). This is engaged as an input to our active detector, where for each initial pedestrian proposal, a sequence of coordinate transformation actions is carried out to determine its proper x-y 2D location and the layer of feature representation, or eventually terminated as being background. Empirically our approach is demonstrated to produce overall lower detection errors on widely used benchmarks, and it works particularly well with far-scale pedestrians. For example, compared with 60.51% log-average miss rate of the state-of-the-art MS-CNN for far-scale pedestrians (those below 80 pixels in bounding-box height) of the Caltech benchmark, the miss rate of our approach is 41.85%, with a notable reduction of 18.66%.
Xiaowei Zhang 0003, Li Cheng 0001, Bo Li 0006, Hai-Miao Hu
IEEE Trans. Image Process.2
2017 TMRCP: A Trend-Matching Resources Coupled Prediction Method over Data Stream
Runfan Wu, Yijie Wang 0001, Xingkong Ma, Li Cheng 0001
ICONIP (5)4
2017 Multiview and Multimodal Pervasive Indoor Localization
abstract
Pervasive indoor localization (PIL) aims to locate an indoor mobile-phone user without any infrastructure assistance. Conventional PIL approaches employ a single probe (i.e., target) measurement to localize by identifying its best match out of a fingerprint gallery. However, a single measurement usually captures limited and inadequate location features. More importantly, the reliance on a single measurement bears the inherent risk of being inaccurate and unreliable, due to the fact that the measurement could be noisy and even corrupted.
Zhenguang Liu, Li Cheng 0001, Anan Liu, Xiangnan He 0001, Roger Zimmermann
ACM Multimedia2
2017 Lie-X: Depth Image Based Articulated Object Pose Estimation, Tracking, and Action Recognition on Lie Groups
Chi Xu 0002, Lakshmi Narasimhan Govindarajan, Yu Zhang 0004, Li Cheng 0001
Int. J. Comput. Vis.4
2017 Pose Estimation from Line Correspondences: A Complete Analysis and a Series of Solutions
abstract
In this paper we deal with the camera pose estimation problem from a set of 2D/3D line correspondences, which is also known as PnL (Perspective-n-Line) problem. We carry out our study by comparing PnL with the well-studied PnP (Perspective-n-Point) problem, and our contributions are three-fold: (1) We provide a complete 3D configuration analysis for P3L, which includes the well-known P3P problem as well as several existing analyses as special cases. (2) By exploring the similarity between PnL and PnP, we propose a new subset-based PnL approach as well as a series of linear-formulation-based PnL approaches inspired by their PnP counterparts. (3) The proposed linear-formulation-based methods can be easily extended to deal with the line and point features simultaneously.
Chi Xu 0002, Lilian Zhang, Li Cheng 0001, Reinhard Koch
IEEE Trans. Pattern Anal. Mach. Intell.3
2017 Hand action detection from ego-centric depth sequences with error-correcting Hough transform
Chi Xu 0002, Lakshmi Narasimhan Govindarajan, Li Cheng 0001
Pattern Recognit.3
2017 Segment 2D and 3D Filaments by Learning Structured and Contextual Features
abstract
We focus on the challenging problem of filamentary structure segmentation in both 2D and 3D images, including retinal vessels and neurons, among others. Despite the increasing amount of efforts in learning based methods to tackle this problem, there still lack proper data-driven feature construction mechanisms to sufficiently encode contextual labelling information, which might hinder the segmentation performance. This observation prompts us to propose a data-driven approach to learn structured and contextual features in this paper. The structured features aim to integrate local spatial label patterns into the feature space, thus endowing the follow-up tree classifiers capability to grouping training examples with similar structure into the same leaf node when splitting the feature space, and further yielding contextual features to capture more of the global contextual information. Empirical evaluations demonstrate that our approach outperforms state-of-the-arts on well-regarded testbeds over a variety of applications. Our code is also made publicly available in support of the open-source research activities.
Lin Gu 0003, Xiaowei Zhang 0002, He Zhao 0002, Huiqi Li, Li Cheng 0001
IEEE Trans. Medical Imaging5
2017 Fusion of Magnetic and Visual Sensors for Indoor Localization: Infrastructure-Free and More Effective
abstract
Accurate and infrastructure-free indoor positioning can be very useful in a variety of applications. However, most existing approaches (e.g., WiFi and infrared-based methods) for indoor localization heavily rely on infrastructure, which is neither scalable nor pervasively available. In this paper, we propose a novel indoor localization and tracking approach, termed VMag, that does not require any infrastructure assistance. The user can be localized while simply holding a smartphone. To the best of our knowledge, the proposed method is the first exploration of fusing geomagnetic and visual sensing for indoor localization. More specifically, we conduct an in-depth study on both the advantageous properties and the challenges in leveraging the geomagnetic field and visual images for indoor localization. Based on these studies, we design a context-aware particle filtering framework to track the user with the goal of maximizing the positioning accuracy. We also introduce a neural-network-based method to extract deep features for the purpose of indoor positioning. We have conducted extensive experiments on four different indoor settings including a laboratory, a garage, a canteen, and an office building. Experimental results demonstrate the superior performance of VMag over the state of the art with these four indoor settings.
Zhenguang Liu, Qi Liu 0049, Yifang Yin, Li Cheng 0001, Roger Zimmermann
IEEE Trans. Multim.5
2016 Recognizing Complex Activities by a Probabilistic Interval-Based Model
abstract
A key challenge in complex activity recognition is the fact that a complex activity can often be performed in several different ways, with each consisting of its own configuration of atomic actions and their temporal dependencies. This leads us to define an atomic activity-based probabilistic framework that employs Allen's interval relations to represent local temporal dependencies. The framework introduces a latent variable from the Chinese Restaurant Process to explicitly characterize these unique internal configurations of a particular complex activity as a variable number of tables.It can be analytically shown that the resulting interval network satisfies the transitivity property, and as a result, all local temporal dependencies can be retained and are globally consistent.Empirical evaluations on benchmark datasets suggest our approach significantly outperforms the state-of-the-art methods.
Li Liu 0001, Li Cheng 0001, Ye Liu 0002, Yongpo Jia, David S. Rosenblum
AAAI2
2016 Estimate Hand Poses Efficiently from Single Depth Images
abstract
This paper aims to tackle the practically very challenging problem of efficient and accurate hand pose estimation from single depth images. A dedicated two-step regression forest pipeline is proposed: given an input hand depth image, step one involves mainly estimation of 3D location and in-plane rotation of the hand using a pixel-wise regression forest. This is utilized in step two which delivers final hand estimation by a similar regression forest model based on the entire hand image patch. Moreover, our estimation is guided by internally executing a 3D hand kinematic chain model. For an unseen test image, the kinematic model parameters are estimated by a proposed dynamically weighted scheme. As a combined effect of these proposed building blocks, our approach is able to deliver more precise estimation of hand poses. In practice, our approach works at 15.6 frame-per-second (FPS) on an average laptop when implemented in CPU, which is further sped-up to 67.2 FPS when running on GPU. In addition, we introduce and make publicly available a data-glove annotated depth image dataset covering various hand shapes and gestures, which enables us conducting quantitative analyses on real-world hand images. The effectiveness of our approach is verified empirically on both synthetic and the annotated real-world datasets for hand pose estimation, as well as related applications including part-based labeling and gesture classification. In addition to empirical studies, the consistency property of our approach is also theoretically analyzed.
Chi Xu 0002, Ashwin Nanjappa, Xiaowei Zhang 0002, Li Cheng 0001
Int. J. Comput. Vis.4
2016 Action Recognition in Still Images With Minimum Annotation Efforts
abstract
We focus on the problem of still image-based human action recognition, which essentially involves making prediction by analyzing human poses and their interaction with objects in the scene. Besides image-level action labels (e.g., riding, phoning), during both training and testing stages, existing works usually require additional input of human bounding boxes to facilitate the characterization of the underlying human-object interactions. We argue that this additional input requirement might severely discourage potential applications and is not very necessary. To this end, a systematic approach was developed in this paper to address this challenging problem of minimum annotation efforts, i.e., to perform recognition in the presence of only image-level action labels in the training stage. Experimental results on three benchmark data sets demonstrate that compared with the state-of-the-art methods that have privileged access to additional human bounding-box annotations, our approach achieves comparable or even superior recognition accuracy using only action annotations in training. Interestingly, as a by-product in many cases, our approach is able to segment out the precise regions of underlying human-object interactions.
Yu Zhang 0004, Li Cheng 0001, Jianxin Wu 0001, Jianfei Cai 0001, Minh N. Do, Jiangbo Lu
IEEE Trans. Image Process.2
2016 A Graph-Theoretical Approach for Tracing Filamentary Structures in Neuronal and Retinal Images
abstract
The aim of this study is about tracing filamentary structures in both neuronal and retinal images. It is often crucial to identify single neurons in neuronal networks, or separate vessel tree structures in retinal blood vessel networks, in applications such as drug screening for neurological disorders or computer-aided diagnosis of diabetic retinopathy. Both tasks are challenging as the same bottleneck issue of filament crossovers is commonly encountered, which essentially hinders the ability of existing systems to conduct large-scale drug screening or practical clinical usage. To address the filament crossovers' problem, a two-step graph-theoretical approach is proposed in this paper. The first step focuses on segmenting filamentary pixels out of the background. This produces a filament segmentation map used as input for the second step, where they are further separated into disjointed filaments. Key to our approach is the idea that the problem can be reformulated as label propagation over directed graphs, such that the graph is to be partitioned into disjoint sub-graphs, or equivalently, each of the neurons (vessel trees) is separated from the rest of the neuronal (vessel) network. This enables us to make the interesting connection between the tracing problem and the digraph matrix-forest theorem in algebraic graph theory for the first time. Empirical experiments on neuronal and retinal image datasets demonstrate the superior performance of our approach over existing methods.
Jaydeep De, Li Cheng 0001, Xiaowei Zhang 0002, Feng Lin 0002, Huiqi Li, Ong Kok Haur, Weimiao Yu, Yuanhong Yu 0002, Sohail Ahmed
IEEE Trans. Medical Imaging2
2015 Robust Multivariate Regression with Grossly Corrupted Observations and Its Application to Personality Prediction
Xiaowei Zhang 0002, Li Cheng 0001, Tingshao Zhu
ACML2
2015 Integrated Foreground Segmentation and Boundary Matting for Live Videos
abstract
The objective of foreground segmentation is to extract the desired foreground object from input videos. Over the years, there have been significant amount of efforts on this topic. Nevertheless, there still lacks a simple yet effective algorithm that can process live videos of objects with fuzzy boundaries (e.g., hair) captured by freely moving cameras. This paper presents an algorithm toward this goal. The key idea is to train and maintain two competing one-class support vector machines at each pixel location, which model local color distributions for both foreground and background, respectively. The usage of two competing local classifiers, as we have advocated, provides higher discriminative power while allowing better handling of ambiguities. By exploiting this proposed machine learning technique, and by addressing both foreground segmentation and boundary matting problems in an integrated manner, our algorithm is shown to be particularly competent at processing a wide range of videos with complex backgrounds from freely moving cameras. This is usually achieved with minimum user interactions. Furthermore, by introducing novel acceleration techniques and by exploiting the parallel structure of the algorithm, near real-time processing speed (14 frames/s without matting and 8 frames/s with matting on a midrange PC & GPU) is achieved for VGA-sized videos.
Minglun Gong, Yiming Qian, Li Cheng 0001
IEEE Trans. Image Process.3
2014 Tracing Retinal Blood Vessels by Matrix-Forest Theorem of Directed Graphs
Li Cheng 0001, Jaydeep De, Xiaowei Zhang 0002, Feng Lin 0002, Huiqi Li
MICCAI (1)1
2014 Tracing retinal vessel trees by transductive inference
abstract
BACKGROUND: Structural study of retinal blood vessels provides an early indication of diseases such as diabetic retinopathy, glaucoma, and hypertensive retinopathy. These studies require accurate tracing of retinal vessel tree structure from fundus images in an automated manner. However, the existing work encounters great difficulties when dealing with the crossover issue commonly-seen in vessel networks. RESULTS: In this paper, we consider a novel graph-based approach to address this tracing with crossover problem: After initial steps of segmentation and skeleton extraction, its graph representation can be established, where each segment in the skeleton map becomes a node, and a direct contact between two adjacent segments is translated to an undirected edge of the two corresponding nodes. The segments in the skeleton map touching the optical disk area are considered as root nodes. This determines the number of trees to-be-found in the vessel network, which is always equal to the number of root nodes. Based on this undirected graph representation, the tracing problem is further connected to the well-studied transductive inference in machine learning, where the goal becomes that of properly propagating the tree labels from those known root nodes to the rest of the graph, such that the graph is partitioned into disjoint sub-graphs, or equivalently, each of the trees is traced and separated from the rest of the vessel network. This connection enables us to address the tracing problem by exploiting established development in transductive inference. Empirical experiments on public available fundus image datasets demonstrate the applicability of our approach. CONCLUSIONS: We provide a novel and systematic approach to trace retinal vessel trees with the present of crossovers by solving a transductive learning problem on induced undirected graphs.
Jaydeep De, Huiqi Li, Li Cheng 0001
BMC Bioinform.3
2014 Recognizing flu-like symptoms from videos
abstract
BACKGROUND: Vision-based surveillance and monitoring is a potential alternative for early detection of respiratory disease outbreaks in urban areas complementing molecular diagnostics and hospital and doctor visit-based alert systems. Visible actions representing typical flu-like symptoms include sneeze and cough that are associated with changing patterns of hand to head distances, among others. The technical difficulties lie in the high complexity and large variation of those actions as well as numerous similar background actions such as scratching head, cell phone use, eating, drinking and so on. RESULTS: In this paper, we make a first attempt at the challenging problem of recognizing flu-like symptoms from videos. Since there was no related dataset available, we created a new public health dataset for action recognition that includes two major flu-like symptom related actions (sneeze and cough) and a number of background actions. We also developed a suitable novel algorithm by introducing two types of Action Matching Kernels, where both types aim to integrate two aspects of local features, namely the space-time layout and the Bag-of-Words representations. In particular, we show that the Pyramid Match Kernel and Spatial Pyramid Matching are both special cases of our proposed kernels. Besides experimenting on standard testbed, the proposed algorithm is evaluated also on the new sneeze and cough set. Empirically, we observe that our approach achieves competitive performance compared to the state-of-the-arts, while recognition on the new public health dataset is shown to be a non-trivial task even with simple single person unobstructed view. CONCLUSIONS: Our sneeze and cough video dataset and newly developed action recognition algorithm is the first of its kind and aims to kick-start the field of action recognition of flu-like symptoms from videos. It will be challenging but necessary in future developments to consider more complex real-life scenario of detecting these actions simultaneously from multiple persons in possibly crowded environments.
Tuan Hue Thi, Li Wang 0033, Jian Zhang 0002, Sebastian Maurer-Stroh, Li Cheng 0001
BMC Bioinform.6
2014 Semi-supervised Domain Adaptation on Manifolds
abstract
In real-life problems, the following semi-supervised domain adaptation scenario is often encountered: we have full access to some source data, which is usually very large; the target data distribution is under certain unknown transformation of the source data distribution; meanwhile, only a small fraction of the target instances come with labels. The goal is to learn a prediction model by incorporating information from the source domain that is able to generalize well on the target test instances. We consider an explicit form of transformation functions and especially linear transformations that maps examples from the source to the target domain, and we argue that by proper preprocessing of the data from both source and target domains, the feasible transformation functions can be characterized by a set of rotation matrices. This naturally leads to an optimization formulation under the special orthogonal group constraints. We present an iterative coordinate descent solver that is able to jointly learn the transformation as well as the model parameters, while the geodesic update ensures the manifold constraints are always satisfied. Our framework is sufficiently general to work with a variety of loss functions and prediction problems. Empirical evaluations on synthetic and real-world experiments demonstrate the competitive performance of our method with respect to the state-of-the-art.
Li Cheng 0001, Sinno Jialin Pan
IEEE Trans. Neural Networks Learn. Syst.1
2013 Efficient Hand Pose Estimation from a Single Depth Image
abstract
We tackle the practical problem of hand pose estimation from a single noisy depth image. A dedicated three-step pipeline is proposed: Initial estimation step provides an initial estimation of the hand in-plane orientation and 3D location, Candidate generation step produces a set of 3D pose candidate from the Hough voting space with the help of the rotational invariant depth features, Verification step delivers the final 3D hand pose as the solution to an optimization problem. We analyze the depth noises, and suggest tips to minimize their negative impacts on the overall performance. Our approach is able to work with Kinect-type noisy depth images, and reliably produces pose estimations of general motions efficiently (12 frames per second). Extensive experiments are conducted to qualitatively and quantitatively evaluate the performance with respect to the state-of-the-art methods that have access to additional RGB images. Our approach is shown to deliver on par or even better results.
Chi Xu 0002, Li Cheng 0001
ICCV2
2013 Finding Distinctive Shape Features for Automatic Hematoma Classification in Head CT Images from Traumatic Brain Injuries
abstract
Computer aided diagnosis (CAD) in medical imaging is of growing interest in recent years. Our proposed CAD system aims to enhance diagnosis and prognosis of traumatic brain injury (TBI) patients with hematomas. Hematoma caused by blood vessel rupture is the major lesion in TBI cases and is usually assessed using head computed tomography (CT). In our CAD system, we segment the hematoma region from each slice of a CT series, extract features from the hematoma segments, and automatically classify the hematoma types using machine learning methods. We propose two sets of shape based features for each segmented hematoma region. The first set contains primitive features describing the overall shape of a hematoma region. The features in the second set are based on the dissimilarities of the shapes of hematoma regions measured by geodesic distances. After feature extraction, we classify the hematoma regions into three types -- epidural hematoma, sub-dural hematoma, and intracerebral hematoma, using random forest. Each tree of the random forest votes one class for each hematoma, and the random forest takes the class label with the majority votes for the hematoma. As hematomas are volumetric in nature, some hematomas are observed across several consecutive slices in the same CT series. For each class, we add the votes from each hematoma slice that comprises the volumetric hematoma in that class, then we take the class with the majority of the summed votes as the class label for that volumetric hematoma. The overall classification accuracies for hematoma region from each CT slice are 80.7%, 81.3%, and 81.1% using primitive features only, geodesic distance features only, or both sets of features, respectively. For volumetric hematoma classification, the overall accuracies are 80.9%, 81.5%, and 81.5% respectively. The results are promising to radiologists and neurosurgeons specialized in this field of research.
Tianxia Gong, Nengli Lim, Li Cheng 0001, Hwee Kuan Lee, Bolan Su, Chew Lim Tan, Shimiao Li, C. C. Tchoyoson Lim, Boon Chuan Pang, Cheng Kiang Lee
ICTAI3
2013 Exploiting Syntactic, Semantic, and Lexical Regularities in Language Modeling via Directed Markov Random Fields
abstract
We present a directed Markov random field (MRF) model that combinesn‐gram models, probabilistic context‐free grammars (PCFGs), and probabilistic latent semantic analysis (PLSA) for the purpose of statistical language modeling. Even though the composite directed MRF model potentially has an exponential number of loops and becomes a context‐sensitive grammar, we are nevertheless able to estimate its parameters in cubic time using an efficient modified Expectation‐Maximization (EM) method,the generalized inside–outside algorithm, which extends the inside–outside algorithm to incorporate the effects of then‐gram and PLSA language models. We generalize various smoothing techniques to alleviate the sparseness ofn‐gram counts in cases where there are hidden variables. We also derive an analogous algorithm to find the most likely parse of a sentence and to calculate the probability of initial subsequence of a sentence, all generated by the composite language model. Our experimental results on theWall Street Journalcorpus show that we obtain significant reductions in perplexity compared to the state‐of‐the‐art baseline trigram model with Good–Turing and Kneser–Ney smoothing techniques.
Shaomin Wang, Li Cheng 0001, Russell Greiner, Dale Schuurmans
Comput. Intell.3
2013 Machine learning in motion analysis: New advances
Matti Pietikäinen, Matthew Turk 0001, Liang Wang 0001, Guoying Zhao 0001, Li Cheng 0001
Image Vis. Comput.5
2012 Integrating local action elements for action analysis
Tuan Hue Thi, Li Cheng 0001, Jian Zhang 0002, Li Wang 0033, Shin'ichi Satoh 0001
Comput. Vis. Image Underst.2
2012 Structured learning of local features for human action classification and localization
Tuan Hue Thi, Li Cheng 0001, Jian Zhang 0002, Li Wang 0033, Shin'ichi Satoh 0001
Image Vis. Comput.2
2011 Incorporating estimated motion in real-time background subtraction
abstract
Many existing background subtraction approaches model background color only and detect foreground as outliers, and hence may confuse background changes or noises with true foreground. We present a novel algorithm that utilizes motion cues computed from an optical flow algorithm. The additional motion information allows aligning moving foreground objects over time so that models can be built for foreground as well. It also facilities background (and foreground) modeling since both color and motion cues can be utilized. In practice, our GPU implementation is able to process QVGA-sized video sequences at 39.3 FPS on a laptop. Quantitative evaluation on standard testbeds demonstrate the competitive performance of our approach.
Minglun Gong, Li Cheng 0001
ICIP2
2011 Discriminative Segmentation of Microscopic Cellular Images
Li Cheng 0001, Weimiao Yu, Andre Cheah
MICCAI (1)1
2011 Human Action Segmentation and Recognition Using Discriminative Semi-Markov Models
Qinfeng Shi, Li Cheng 0001, Li Wang 0033, Alexander J. Smola
Int. J. Comput. Vis.2
2011 Real-Time Discriminative Background Subtraction
abstract
The authors examine the problem of segmenting foreground objects in live video when background scene textures change over time. In particular, we formulate background subtraction as minimizing a penalized instantaneous risk functional--yielding a local online discriminative algorithm that can quickly adapt to temporal changes. We analyze the algorithm's convergence, discuss its robustness to nonstationarity, and provide an efficient nonlinear extension via sparse kernels. To accommodate interactions among neighboring pixels, a global algorithm is then derived that explicitly distinguishes objects versus background using maximum a posteriori inference in a Markov random field (implemented via graph-cuts). By exploiting the parallel nature of the proposed algorithms, we develop an implementation that can run efficiently on the highly parallel graphics processing unit (GPU). Empirical studies on a wide variety of datasets demonstrate that the proposed approach achieves quality that is comparable to state-of-the-art offline methods, while still being suitable for real-time video analysis ( ≥ 75 fps on a mid-range GPU).
Li Cheng 0001, Minglun Gong, Dale Schuurmans, Terry Caelli
IEEE Trans. Image Process.1
2011 Elastic Sequence Correlation for Human Action Analysis
abstract
This paper addresses the problem of automatically analyzing and understanding human actions from video footage. An "action correlation" framework, elastic sequence correlation (ESC), is proposed to identify action subsequences from a database of (possibly long) video sequences that are similar to a given query video action clip. In particular, we show that two well-known algorithms, namely approximate pattern matching in computer and information sciences and dynamic time warping (DTW) method in signal processing, are special cases of our ESC framework. The proposed framework is applied to two important real-world applications: action pattern retrieval, as well as action segmentation and recognition, where, on average, its run time speed (in matlab) is about 3.3 frames per second. In addition, comparing with the state-of-the-art algorithms on a number of challenging data sets, our approach is demonstrated to perform competitively.
Li Wang 0033, Li Cheng 0001, Liang Wang 0001
IEEE Trans. Image Process.2
2010 Human Action Recognition and Localization in Video Using Structured Learning of Local Space-Time Features
abstract
This paper presents a unified framework for human action classification and localization in video using structured learning of local space-time features. Each human action class is represented by a set of its own compact set of local patches. In our approach, we first use a discriminative hierarchical Bayesian classifier to select those space-time interest points that are constructive for each particular action. Those concise local features are then passed to a Support Vector Machine with Principal Component Analysis projection for the classification task. Meanwhile, the action localization is done using Dynamic Conditional Random Fields developed to incorporate the spatial and temporal structure constraints of superpixels extracted around those features. Each superpixel in the video is defined by the shape and motion information of its corresponding feature region. Compelling results obtained from experiments on KTH [22], Weizmann [1], HOHA [13] and TRECVid [23] datasets have proven the efficiency and robustness of our framework for the task of human action recognition and localization in video.
Tuan Hue Thi, Jian Zhang 0002, Li Cheng 0001, Li Wang 0033, Shin'ichi Satoh 0001
AVSS3
2010 Implicit Motion-Shape Model: A generic approach for action matching
abstract
We develop a robust technique to find similar matches of human actions in video. Given a query video, Motion History Images (MHI) are constructed for consecutive keyframes. This is followed by dividing the MHI into local Motion-Shape regions, which allows us to analyze the action as a set of sparse space-time patches in 3D. Inspired by the idea of Generalized Hough Transform, we develop the Implicit Motion-Shape Model that allows the integration of these local patches to describe the dynamic characteristics of the query action. In the same way we retrieve motion segments from video candidates, then project them onto the Hough Space built by the query model. This produces the matching score by running Parzen window density estimation under different scales. Empirical experiments on popular datasets demonstrate the efficiency of this approach, where highly accurate matches are returned within acceptable processing time.
Tuan Hue Thi, Li Cheng 0001, Jian Zhang 0002, Li Wang 0033
ICIP2
2010 Efficient Learning to Label Images
abstract
Conditional random field methods (CRFs) have gained popularity for image labeling tasks in recent years. In this paper, we describe an alternative discriminative approach, by extending the large margin principle to incorporate spatial correlations among neighboring pixels. In particular, by explicitly enforcing the sub modular condition, graph-cuts is conveniently integrated as the inference engine to attain the optimal label assignment efficiently. Our approach allows learning a model with thousands of parameters, and is shown to be capable of readily incorporating higher-order scene context. Empirical studies on a variety of image datasets suggest that our approach performs competitively compared to the state-of-the-art scene labeling methods.
Ke Jia, Li Cheng 0001, Nianjun Liu, Lei Wang 0001
ICPR2
2010 Weakly Supervised Action Recognition Using Implicit Shape Models
abstract
In this paper, we present a robust framework for action recognition in video, that is able to perform competitively against the state-of-the-art methods, yet does not rely on sophisticated background subtraction preprocess to remove background features. In particular, we extend the Implicit Shape Modeling (ISM) of [10] for object recognition to 3D to integrate local spatiotemporal features, which are produced by a weakly supervised Bayesian kernel filter. Experiments on benchmark datasets (including KTH and Weizmann) verifies the effectiveness of our approach.
Tuan Hue Thi, Li Cheng 0001, Jian Zhang 0002, Li Wang 0033, Shin'ichi Satoh 0001
ICPR2
2009 Human Body Articulation for Action Recognition in Video Sequences
abstract
This paper presents a new technique for action recognition in video using human body part-based approach, combining both local feature description of each body part, and global graphical model structure of the human action. The human body is divided into elementary points from which a Decomposable Triangulated Graph will be built. The temporal variation of human activity is encoded in the velocity distribution of each node in the graph, while the graph structure shows the spatial configuration of all the nodes in the action. Tracking trajectories of unlabeled good feature points are correctly labeled using Maximum a Posterior probability. Dynamic Programming is then implemented to boost up the exhaustive search for the optimal labeling of unknown body parts and the best possible action. A simple and efficient technique for building the optimal structure of the human action graph is also implemented. Experimental results on the KTH dataset proves the success and potential applications of this proposed technique.
Tuan Hue Thi, Sijun Lu, Jian Zhang 0002, Li Cheng 0001, Li Wang 0033
AVSS4
2009 Inference of the structural credit risk model using MLE
abstract
Credit risk analysis is not only an important research topic in finance, but also of interest in everyday life. Unfortunately, the non-linear nature of the widely accepted Black-Scholes option price model, which sits at the very heart of the structural credit risk model, causes great difficulty when inferring the latent asset value sequence from observed data. The main contribution of this paper is to address this problem by pursuing maximum likelihood state estimation (MLE) instead of the usual particle filtering approach. Experiments demonstrate the competitiveness of the proposed MLE approach: it achieves a much lower inference error and a much lower running time than particle filtering methods. This work has merit for the general problem of inferring latent values for probabilistic time-series.
Li Cheng 0001, Dale Schuurmans
CIFEr2
2009 Realtime background subtraction from dynamic scenes
abstract
This paper examines the problem of moving object detection. More precisely, it addresses the difficult scenarios where background scene textures in the video might change over time. In this paper, we formulate the problem mathematically as minimizing a constrained risk functional motivated from the large margin principle. It is a generalization of the one class support vector machines (1-SVMs) to accommodate spatial interactions, which is further incorporated into an online learning framework to track temporal changes. As a result it yields a closed-form update formula, a central component of the proposed algorithm to enable prompt adaptation to spatio-temporal changes. We also analyze the mistake bound and discuss issues such as dealing with non-stationary distributions, making use of kernels and efficient inference by a variant of dynamic programming. By exploiting the inherently concurrent structure, the proposed approach is designed to work with the highly parallel graphics processors (GPUs) to facilitate realtime analysis. Our empirical study demonstrates that the proposed approach works in realtime (over 80 frames per second) and at the same time performs competitively against state-of-the-art offline and quasi-realtime methods.
Li Cheng 0001, Minglun Gong
ICCV1
2009 Discriminative Maximum Margin Image Object Categorization with Exact Inference
abstract
Categorizing multiple objects in images is essentially a structured prediction problem: the label of an object is in general dependent on the labels of other objects in the image. We explicitly model object dependencies in a sparse graphical topology induced by the adjacency of objects in the image, which benefits inference, and then use maximum margin principle to learn the model discriminatively. Moreover, we propose a novel exact inference method, which is used in training to find the most violated constraint required by cutting plane method. A slightly modified inference method is used in testing when the target labels are unseen. Experiment results on both synthetic and real datasets demonstrate the improvement of the proposed approach over the state-of-the-art methods.
Qinfeng Shi, Luping Zhou, Li Cheng 0001, Dale Schuurmans
ICIG3
2009 Learning-based multiview video coding
abstract
In the past decade, machine learning techniques have made great progress. Inspired by the recent advancement on semi-supervised learning techniques, we propose a novel learning-based multiview video compression framework. Our scheme can efficiently compress the multiview video represented by multiview-video-plus-depth (MVD) format.We model the multiview video compression problem as a semi-supervised learning problem and design sophisticated mechanisms to achieve high compression efficiency. Our approach is significantly different from the traditional hybrid coding scheme such as H.264-based multiview video coding methods. The preliminary results show promising compression performance.
Baochun Bai, Li Cheng 0001, Pierre Boulanger, Janelle J. Harms
PCS2
2009 Learning Graph Matching
abstract
As a fundamental problem in pattern recognition, graph matching has applications in a variety of fields, from computer vision to computational biology. In graph matching, patterns are modeled as graphs and pattern recognition amounts to finding a correspondence between the nodes of different graphs. Many formulations of this problem can be cast in general as a quadratic assignment problem, where a linear term in the objective function encodes node compatibility and a quadratic term encodes edge compatibility. The main research focus in this theme is about designing efficient algorithms for approximately solving the quadratic assignment problem, since it is NP-hard. In this paper we turn our attention to a different question: how to estimate compatibility functions such that the solution of the resulting graph matching problem best matches the expected solution that a human would manually provide. We present a method for learning graph matching: the training examples are pairs of graphs and the 'labels' are matches between them. Our experimental results reveal that learning can substantially improve the performance of standard graph matching algorithms. In particular, we find that simple linear assignment with such a learning scheme outperforms Graduated Assignment with bistochastic normalisation, a state-of-the-art quadratic assignment relaxation algorithm.
Tibério S. Caetano, Julian J. McAuley, Li Cheng 0001, Quoc V. Le, Alexander J. Smola
IEEE Trans. Pattern Anal. Mach. Intell.3
2008 Prediction and Change Detection in Sequential Data for Interactive Applications
Jun Zhou 0001, Li Cheng 0001, Walter F. Bischof
AAAI2
2008 Consistent image analogies using semi-supervised learning
abstract
In this paper we study the following problem: given two source images A and Apsila, and a target image B, can we learn to synthesize a new image Bpsila which relates to B in the same way that Apsila relates to A? We propose an algorithm which a) uses a semi-supervised component to exploit the fact that the target image B is available apriori, b) uses inference on a Markov random field (MRF) to ensure global consistency, and c) uses image quilting to ensure local consistency. Our algorithm can also deal with the case when A is only partially labeled, that is, only small parts of Apsila are available for training. Empirical evaluation shows that our algorithm consistently produces visually pleasing results, outperforming the state of the art.
Li Cheng 0001, S. V. N. Vishwanathan
CVPR1
2008 Discriminative human action segmentation and recognition using semi-Markov model
abstract
Given an input video sequence of one person conducting a sequence of continuous actions, we consider the problem of jointly segmenting and recognizing actions. We propose a discriminative approach to this problem under a semi-Markov model framework, where we are able to define a set of features over input-output space that captures the characteristics on boundary frames, action segments and neighboring action segments, respectively. In addition, we show that this method can also be used to recognize the person who performs in this video sequence. A Viterbi-like algorithm is devised to help efficiently solve the induced optimization problem. Experiments on a variety of datasets demonstrate the effectiveness of the proposed method.
Qinfeng Shi, Li Wang 0033, Li Cheng 0001, Alexander J. Smola
CVPR3
2008 Real-time foreground segmentation on GPUs using local online learning and global graph cut optimization
abstract
This paper is to address the problem of foreground separation from the background modeling perspective. In particular, we deal with the difficult scenarios where the background texture might change spatially and temporally. A novel approach is proposed that incorporates a pixel-based online learning method to adapt to temporal background changes promptly, together with a graph cuts method to propagate per-pixel evaluation results over nearby pixels. Empirical experiments on a variety of datasets demonstrate the competitiveness of the proposed approach, which is also able to work in real-time on the Graphics Processing Unit (GPU) of programmable graphics cards.
Minglun Gong, Li Cheng 0001
ICPR2
2007 Learning Graph Matching
abstract
As a fundamental problem in pattern recognition, graph matching has found a variety of applications in the field of computer vision. In graph matching, patterns are modeled as graphs and pattern recognition amounts to finding a correspondence between the nodes of different graphs. There are many ways in which the problem has been formulated, but most can be cast in general as a quadratic assignment problem, where a linear term in the objective function encodes node compatibility functions and a quadratic term encodes edge compatibility functions. The main research focus in this theme is about designing efficient algorithms for solving approximately the quadratic assignment problem, since it is NP-hard. In this paper, we turn our attention to the complementary problem: how to estimate compatibility functions such that the solution of the resulting graph matching problem best matches the expected solution that a human would manually provide. We present a method for learning graph matching: the training examples are pairs of graphs and the "labels" are matchings between pairs of graphs. We present experimental results with real image data which give evidence that learning can improve the performance of standard graph matching algorithms. In particular, it turns out that linear assignment with such a learning scheme may improve over state-of-the-art quadratic assignment relaxations. This finding suggests that for a range of problems where quadratic assignment was thought to be essential for securing good results, linear assignment, which is far more efficient, could be just sufficient if learning is performed.
Tibério S. Caetano, Li Cheng 0001, Quoc V. Le, Alexander J. Smola
ICCV2
2007 Learning to compress images and videos
abstract
We present an intuitive scheme for lossy color-image compression: Use the color information from a few representative pixels to learn a model which predicts color on the rest of the pixels. Now, storing the representative pixels and the image in grayscale suffice to recover the original image. A similar scheme is also applicable for compressing videos, where a single model can be used to predict color on many consecutive frames, leading to better compression. Existing algorithms for colorization -- the process of adding color to a grayscale image or video sequence -- are tedious, and require intensive human-intervention. We bypass these limitations by using a graph-based inductive semi-supervised learning module for colorization, and a simple active learning strategy to choose the representative pixels. Experiments on a wide variety of images and video sequences demonstrate the efficacy of our algorithm.
Li Cheng 0001, S. V. N. Vishwanathan
ICML1
2007 Bayesian stereo matching
Li Cheng 0001, Terry Caelli
Comput. Vis. Image Underst.1
2007 Online Learning With Novelty Detection in Human-Guided Road Tracking
abstract
Current image processing and pattern recognition algorithms are not robust enough to make automated remote sensing image interpretation feasible. For this reason, we need to develop image interpretation systems that rely on human guidance. In this paper, we tackle the problem of semiautomatic road tracking in aerial photos. We propose an online learning approach that naturally integrates inputs from human experts with computational algorithms to learn road tracking. Human inputs provide the online learner with training examples to generate road predictors. An ensemble of road predictors is learned incrementally and used to automatically track roads. When novel situations are encountered, control is returned back to the human expert to initialize a new training and tracking iteration. Our approach is computationally efficient, and it can rapidly adapt to dynamic situations where the image feature distributions change. Experimental results confirm that our approach is effective and superior to existing methods.
Jun Zhou 0001, Li Cheng 0001, Walter F. Bischof
IEEE Trans. Geosci. Remote. Sens.2
2006 An Online Discriminative Approach to Background Subtraction
abstract
We present a simple, principled approach to detecting foreground objects in video sequences in real-time. Our method is based on an on-line discriminative learning technique that is able to cope with illumination changes due to discontinuous switching, or illumination drifts caused by slower processes such as varying time of the day. Starting from a discriminative learning principle, we derive a training algorithm that, for each pixel, computes a weighted linear combination of selected past observations with time-decay. We present experimental results that show the proposed approach outperforms existing methods on both synthetic sequencse and real video data.
Li Cheng 0001, Dale Schuurmans, Terry Caelli, S. V. N. Vishwanathan
AVSS1
2006 implicit Online Learning with Kernels
abstract
We present two new algorithms for online learning in reproducing kernel Hilbert spaces. Our first algorithm, ILK (implicit online learning with kernels), employs a new, implicit update technique that can be applied to a wide variety of convex loss functions. We then introduce a bounded memory version, SILK (sparse ILK), that maintains a compact representation of the predictor without compromising solution quality, even in non-stationary environments. We prove loss bounds and analyze the convergence rate of both. Experimental evidence shows that our proposed algorithms outperform current methods on synthetic and real data.
Li Cheng 0001, S. V. N. Vishwanathan, Dale Schuurmans, Terry Caelli
NIPS1
2006 Component Optimization for Image Understanding: A Bayesian Approach
abstract
In this paper, the optimizations of three fundamental components of image understanding: segmentation/annotation, 3D sensing (stereo) and 3D fitting, are posed and integrated within a Bayesian framework. This approach benefits from recent advances in statistical learning which have resulted in greatly improved flexibility and robustness. The first two components produce annotation (region labeling) and depth maps for the input images, while the third module integrates and resolves the inconsistencies between region labels and depth maps to fit most likely 3D models. To illustrate the application of these ideas, we have focused on the difficult problem of fitting individual tree models to tree stands which is a major challenge for vision-based forestry inventory systems.
Li Cheng 0001, Terry Caelli, G. Arturo Sanchez-Azofeifa
IEEE Trans. Pattern Anal. Mach. Intell.1
2005 Variational Bayesian image modelling
abstract
We present a variational Bayesian framework for performing inference, density estimation and model selection in a special class of graphical models---Hidden Markov Random Fields (HMRFs). HMRFs are particularly well suited to image modelling and in this paper, we apply them to the problem of image segmentation. Unfortunately, HMRFs are notoriously hard to train and use because the exact inference problems they create are intractable. Our main contribution is to introduce an efficient variational approach for performing approximate inference of the Bayesian formulation of HMRFs, which we can then apply to the density estimation and model selection problems that arise when learning image models from data. With this variational approach, we can conveniently tackle the problem of image segmentation. We present experimental results which show that our technique outperforms recent HMRF-based segmentation methods on real world images.
Li Cheng 0001, Feng Jiao, Dale Schuurmans
ICML1
2005 Exploiting syntactic, semantic and lexical regularities in language modeling via directed Markov random fields
abstract
We present a directed Markov random field (MRF) model that combines n-gram models, probabilistic context free grammars (PCFGs) and probabilistic latent semantic analysis (PLSA) for the purpose of statistical language modeling. Even though the composite directed MRF model potentially has an exponential number of loops and becomes a context sensitive grammar, we are nevertheless able to estimate its parameters in cubic time using an efficient modified EM method, the generalized inside-outside algorithm, which extends the inside-outside algorithm to incorporate the effects of the n-gram and PLSA language models. We generalize various smoothing techniques to alleviate the sparseness of n-gram counts in cases where there are hidden variables. We also derive an analogous algorithm to calculate the probability of initial subsequence of a sentence, generated by the composite language model. Our experimental results on the Wall Street Journal corpus show that we obtain significant reductions in perplexity compared to the state-of-the-art baseline trigram model with Good-Turing and Kneser-Ney smoothings.
Shaomin Wang, Russell Greiner, Dale Schuurmans, Li Cheng 0001
ICML5
2003 Doubly-MRF stereo matching
abstract
We examine a new double-layered Markov random field probabilistic framework for stereo matching and use belief propagation for approximate inference. Our initial experimental results are promising and future developments are discussed.
Li Cheng 0001, Terry Caelli
ICASSP (3)1