Jianfang Hu

dblp:10/10388 · also Jian-Fang Hu · DBLP profile ↗
← Back
57ranked-venue papers
9as first author
36since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 42 · 5 first-author · 27 since 2021Artificial intelligence and machine learning · 38 · 8 first-author · 24 since 2021
YearPublicationVenuePosition
2026 TubeRMC: Tube-conditioned Reconstruction with Mutual Constraints for Weakly-supervised Spatio-Temporal Video Grounding
abstract
Spatio-Temporal Video Grounding (STVG) aims to localize a spatio-temporal tube that corresponds to a given language query in an untrimmed video. This is a challenging task since it involves complex vision-language understanding and spatiotemporal reasoning. Recent works have explored weakly-supervised setting in STVG to eliminate reliance on fine-grained annotations like bounding boxes or temporal stamps. However, they typically follow a simple late-fusion manner, which generates tubes independent of the text description, often resulting in failed target identification and inconsistent target tracking. To address this limitation, we propose a Tube-conditioned Reconstruction with Mutual Constraints (TubeRMC) framework that generates text-conditioned candidate tubes with pre-trained visual grounding models and further refine them via tube-conditioned reconstruction with spatio-temporal constraints. Specifically, we design three reconstruction strategies from temporal, spatial, and spatio-temporal perspectives to comprehensively capture rich tube-text correspondences. Each strategy is equipped with a Tube-conditioned Reconstructor, utilizing spatio-temporal tubes as condition to reconstruct the key clues in the query. We further introduce mutual constraints between spatial and temporal proposals to enhance their quality for reconstruction. TubeRMC outperforms existing methods on two public benchmarks VidSTG and HCSTVG. Further visualization shows that TubeRMC effectively mitigates both target identification errors and inconsistent tracking.
Jinxuan Li, Jianfang Hu, Chaolei Tan, Tianming Liang, Beihao Xia
AAAI3
2026 Human Motion Prediction via Continual Prior Compensation
abstract
Human Motion Prediction (HMP) aims to predict future human poses at different moments according to observed past motion sequences. Previous approaches mainly treated the prediction of different temporal moments as a single prediction task and learned the predictions of varied moments simultaneously, which would encounter a main limitation: the learning of short-term predictions (referring to "near-future" prediction) could be hindered by the predictions of long-term (referring to "far-future" prediction) motions. In this paper, we develop a novel temporal continual learning framework called Continual Prior Compensation (CPC) to progressively train HMP models, in which we divide the prediction task of motions corresponding to varied temporal moments into several subtasks and train the model in a multi-stage manner. To mitigate the prior information forgetting in the progressive training, we further introduce a learnable random variable Prior Compensation Factor (PCF) to explicitly measure the prior knowledge loss. We theoretically show that the PCF can be efficiently learned together with the model parameters by minimizing a reasonable upper bound of the objective function. The proposed CPC is further enhanced to estimate the prior information loss for each subtask and a new framework called Continual Prior Compensation++ (CPC++) with Fine-Grained Prior Compensation Factor (FGPCF) is finally developed. Our CPC and CPC++ frameworks are quite flexible and can be easily integrated with different HMP backbone models and adapted to various datasets and applications. Extensive experiments on three HMP benchmark datasets using multiple SOTA HMP backbones (PGBIG, siMLPe, MotionMixer, and LTD) demonstrate the effectiveness and flexibility of our frameworks.
Jianwei Tang, Jianfang Hu, Tianming Liang, Xiaotong Lin 0002, Jiangxin Sun, Wei-Shi Zheng 0001, Jian-Huang Lai
IEEE Trans. Pattern Anal. Mach. Intell.2
2026 Arbitrary-scale atmospheric downscaling with mixture of implicit neural networks trained on fixed-scale data
Teng-Yue Chen, Jie-Lan Xie, Wei Zhou 0097, Jianfang Hu, Peng-Qin Yao, Tianming Liang, Wei-Shi Zheng 0001, Pak Wai Chan
Pattern Recognit.4
2026 A training-free framework for text-to-image person re-identification via query-prototype matching
Jianfang Hu, Jian-Huang Lai
Pattern Recognit.3
2025 CLIP-RestoreX: Restore Image Structure and Perception in Exposure Correction
abstract
Exposure correction aims to adjust the exposure of an under- and over-exposed image to enhance its overall visual quality. The core challenge of this task lies in that it requires to faithfully restore both the structure and perception information. In this work, we present a novel exposure correction method, referred to as CLIP-RestoreX, that leverages structural and perceptual priors from CLIP to tackle exposure correction. Specifically, we in CLIP-RestoreX propose to perform exposure correction by aligning CLIP-based structural and perceptual feature of the impaired image with its ground-truth image. To better restore the damaged structural information and perceptual information, we further design a frequency-domain based feature enhancement diffusion model, where we utilize the globality of Fourier transform to help reveal potential the relationship within the features. We conduct extensive experiments on several benchmark datasets. The results demonstrate that the proposed CLIP-RestoreX outperforms state-of-the-art exposure correction methods.
Qing Zhang 0006, Jianfang Hu, Wei-Shi Zheng 0001
AAAI3
2025 SAUGE: Taming SAM for Uncertainty-Aligned Multi-Granularity Edge Detection
abstract
Edge labels are typically at various granularity levels owing to the varying preferences of annotators, thus handling the subjectivity of per-pixel labels has been a focal point for edge detection. Previous methods often employ a simple voting strategy to diminish such label uncertainty or impose a strong assumption of labels with a pre-defined distribution, e.g., Gaussian. In this work, we unveil that the segment anything model (SAM) provides strong prior knowledge to model the uncertainty in edge labels. Our key insight is that the intermediate SAM features inherently correspond to object edges at various granularities, which reflects different edge options due to uncertainty. Therefore, we attempt to align uncertainty with granularity by regressing intermediate SAM features from different layers to object edges at multi-granularity levels. In doing so, the model can fully and explicitly explore diverse ``uncertainties'' in a data-driven fashion. Specifically, we inject a lightweight module (~ 1.5% additional parameters) into the frozen SAM to progressively fuse and adapt its intermediate features to estimate edges from coarse to fine. It is crucial to normalize the granularity level of human edge labels to match their innate uncertainty. For this, we simply perform linear blending to the real edge labels at hand to create pseudo labels with varying granularities. Consequently, our uncertainty-aligned edge detector can flexibly produce edges at any desired granularity (including an optimal one). Thanks to SAM, our model uniquely demonstrates strong generalizability for cross-dataset edge detection. Extensive experimental results on BSDS500, Muticue and NYUDv2 validate our model's superiority.
Xing Liufu, Chaolei Tan, Xiaotong Lin 0002, Yonggang Qi, Jinxuan Li, Jianfang Hu
AAAI6
2025 Modeling Multiple Normal Action Representations for Error Detection in Procedural Tasks
abstract
Error detection in procedural activities is essential for consistent and correct outcomes in AR-assisted and robotic systems. Existing methods often focus on temporal ordering errors or rely on static prototypes to represent normal actions. However, these approaches typically overlook the common scenario where multiple, distinct actions are valid following a given sequence of executed actions. This leads to two issues: (1) the model cannot effectively detect errors using static prototypes when the inference environment or action execution distribution differs from training; and (2) the model may also use the wrong prototypes to detect errors if the ongoing action label is not the same as the predicted one. To address this problem, we propose an Adaptive Multiple Normal Action Representation (AMNAR) framework. AMNAR predicts all valid next actions and reconstructs their corresponding normal action representations, which are compared against the ongoing action to detect errors. Extensive experiments demonstrate that AMNAR achieves state-of-the-art performance, highlighting the effectiveness of AMNAR and the importance of modeling multiple valid next actions in error detection. The code is available at https://github.com/iSEE-Laboratory/AMNAR.
Wei-Jin Huang, Yuan-Ming Li, Zhi-Wei Xia, Yu-Ming Tang, Kun-Yu Lin, Jianfang Hu, Wei-Shi Zheng 0001
CVPR6
2025 Stochastic Human Motion Prediction with Memory of Action Transition and Action Characteristic
abstract
Action-driven stochastic human motion prediction aims to generate future motion sequences of a pre-defined target action based on given past observed sequences performing non-target actions. This task primarily presents two challenges. Firstly, generating smooth transition motions is hard due to the varying transition speeds of different actions. Secondly, the action characteristic is difficult to be learned because of the similarity of some actions. These issues cause the predicted results to be unreasonable and inconsistent. As a result, we propose two memory banks, the Soft-transition Action Bank (STAB) and Action Characteristic Bank (ACB), to tackle the problems above. The STAB stores the action transition information. It is equipped with the novel soft searching approach, which encourages the model to focus on multiple possible action categories of observed motions. The ACB records action characteristic, which produces more prior information for predicting certain actions. To fuse the features retrieved from the two banks better, we further propose the Adaptive Attention Adjustment (AAA) strategy. Extensive experiments on four motion prediction datasets demonstrate that our approach consistently outperforms the previous state-of-the-art. The demo and code are available at https://hyqlat.github.io/STABACB.github.io/.
Jianwei Tang, Tengyue Chen, Jianfang Hu
CVPR4
2025 Panorama Generation From NFoV Image Done Right
abstract
Generating 360-degree panoramas from narrow field of view (NFoV) image is a promising computer vision task for Virtual Reality (VR) applications. Existing methods mostly assess the generated panoramas with InceptionNet or CLIP based metrics, which tend to perceive the image quality and is not suitable for evaluating the distortion. In this work, we first propose a distortion-specific CLIP, named Distort-CLIP to accurately evaluate the panorama distortion and discover the "visual cheating" phenomenon in previous works (i.e., tending to improve the visual results by sacrificing distortion accuracy). This phenomenon arises because prior methods employ a single network to learn the distinct panorama distortion and content completion at once, which leads the model to prioritize optimizing the latter. To address the phenomenon, we propose PanoDecouple, a decoupled diffusion model framework, which decouples the panorama generation into distortion guidance and content completion, aiming to generate panoramas with both accurate distortion and visual appeal. Specifically, we design a DistortNet for distortion guidance by imposing panorama-specific distortion prior and a modified condition registration mechanism; and a ContentNet for content completion by imposing perspective image information. Additionally, a distortion correction loss function with Distort-CLIP is introduced to constrain the distortion explicitly. The extensive experiments validate that PanoDecouple surpasses existing methods both in distortion and visual metrics.
Dian Zheng, Xiao-Ming Wu 0002, Cao Li, Chengfei Lv, Jianfang Hu, Wei-Shi Zheng 0001
CVPR6
2025 ViSpeak: Visual Instruction Feedback in Streaming Videos
Shenghao Fu, Qize Yang, Yuan-Ming Li, Yi-Xing Peng, Kun-Yu Lin, Xihan Wei, Jianfang Hu, Xiaohua Xie, Wei-Shi Zheng 0001
ICCV7
2025 ReferDINO: Referring Video Object Segmentation with Visual Grounding Foundations
abstract
Referring video object segmentation (RVOS) aims to segment target objects throughout a video based on a text description. This is challenging as it involves deep vision-language understanding, pixel-level dense prediction and spatiotemporal reasoning. Despite notable progress in recent years, existing methods still exhibit a noticeable gap when considering all these aspects. In this work, we propose \textbf{ReferDINO}, a strong RVOS model that inherits region-level vision-language alignment from foundational visual grounding models, and is further endowed with pixel-level dense perception and cross-modal spatiotemporal reasoning. In detail, ReferDINO integrates two key components: 1) a grounding-guided deformable mask decoder that utilizes location prediction to progressively guide mask prediction through differentiable deformation mechanisms; 2) an object-consistent temporal enhancer that injects pretrained time-varying text features into inter-frame interaction to capture object-aware dynamic changes. Moreover, a confidence-aware query pruning strategy is designed to accelerate object decoding without compromising model performance. Extensive experimental results on five benchmarks demonstrate that our ReferDINO significantly outperforms previous methods (e.g., +3.9% (\mathcal{J}&\mathcal{F}) on Ref-YouTube-VOS) with real-time inference speed (51 FPS).
Tianming Liang, Kun-Yu Lin, Chaolei Tan, Jianguo Zhang 0001, Wei-Shi Zheng 0001, Jianfang Hu
ICCV6
2025 Recovering Human Mesh from Videos by 2D and 3D Deformable Attentions
abstract
Existing methods for 3D human mesh recovery from video rely mainly on Recurrent Neural Networks (RNNs) or Transformers. However, due to the limitations of RNNs in temporal modeling and the dense strategies of traditional attention mechanisms, these methods struggle to efficiently model human motion in videos. To address this issue, we propose a novel method that exploits sparse deformable attention mechanisms, efficiently extracting critical spatio-temporal mesh information from the input videos. Specifically, our method consists of two novel modules: the 3D Deformable Mesh Attention (3D-DMA) module and the 2D Deformable Mesh Attention (2D-DMA) module. The 3D-DMA module adaptively extracts human mesh information by sparsely aggregating the features at varied spatio-temporal locations with attentions, while the 2D-DMA module captures the mesh contexts with the attentions of features at sparsely sampled spatial locations. By fusing the outputs of the two modules, we obtain accurate and smooth human body estimations. Extensive experiments show that our model outperforms previous state-of-the-art methods on three widely used benchmarks.
Yulei Kang, Teng-Yue Chen, Xiaotong Lin 0002, Siyu Jiang, Jianfang Hu
ICME5
2025 Efficient Text-to-Motion via Multi-Head Generative Masked Modeling
abstract
Text-to-motion generation has attracted increasing attention in recent years. Existing methods primarily employ Vector Quantized Variational AutoEncoder (VQ-VAE) as the tokenizer for motion representation. However, such vector quantization maps the continuous space into limited discrete tokens, which inevitably leads to significant information loss. To address this limitation, we propose a simple yet effective approach to expand the capacity of discrete motion representation space, effectively reducing the information loss without incurring additional overhead. Specifically, we present a Multi-Head Generative Masked model which a) exploits multi-head mechanism into quantization for a high-fidelity motion tokenizer, and b) simultaneously generates tokens across different heads using a bi-directional transformer. In this way, the motion representation space can be expanded exponentially at the cost of negligible time and parameter overhead. Extensive experiments demonstrate that our method achieves state-of-the-art performance on both HumanML3D and KIT-ML datasets.
Heng Li 0015, Xing Liufu, Xiaotong Lin 0002, Jianfang Hu
ICME5
2025 Context Consistency Learning via Sentence Removal for Semi-Supervised Video Paragraph Grounding
abstract
Semi-Supervised Video Paragraph Grounding (SSVPG) aims to localize multiple sentences in a paragraph from an untrimmed video with limited temporal annotations. Existing methods focus on teacher-student consistency learning and video-level contrastive loss, but they overlook the importance of perturbing query contexts to generate strong supervisory signals. In this work, we propose a novel Context Consistency Learning (CCL) framework that unifies the paradigms of consistency regularization and pseudo-labeling to enhance semi-supervised learning. Specifically, we first conduct teacher-student learning where the student model takes as inputs strongly-augmented samples with sentences removed and is enforced to learn from the adequately strong supervisory signals from the teacher model. Afterward, we conduct model retraining based on the generated pseudo labels, where the mutual agreement between the original and augmented views’ predictions is utilized as the label confidence. Extensive experiments show that CCL outperforms existing methods by a large margin.
Yaokun Zhong, Siyu Jiang, Jianfang Hu
ICME4
2025 Weakly-Supervised Temporal Action Localization by Progressive Complementary Learning
abstract
Weakly-Supervised Temporal Action Localization (WSTAL) aims to localize and classify action instances in long untrimmed videos with only video-level category labels as supervision. A critical challenge of WSTAL is the large gap between video-level supervision and unavailable snippet-level supervision. Prevailing methods typically assign pseudo labels to snippets, but these methods suffer from significant noise caused by the pseudo snippet-level labels. In this work, we address the WSTAL from a novel category exclusion perspective, which gradually enhances the snippet-level supervision to bridge the gap. Our proposed Progressive Complementary Learning (ProCL) is inspired by the fact that, video-level labels precisely indicate the categories that all snippets surely do not belong to, which is ignored by previous works. Accordingly, we first exclude these surely non-existent categories by the deterministic complementary learning. And then, we introduce the entropy-based pseudo complementary learning that is able to exclude more categories for snippets of less ambiguity. Furthermore, for the remaining ambiguous snippets, we attempt to reduce the ambiguity by distinguishing foreground actions from the background. Extensive experimental results show that our method achieves new state-of-the-art performance on THUMOS14, ActivityNet1.3, and MultiTHUMOS benchmarks.
Jia-Run Du, Jia-Chang Feng, Kun-Yu Lin, Fa-Ting Hong, Zhongang Qi, Ying Shan, Jianfang Hu, Wei-Shi Zheng 0001
IEEE Trans. Circuits Syst. Video Technol.7
2025 Progressive Human Motion Generation Based on Text and Few Motion Frames
abstract
Although existing text-to-motion (T2M) methods can produce realistic human motion from text description, it is still difficult to align the generated motion with the desired postures since using text alone is insufficient for precisely describing diverse postures. To achieve more controllable generation, an intuitive way is to allow the user to input a few motion frames describing precise desired postures. Thus, we explore a new Text-Frame-to-Motion (TF2M) generation task that aims to generate motions from text and very few given frames. Intuitively, the closer a frame is to a given frame, the lower the uncertainty of this frame is when conditioned on this given frame. Hence, we propose a novel Progressive Motion Generation (PMG) method to progressively generate a motion from the frames with low uncertainty to those with high uncertainty in multiple stages. During each stage, new frames are generated by a Text-Frame Guided Generator conditioned on frame-aware semantics of the text, given frames, and frames generated in previous stages. Additionally, to alleviate the train-test gap caused by multi-stage accumulation of incorrectly generated frames during testing, we propose a Pseudo-frame Replacement Strategy for training. Experimental results show that our PMG outperforms existing T2M generation methods by a large margin with even one given frame, validating the effectiveness of our PMG. Code is available here.
Ling-An Zeng, Gaojie Wu, Ancong Wu, Jianfang Hu, Wei-Shi Zheng 0001
IEEE Trans. Circuits Syst. Video Technol.4
2025 Rethinking Temporal Context in Video-QA: A Comprehensive Study of Single-Frame Static Bias
abstract
Video question answering (Video-QA) has emerged as a core task in the vision-language domain, which requires the models to understand a given video and answer textual questions related to the video. Compared to conventional image-language tasks, Video-QA is designed for improving the models' capacity of memorizing and integrating multi-frame temporal cues associated with the questions. While significant performance improvements have recently been witnessed on public benchmarks, in this work, we rethink whether these improvements truly stem from better understanding of video temporal context as expected. To this end, we accomplish a strong single-frame baseline model trained with knowledge distillation. With this model, we surprisingly find that visiting only one single frame, without incorporating multi-frame and temporal information, is sufficient to achieve state-of-the-art (SOTA) performance on multiple mainstream benchmarks. This finding reveals the prevalence of single-frame bias in current benchmarks for the first time. Around the single-frame bias, we conduct an in-depth analysis on multiple popular benchmarks, which demonstrate that: (i) merely relying on one frame is able to achieve comparable performance with SOTA temporal Video-QA models; (ii) simply ensembling the prediction scores of only 3 separate frames is able to surpass temporal SOTAs. Furthermore, we observe that most of the benchmarks are biased towards central segments, and even the latest benchmarks tailored for temporal reasoning still suffer from severe single-frame bias. In case study, we find two key properties of low-bias instances: the question emphasizes temporal dependency and contextual understanding, and the associated video content presents significant variability in scenes, actions or interactions. Through further analysis on compositional reasoning datasets, we find that constructing explicit object/event interactions upon videos to fill in well-designed temporal question templates can effectively reduce the single-frame bias during annotation. We hope our analysis helps facilitate future efforts in the field towards mitigating static bias and highlighting temporal reasoning.
Tianming Liang, Jianfang Hu, Xiangyang Yu, Wei-Shi Zheng 0001, Jian-Huang Lai
IEEE Trans. Multim.3
2025 Lightweight image super-resolution via an efficient local-global transformer network with adaptive attention windows
Simin Zheng, Peiyao Chen, Jianfang Hu, Ruichu Cai
Vis. Comput.5
2024 Ranking Distillation for Open-Ended Video Question Answering with Insufficient Labels
abstract
This paper focuses on open-ended video question answering, which aims to find the correct answers from a large answer set in response to a video-related question. This is essentially a multi-label classification task, since a question may have multiple answers. However, due to annotation costs, the labels in existing benchmarks are always extremely insufficient, typically one answer per question. As a result, existing works tend to directly treat all the unlabeled answers as negative labels, leading to limited ability for generalization. In this work, we introduce a simple yet effective ranking distillation framework (RADI) to mitigate this problem without additional manual annotation. RADI employs a teacher model trained with incomplete labels to generate rankings for potential answers, which contain rich knowledge about label priority as well as label-associated visual cues, thereby enriching the insufficient labeling information. To avoid overconfidence in the imperfect teacher model, we further present two robust and parameter-free ranking distillation approaches: a pairwise approach which introduces adaptive soft margins to dynamically refine the optimization constraints on various pairwise rankings, and a listwise approach which adopts sampling-based partial listwise learning to resist the bias in teacher ranking. Extensive experiments on five popular benchmarks consistently show that both our pairwise and listwise RADIs outperform state-of-the-art methods. Further analysis demonstrates the effectiveness of our methods on the insufficient labeling problem.
Tianming Liang, Chaolei Tan, Beihao Xia, Wei-Shi Zheng 0001, Jianfang Hu
CVPR5
2024 Siamese Learning with Joint Alignment and Regression for Weakly-Supervised Video Paragraph Grounding
abstract
Video Paragraph Grounding (VPG) is an emerging task in video-language understanding, which aims at localizing multiple sentences with semantic relations and temporal or-der from an untrimmed video. However, existing VPG approaches are heavily reliant on a considerable number of temporal labels that are laborious and time-consuming to acquire. In this work, we introduce and explore Weakly-Supervised Video Paragraph Grounding (WSVPG) to elim-inate the need of temporal annotations. Different from pre-vious weakly-supervised grounding frameworks based on multiple instance learning or reconstruction learning for two-stage candidate ranking, we propose a novel siamese learning framework that jointly learns the cross-modal feature alignment and temporal coordinate regression without timestamp labels to achieve concise one-stage localization for WSVPG. Specifically, we devise a Siamese Grounding TRansformer (SiamGTR) consisting of two weight-sharing branches for learning complementary supervision. An Aug-mentation Branch is utilized for directly regressing the tem-poral boundaries of a complete paragraph within a pseudo video, and an Inference Branch is designed to capture the order-guided feature correspondence for localizing multi-ple sentences in a normal video. We demonstrate by exten-sive experiments that our paradigm has superior practica-bility and flexibility to achieve efficient weakly-supervised or semi-supervised learning, outperforming state-of-the-art methods trained with the same or stronger supervision.
Chaolei Tan, Jian-Huang Lai, Wei-Shi Zheng 0001, Jianfang Hu
CVPR4
2024 Selective Hourglass Mapping for Universal Image Restoration Based on Diffusion Model
abstract
Universal image restoration is a practical and poten-tial computer vision task for real-world applications. The main challenge of this task is handling the different degra-dation distributions at once. Existing methods mainly utilize task-specific conditions (e.g., prompt) to guide the model to learn different distributions separately, named multi-partite mapping. However, it is not suitable for universal model learning as it ignores the shared information between different tasks. In this work, we propose an advanced selective hourglass mapping strategy based on diffusion model, termed DiffUIR. Two novel considerations make our Dif-fUIR non-trivial. Firstly, we equip the model with strong condition guidance to obtain accurate generation direction of diffusion model (selective). More importantly, DiffUIR integrates a flexible shared distribution term (SDT) into the diffusion algorithm elegantly and naturally, which gradually maps different distributions into a shared one. In the reverse process, combined with SDT and strong condition guidance, DiffUIR iteratively guides the shared distribution to the task-specific distribution with high image quality (hourglass). Without bells and whistles, by only modifying the mapping strategy, we achieve state-of-the-art performance on five image restoration tasks, 22 benchmarks in the universal setting and zero-shot generalization setting. Surprisingly, by only using a lightweight model (only 0.89M), we could achieve outstanding performance. The source code and pre-trained models are available at https://github.com/iSEE-Laboratory/DiffUIR.
Dian Zheng, Xiao-Ming Wu 0002, Shuzhou Yang, Jianfang Hu, Wei-Shi Zheng 0001
CVPR5
2024 Progressive Pretext Task Learning for Human Trajectory Prediction
Xiaotong Lin 0002, Tianming Liang, Jian-Huang Lai, Jianfang Hu
ECCV (30)4
2024 Out-of-Distribution Detection by Principal Component Correspondence
abstract
Out-of-distribution (OOD) detection is vital for the safe application of intelligent systems in real-world scenarios. This paper proposes an enhancement to OOD detection by leveraging the consistency in cognition between two models, both pretrained on in-distribution (ID) data. Specifically, for a given test sample, we first apply Principal Component Analysis (PCA)-based projection on the feature vectors from each model. These obtained feature vectors (with correlation between dimensions decoupled by PCA projection) are then aligned using a multiple linear mapping, which is fitted using the least squares method on the training data. We hypothesize that the regression error for OOD data will be larger than that for ID data, making it a useful metric for OOD detection. Our experimental results demonstrate the effectiveness of this method. When combined with existing robust baselines, our approach achieves state-of-the-art performance in OOD detection.
Xiaoyuan Guan, Zhiyong Gan, Ling Deng, Jiankang Chen, Shenshen Bu, Chunliang Zhao, Jianfang Hu, Wei-Shi Zheng 0001
ICME8
2024 SynopGround: A Large-Scale Dataset for Multi-Paragraph Video Grounding from TV Dramas and Synopses
Chaolei Tan, Zihang Lin, Junfu Pu, Zhongang Qi, Wei-Yi Pei, Yexin Wang, Ying Shan, Wei-Shi Zheng 0001, Jianfang Hu
ACM Multimedia10
2024 AdvAD: Exploring Non-Parametric Diffusion for Imperceptible Adversarial Attacks
abstract
Imperceptible adversarial attacks aim to fool DNNs by adding imperceptible perturbation to the input data. Previous methods typically improve the imperceptibility of attacks by integrating common attack paradigms with specifically designed perception-based losses or the capabilities of generative models. In this paper, we propose Adversarial Attacks in Diffusion (AdvAD), a novel modeling framework distinct from existing attack paradigms. AdvAD innovatively conceptualizes attacking as a non-parametric diffusion process by theoretically exploring basic modeling approach rather than using the denoising or generation abilities of regular diffusion models requiring neural networks. At each step, much subtler yet effective adversarial guidance is crafted using only the attacked model without any additional network, which gradually leads the end of diffusion process from the original image to a desired imperceptible adversarial example. Grounded in a solid theoretical foundation of the proposed non-parametric diffusion process, AdvAD achieves high attack efficacy and imperceptibility with intrinsically lower overall perturbation strength. Additionally, an enhanced version AdvAD-X is proposed to evaluate the extreme of our novel framework under an ideal scenario. Extensive experiments demonstrate the effectiveness of the proposed AdvAD and AdvAD-X. Compared with state-of-the-art imperceptible attacks, AdvAD achieves an average of 99.9% (+17.3%) ASR with 1.34 (-0.97) $l_2$ distance, 49.74 (+4.76) PSNR and 0.9971 (+0.0043) SSIM against four prevalent DNNs with three different architectures on the ImageNet-compatible dataset. Code is available at https://github.com/XianguiKang/AdvAD.
Ziqiang He, Anwei Luo, Jianfang Hu, Z. Jane Wang 0001, Xiangui Kang
NeurIPS4
2024 Time-Frequency Mutual Learning for Moment Retrieval and Highlight Detection
Yaokun Zhong, Tianming Liang, Jianfang Hu
PRCV (5)3
2024 Beyond Minimum-of-N: Rethinking the Evaluation and Methods of Pedestrian Trajectory Prediction
abstract
Pedestrian trajectory prediction is an essential task in real-world applications, aimed at predicting plausible future trajectories based on limited observations. In this work, we rethink the standard evaluation metric of the pedestrian trajectory prediction task: Minimum-of-N Average Displacement Error (MoN-ADE). As for multi-modal prediction models that generate multiple trajectories for each pedestrian, this metric typically evaluates the model by only considering the one that is closest to the ground-truth trajectory. However, such an evaluation protocol cannot comprehensively evaluate the predictive ability of the model, and potentially encourage models to generate high-variance and dispersed trajectory distributions. This is quite impractical especially for many real-world scenes like autonomous driving that require precise and convergent trajectory predictions. To address these limitations, we design a novel metric towards comprehensive evaluation in pedestrian trajectory prediction, which moves beyond the traditional reliance on the closest prediction. Specifically, we replace the Minimum-of-N strategy with an insightful Random-Sampling-K strategy to calculate the expectations of the minimum ADE and formulate a novel metric: Area Under the Curve (AUC). Furthermore, motivated by the proposed metric, we introduce a novel objective function named K-Ensemble Loss, which guides the state-of-the-art models to optimize the whole prediction distribution and reduce the uncertainty caused by the high-variance predictions. Extensive experiments on three real-world datasets demonstrate that the proposed metric and objective function are provided with significant effectiveness and flexibility.
Xiaotong Lin 0002, Yejia Huang, Zizhen Zhang, Jianfang Hu
IEEE Trans. Circuits Syst. Video Technol.5
2023 Collaborative Static and Dynamic Vision-Language Streams for Spatio-Temporal Video Grounding
abstract
Spatio-Temporal Video Grounding (STVG) aims to localize the target object spatially and temporally according to the given language query. It is a challenging task in which the model should well understand dynamic visual cues (e.g., motions) and static visual cues (e.g., object appearances) in the language description, which requires effective joint modeling of spatiotemporal visuallinguistic dependencies. In this work, we propose a novel framework in which a static vision-language stream and a dynamic vision-language stream are developed to collaboratively reason the target tube. The static stream performs cross-modal understanding in a single frame and learns to attend to the target object spatially according to intraframe visual cues like object appearances. The dynamic stream models visual-linguistic dependencies across multiple consecutive frames to capture dynamic cues like motions. We further design a novel cross-stream collaborative block between the two streams, which enables the static and dynamic streams to transfer useful and complementary information from each other to achieve collaborative reasoning. Experimental results show the effectiveness of the collaboration of the two streams and our overall frame-work achieves new state-of-the-art performance on both HCSTVG and VidSTG datasets.
Zihang Lin, Chaolei Tan, Jianfang Hu, Zhi Jin 0002, Tiancai Ye, Wei-Shi Zheng 0001
CVPR3
2023 Hierarchical Semantic Correspondence Networks for Video Paragraph Grounding
abstract
Video Paragraph Grounding (VPG) is an essential yet challenging task in vision-language understanding, which aims to jointly localize multiple events from an untrimmed video with a paragraph query description. One of the critical challenges in addressing this problem is to comprehend the complex semantic relations between visual and textual modalities. Previous methods focus on modeling the contextual information between the video and text from a single-level perspective (i.e., the sentence level), ignoring rich visual-textual correspondence relations at different semantic levels, e.g., the video-word and video-paragraph correspondence. To this end, we propose a novel Hierarchical Semantic Correspondence Network (HSCNet), which explores multi-level visual-textual correspondence by learning hierarchical semantic alignment and utilizes dense supervision by grounding diverse levels of queries. Specifically, we develop a hierarchical encoder that encodes the multi-modal inputs into semantics-aligned representations at different levels. To exploit the hierarchical semantic correspondence learned in the encoder for multi-level supervision, we further design a hierarchical decoder that progressively performs finer grounding for lower-level queries conditioned on higher-level semantics. Extensive experiments demonstrate the effectiveness of HSCNet and our method significantly outstrips the state-of-the-arts on two challenging benchmarks, i.e., ActivityNet-Captions and TACoS.
Chaolei Tan, Zihang Lin, Jianfang Hu, Wei-Shi Zheng 0001, Jian-Huang Lai
CVPR3
2023 Learning Discriminative Proposal Representation for Multi-object Tracking
Yejia Huang, Xianqin Liu, Jianfang Hu
ICIG (2)4
2023 Temporal Continual Learning with Prior Compensation for Human Motion Prediction
abstract
Human Motion Prediction (HMP) aims to predict future poses at different moments according to past motion sequences. Previous approaches have treated the prediction of various moments equally, resulting in two main limitations: the learning of short-term predictions is hindered by the focus on long-term predictions, and the incorporation of prior information from past predictions into subsequent predictions is limited. In this paper, we introduce a novel multi-stage training framework called Temporal Continual Learning (TCL) to address the above challenges. To better preserve prior information, we introduce the Prior Compensation Factor (PCF). We incorporate it into the model training to compensate for the lost prior information. Furthermore, we derive a more reasonable optimization objective through theoretical derivation. It is important to note that our TCL framework can be easily integrated with different HMP backbone models and adapted to various datasets and applications. Extensive experiments on four HMP benchmark datasets demonstrate the effectiveness and flexibility of TCL. The code is available at https://github.com/hyqlat/TCL.
Jianwei Tang, Jiangxin Sun, Xiaotong Lin 0002, Wei-Shi Zheng 0001, Jianfang Hu
NeurIPS6
2023 Reconstruction with robustness: A semantic prior guided face super-resolution framework for multiple degradations
Hongjun Wu 0003, Huanrong Zhang, Zhi Jin 0002, Driton Salihu, Jianfang Hu
Image Vis. Comput.6
2022 You Never Stop Dancing: Non-freezing Dance Generation via Bank-constrained Manifold Projection
abstract
One of the most overlooked challenges in dance generation is that the auto-regressive frameworks are prone to freezing motions due to noise accumulation. In this paper, we present two modules that can be plugged into the existing models to enable them to generate non-freezing and high fidelity dances. Since the high-dimensional motion data are easily swamped by noise, we propose to learn a low-dimensional manifold representation by an auto-encoder with a bank of latent codes, which can be used to reduce the noise in the predicted motions, thus preventing from freezing. We further extend the bank to provide explicit priors about the future motions to disambiguate motion prediction, which helps the predictors to generate motions with larger magnitude and higher fidelity than possible before. Extensive experiments on AIST++, a public large-scale 3D dance motion benchmark, demonstrate that our method notably outperforms the baselines in terms of quality, diversity and time length.
Jiangxin Sun, Huang Hu, Hanjiang Lai, Zhi Jin 0002, Jianfang Hu
NeurIPS6
2022 APANet: Auto-Path Aggregation for Future Instance Segmentation Prediction
abstract
Despite the remarkable progress achieved in conventional instance segmentation, the problem of predicting instance segmentation results for unobserved future frames remains challenging due to the unobservability of future data. Existing methods mainly address this challenge by forecasting features of future frames. However, these methods always treat features of multiple levels (e.g., coarse-to-fine pyramid features) independently and do not exploit them collaboratively, which results in inaccurate prediction for future frames; and moreover, such a weakness can partially hinder self-adaption of a future segmentation prediction model for different input samples. To solve this problem, we propose an adaptive aggregation approach called Auto-Path Aggregation Network (APANet), where the spatio-temporal contextual information obtained in the features of each individual level is selectively aggregated using the developed "auto-path". The "auto-path" connects each pair of features extracted at different pyramid levels for task-specific hierarchical contextual information aggregation, which enables selective and adaptive aggregation of pyramid features in accordance with different videos/frames. Our APANet can be further optimized jointly with the Mask R-CNN head as a feature decoder and a Feature Pyramid Network (FPN) feature encoder, forming a joint learning system for future instance segmentation prediction. We experimentally show that the proposed method can achieve state-of-the-art performance on three video-based instance segmentation benchmarks for future instance segmentation prediction.
Jianfang Hu, Jiangxin Sun, Zihang Lin, Jian-Huang Lai, Wenjun Zeng 0001, Wei-Shi Zheng 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2021 Predictive Feature Learning for Future Segmentation Prediction
abstract
Future segmentation prediction aims to predict the segmentation masks for unobserved future frames. Most existing works addressed it by directly predicting the intermediate features extracted by existing segmentation models. However, these segmentation features are learned to be local discriminative (with rich details) and are always of high resolution/dimension. Hence, the complicated spatiotemporal variations of these features are difficult to predict, which motivates us to learn a more predictive representation. In this work, we develop a novel framework called Predictive Feature Autoencoder. In the proposed framework, we construct an autoencoder which serves as a bridge between the segmentation features and the predictor. In the latent feature learned by the autoencoder, global structures are enhanced and local details are suppressed so that it is more predictive. In order to reduce the risk of vanishing the suppressed details during recurrent feature prediction, we further introduce a reconstruction constraint in the prediction module. Extensive experiments show the effectiveness of the proposed approach and our method outperforms state-of-the-arts by a considerable margin.
Zihang Lin, Jiangxin Sun, Jianfang Hu, Qi-Zhi Yu, Jian-Huang Lai, Wei-Shi Zheng 0001
ICCV3
2021 Action-guided 3D Human Motion Prediction
abstract
The ability of forecasting future human motion is important for human-machine interaction systems to understand human behaviors and make interaction. In this work, we focus on developing models to predict future human motion from past observed video frames. Motivated by the observation that human motion is closely related to the action being performed, we propose to explore action context to guide motion prediction. Specifically, we construct an action-specific memory bank to store representative motion dynamics for each action category, and design a query-read process to retrieve some motion dynamics from the memory bank. The retrieved dynamics are consistent with the action depicted in the observed video frames and serve as a strong prior knowledge to guide motion prediction. We further formulate an action constraint loss to ensure the global semantic consistency of the predicted motion. Extensive experiments demonstrate the effectiveness of the proposed approach, and we achieve state-of-the-art performance on 3D human motion prediction.
Jiangxin Sun, Zihang Lin, Xintong Han, Jianfang Hu, Jia Xu 0011, Wei-Shi Zheng 0001
NeurIPS4
2020 Aggregating Spatio-temporal Context for Video Object Segmentation
Jianfang Hu, Wei-Shi Zheng 0001
PRCV (1)2
2020 Fast Collective Activity Recognition Under Weak Supervision
abstract
Collective activity recognition, which tells what activity a group of people is performing, is a cutting-edge research topic in computer vision. Different from action performed by individuals, collective activity needs to consider the complex interactions among different people. However, most previous works require exhaustive annotations such as accurate label information of individual actions, pairwise interactions, and poses, which could not be easily available in practice. Moreover, most of them treat human detection as a decoupled task before collective activity recognition and leverage all detected persons. This not only ignores the mutual relation between the two tasks, which makes it hard for filtering out irrelevant people, but also probably increases the computation burden when reasoning the collective activities. In this paper, we propose a fast weakly supervised deep learning architecture for collective activity recognition. For fast inference, we propose to make the actor detection and weakly supervised collective activity reasoning collaborate in an end-to-end framework by sharing convolutional layers between them. The joint learning makes the two tasks united and reinforced each other, so that it is more effective to filter out the outliers who are not involved in the activity. For the weakly supervised learning, we propose a latent embedding scheme for mining person-group interactive relationship to get rid of the use of any pairwise relation between people and the individual action labels as well. The experimental results show that the proposed framework achieves comparable or even better performance as compared to the state-of-the-art on three datasets. Our joint modelling reasons collective activities at the speed of 22.65 fps, which is the fastest ever known and substantially makes collective activity recognition more towards real-time applications.
Peizhen Zhang, Yongyi Tang, Jianfang Hu, Wei-Shi Zheng 0001
IEEE Trans. Image Process.3
2019 Action Knowledge Transfer for Action Prediction with Partial Videos
abstract
Predicting action class from partially observed videos, which is known as action prediction, is an important task in computer vision field with many applications. The challenge for action prediction mainly lies in the lack of discriminative action information for the partially observed videos. To tackle this challenge, in this work, we propose to transfer action knowledge learned from fully observed videos for improving the prediction of partially observed videos. Specifically, we develop a two-stage learning framework for action knowledge transfer. At the first stage, we learn feature embeddings and discriminative action classifier from full videos. The knowledge in the learned embeddings and classifier is then transferred to the partial videos at the second stage. Our experiments on the UCF-101 and HMDB-51 datasets show that the proposed action knowledge transfer method can significantly improve the performance of action prediction, especially for the actions with small observation ratios (e.g., 10%). We also experimentally illustrate that our method outperforms all the state-of-the-art action prediction systems.
Yijun Cai, Haoxin Li, Jianfang Hu, Wei-Shi Zheng 0001
AAAI3
2019 Progressive Teacher-Student Learning for Early Action Prediction
abstract
The goal of early action prediction is to recognize actions from partially observed videos with incomplete action executions, which is quite different from action recognition. Predicting early actions is very challenging since the partially observed videos do not contain enough action information for recognition. In this paper, we aim at improving early action prediction by proposing a novel teacherstudent learning framework. Our framework involves a teacher model for recognizing actions from full videos, a student model for predicting early actions from partial videos, and a teacher-student learning block for distilling progressive knowledge from teacher to student, crossing different tasks. Extensive experiments on three public action datasets show that the proposed progressive teacher-student learning framework can consistently improve performance of early action prediction model. We have also reported the state-of-the-art performances for early action prediction on all of these sets.
Xionghui Wang, Jianfang Hu, Jian-Huang Lai, Jianguo Zhang 0001, Wei-Shi Zheng 0001
CVPR2
2019 DBDNet: Learning Bi-directional Dynamics for Early Action Prediction
abstract
Predicting future actions from observed partial videos is very challenging as the missing future is uncertain and sometimes has multiple possibilities. To obtain a reliable future estimation, a novel encoder-decoder architecture is proposed for integrating the tasks of synthesizing future motions from observed videos and reconstructing observed motions from synthesized future motions in an unified framework, which can capture the bi-directional dynamics depicted in partial videos along the temporal (past-to-future) direction and reverse chronological (future-back-to-past) direction. We then employ a bi-directional long short-term memory (Bi-LSTM) architecture to exploit the learned bi-directional dynamics for predicting early actions. Our experiments on two benchmark action datasets show that learning bi-directional dynamics benefits the early action prediction and our system clearly outperforms the state-of-the-art methods.
Guoliang Pang, Xionghui Wang, Jianfang Hu, Qing Zhang 0006, Wei-Shi Zheng 0001
IJCAI3
2019 Predicting Future Instance Segmentation with Contextual Pyramid ConvLSTMs
abstract
Despite the remarkable progress in instance segmentation, the problem of predicting future instance segmentation remains challenging due to the unobservability of future data. Existing methods mainly address this challenge by forecasting pyramid features to represent unobserved future frames. However, they mainly predict features for each pyramid level independently, and ignore the underlying structural relationship between features of different levels.
Jiangxin Sun, Jiafeng Xie, Jianfang Hu, Zihang Lin, Jian-Huang Lai, Wenjun Zeng 0001, Wei-Shi Zheng 0001
ACM Multimedia3
2019 Early Action Prediction by Soft Regression
abstract
We propose a novel approach for predicting on-going action with the assistance of a low-cost depth camera. Our approach introduces a soft regression-based early prediction framework. In this framework, we estimate soft labels for the subsequences at different progress levels, jointly learned with an action predictor. Our formulation of soft regression framework 1) overcomes a usual assumption in existing early action prediction systems that the progress level of on-going sequence is given in the testing stage; and 2) presents a theoretical framework to better resolve the ambiguity and uncertainty of subsequences at early performing stage. The proposed soft regression framework is further enhanced in order to take the relationships among subsequences and the discrepancy of soft labels over different classes into consideration, so that a Multiple Soft labels Recurrent Neural Network (MSRNN) is finally developed. For real-time performance, we also introduce a new RGB-D feature called "local accumulative frame feature (LAFF)", which can be computed efficiently by constructing an integral feature map. Our experiments on three RGB-D benchmark datasets and an unconstrained RGB action set demonstrate that the proposed regression-based early action prediction model outperforms existing models significantly and also show that the early action prediction on RGB-D sequence is more accurate than that on RGB channel.
Jianfang Hu, Wei-Shi Zheng 0001, Lianyang Ma, Gang Wang 0012, Jian-Huang Lai, Jianguo Zhang 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2018 Improving Fast Segmentation With Teacher-Student Learning
Jiafeng Xie, Bing Shuai, Jianfang Hu, Wei-Shi Zheng 0001
BMVC3
2018 Deep Bilinear Learning for RGB-D Action Recognition
Jianfang Hu, Wei-Shi Zheng 0001, Jian-Huang Lai, Jianguo Zhang 0001
ECCV (7)1
2018 GL-PAM RGB-D Gesture Recognition
abstract
The existing approaches for RGB-D gesture recognition mainly developed their systems based on the global features extracted from full sequences, which makes them unreliable for capturing some important movements. In this paper, we propose to combine the global and local context information extracted from posture, appearance, and motion sequences. Our experimental results on a large scale RGB-D gesture dataset show that the proposed global and local contexts can complement well with each other for efficiently characterizing gestures, and thus achieve the 2nd place in the ChaLearn LAP Large-scale Isolated Gesture Recognition Challenge (Round 2).
Benchao Li, Wanhua Li 0001, Yongyi Tang, Jianfang Hu, Wei-Shi Zheng 0001
ICIP4
2018 Global-Local Temporal Saliency Action Prediction
abstract
Action prediction on a partially observed action sequence is a very challenging task. To address this challenge, we first design a global-local distance model, where a global-temporal distance compares subsequences as a whole and local-temporal distance focuses on individual segment. Our distance model introduces temporal saliency for each segment to adapt its contribution. Finally, a global-local temporal action prediction model is formulated in order to jointly learn and fuse these two types of distances. Such a prediction model is capable of recognizing action of: 1) an on-going sequence and 2) a sequence with arbitrarily frames missing between the beginning and end (known as gap-filling). Our proposed model is tested and compared with related action prediction models on BIT, UCF11, and HMDB data sets. The results demonstrated the effectiveness of our proposal. In particular, we showed the benefit of our proposed model on predicting unseen action types and the advantage on addressing the gapfilling problem as compared with recently developed action prediction models.
Shaofan Lai, Wei-Shi Zheng 0001, Jianfang Hu, Jianguo Zhang 0001
IEEE Trans. Image Process.3
2017 Latent embeddings for collective activity recognition
abstract
Rather than simply recognizing the action of a person individually, collective activity recognition aims to find out what a group of people is acting in a collective scene. Previous state-of-the-art methods using hand-crafted potentials in conventional graphical model which can only define a limited range of relations. Thus, the complex structural dependencies among individuals involved in a collective scenario cannot be fully modeled. In this paper, we overcome these limitations by embedding latent variables into feature space and learning the feature mapping functions in a deep learning framework. The embeddings of latent variables build a global relation containing person-group interactions and richer contextual information by jointly modeling broader range of individuals. Besides, we assemble attention mechanism during embedding for achieving more compact representations. We evaluate our method on three collective activity datasets, where we contribute a much larger dataset in this work. The proposed model has achieved clearly better performance as compared to the state-of-the-art methods in our experiments.
Yongyi Tang, Peizhen Zhang, Jianfang Hu, Wei-Shi Zheng 0001
AVSS3
2017 Jointly Learning Heterogeneous Features for RGB-D Activity Recognition
abstract
In this paper, we focus on heterogeneous features learning for RGB-D activity recognition. We find that features from different channels (RGB, depth) could share some similar hidden structures, and then propose a joint learning model to simultaneously explore the shared and feature-specific components as an instance of heterogeneous multi-task learning. The proposed model formed in a unified framework is capable of: 1) jointly mining a set of subspaces with the same dimensionality to exploit latent shared features across different feature channels, 2) meanwhile, quantifying the shared and feature-specific components of features in the subspaces, and 3) transferring feature-specific intermediate transforms (i-transforms) for learning fusion of heterogeneous features across datasets. To efficiently train the joint model, a three-step iterative optimization algorithm is proposed, followed by a simple inference model. Extensive experimental results on four activity datasets have demonstrated the efficacy of the proposed method. A new RGB-D activity dataset focusing on human-object interaction is further contributed, which presents more challenges for RGB-D activity benchmarking.
Jianfang Hu, Wei-Shi Zheng 0001, Jian-Huang Lai, Jianguo Zhang 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2017 Multi-task mid-level feature learning for micro-expression recognition
Jiachi He, Jianfang Hu, Wei-Shi Zheng 0001
Pattern Recognit.2
2017 Sparse transfer for facial shape-from-shading
Jianfang Hu, Wei-Shi Zheng 0001, Xiaohua Xie, Jian-Huang Lai
Pattern Recognit.1
2016 Real-Time RGB-D Activity Prediction by Soft Regression
Jianfang Hu, Wei-Shi Zheng 0001, Lianyang Ma, Gang Wang 0012, Jian-Huang Lai
ECCV (1)1
2016 Facial skin beautification via sparse representation over learned layer dictionary
abstract
In this paper, we propose a facial skin beautification framework to remove facial spots based on layer dictionary learning and sparse representation. More precisely, we first decompose the face image into three layers: lighting layer, detail layer and color layer. The corresponding detail layer dictionary are learned by using 60 thousands beauty images collected from the Internet. Thereafter, the detail layer of the image is reconstructed by using sparse representation. Moreover, a binary mask obtained from the learned layer is used to transform detail information from original detail layer to the learned one. The experiment results demonstrate that the proposed method is more effective in eliminating moles, flaws and wrinkles in face image compared with representative commercial systems like PicTreat, Portrait+, Portraitrue and MeituPic.
Xiaobin Chang, Xiaohua Xie, Jianfang Hu, Wei-Shi Zheng 0001
IJCNN4
2016 One-pass online learning: A local approach
Zhaoze Zhou, Wei-Shi Zheng 0001, Jianfang Hu, Yong Xu 0001, Jane You
Pattern Recognit.3
2016 Exemplar-Based Recognition of Human-Object Interactions
abstract
Human action can be recognized from a single still image by modeling human-object interactions (HOIs), which infers the mutual spatial structure information between human and the manipulated object as well as their appearance. Existing approaches rely heavily on accurate detection of human and object and estimation of human pose; they are thus sensitive to large variations of human poses, occlusion, and unsatisfactory detection of small size objects. To overcome this limitation, a novel exemplar-based approach is proposed in this paper. Our approach learns a set of spatial pose-object interaction exemplars, which are probabilistic density functions describing spatially how a person is interacting with a manipulated object for different activities. Specifically, a new framework consisting of an exemplar-based HOI descriptor and an associated matching model is formulated for robust human action recognition in still images. In addition, the framework is extended to perform HOI recognition in videos, where the proposed exemplar representation is used for implicit frame selection to negate irrelevant or noisy frames by temporal structured HOI modeling. Extensive experiments are carried out on two image action datasets and two video action datasets. The results demonstrate the effectiveness of our proposed methods and show that our approach is able to achieve state-of-the-art performance, compared with several recently proposed competitors.
Jianfang Hu, Wei-Shi Zheng 0001, Jian-Huang Lai, Shaogang Gong, Tao Xiang 0002
IEEE Trans. Circuits Syst. Video Technol.1
2015 Jointly learning heterogeneous features for RGB-D activity recognition
abstract
In this paper, we focus on heterogeneous feature learning for RGB-D activity recognition. Considering that features from different channels could share some similar hidden structures, we propose a joint learning model to simultaneously explore the shared and feature-specific components as an instance of heterogenous multi-task learning. The proposed model in an unified framework is capable of: 1) jointly mining a set of subspaces with the same dimensionality to enable the multi-task classifier learning, and 2) meanwhile, quantifying the shared and feature-specific components of features in the subspaces. To efficiently train the joint model, a three-step iterative optimization algorithm is proposed, followed by two inference models. Extensive results on three activity datasets have demonstrated the efficacy of the proposed method. In addition, a novel RGB-D activity dataset focusing on human-object interaction is collected for evaluating the proposed method, which will be made available to the community for RGB-D activity benchmarking and analysis.
Jianfang Hu, Wei-Shi Zheng 0001, Jian-Huang Lai, Jianguo Zhang 0001
CVPR1
2013 Recognising Human-Object Interaction via Exemplar Based Modelling
abstract
Human action can be recognised from a single still image by modelling Human-object interaction (HOI), which infers the mutual spatial structure information between human and object as well as their appearance. Existing approaches rely heavily on accurate detection of human and object, and estimation of human pose. They are thus sensitive to large variations of human poses, occlusion and unsatisfactory detection of small size objects. To overcome this limitation, a novel exemplar based approach is proposed in this work. Our approach learns a set of spatial pose-object interaction exemplars, which are density functions describing how a person is interacting with a manipulated object for different activities spatially in a probabilistic way. A representation based on our HOI exemplar thus has great potential for being robust to the errors in human/object detection and pose estimation. A new framework consists of a proposed exemplar based HOI descriptor and an activity specific matching model that learns the parameters is formulated for robust human activity recognition. Experiments on two benchmark activity datasets demonstrate that the proposed approach obtains state-of-the-art performance.
Jianfang Hu, Wei-Shi Zheng 0001, Jian-Huang Lai, Shaogang Gong, Tao Xiang 0002
ICCV1