VLDB 2026 Research / reviewers in the wild / expert
Hongxun Yao
dblp:y/HongxunYao
· DBLP profile ↗
275ranked-venue papers
2as first author
69since 2021 · last 2026
0000-0003-3298-2574ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 194 · 1 first-author · 42 since 2021Artificial intelligence and machine learning · 101 · 1 first-author · 40 since 2021Databases, data management, data science and information retrieval · 7 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 since 2021Human-computer interaction and ubiquitous computing · 6Computer networks · 1Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Content-aware Information Compression and Selection for Whole Slide Image AnalysisabstractRecent advances in multi-instance learning (MIL) have demonstrated impressive performance in whole slide image (WSI) analysis. However, current methods search for cues and draw conclusions from all instances or regions, resulting in excessive redundant computation and suboptimal representation quality due to irrelevant and uninformative feature interference. To address these issues, we propose CICS, an efficient and general framework that performs compact information compression and selection for high-efficiency WSI analysis. In particular, CICS features two key components: (1) context-aware compression (CAC), which partitions the instance space into sub-regions and applies learnable compression to discard irrelevant components, reduce computational complexity while facilitating information selection, and (2) global-proximity selective attention (GPSA), which cherry-picks the most informative representation with a proximity-assisted global dynamic selection strategy. Building upon these innovations, CICS forms a plug-and-play module that reduces computational complexity through compact instance representations while improving feature quality by preserving the most informative cues. Extensive experiments on six WSI classification and survival prediction datasets show that CICS consistently improves the performance of multiple representative MIL methods. It achieves 2.5%, 7.7%, and 3.9% accuracy gain over the state-of-the-art Transformer-based TransMIL, Mamba-based MambaMIL, and graph-based WIKG methods on the ESCA dataset. Hongxun Yao, Sicheng Zhao, Yi Xiao 0003 |
AAAI | 2 |
| 2026 | Retrieval-Augmented Camera Control for Video DiffusionabstractVideo Diffusion Models (VDMs) have demonstrated unprecedented creativity and realism in generating visual content. However, taming these models to strictly adhere to specific camera trajectories remains a persistent challenge. Existing approaches predominantly follow a "Warp-and-Prediction" paradigm, which forces models to learn complex geometric inpainting through extensive fine-tuning. This often compromises the model’s native generative diversity. To address this, we propose a novel training-free framework that shifts the paradigm to ”Retrieve-and-Refine”. First, we introduce Retrieval Augmented Patch Inpainting. By leveraging hybrid priors from foundation models (e.g., DINOv3 and VGGT), this module retrieves source patches to construct a trajectory-aligned "Draft Video", transforming the ill-posed inpainting problem into a draft video refine task. Subsequently, our Coupled Dual-Path Refinement elevates this draft into a photorealistic sequence. By dynamically injecting generative priors into a parallel control branch, this mechanism heals stitching artifacts and synthesizes coherent content for out-of-view regions. Extensive experiments demonstrate that our method achieves state-of-the-art performance, surpassing both training-free baselines and fully optimized methods in visual fidelity and camera control accuracy. Lining Wang, Hongxun Yao |
ICMR | 2 |
| 2026 | Augmenting and contrasting distortion for open panoramic segmentation
Sicheng Zhao, Jiankun Zhu, Xi Chen 0110, Hongxun Yao |
Sci. China Inf. Sci. | 6 |
| 2026 | SfMamba: Efficient source-free domain adaptation via selective scan modeling
Xi Chen 0110, Hongxun Yao, Sicheng Zhao, Jiankun Zhu, Kui Jiang |
Expert Syst. Appl. | 2 |
| 2026 | Noisy correspondence decomposition for robust composed image retrieval
Zhaopan Xu, Chenrui Zhou, Xi Chen 0110, Wangbo Zhao, Jianning Zhang, Xiaojiang Peng, Hongxun Yao, Kaipeng Zhang |
Expert Syst. Appl. | 9 |
| 2026 | CMPF: Harmonizing Cross-Model Prior Fusion for Open-Vocabulary Segmentation
Sicheng Zhao, Xi Chen 0110, Hongxun Yao, Haosen Yang 0003, Yanhao Zhang 0001, Sheng Jin 0002, Xiatian Zhu, Haonan Lu, Kui Jiang, Guiguang Ding |
Int. J. Comput. Vis. | 3 |
| 2026 | V3D: Enhancing text-to-3D synthesis through a view-consistent multi-view diffusion model
Lining Wang, Hongxun Yao |
Neurocomputing | 2 |
| 2026 | Dual-prompt-based binary matching for open set domain adaptation
Jidong Yang, Shouxu Jiang, Hongxun Yao, Lingji Xu, Sheng Jin 0002, Huicong Zhang, Zhaopan Xu |
Neurocomputing | 3 |
| 2026 | Noisy Correspondence Rectification in Multimodal Clustering Space for Cross-Modal MatchingabstractAs one of the most fundamental techniques in multimodal learning, cross-modal matching aims to project various sensory modalities into a shared feature space. To achieve this, massive and correctly aligned data pairs are required for model training. However, unlike unimodal datasets, multimodal datasets are extremely harder to collect and annotate precisely. As an alternative, the co-occurred data pairs (e.g., image-text pairs) collected from the Internet have been widely exploited in the area. Unfortunately, the cheaply collected dataset unavoidably contains many mismatched data pairs, which have been proven to be harmful to the model's performance. To address this, we propose BiCro++ (Improved Bidirectional Cross-modal Similarity Consistency). This module can be integrated into existing cross-modal matching models, enhancing their robustness against noisy data through self-adaptive soft labels that dynamically reflect the true correspondence of data pairs. The basic idea of BiCro++ is motivated by that - taking image-text matching as an example - similar images should have similar textual descriptions and vice versa. This bidirectional similarity consistency can be directly translated into soft labels as a self-supervision signal to train the matching model. To further refine soft label quality, BiCro++ first introduces a Diagonal-Dominance Purification process to identify reliable anchor points from noisy dataset as the reference for soft label estimation. Then it employs a Hybrid-level Codebook Alignment mechanism that establishes enhanced consistency in bidirectional cross-modal similarity. The experiments on three popular cross-modal matching datasets show that our method significantly improves the noise-robustness of various matching models, and surpasses the state-of-the-art method by an average of 5.3%, 3.1% and 6.4% in terms of recall, respectively. Shuo Yang 0006, Yancheng Long, Zeke Xie, Hongxun Yao, Min Xu 0001, Liqiang Nie |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2026 | Emotional conflict adaptation for multimodal sentiment analysis
Tingting Han 0003, Lingyun Yu 0004, Min Tan 0005, Zhou Yu 0001, Hongxun Yao |
Pattern Recognit. | 5 |
| 2026 | Training-Free Noisy Correspondence Rectification With Multimodal Conceptual Knowledge
Zhaopan Xu, Wangbo Zhao, Xi Chen 0110, Sheng Jin 0002, Hongxun Yao |
IEEE Signal Process. Lett. | 5 |
| 2026 | Learning With Dual Noisy Labels for Text-to-Image Person Re-IdentificationabstractText-to-image person re-identification (TIReID) aims to identify a target person from a given textual description. Although recent work has made significant progress, most of it implicitly assumes that the sample annotations are correct and that the cross-modal correspondence in each image-text pair is well aligned. However, such an assumption requires elaborately annotated datasets, which are expensive and even impossible to obtain in practice. To alleviate this issue, in this letter, we explore a new TIReID setting, termed learning with dual noisy labels, in which the model learns from data with both noisy identity labels and noisy correspondence. We propose a general framework called TDTD (Two stage framework forDual noise ofTIReID) to achieve this. In the first stage, a Noise-Aware Preliminary Learning (NAPL) strategy selects “easy” triplets to train a noise-tolerant initial model. In the second stage, the model leverages reliable representations from NAPL to automatically correct both identity and correspondence errors via soft-label estimation and is then fine-tuned on the entire dataset using a dual noise-robust triplet loss. Extensive experiments on three public benchmarks, CUHK-PEDES, ICFG-PEDES, and RSTPReID, demonstrate the performance and robustness of TDTD, achieving state-of-the-art results under dual noise conditions. Zhaopan Xu, Wangbo Zhao, Xiaojiang Peng, Hongxun Yao |
IEEE Signal Process. Lett. | 6 |
| 2026 | HEART: Emotionally Grounded Video Captioning via Hierarchical Emotion-Aligned RepresentationabstractEmotional Video Captioning (EVC) seeks to generate video descriptions that are both factually accurate and emotionally expressive. However, existing approaches often lack structured semantic grounding and fine-grained temporal modeling, leading to incomplete or emotionally inconsistent captions. To address these issues, we proposeHEART(HierarchicalEmotion-AlignedRepresentation withTemporal structure), a unified framework that jointly models hierarchical visual semantics and multi-scale temporal context. Specifically, HEART introduces a Hierarchical Semantic Extraction Module that decomposes visual content into entity-, action-, and event-level representations, providing a rich foundation for multi-level emotional alignment. A Temporal Pyramid Module captures short- and long-range temporal dependencies through multi-scale convolution, enabling temporally coherent captioning. Together, these components enable HEART to generate captions that are both emotionally grounded and temporally complete. To support this framework, we construct EmoStruct, a new benchmark dataset with fine-grained emotional annotations at the subject and predicate levels. Experiments on EmoStruct and public datasets demonstrate that HEART significantly outperforms prior methods in both semantic and emotional dimensions. Tingting Han 0003, Yuxuan Gong, Sicheng Zhao, Min Tan 0005, Zhou Yu 0001, Hongxun Yao |
IEEE Trans. Affect. Comput. | 6 |
| 2026 | PROMIND: Privacy-Protected Mental Health Intelligence with Noise Defense for Depression Recognition
Wuxin Shen, Hongxun Yao |
IEEE Trans. Affect. Comput. | 3 |
| 2026 | M2DAO-Talker: Harmonizing Multi-Granular Motion Decoupling and Alternating Optimization for Talking-Head GenerationabstractAudio-driven talking head generation holds significant potential for film production. While existing 3D methods have advanced motion modeling and content synthesis, they often produce rendering artifacts, such as motion blur, temporal jitter, and local penetration, due to limitations in representing stable, fine-grained motion fields. Through systematic analysis, we reformulate talking head generation into a unified framework comprising three steps: video preprocessing, motion representation, and rendering reconstruction. This framework underpins our proposed M2DAO-Talker, which addresses current limitations via multi-granular motion decoupling and alternating optimization. Specifically, we devise a novel 2D portrait preprocessing pipeline to extract frame-wise deformation control conditions (motion region segmentation masks, and camera parameters) to facilitate motion representation. To ameliorate motion modeling, we elaborate a multi-granular motion decoupling strategy, which independently models non-rigid (oral and facial) and rigid (head) motions for improved reconstruction accuracy. Meanwhile, a motion consistency constraint is developed to ensure head-torso kinematic consistency, thereby mitigating penetration artifacts caused by motion aliasing. In addition, an alternating optimization strategy is designed to iteratively refine facial and oral motion parameters, enabling more realistic video generation. Experiments across multiple datasets show that M2DAO-Talker achieves state-of-the-art performance, with the 2.43 dB PSNR improvement in generation quality and 0.64 gain in user-evaluated video realness versus TalkingGaussian while with 150 FPS inference speed. Our project homepage is https://m2dao-talker.github.io/M2DAO-Talk.github.io. Kui Jiang, Junjun Jiang, Hongxun Yao, Xiaopeng Fan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | PH-Mamba: Enhancing Mamba With Position Encoding and Harmonized Attention for Image Deraining and BeyondabstractMamba and its variants excel at modeling long-range dependencies with linear computational complexity, making them effective for diverse vision tasks. However, Mamba's reliance on unfolding 1D sequential representations necessitates multiple directional scans to recover lost spatial dependencies. This introduces significant computational overhead, redundant token traversal, and inefficiencies that compromise accuracy in real-world applications. To this end, we propose PH-Mamba, a novel framework integrating position encoding and harmonized attention for image deraining and beyond. PH-Mamba transforms Mamba's scanning process into a position-guided, unidirectional scanning that selectively prioritizes degradation-relevant tokens. Specifically, we devise a position-guided hybrid Mamba module (PHMM) that jointly encodes perturbation features alongside their spatial coordinates and harmonized representation to model consistent degradation patterns. Within PHMM, a harmonized Transformer is developed to focus on uncertain regions while suppressing noise interference, thereby improving spatial modeling fidelity. Additionally, we employ a vector decomposition and synthesis strategy to enable the unified representation layout to global degradation by directional scanning while minimizing redundancy. By cascading multiple PHMM blocks, PH-Mamba combines global positional guidance with local differential features to strengthen contextual learning. Extensive experiments demonstrate the superiority of PH-Mamba across low-level image restoration benchmarks. For example, compared to NeRD, PH-Mamba achieves a 0.60 dB PSNR improvement while requiring 88.9% fewer parameters, 36.2% less computation, and 63.0% faster inference time. Kui Jiang, Junjun Jiang, Xianming Liu 0005, Hongxun Yao, Chia-Wen Lin |
IEEE Trans. Image Process. | 4 |
| 2025 | OODML: Whole Slide Image Classification Meets Online Pseudo-Supervision and Dynamic Mutual LearningabstractBag-label-based multi-instance learning (MIL) has demonstrated significant performance in whole slide image (WSI) analysis, particularly in pseudo-label-based learning schemes. However, due to inaccurate feature representation and interference, existing MIL methods often yield unreliable pseudo-labels, which spawn undesired predictions. To address these issues, we propose an Online Pseudo-Supervision and Dynamic Mutual Learning (OODML) framework that enhances pseudo-label generation and feature representation while exploring their mutual learning to improve bag-level prediction. Specifically, we design an Adaptive Memory Bank (AMB) to collect the most informative components of the current WSI. We also introduce a Self-Progressive Feature Fusion (SPFF) module that integrates label-related historical information from the AMB with current semantic variations, thereby enhancing the representation of pseudo-bag tokens. Furthermore, we propose a Decision Revision Pseudo-Label (DRPL) generation scheme to explore intrinsic connections between pseudo-bag representations and bag-label predictions, resulting in more reliable pseudo-label generation. To alleviate redundant and ambiguous representations, the class-wise prior of pseudo-label prediction is borrowed to facilitate label-related feature learning and to update the AMB, forming a mutual refinement between feature representation and pseudo-label generation. Additionally, a Dynamic Decision-Making (DDM) module is developed to harmonize explicit and implicit representations of bag information for more robust decision-making. Extensive experiments on four datasets demonstrate that our OODML surpasses the state-of-the-art by 3.3% and 6.9% on the CAMELYON16 and TCGA Lung datasets. Kui Jiang, Hongxun Yao, Yi Xiao 0003, Zhongyuan Wang 0001 |
AAAI | 3 |
| 2025 | Bridge Then Begin Anew: Generating Target-Relevant Intermediate Model for Source-Free Visual Emotion AdaptationabstractVisual emotion recognition (VER), which aims at understanding humans' emotional reactions toward different visual stimuli, has attracted increasing attention. Given the subjective and ambiguous characteristics of emotion, annotating a reliable large-scale dataset is hard. For reducing reliance on data labeling, domain adaptation offers an alternative solution by adapting models trained on labeled source data to unlabeled target data. Conventional domain adaptation methods require access to source data. However, due to privacy concerns, source emotional data may be inaccessible. To address this issue, we propose an unexplored task: source-free domain adaptation (SFDA) for VER, which does not have access to source data during the adaptation process. To achieve this, we propose a novel framework termed Bridge then Begin Anew (BBA), which consists of two steps: domain-bridged model generation (DMG) and target-related model adaptation (TMA). First, the DMG bridges cross-domain gaps by generating an intermediate model, avoiding direct alignment between two VER datasets with significant differences. Then, the TMA begins training the target model anew to fit the target structure, avoiding the influence of source-specific knowledge. Extensive experiments are conducted on six SFDA settings for VER. The results demonstrate the effectiveness of BBA, which achieves remarkable performance gains compared with state-of-the-art SFDA methods and outperforms representative unsupervised domain adaptation approaches. Jiankun Zhu, Sicheng Zhao, Wenbo Tang, Zhaopan Xu, Tingting Han 0003, Pengfei Xu 0001, Hongxun Yao |
AAAI | 8 |
| 2025 | M3amba: Memory Mamba is All You Need for Whole Slide Image ClassificationabstractMulti-instance learning (MIL) has demonstrated impressive performance in whole slide image (WSI) analysis. However, existing approaches struggle with undesirable results and unbearable computational overhead due to the quadratic complexity of Transformers. Recently, Mamba has offered a feasible solution for modeling long-range dependencies with linear complexity. However, vanilla Mamba inherently suffers from contextual forgetting issues, making it ill-suited for capturing global dependencies across instances in large-scale WSIs. To address this, we propose a memory-driven Mamba network, dubbed M3amba, to fully explore the global latent relations among instances. Specifically, M3amba retains and iteratively updates historical information with a dynamic memory bank (DMB), thus overcoming the catastrophic forgetting defects of Mamba for long-term context representation. For better feature representation, M3amba involves an intra-group bidirectional Mamba (BiMamba) block to refine local interactions within groups. Meanwhile, we additionally perform cross-attention fusion to incorporate relevant historical information across groups, facilitating richer inter-group connections. The joint learning of inter- and intra-group representations with memory merits enables M3amba with a more powerful capability for achieving accurate and comprehensive WSI representation. Extensive experiments on four datasets demonstrate that M3amba outperforms the state-of-the-art by 6.2% and 7.0% in accuracy on the TCGA BRCA and TCGA Lung datasets while maintaining low computational costs. Kui Jiang, Yi Xiao 0003, Sicheng Zhao, Hongxun Yao |
CVPR | 5 |
| 2025 | Gaussian Constrained Diffeomorphic Deformation Network for Panoramic Semantic SegmentationabstractPanoramic semantic segmentation has garnered increasing attention due to its ability to provide comprehensive environmental perception. However, it requires a large number of annotated panoramic images to achieve satisfactory performance, which is costly. Recently, Domain Adaptation for Panoramic Semantic Segmentation (DA4PASS) has been proposed to reduce the reliance on annotated data by transferring segmentation models trained on annotated pinhole images to unlabelled panoramic images. Previous DA4PASS methods mainly focus on aligning features between pinhole and panoramic images, overlooking the unique appearance characteristics of panoramic images, particularly object distortion. To address the appearance discrepancies between pinhole and panoramic images, we propose Gaussian Constrained Diffeomorphic Deformation Network (GCDDN), which applies a panoramic deformation transformation obtained by Gaussian kernels to the annotated pinhole images. Specifically, GCDDN predicts multiple Gaussian kernels and performs first-order horizontal/vertical differences to obtain a naturally smooth and reversible panoramic deformation field, which is diffeomorphic. Due to its universality, GCDDN can be integrated into any domain adaptation (DA) method. Extensive experimental results demonstrate that integrating GCDDN leads to substantial improvements in both DA methods for pinhole images and those specifically designed for panoramic images, with a maximum gain of 1.80% in outdoor scenarios. Code is available at https://github.com/jingjiang02/GCDDN. Jiankun Zhu, Zhaopan Xu, Xi Chen 0110, Sicheng Zhao, Hongxun Yao |
ICASSP | 6 |
| 2025 | Multimodal and Multiple Prompts for BiologyabstractPrompt learning for vision-language models, e.g. CLIP, is a rapidly advancing topic, focusing on leveraging pre-trained foundational models to generalize to downstream tasks with minimal training samples. Current prompt-tuning methods primarily optimize on general datasets, yet their application in biological datasets remains limited due to the field’s specificity. To address this gap, we introduce Multimodal and Multiple Prompts for Biology (MMPB). Our approach incorporates a Vision Prompt Injection module, which injects diverse biological visual information from the vision branch into the language branch. Additionally, we propose a multiple prompts scheme tailored to the fine-grained naming conventions in biology. We assess the method’s effectiveness on generation to novel classes, and results demonstrate that MMPB significantly outperforms existing approaches across eight biological datasets. Specifically, MMPB achieves a 14.90% absolute improvement in novel classes performance and an 11.71% gain in harmonic mean compared to the state-of-the-art. Chonghuinan Wang, Xi Chen 0110, Hongxun Yao |
ICASSP | 3 |
| 2025 | Open-Vocabulary Visual Emotion Adaptation via Prompt LearningabstractVisual Emotion Recognition (VER) aims to identify emotions from visual content and has garnered significant attention in recent years due to its wide-ranging applications. Although deep learning-based methods have shown success in VER, they require extensive labeled data, which is costly. Unsupervised Domain Adaptation (UDA) methods can reduce reliance on annotated data by transferring models trained on labeled datasets to unlabeled data. However, these methods assume that the source and target domains share the same label space. In practice, this assumption is often violated due to the inherent ambiguity and subjectivity in emotion labeling. To address this limitation, we propose a novel prompt learning paradigm for open-vocabulary visual emotion UDA, termed Domain-specific Ensemble Prompting (DSEP). DSEP leverages psychological emotion models to unify emotion labels into a common space in an ensemble manner, enhancing the open-vocabulary capabilities of UDA. It then combines ensemble label prompts with domain-specific content prompts to achieve open-vocabulary UDA. To our knowledge, we are the first to explore open-vocabulary adaptation for VER. Extensive experiments demonstrate that DSEP consistently outperforms state-of-the-art methods across four public benchmarks. Zhaopan Xu, Sicheng Zhao, Xiaojiang Peng, Hongxun Yao |
ICASSP | 4 |
| 2025 | Learning Class Prototypes for Visual Emotion RecognitionabstractVisual emotion recognition (VER), which aims at understanding humans’ emotional reactions toward different visual stimuli, has attracted increasing attention. However, because of the subjectivity and complex nature of emotion, existing VER methods suffer from one or more of the following problems: 1) semantic gap: the large affective gap between visual clues and the emotional expressions; 2) overfitting: the lack of model robustness due to unclear features in the emotional category samples; 3) label ambiguity: the overlap between categories caused by diverse emotional responses. To address these limitations, we present a novel VER method named ProtoEmotion (PoE), exploring discriminative emotional representations by jointly learning prototypes of textual emotional expressions and visual features. Specifically, text prototypes build explicit textual features for each emotion category by extracting prototypes of learnable prompts from multiple aspects, reducing semantic differences. The visual prototypes capture the most defining image features of each category, providing a more robust and discriminative feature representation, while bringing samples closer together to reduce overfitting. In addition, to alleviate the label ambiguity, we propose a label smoothing algorithm based on the prototype distance. Extensive experiments demonstrate the effectiveness of PoE, which outperforms the state-of-the-art by 1.37% on FI and 1.52% on EmotionROI datasets. Jiankun Zhu, Sicheng Zhao, Zhaopan Xu, Wenbo Tang, Hongxun Yao |
ICASSP | 6 |
| 2025 | GMMamba: Group Masking Mamba for Whole Slide Image Classification
Hongxun Yao, Kui Jiang, Yi Xiao 0003, Sicheng Zhao |
ICCV | 2 |
| 2025 | LV-VTON: Long-Video Virtual Try-On via Enhanced Visual Autoregressive ModelingabstractVideo virtual try-on (VTON) aims to dress a target person in a desired garment while preserving the motion and identity in the original video. Generating long-duration VTON videos exacerbates challenges in achieving temporal coherence and visual fidelity. Existing image-based methods struggle with temporal consistency due to their frame-by-frame manner, while video-based approaches often sacrifice details for global coherence, resulting in blurry results. To address these limitations, we propose the first VTON framework tailored for long-video generation, namely LV-VTON, based on UNet Diffusion Transformer (UDiT) and Visual Autoregressive Modeling (VAR). Our LV-VTON framework includes three key components, i.e., Garment Appearance Module, Temporal Consistency Module, and Extended Video Synthesis Module. Notably, we collected a diverse and high-quality dataset, namely LongTry, to advance long video virtual try-on research. Extensive experiments demonstrate that LV-VTON excels at synthesizing various long VTON videos, outperforming state-of-the-art methods in detail preservation, temporal consistency, and long-video synthesis. Lulu Tian, Hongxun Yao, Ming Li 0073 |
ICME | 2 |
| 2025 | DreamAnimate: Temporal Consistency and Detail Preservation for Character AnimationabstractCharacter animation aims to generate realistic, high-quality videos from a reference image and target frames. However, existing methods struggle to balance fine-grained detail preservation with temporal consistency. This limitation results in artifacts like flickering and unrealistic deformations, especially in facial and hand regions. To address these challenges, we propose DreamAnimate, a novel framework that synthesizes temporally consistent and detail-rich animations. DreamAnimate integrates three modules: the Progressive Motion Estimation module ensures accurate motion alignment and temporal stability by refining keypoint heatmaps, the Global Affine Transformation module generates dense motion flows to handle complex motions and occlusions, and the Character Animation Fusion module combines intermediate synthesis using a UNet architecture and an Animation Fusion Network to produce high-quality animations. Extensive experiments demonstrate that DreamAnimate outperforms state-of-the-art methods, achieving superior fidelity and effectively capturing intricate facial expressions and hand movements. Code and models will be released at https://github.com/cslltian/DreamAnimate in the near future. Lulu Tian, Hongxun Yao, Zhaopan Xu, Jiankun Zhu, Xi Chen 0110, Yuxin Hou |
ICME | 2 |
| 2025 | RetrievFace: Retrieval-Enhanced Diffusion for Controllable Text-Guided Face Editing
Lulu Tian, Hongxun Yao |
ICMR | 2 |
| 2025 | Emotion in a Bottle: Information Bottleneck Guided Disentanglement for Emotion Domain AdaptationabstractVisual emotion recognition (VER), which aims to understand human emotional reactions to different visual stimuli, has garnered increasing attention. However, the inherent ambiguity of emotional features presents significant challenges for data annotation in supervised learning paradigms. To address this limitation, emotion domain adaptation (EDA) facilitates knowledge transfer from labeled source domains to unlabeled target domains. Recently, large visual-language models such as CLIP have demonstrated impressive transfer performance on traditional UDA tasks. However, when generalizing to more abstract concepts such as emotion, the misalignment between CLIP and emotion spaces greatly affects the model performance. To address these challenges, we propose a CLIP-based emotion disentanglement (EmoD) framework designed for EDA. Leveraging perspectives from information bottleneck theory, EmoD implements a disentangler network that extracts emotion-specific features while removing redundant emotion-agnostic information. It also incorporates cross-domain feature alignment to reduce the affective gap between domains. Experimental evaluations in six EDA settings demonstrate that EmoD achieves state-of-the-art performance, surpassing traditional CLIP-based UDA methods by an average of 2.53%. Jiankun Zhu, Sicheng Zhao, Lulu Tian, Xi Chen 0110, Hongxun Yao |
ACM Multimedia | 6 |
| 2025 | 2D Semantic-Guided Semantic Scene Completion
Xianzhu Liu, Haozhe Xie, Shengping Zhang, Hongxun Yao, Rongrong Ji, Liqiang Nie, Dacheng Tao |
Int. J. Comput. Vis. | 4 |
| 2025 | Drop inherent biases: Multi-level attention calibration for robust cross-domain few-shot classification
Hongxun Yao |
Neurocomputing | 3 |
| 2025 | FMVP: Fine-grained Meta Visual Prompt enabled domain-specific few-shot classification
Hongxun Yao |
Neurocomputing | 2 |
| 2025 | Bias-in-debias-out: Hierarchical channel-spatial bias calibration for cross-domain few-shot classification
Hongxun Yao |
Knowl. Based Syst. | 2 |
| 2025 | Collaborative optimization for whole slide image classification
Hongxun Yao, Kui Jiang, Yi Xiao 0003 |
Knowl. Based Syst. | 2 |
| 2025 | GraphMamba: Whole slide image classification meets graph-driven selective state space model
Hongxun Yao, Sicheng Zhao, Kui Jiang, Yi Xiao 0003 |
Pattern Recognit. | 2 |
| 2025 | Patch-Based Spatio-Temporal Deformable Attention BiRNN for Video DeblurringabstractSuccessful video deblurring relies on effectively using sharp pixels from other frames to recover the blurry pixels of the current frame. However, mainstream methods only use estimated optical flows to align and fuse features from adjacent frames without considering the pixel-wise blur levels, leading to the introduction of blurry pixels from adjacent frames. Furthermore, these methods fail to effectively exploit information from the entire input video. To address these limitations, we propose STDANet++, which redesigns the state-of-the-art method STDANet by introducing patch-based spatio-temporal deformable attention (PSTDA) module and long-term frame fusion (LTFF) module to the BiRNN-based structure. By effectively utilizing sharp information across the entire video, the proposed method outperforms state-of-the-art methods on the GoPro, DVD and BSD datasets, according to our experimental results. The source code is available at https://github.com/huicongzhang/STDANetPP. Huicong Zhang, Haozhe Xie, Shengping Zhang, Hongxun Yao |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Guest Editorial: Special Issue on Fuzzy Affective Computing Systems
Sicheng Zhao, Hongxun Yao, Xinde Li, James Z. Wang 0001, Björn W. Schuller |
IEEE Trans. Fuzzy Syst. | 2 |
| 2024 | MotionMAE: Self-supervised Video Representation Learning with Motion-Aware Masked Autoencoders
Haosen Yang 0003, Deng Huang, Jiannan Wu, Hongxun Yao, Yi Jiang 0009, Xiatian Zhu, Zehuan Yuan |
BMVC | 5 |
| 2024 | Blur-Aware Spatio-Temporal Sparse Transformer for Video DeblurringabstractVideo deblurring relies on leveraging information from other frames in the video sequence to restore the blurred regions in the current frame. Mainstream approaches employ bidirectional feature propagation, spatio-temporal Transformers, or a combination of both to extract information from the video sequence. However, limitations in memory and computational resources constraints the temporal window length of the spatio-temporal transformer, preventing the extraction of longer temporal contextual information from the video sequence. Additionally, bidirectional feature propagation is highly sensitive to inaccurate optical flow in blurry frames, leading to error accumulation during the propagation process. To address these issues, we propose BSSTNet, Blur-aware Spatio-temporal Sparse Transformer Network. It introduces the blur map, which converts the originally dense attention into a sparse form, enabling a more extensive utilization of information throughout the entire video sequence. Specifically, BSSTNet (1) uses a longer temporal window in the transformer, lever-aging information from more distant frames to restore the blurry pixels in the current frame. (2) introduces bidirectional feature propagation guided by blur maps, which reduces error accumulation caused by the blur frame. The experimental results demonstrate the proposed BSSTNet out-performs the state-of-the-art methods on the GoPro and DVD datasets. Huicong Zhang, Haozhe Xie, Hongxun Yao |
CVPR | 3 |
| 2024 | Dynamic Policy-Driven Adaptive Multi-Instance Learning for Whole Slide Image ClassificationabstractMulti-Instance Learning (MIL) has shown impressive performance for histopathology whole slide image (WSI) analysis using bags or pseudo-bags. It involves instance sampling, feature representation, and decision-making. However, existing MIL-based technologies at least suffer from one or more of the following problems: 1) requiring high storage and intensive preprocessing for numerous instances (sampling); 2) potential over-fitting with limited knowledge to predict bag labels (feature representation); 3) pseudo-bag counts and prior biases affect model robustness and generalizability (decision-making). Inspired by clinical diagnostics, using the past sampling instances can facili-tate the final WSI analysis, but it is barely explored in prior technologies. To break free these limitations, we integrate the dynamic instance sampling and reinforcement learning into a unified framework to improve the instance selection and feature aggregation, forming a novel Dynamic Policy Instance Selection (DPIS) scheme for better and more cred-ible decision-making. Specifically, the measurement of feature distance and reward function are employed to boost continuous instance sampling. To alleviate the over-fitting, we explore the latent global relations among instances for more robust and discriminative feature representation while establishing reward and punishment mechanisms to correct biases in pseudo-bags using contrastive learning. These strategies form the final Dynamic Policy-Driven Adaptive Multi-Instance Learning (PAMIL) method for WSI tasks. Extensive experiments reveal that our PAMIL method outperforms the state-of-the-art by 3.8% on CAMELYON16 and 4.4% on TCGA lung cancer datasets. Kui Jiang, Hongxun Yao |
CVPR | 3 |
| 2024 | Dataset Growth
Ziheng Qin, Zhaopan Xu, Zangwei Zheng, Zebang Cheng, Hao Tang 0005, Baigui Sun, Xiaojiang Peng, Radu Timofte, Hongxun Yao, Kai Wang 0036, Yang You 0001 |
ECCV (9) | 11 |
| 2024 | Uncertainty-aware pseudo-label filtering for source-free unsupervised domain adaptation
Xi Chen 0110, Haosen Yang 0003, Huicong Zhang, Hongxun Yao, Xiatian Zhu |
Neurocomputing | 4 |
| 2024 | Hierarchical pose net: spatial hierarchical body tree driven multi-person pose estimation
Haoran Li 0012, Hongxun Yao, Yuxin Hou |
Multim. Tools Appl. | 2 |
| 2024 | Artistic image synthesis with tag-guided correlation matching
Dilin Liu, Hongxun Yao |
Multim. Tools Appl. | 2 |
| 2024 | Artistic image synthesis from unsupervised segmentation maps
Dilin Liu, Hongxun Yao, Xiusheng Lu |
Multim. Tools Appl. | 2 |
| 2024 | In-use calibration: improving domain-specific fine-grained few-shot recognition
Hongxun Yao |
Neural Comput. Appl. | 2 |
| 2024 | Stereo Image Restoration via Attention-Guided Correspondence LearningabstractAlthough stereo image restoration has been extensively studied, most existing work focuses on restoring stereo images with limited horizontal parallax due to the binocular symmetry constraint. Stereo images with unlimited parallax (e.g., large ranges and asymmetrical types) are more challenging in real-world applications and have rarely been explored so far. To restore high-quality stereo images with unlimited parallax, this paper proposes an attention-guided correspondence learning method, which learns both self- and cross-views feature correspondence guided by parallax and omnidirectional attention. To learn cross-view feature correspondence, a Selective Parallax Attention Module (SPAM) is proposed to interact with cross-view features under the guidance of parallax attention that adaptively selects receptive fields for different parallax ranges. Furthermore, to handle asymmetrical parallax, we propose a Non-local Omnidirectional Attention Module (NOAM) to learn the non-local correlation of both self- and cross-view contexts, which guides the aggregation of global contextual features. Finally, we propose an Attention-guided Correspondence Learning Restoration Network (ACLRNet) upon SPAMs and NOAMs to restore stereo images by associating the features of two views based on the learned correspondence. Extensive experiments on five benchmark datasets demonstrate the effectiveness and generalization of the proposed method on three stereo image restoration tasks including super-resolution, denoising, and compression artifact reduction. Shengping Zhang, Wei Yu 0004, Feng Jiang 0001, Liqiang Nie, Hongxun Yao, Qingming Huang, Dacheng Tao |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | End-to-End Human Instance MattingabstractHuman instance matting aims to estimate an alpha matte for each human instance in an image, which is extremely challenging and has rarely been studied so far. Despite some efforts to use instance segmentation to generate a trimap for each instance and apply trimap-based matting methods, the resulting alpha mattes are often inaccurate due to inaccurate segmentation. In addition, this approach is computationally inefficient due to multiple executions of the matting method. To address these problems, this paper proposes a novel End-to-End Human Instance Matting (E2E-HIM) framework for simultaneous multiple instance matting in a more efficient manner. Specifically, a general perception network first extracts image features and decodes instance contexts into latent codes. Then, a united guidance network exploits spatial attention and semantics embedding to generate united semantics guidance, which encodes the locations and semantic correspondences of all instances. Finally, an instance matting network decodes the image features and united semantics guidance to predict all instance-level alpha mattes. In addition, we construct a large-scale human instance matting dataset (HIM-100K) comprising over 100,000 human images with instance alpha matte labels. Experiments on HIM-100K demonstrate the proposed E2E-HIM outperforms the existing methods on human instance matting with 50% lower errors and 5× faster speed (6 instances in a 640 × 640 image). Experiments on the PPM-100, RWP-636, and P3M datasets demonstrate that E2E-HIM also achieves competitive performance on traditional human matting. Qinglin Liu, Shengping Zhang, Quanling Meng, Bineng Zhong 0001, Peiqiang Liu, Hongxun Yao |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Limb-Aware Virtual Try-On Network With Progressive Clothing WarpingabstractImage-based virtual try-on aims to transfer an in-shop clothing image to a person image. Most existing methods adopt a single global deformation to perform clothing warping directly, which lacks fine-grained modeling of in-shop clothing and leads to distorted clothing appearance. In addition, existing methods usually fail to generate limb details well because they are limited by the used clothing-agnostic person representation without referring to the limb textures of the person image. To address these problems, we propose Limb-aware Virtual Try-on Network named PL-VTON, which performs fine-grained clothing warping progressively and generates high-quality try-on results with realistic limb details. Specifically, we present Progressive Clothing Warping (PCW) that explicitly models the location and size of in-shop clothing and utilizes a two-stage alignment strategy to progressively align the in-shop clothing with the human body. Moreover, a novel gravity-aware loss that considers the fit of the person wearing clothing is adopted to better handle the clothing edges. Then, we design Person Parsing Estimator (PPE) with a non-limb target parsing map to semantically divide the person into various regions, which provides structural constraints on the human body and therefore alleviates texture bleeding between clothing and body regions. Finally, we introduce Limb-aware Texture Fusion (LTF) that focuses on generating realistic details in limb regions, where a coarse try-on result is first generated by fusing the warped clothing image with the person image, then limb textures are further fused with the coarse result under limb-aware guidance to refine limb details. Extensive experiments demonstrate that our PL-VTON outperforms the state-of-the-art methods both qualitatively and quantitatively. Shengping Zhang, Weigang Zhang, Xiangyuan Lan, Hongxun Yao, Qingming Huang |
IEEE Trans. Multim. | 5 |
| 2023 | BiCro: Noisy Correspondence Rectification for Multi-modality Data via Bi-directional Cross-modal Similarity ConsistencyabstractAs one of the most fundamental techniques in multi-modal learning, cross-modal matching aims to project various sensory modalities into a shared feature space. To achieve this, massive and correctly aligned data pairs are required for model training. However, unlike unimodal datasets, multimodal datasets are extremely harder to collect and annotate precisely. As an alternative, the co-occurred data pairs (e.g., image-text pairs) collected from the Internet have been widely exploited in the area. Unfortunately, the cheaply collected dataset unavoidably contains many mismatched data pairs, which have been proven to be harmful to the model's performance. To address this, we propose a general framework called BiCro (Bidirectional Cross-modal similarity consistency), which can be easily integrated into existing cross-modal matching models and improve their robustness against noisy data. Specifically, BiCro aims to estimate soft labels for noisy data pairs to reflect their true correspondence degree. The basic idea of BiCro is motivated by that – taking image-text matching as an example – similar images should have similar textual descriptions and vice versa. Then the consistency of these two similarities can be recast as the estimated soft labels to train the matching model. The experiments on three popular cross-modal matching datasets demonstrate that our method significantly improves the noise-robustness of various matching models, and surpass the state-of-the-art by a clear margin. The code is available at https://github.com/xu5zhao/BiCro. Shuo Yang 0006, Zhaopan Xu, Kai Wang 0036, Yang You 0001, Hongxun Yao, Tongliang Liu, Min Xu 0001 |
CVPR | 5 |
| 2023 | BRMR: TAL Based on Boundary Refinement and Multi-scale Regression
Jiankun Zhu, Lining Wang, Hongxun Yao |
ICIG (4) | 4 |
| 2023 | Graph Convolutional GRU for Music-Oriented Dance Choreography GenerationabstractMusic-oriented dance choreography is a complex art form that requires consideration of various factors in music understanding and physical aesthetics. In addition, learning model to generate dance sequences automatically needs to overcome not only the challenges from choreography but also the complexity of human skeleton structure. In this paper, we propose a graph convolutional GRU (GraphGRU) model for music-oriented dance choreography by integrating the graph convolutional networks into the gated recurrent unit. The GraphGRU can model the spatial relationship among the human skeleton joints and the temporal relationship of the dancing sequences. Besides, we define multiple dance-driven kinematic graphs for GraphGRU and establish a path attention block to fuse the graph features. Moreover, we design a reverse-generation loss to ensure consistency between the generated dance and the music. Extensive experiments demonstrate that the proposed model GraphGRU can effectively solve the task of dance choreography. Yuxin Hou, Hongxun Yao, Haoran Li 0012 |
ICME | 2 |
| 2023 | Learning Enriched Hop-Aware Correlation for Robust 3D Human Pose Estimation
Shengping Zhang, Chenyang Wang 0002, Liqiang Nie, Hongxun Yao, Qingming Huang, Qi Tian 0001 |
Int. J. Comput. Vis. | 4 |
| 2023 | Correction to: Learning Enriched Hop-Aware Correlation for Robust 3D Human Pose Estimation
Shengping Zhang, Chenyang Wang 0002, Liqiang Nie, Hongxun Yao, Qingming Huang, Qi Tian 0001 |
Int. J. Comput. Vis. | 4 |
| 2023 | MIFNet: Multiple instances focused temporal action proposal generation
Lining Wang, Hongxun Yao, Haosen Yang 0003, Sibo Wang 0012, Sheng Jin 0002 |
Neurocomputing | 2 |
| 2023 | Focus nuance and toward diversity: exploring domain-specific fine-grained few-shot recognition
Hongxun Yao |
Neural Comput. Appl. | 2 |
| 2023 | FakePoI: A Large-Scale Fake Person of Interest Video Detection Benchmark and a Strong BaselineabstractDeepfake technique can synthesize realistic images, audios, and videos, facilitating the thriving of entertainment, education, healthcare, and other industries. However, its abuse may pose potential threats to personal privacy, social stability, and even national security. Therefore, the development of deepfake detection methods is attracting more and more attention. Existing works mainly focus on the detection of common videos for entertainment purposes. In contrast, fake videos maliciously synthesized for Person of Interest (PoI, i.e., who is in an authoritative position and has broadly public influences) are much more harmful to society because of celebrity endorsement. However, there is no particular benchmark for driving related research in the community. Motivated by this observation, we present the first large-scale benchmark dataset, named FakePoI, to enable the research on fake PoI detection. It contains numerous fake videos of important people from all walks of life, e.g., police chiefs, city mayors, famous artists, and well-known Internet bloggers. In summary, our FakePoI includes 11092 synthesized videos where only a few clips rather than the entire are fake. Previous fake detection algorithms deteriorate heavily or even fail on our FakePoI due to two main challenges. On the one hand, the rich diversity of our fake videos makes it pretty difficult to find universally applicable patterns for detection. On the other hand, the high credibility contributed by the presence of real frames easily confuses a common detector. To tackle these challenges, we present an amplifier framework, highlighting the feature gap between real and generated video frames. Specifically, we present a quadruplet loss to narrow the distance of all real PoIs and meanwhile push away each real and fake PoI in embedding space. We implement our framework and conduct extensive experiments on the proposed benchmark. The quantitative results demonstrate that our approach outperforms existing methods significantly, setting a strong baseline on FakePoI. The qualitative analysis also shows its superiority. We will release our dataset and code athttps://github.com/cslltian/deepfake-detectionto encourage future research on this valuable area. Lulu Tian, Hongxun Yao, Ming Li 0073 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Unsupervised Low-Light Video Enhancement With Spatial-Temporal Co-Attention TransformerabstractExisting low-light video enhancement methods are dominated by Convolution Neural Networks (CNNs) that are trained in a supervised manner. Due to the difficulty of collecting paired dynamic low/normal-light videos in real-world scenes, they are usually trained on synthetic, static, and uniform motion videos, which undermines their generalization to real-world scenes. Additionally, these methods typically suffer from temporal inconsistency (e.g., flickering artifacts and motion blurs) when handling large-scale motions since the local perception property of CNNs limits them to model long-range dependencies in both spatial and temporal domains. To address these problems, we propose the first unsupervised method for low-light video enhancement to our best knowledge, named LightenFormer, which models long-range intra- and inter-frame dependencies with a spatial-temporal co-attention transformer to enhance brightness while maintaining temporal consistency. Specifically, an effective but lightweight S-curve Estimation Network (SCENet) is first proposed to estimate pixel-wise S-shaped non-linear curves (S-curves) to adaptively adjust the dynamic range of an input video. Next, to model the temporal consistency of the video, we present a Spatial-Temporal Refinement Network (STRNet) to refine the enhanced video. The core module of STRNet is a novel Spatial-Temporal Co-attention Transformer (STCAT), which exploits multi-scale self- and cross-attention interactions to capture long-range correlations in both spatial and temporal domains among frames for implicit motion estimation. To achieve unsupervised training, we further propose two non-reference loss functions based on the invertibility of the S-curve and the noise independence among frames. Extensive experiments on the SDSD and LLIV-Phone datasets demonstrate that our LightenFormer outperforms state-of-the-art methods. Xiaoqian Lv, Shengping Zhang, Chenyang Wang 0002, Weigang Zhang, Hongxun Yao, Qingming Huang |
IEEE Trans. Image Process. | 5 |
| 2022 | Temporal Action Proposal Generation with Background ConstraintabstractTemporal action proposal generation (TAPG) is a challenging task that aims to locate action instances in untrimmed videos with temporal boundaries. To evaluate the confidence of proposals, the existing works typically predict action score of proposals that are supervised by the temporal Intersection-over-Union (tIoU) between proposal and the ground-truth. In this paper, we innovatively propose a general auxiliary Background Constraint idea to further suppress low-quality proposals, by utilizing the background prediction score to restrict the confidence of proposals. In this way, the Background Constraint concept can be easily plug-and-played into existing TAPG methods (BMN, GTAD). From this perspective, we propose the Background Constraint Network (BCNet) to further take advantage of the rich information of action and background. Specifically, we introduce an Action-Background Interaction module for reliable confidence evaluation, which models the inconsistency between action and background by attention mechanisms at the frame and clip levels. Extensive experiments are conducted on two popular benchmarks, ActivityNet-1.3 and THUMOS14. The results demonstrate that our method outperforms state-of-the-art methods. Equipped with the existing action classifier, our method also achieves remarkable performance on the temporal action localization task. Haosen Yang 0003, Lining Wang, Sheng Jin 0002, Boyang Xia, Hongxun Yao, Hujie Huang |
AAAI | 6 |
| 2022 | Spatio-Temporal Deformable Attention Network for Video Deblurring
Huicong Zhang, Haozhe Xie, Hongxun Yao |
ECCV (16) | 3 |
| 2021 | Asynchronous Teacher Guided Bit-wise Hard Mining for Online HashingabstractOnline hashing for streaming data has attracted increasing attention recently. However, most existing algorithms focus on batch inputs and instance-balanced optimization, which is limited in the single datum input case and does not match the dynamic training in online hashing. Furthermore, constantly updating the online model with new-coming samples will inevitably lead to the catastrophic forgetting problem. In this paper, we propose a novel online hashing method to handle the above-mentioned issues jointly, termed Asynchronus Teacher-Guided Bit-wise Hard Mining for Online Hashing. Firstly, to meet the needs of datum-wise online hashing, we design a novel binary codebook that is discriminative to separate different classes. Secondly, we propose a novel semantic loss (termed bit-wise attention loss) to dynamically focus on hard samples of each bit during training. Last but not least, we design a asynchronous knowledge distillation scheme to alleviate the catastrophic forgetting problem, where the teacher model is delaying updated to maintain the old knowledge, guiding the student model learning. Extensive experiments conducted on two public benchmarks demonstrate the favorable performance of our method over the state-of-the-arts. Sheng Jin 0002, Qin Zhou 0002, Hongxun Yao, Yao Liu 0014, Xian-Sheng Hua 0001 |
AAAI | 3 |
| 2021 | Efficient Regional Memory Network for Video Object SegmentationabstractRecently, several Space-Time Memory based networks have shown that the object cues (e.g. video frames as well as the segmented object masks) from the past frames are useful for segmenting objects in the current frame. However, these methods exploit the information from the memory by global-to-global matching between the current and past frames, which lead to mismatching to similar objects and high computational complexity. To address these problems, we propose a novel local-to-local matching solution for semi-supervised VOS, namely Regional Memory Network (RMNet). In RMNet, the precise regional memory is constructed by memorizing local regions where the target objects appear in the past frames. For the current query frame, the query regions are tracked and predicted based on the optical flow estimated from the previous frame. The proposed local-to-local matching effectively alleviates the ambiguity of similar objects in both memory and query frames, which allows the information to be passed from the regional memory to the query region efficiently and effectively. Experimental results indicate that the proposed RM-Net performs favorably against state-of-the-art methods on the DAVIS and YouTube-VOS datasets. Haozhe Xie, Hongxun Yao, Shangchen Zhou, Shengping Zhang, Wenxiu Sun |
CVPR | 2 |
| 2021 | Adaptive Spatio-Temporal Convolutional Network for Video Deblurring
Fengzhi Duan, Hongxun Yao |
ICIG (3) | 2 |
| 2021 | 3D Reconstruction from Single-View Image Using Feature Selection
Hongxun Yao |
ICIG (3) | 2 |
| 2021 | Visual Chirality Meets Freehand SketchesabstractVisual chirality measures the distribution variation of visual data under transformation, while it has not been explored in freehand sketches yet. In this paper, we investigate the vertical flipping associated with visual chirality in freehand sketches. Our analysis of investigation results reveals that the vertical flipping shows a high degree of visual chirality. To utilize the high-level cues automatically discovered by predicting the vertical flipping, we propose a Visual Chirality Attention (VCA) module for deep CNNs, which consists of two sequential sub-modules: channel and chirality attention. Experimental results of sketch recognition on TU-Berlin dataset show that our method performs more favorably against state-of-the-art attention-based methods. Our code can be found at https://github.com/zhengyinghit/VCANet. Ying Zheng 0009, Yiyi Zhang 0001, Xiaogang Xu 0001, Jun Wang 0134, Hongxun Yao |
ICIP | 5 |
| 2021 | Image editing with varying intensities of processing
Yasi Wang, Yuankai Qi, Hongxun Yao, Dong Gong, Qi Wu 0001 |
Comput. Vis. Image Underst. | 3 |
| 2021 | Toward 3D object reconstruction from stereo images
Haozhe Xie, Hongxun Yao, Shangchen Zhou, Shengping Zhang, Xiaojun Tong, Wenxiu Sun |
Neurocomputing | 2 |
| 2021 | Sketch-specific data augmentation for freehand sketch recognition
Ying Zheng 0009, Hongxun Yao, Xiaoshuai Sun, Shengping Zhang, Sicheng Zhao, Fatih Porikli |
Neurocomputing | 2 |
| 2021 | Unsupervised Discrete Hashing With Affinity SimilarityabstractIn recent years, supervised hashing has been validated to greatly boost the performance of image retrieval. However, the label-hungry property requires massive label collection, making it intractable in practical scenarios. To liberate the model training procedure from laborious manual annotations, some unsupervised methods are proposed. However, the following two factors make unsupervised algorithms inferior to their supervised counterparts: (1) Without manually-defined labels, it is difficult to capture the semantic information across data, which is of crucial importance to guide robust binary code learning. (2) The widely adopted relaxation on binary constraints results in quantization error accumulation in the optimization procedure. To address the above-mentioned problems, in this paper, we propose a novel Unsupervised Discrete Hashing method (UDH). Specifically, to capture the semantic information, we propose a balanced graph-based semantic loss which explores the affinity priors in the original feature space. Then, we propose a novel self-supervised loss, termed orthogonal consistent loss, which can leverage semantic loss of instance and impose independence of codes. Moreover, by integrating the discrete optimization into the proposed unsupervised framework, the binary constraints are consistently preserved, alleviating the influence of quantization errors. Extensive experiments demonstrate that UDH outperforms state-of-the-art unsupervised methods for image retrieval. Sheng Jin 0002, Hongxun Yao, Qin Zhou 0002, Yao Liu 0014, Jianqiang Huang 0001, Xian-Sheng Hua 0001 |
IEEE Trans. Image Process. | 2 |
| 2021 | Deep Semantic Parsing of Freehand Sketches With Homogeneous Transformation, Soft-Weighted Loss, and Staged LearningabstractIn this paper, we propose a novel deep framework for part-level semantic parsing of freehand sketches, which makes three main contributions that are experimentally shown to have substantial practical merit. First, we propose a homogeneous transformation method to address the problem of domain adaptation. For the task of sketch parsing, there is no available data of labeled freehand sketches that can be directly used for model training. An alternative solution is to learn from datasets of real image parsing, while the domain adaptation is an inevitable problem. Unlike existing methods that utilize the edge maps of real images to approximate freehand sketches, the proposed homogeneous transformation method transforms the data from domains of real images and freehand sketches into a homogeneous space to minimize the semantic gap. Second, we design a soft-weighted loss function as guidance for the training process, which gives attention to both the ambiguous label boundary and class imbalance. Third, we present a staged learning strategy to improve the parsing performance of the trained model, which takes advantage of the shared information and specific characteristic from different sketch categories. Extensive experimental results demonstrate the effectiveness of the above three methods. Specifically, to evaluate the generalization ability of our homogeneous transformation method, additional experiments for the task of sketch-based image retrieval are conducted on the QMUL FG-SBIR dataset. Finally, by integrating the proposed three methods into a unified framework of deep semantic sketch parsing (DeepSSP), we achieve the state-of-the-art on the public SketchParse dataset. Ying Zheng 0009, Hongxun Yao, Xiaoshuai Sun |
IEEE Trans. Multim. | 2 |
| 2020 | SSAH: Semi-Supervised Adversarial Deep Hashing with Self-Paced Hard Sample GenerationabstractDeep hashing methods have been proved to be effective and efficient for large-scale Web media search. The success of these data-driven methods largely depends on collecting sufficient labeled data, which is usually a crucial limitation in practical cases. The current solutions to this issue utilize Generative Adversarial Network (GAN) to augment data in semi-supervised learning. However, existing GAN-based methods treat image generations and hashing learning as two isolated processes, leading to generation ineffectiveness. Besides, most works fail to exploit the semantic information in unlabeled data. In this paper, we propose a novel Semi-supervised Self-pace Adversarial Hashing method, named SSAH to solve the above problems in a unified framework. The SSAH method consists of an adversarial network (A-Net) and a hashing network (H-Net). To improve the quality of generative images, first, the A-Net learns hard samples with multi-scale occlusions and multi-angle rotated deformations which compete against the learning of accurate hashing codes. Second, we design a novel self-paced hard generation policy to gradually increase the hashing difficulty of generated samples. To make use of the semantic information in unlabeled ones, we propose a semi-supervised consistent loss. The experimental results show that our method can significantly improve state-of-the-art models on both the widely-used hashing datasets and fine-grained datasets. Sheng Jin 0002, Shangchen Zhou, Yao Liu 0014, Chao Chen 0026, Xiaoshuai Sun, Hongxun Yao, Xian-Sheng Hua 0001 |
AAAI | 6 |
| 2020 | GRNet: Gridding Residual Network for Dense Point Cloud Completion
Haozhe Xie, Hongxun Yao, Shangchen Zhou, Jiageng Mao, Shengping Zhang, Wenxiu Sun |
ECCV (9) | 2 |
| 2020 | PRF-Ped: Multi-scale Pedestrian Detector with Prior-based Receptive FieldabstractMulti-scale feature representation is a common strategy to handle the scale variation in pedestrian detection. Existing methods simply utilize the convolutional pyramidal features for multi-scale representation. However, they rarely pay attention to the differences among different feature scales and extract multi-scale features from a single feature map, which may make the detectors sensitive to scale-variance in multi-scale pedestrian detection. In this paper, we introduce a bidirectional feature enhancement module (BFEM) to augment the semantic information of low-level features and the localization information of high-level features. In addition, we propose a prior-based receptive field block (PRFB) for multi-scale pedestrian feature extraction, where the receptive field is closer to the aspect ratio of the pedestrian target. Consequently, it is less affected by the surrounding background when extracting features. Experimental results indicate that the proposed method outperforms the state-of-the-art methods on the CityPersons and Caltech datasets. Yuzhi Tan, Hongxun Yao, Haoran Li 0012, Xiusheng Lu, Haozhe Xie |
ICPR | 2 |
| 2020 | An Effective Way to Boost Black-Box Adversarial Attack
Xinjie Feng, Hongxun Yao, Wenbin Che, Shengping Zhang |
MMM (1) | 2 |
| 2020 | Pix2Vox++: Multi-scale Context-aware 3D Object Reconstruction from Single and Multiple Images
Haozhe Xie, Hongxun Yao, Shengping Zhang, Shangchen Zhou, Wenxiu Sun |
Int. J. Comput. Vis. | 2 |
| 2020 | Actionness-pooled Deep-convolutional Descriptor for fine-grained action recognition
Tingting Han 0003, Hongxun Yao, Xiaoshuai Sun, Wenlong Xie, Sicheng Zhao, Wei Yu 0004 |
Neurocomputing | 2 |
| 2020 | TVENet: Temporal variance embedding network for fine-grained action representation
Tingting Han 0003, Hongxun Yao, Wenlong Xie, Xiaoshuai Sun, Sicheng Zhao, Jun Yu 0002 |
Pattern Recognit. | 2 |
| 2020 | Conditional GAN based individual and global motion fusion for multiple object tracking in UAV videos
Hongyang Yu 0001, Guorong Li, Li Su 0003, Bineng Zhong 0001, Hongxun Yao, Qingming Huang |
Pattern Recognit. Lett. | 5 |
| 2020 | Discrete Probability Distribution Prediction of Image Emotions with Shared Sparse LearningabstractComputationally modelling the affective content of images has been extensively studied recently because of its wide applications in entertainment, advertisement, and education. Significant progress has been made on designing discriminative features to bridge the affective gap. Assuming that viewers can reach a consensus on the emotion of images, most existing works focused on assigning the dominant emotion category or the average dimension values to an image. However, the image emotions perceived by viewers are subjective by nature with the influence of personal and situational factors. In this paper, we propose a novel machine learning approach that characterizes the categorical image emotions as a discrete probability distribution (DPD). To associate emotion with the visual features extracted from images, we present shared sparse learning to learn the combination coefficients, with which the DPD of an unseen image is predicted by linearly combining the DPDs of the training images. Furthermore, we extend our method to the setup where multi-features are available and learn the optimal weights for each feature to reflect the importance of different features. Extensive experiments are carried out on Abstract, Emotion6 and IESN datasets and the results demonstrate the superiority of the proposed method, as compared to the state-of-the-art approaches. Sicheng Zhao, Guiguang Ding, Yue Gao 0002, Xin Zhao 0020, Youbao Tang, Jungong Han, Hongxun Yao, Qingming Huang |
IEEE Trans. Affect. Comput. | 7 |
| 2020 | Deep Saliency Hashing for Fine-Grained RetrievalabstractIn recent years, hashing methods have been proved to be effective and efficient for large-scale Web media search. However, the existing general hashing methods have limited discriminative power for describing fine-grained objects that share similar overall appearance but have a subtle difference. To solve this problem, we for the first time introduce the attention mechanism to the learning of fine-grained hashing codes. Specifically, we propose a novel deep hashing model, named deep saliency hashing (DSaH), which automatically mines salient regions and learns semantic-preserving hashing codes simultaneously. DSaH is a two-step end-to-end model consisting of an attention network and a hashing network. Our loss function contains three basic components, including the semantic loss, the saliency loss, and the quantization loss. As the core of DSaH, the saliency loss guides the attention network to mine discriminative regions from pairs of images.We conduct extensive experiments on both fine-grained and general retrieval datasets for performance evaluation. Experimental results on fine-grained datasets, including Oxford Flowers, Stanford Dogs, and CUB Birds demonstrate that our DSaH performs the best for the fine-grained retrieval task and beats the strongest competitor (DTQ) by approximately 10% on both Stanford Dogs and CUB Birds. DSaH is also comparable to several state-of-the-art hashing methods on CIFAR-10 and NUS-WIDE. Sheng Jin 0002, Hongxun Yao, Xiaoshuai Sun, Shangchen Zhou, Lei Zhang 0006, Xian-Sheng Hua 0001 |
IEEE Trans. Image Process. | 2 |
| 2019 | Pix2Vox: Context-Aware 3D Reconstruction From Single and Multi-View ImagesabstractRecovering the 3D representation of an object from single-view or multi-view RGB images by deep neural networks has attracted increasing attention in the past few years. Several mainstream works (e.g., 3D-R2N2) use recurrent neural networks (RNNs) to fuse multiple feature maps extracted from input images sequentially. However, when given the same set of input images with different orders, RNN-based approaches are unable to produce consistent reconstruction results. Moreover, due to long-term memory loss, RNNs cannot fully exploit input images to refine reconstruction results. To solve these problems, we propose a novel framework for single-view and multi-view 3D reconstruction, named Pix2Vox. By using a well-designed encoder-decoder, it generates a coarse 3D volume from each input image. Then, a context-aware fusion module is introduced to adaptively select high-quality reconstructions for each part (e.g., table legs) from different coarse 3D volumes to obtain a fused 3D volume. Finally, a refiner further refines the fused 3D volume to generate the final output. Experimental results on the ShapeNet and Pix3D benchmarks indicate that the proposed Pix2Vox outperforms state-of-the-arts by a large margin. Furthermore, the proposed method is 24 times faster than 3D-R2N2 in terms of backward inference time. The experiments on ShapeNet unseen 3D categories have shown the superior generalization abilities of our method. Haozhe Xie, Hongxun Yao, Xiaoshuai Sun, Shangchen Zhou, Shengping Zhang |
ICCV | 2 |
| 2019 | Adaptive Semantic-Visual Tree for Hierarchical EmbeddingsabstractMerchandise categories inherently form a semantic hierarchy with different levels of concept abstraction, especially for fine-grained categories. This hierarchy encodes rich correlations among various categories across different levels, which can effectively regularize the semantic space and thus make prediction less ambiguous. However, previous studies of fine-grained image retrieval primarily focus on semantic similarities or visual similarities. In real application, merely using visual similarity may not satisfy the need of consumers to search merchandise with real-life images, e.g., given a red coat as query image, we might get red suit in recall results only based on visual similarity, since they are visually similar; But the users actually want coat rather than suit even the coat is with different color or texture attributes. We introduce this new problem based on photo shopping in real practice. That's why semantic information are integrated to regularize the margins to make "semantic" prior to "visual". To solve this new problem, we propose a hierarchical adaptive semantic-visual tree (ASVT) to depict the architecture of merchandise categories, which evaluates semantic similarities between different semantic levels and visual similarities within the same semantic class simultaneously. The semantic information satisfies the demand of consumers for similar merchandise with the query while the visual information optimize the correlations within the semantic class. At each level, we set different margins based on the semantic hierarchy and incorporate them as prior information to learn a fine-grained feature embedding. To evaluate our framework, we propose a new dataset named JDProduct, with hierarchical labels collected from actual image queries and official merchandise images on online shopping application. Extensive experimental results on the public CARS196 and CUB-200-2011 datasets demonstrate the superiority of our ASVT framework against compared state-of-the-art methods. Shuo Yang 0003, Wei Yu 0004, Ying Zheng 0009, Hongxun Yao, Tao Mei 0001 |
ACM Multimedia | 4 |
| 2019 | Self-balance Motion and Appearance Model for Multi-object Tracking in UAVabstractUnder the tracking-by-detection framework, multi-object tracking methods try to connect object detections with target trajectories by reasonable policy. Most methods represent objects by the appearance and motion. The inference of the association is mostly judged by a fusion of appearance similarity and motion consistency. However, the fusion ratio between appearance and motion are often determined by subjective setting. In this paper, we propose a novel self-balance method fusing appearance similarity and motion consistency. Extensive experimental results on public benchmarks demonstrate the effectiveness of the proposed method with comparisons to several state-of-the-art trackers. Hongyang Yu 0001, Guorong Li, Weigang Zhang, Hongxun Yao, Qingming Huang |
MMAsia | 4 |
| 2019 | Unsupervised semantic deep hashing
Sheng Jin 0002, Hongxun Yao, Xiaoshuai Sun, Shangchen Zhou |
Neurocomputing | 2 |
| 2019 | Robust visual tracking via scale-and-state-awareness
Yuankai Qi, Shengping Zhang, Qingming Huang, Hongxun Yao |
Neurocomputing | 5 |
| 2019 | Handling missing labels and class imbalance challenges simultaneously for facial action unit recognition
Baoyuan Wu, Yongping Zhao, Hongxun Yao |
Multim. Tools Appl. | 4 |
| 2019 | Action recognition with multi-scale trajectory-pooled 3D convolutional descriptors
Xiusheng Lu, Hongxun Yao, Sicheng Zhao, Xiaoshuai Sun, Shengping Zhang |
Multim. Tools Appl. | 2 |
| 2019 | Gradual recovery based occluded digit images recognition
Yasi Wang, Hongxun Yao, Wei Yu 0004, Dong Wang 0030, Shangchen Zhou, Xiaoshuai Sun |
Multim. Tools Appl. | 2 |
| 2019 | Hedging Deep Features for Visual TrackingabstractConvolutional Neural Networks (CNNs) have been applied to visual tracking with demonstrated success in recent years. Most CNN-based trackers utilize hierarchical features extracted from a certain layer to represent the target. However, features from a certain layer are not always effective for distinguishing the target object from the backgrounds especially in the presence of complicated interfering factors (e.g., heavy occlusion, background clutter, illumination variation, and shape deformation). In this work, we propose a CNN-based tracking algorithm which hedges deep features from different CNN layers to better distinguish target objects and background clutters. Correlation filters are applied to feature maps of each CNN layer to construct a weak tracker, and all weak trackers are hedged into a strong one. For robust visual tracking, we propose a hedge method to adaptively determine weights of weak classifiers by considering both the difference between the historical as well as instantaneous performance, and the difference among all weak trackers over time. In addition, we design a Siamese network to define the loss of each weak tracker for the proposed hedge method. Extensive experiments on large benchmark datasets demonstrate the effectiveness of the proposed algorithm against the state-of-the-art tracking methods. Yuankai Qi, Shengping Zhang, Qingming Huang, Hongxun Yao, Jongwoo Lim, Ming-Hsuan Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2019 | Learning Descriptors With Cube Loss for View-Based 3-D Object Retrievalabstract3-D object retrieval has been a hot research topic in recent years. Within such a field, view-based approaches are attracting increasing attention because of the flexibility of data representation as well as the reported state-of-the-art performance. One of the most important issues related to view-based 3-D object retrieval is how to learn embedding features that are discriminative across classes while being compactly distributed within each class. In this paper, we analyze the difference between the two tasks of classification and retrieval, and propose a novel way to learn a view-pooling feature via a triplet network. In addition, we propose a new loss, named cube loss, which is able to sample a number of triplets equal to the cube of the samples in a batch. With the new loss, both hard-negative and hard-positive pairs can be effectively investigated. The experimental results on the ModelNet benchmark demonstrate that the proposed method achieves superior performance compared to state-of-the-art approaches. Dong Wang 0030, Hongxun Yao, Federico Tombari, Sicheng Zhao, Bin Wang 0032, Hong Liu 0002 |
IEEE Trans. Multim. | 2 |
| 2019 | Discovering Latent Discriminative Patterns for Multi-Mode Event RepresentationabstractRepresentation of videos is essential since it conveys an understanding of video content and enables many higher level tasks to be tackled efficiently. However, it is challenging to propose a rational representation for complex event videos, as most video information is either noisy or redundant. In this paper, we propose a compact event representation method that can concisely describe the inner modes of events. We deem that an optimal event representation scheme should reflect the long-term and high-level visual semantics (visual topics) of events, so different from previous frame-level video semantics representation methods and concept-based video representation methods, we investigate the problem from the perspective of segment-level video representations. We then present three appealing properties of segment-level visual semantics. Based on the observation, we propose different algorithms that rely on a novel deep-visual-word-based video encoding method to discover latent discriminative patterns of events. Finally, our multi-mode event representation is obtained by concatenating the discovered patterns as inner modes. We adopt our event representation for representative event parts mining, which can highlight the visual topics of events and remarkably prune the raw videos. We validate our event representation method based on complex event detection task. Experimental results on two standard benchmarking datasets, MED11 and CCV Dataset, show that the proposed method can significantly outperform the state-of-the-art approaches. Wenlong Xie, Hongxun Yao, Xiaoshuai Sun, Tingting Han 0003, Sicheng Zhao, Tat-Seng Chua |
IEEE Trans. Multim. | 2 |
| 2018 | Cycle-Consistency Based Hierarchical Dense Semantic CorrespondenceabstractThis work aims to estimate dense correspondences between the images from same visual class but with different geometries and visual similarities. This task is particularly challenging because (i) most image pairs have large intra-class variations, and (ii)their visual content is similar only on the high-level structure. To address these problems, this paper proposed a multilevel method to estimate per-pixel correspondences from high-level semantic to low-level structural details by the guidance of cycle-consistency. We utilize CNN feature pyramid to represent images level by level. Meanwhile, we introduce cycle-consistency to measure the reliability of flow vector, which further affects the guidance from higher level to lower level. The proposed method has been extensively evaluated on various challenging benchmarks. The results show that our method significantly outperforms the state-of-the-arts in terms of semantic flow accuracy. Chuang Lin 0003, Hongxun Yao, Wei Yu 0004, Xiaoshuai Sun |
ICIP | 2 |
| 2018 | Local Image Descriptors with Statistical LossesabstractWe present a novel regularization technique for learning local feature descriptors based on statistical information extracted from batches of training samples. With the proposed regularization term, we learn a descriptor distribution in Euclidean space that aims at minimizing the overlap between the distributions of positive pairs and that of negative pairs. The proposed method is able to improve the performance of pairwise and triplet losses with various deep convolution network architectures. This improvement is demonstrated through two different types of architectures, able to obtain state-of-the-art results on the reference benchmark for local feature matching. Dong Wang 0030, Bin Wang 0032, Hongxun Yao, Hong Liu 0002, Federico Tombari |
ICIP | 3 |
| 2018 | Add: Actionness-Pooled Deep-Convolutional DescriptorabstractRecognition of general actions has achieved great breakthroughs in recent years. However, in real-world applications, finer-grained action classification is often needed. The major challenge is that fine-grained actions usually share high similarities in both appearance and motion pattern, making it difficult to distinguish them with existing general action representation. To solve this problem, we introduce visual attention mechanism into the proposed descriptor, termed as Actionness-pooled Deep-convolutional Descriptor (ADD). Instead of pooling features uniformly from the entire video, we aggregate features in sub-regions that are more likely to contain actions according to actionness maps, which endow ADD with the capability of capturing the subtle differences between fine-grained actions. We conduct experiments on HIT Dances dataset, one of the few existing datasets for fine-grained action analysis. Quantitative results have demonstrated that ADD remarkably outperforms traditional two-stream representation. Extensive experiments on two general action benchmarks, JHMDB and UCF101, have additionally proved that combining ADD with end-to-end ConvNet can further boost the recognition performance. Tingting Han 0003, Hongxun Yao, Xiaoshuai Sun, Wenlong Xie, Yanhao Zhang 0001 |
ICME | 2 |
| 2018 | Very High Resolution Image Scene Classification with Semantic Fisher VectorsabstractVery high resolution (VHR) image scene classification is the most challenging of remote sensing data analysis, that has attracted researchers' attention. To improve the precision of VHR image scene classification, we propose a new method based on convolutional features extracted by convolutional neural network (CNN). First, Visual Geometry Group Network (VGG-Net) model is introduced as a feature extractor form the original VHR images. Second, we select the fifth convolutional layer constructed by VGG-Net, which is supposed as convolutional features descriptors. Third, based on Improved Fisher Vector (IFV) coding method, we compute the visual word corresponding to the convolutional features of the image scene. We conduct experiments on the public AID benchmark dataset, which contains 30 different areal categories with sub-meter resolution. Experimental results demonstrate the effectiveness of the proposed method, as compared with several state-of-the-art methods. Souleyman Chaib, Yanfeng Gu, Hongxun Yao, Khaled Belkadi |
IGARSS | 3 |
| 2018 | ASMMC-MMAC 2018: The Joint Workshop of 4th the Workshop on Affective Social Multimedia Computing and first Multi-Modal Affective Computing of Large-Scale Multimedia Data WorkshopabstractAffective social multimedia computing is an emergent research topic for both affective computing and multimedia research communities. Social multimedia is fundamentally changing how we communicate, interact, and collaborate with other people in our daily lives. Social multimedia contains much affective information. Effective extraction of affective information from social multimedia can greatly help social multimedia computing (e.g., processing, index, retrieval, and understanding). Besides, with the rapid development of digital photography and social networks, people get used to sharing their lives and expressing their opinions online. As a result, user-generated social media data, including text, images, audios, and videos, grow rapidly, which urgently demands advanced techniques on the management, retrieval, and understanding of these data. Dong-Yan Huang, Sicheng Zhao, Björn W. Schuller, Hongxun Yao, Jianhua Tao 0001, Min Xu 0001, Lei Xie 0001, Qingming Huang |
ACM Multimedia | 4 |
| 2018 | Hierarchical semantic image matching using CNN feature pyramid
Wei Yu 0004, Xiaoshuai Sun, Kuiyuan Yang, Yong Rui, Hongxun Yao |
Comput. Vis. Image Underst. | 5 |
| 2018 | Online multiple object tracking via exchanging object context
Hongyang Yu 0001, Qingming Huang, Hongxun Yao |
Neurocomputing | 4 |
| 2018 | Rediscover flowers structurally
Hongxun Yao, Xiaoshuai Sun, Sicheng Zhao, Wei Yu 0004 |
Multim. Tools Appl. | 2 |
| 2018 | Exploring part-aware segmentation for fine-grained visual categorization
Hongxun Yao, Xiaoshuai Sun, Sicheng Zhao, Yanhao Zhang 0001 |
Multim. Tools Appl. | 2 |
| 2018 | Off-the-shelf CNN features for 3D object retrieval
Dong Wang 0030, Bin Wang 0032, Sicheng Zhao, Hongxun Yao, Hong Liu 0002 |
Multim. Tools Appl. | 4 |
| 2018 | Event patches: Mining effective parts for event detection and understanding
Wenlong Xie, Hongxun Yao, Sicheng Zhao, Xiaoshuai Sun, Tingting Han 0003 |
Signal Process. | 2 |
| 2018 | Distinctive action sketch for human action recognition
Ying Zheng 0009, Hongxun Yao, Xiaoshuai Sun, Sicheng Zhao, Fatih Porikli |
Signal Process. | 2 |
| 2018 | Predicting Personalized Image Emotion Perceptions in Social NetworksabstractImages can convey rich semantics and induce various emotions to viewers. Most existing works on affective image analysis focused on predicting the dominant emotions for the majority of viewers. However, such dominant emotion is often insufficient in real-world applications, as the emotions that are induced by an image are highly subjective and different with respect to different viewers. In this paper, we propose to predict the personalized emotion perceptions of images for each individual viewer. Different types of factors that may affect personalized image emotion perceptions, including visual content, social context, temporal evolution, and location influence, are jointly investigated. Rolling multi-task hypergraph learning (RMTHG) is presented to consistently combine these factors and a learning algorithm is designed for automatic optimization. For evaluation, we set up a large scale image emotion dataset from Flickr, named Image-Emotion-Social-Net, on both dimensional and categorical emotion representations with over 1 million images and about 8,000 users. Experiments conducted on this dataset demonstrate that the proposed method can achieve significant performance gains on personalized emotion classification, as compared to several state-of-the-art approaches. Sicheng Zhao, Hongxun Yao, Yue Gao 0002, Guiguang Ding, Tat-Seng Chua |
IEEE Trans. Affect. Comput. | 2 |
| 2017 | Non-rigid Object Tracking via Deformable Patches Using Shape-Preserved KCF and Level SetsabstractPart-based trackers are effective in exploiting local details of the target object for robust tracking. In contrast to most existing part-based methods that divide all kinds of target objects into a number of fixed rectangular patches, in this paper, we propose a novel framework in which a set of deformable patches dynamically collaborate on tracking of non-rigid objects. In particular, we proposed a shape-preserved kernelized correlation filter (SP-KCF) which can accommodate target shape information for robust tracking. The SP-KCF is introduced into the level set framework for dynamic tracking of individual patches. In this manner, our proposed deformable patches are target-dependent, have the capability to assume complex topology, and are deformable to adapt to target variations. As these deformable patches properly capture individual target subregions, we exploit their photometric discrimination and shape variation to reveal the trackability of individual target subregions, which enables the proposed tracker to dynamically take advantage of those subregions with good trackability for target likelihood estimation. Finally the shape information of these deformable patches enables accurate object contours to be computed as the tracking output. Experimental results on the latest public sets of challenging sequences demonstrate the effectiveness of the proposed method. Xin Sun 0003, Ngai-Man Cheung, Hongxun Yao, Yiluan Guo |
ICCV | 3 |
| 2017 | The shortest matching path based on novel cycle consistencyabstractCategory-level image matching is extremely challenging due to various intra-class variations. To tackle the large variations, we propose an algorithm to jointly estimate the dense correspondence for image set, which reformulates image set alignment into the problem of shortest path searching. We propose a novel tri-image cycle-consistency to measure the matching “distance” between two image, which is further used to improve the pair-wise dense correspondence. Meanwhile, we utilize CNN feature pyramid to achieve pair-wise image matching hierarchically. Extensive experiments and analysis demonstrate the superiority of our method in matching images with challenging variations. Wei Yu 0004, Hongxun Yao |
ICIP | 3 |
| 2017 | Dancing like a superstar: Action guidance based on pose estimation and conditional pose alignmentabstractAction Guidance (AG) aims at scoring how accurate the action is and giving guidance to the learners on how to correct their actions according to the standard instructive videos. AG has plenty of real-world applications such as sports training, rehabilitation treatment, and dance teaching. However, the problem of assessing the accuracy of action has almost no effective solution. In this paper, we describe a two-stage framework for action guidance. Firstly, we estimate the poses in the test video and standard video using person detection method and convolutional pose machines (CPMs). As for action guidance, we propose a network to compute the essential differences between two poses under different sizes, views, and locations. Extensive experiments with real-world and synthetic datasets demonstrate the effectiveness of our framework. Yuxin Hou, Hongxun Yao, Haoran Li 0012, Xiaoshuai Sun |
ICIP | 2 |
| 2017 | Gated additive skip context connection for object detectionabstractContext information plays an important role in object detection. DeepID concatenates the global context for classification, while Yolo, SSD, and Crafting use local context information for detection. In this paper, we propose a straightforward method to plug the global context information into the Faster RCNN Framework, namely the Skip Context Connection (SCC). We use SCC to inject the global context into the object representation which skips the RoiPooling layer rather than drops it. Therefore, it can not only leverage the context information but also keep the location accuracy from the RCNN framework. We proposed three principles to construct the SCC blocks: effectiveness means fewer parameters, additivity means the features possess the same meaning, and selectable means soft gated addition. We also evaluate several different SCC blocks. The Gated Additive SCC(GA-SCC) which satisfy the three principles get the best performance. Our experiment results on PASCAL VOC 2007 show that GA-SCC can get the steady 1% improvement over the traditional RCNN method. Haoran Li 0012, Hongxun Yao, Yuxin Hou, Xiaoshuai Sun |
ICIP | 2 |
| 2017 | Part-based fine-grained bird image retrieval respecting species correlationabstractMost of the existing works on fine-grained bird image categorization and retrieval focus on finding similar images from the same species and often give little importance to inter-species similarity. In this paper, we devise a new fine-grained retrieval task that searches similar instances from different species. To this end, we propose a two-step strategy. In the first step, we search for visually similar parts to a query image using a deep convolutional neural network (CNN). To improve the quality of the retrieved candidates, we incorporate structural cues into the CNN using a novel part-pooling layer. In the second step, we re-rank the retrieved candidates improving the species diversity. We achieve this by formulating a novel ranking function that balances between the similarity of the candidates to the queried parts, while decreasing the similarity to the query species. We provide experiments on the benchmark CUB200 dataset and demonstrate clear benefits of our schemes. Hongdong Li, Anoop Cherian, Hongxun Yao |
ICIP | 4 |
| 2017 | How many zero crossings? A method for structure-texture image decomposition
Xiaolei Jiang, Hongxun Yao, Shaohui Liu |
Comput. Graph. | 2 |
| 2017 | Text image deblurring via two-tone prior
Xiaolei Jiang, Hongxun Yao, Sicheng Zhao |
Neurocomputing | 2 |
| 2017 | View-based 3D object retrieval with discriminative views
Dong Wang 0030, Bin Wang 0032, Sicheng Zhao, Hongxun Yao, Hong Liu 0002 |
Neurocomputing | 4 |
| 2017 | Actor identification via mining representative actions
Wenlong Xie, Hongxun Yao, Xiaoshuai Sun, Sicheng Zhao, Wei Yu 0004, Shengping Zhang |
Neurocomputing | 2 |
| 2017 | Exploiting the complementary strengths of multi-layer CNN features for image retrieval
Wei Yu 0004, Kuiyuan Yang, Hongxun Yao, Xiaoshuai Sun, Pengfei Xu 0001 |
Neurocomputing | 3 |
| 2017 | Discovering discriminative patches for free-hand sketch analysis
Ying Zheng 0009, Hongxun Yao, Sicheng Zhao, Yasi Wang |
Multim. Syst. | 2 |
| 2017 | Towards more efficient and flexible face image deblurring using robust salient face landmark detection
Yinghao Huang, Hongxun Yao, Sicheng Zhao, Yanhao Zhang 0001 |
Multim. Tools Appl. | 2 |
| 2017 | Anomaly detection based on spatio-temporal sparse representation and visual attention analysis
Chen Wang 0145, Hongxun Yao, Xiaoshuai Sun |
Multim. Tools Appl. | 2 |
| 2017 | Breaking video into pieces for action recognition
Ying Zheng 0009, Hongxun Yao, Xiaoshuai Sun, Xuesong Jiang, Fatih Porikli |
Multim. Tools Appl. | 2 |
| 2017 | Guest Editorial Introduction to the Special Issue on Group and Crowd Behavior Analysis for Intelligent Multicamera Video SurveillanceabstractDespite significant progress in human behavior analysis over the past few years, most of today’s state-of-the-art algorithms focus on analyzing individual behavior in a simple environment monitored by a single camera. Recently, the widespread availability of cameras and a growing need for public safety have shifted the attention of researchers in video surveillance from individual behavior analysis to group and crowd behavior analysis in multicamera networks. Group behavior analysis provides a novel level for describing events, which are semantically more meaningful, highlighting barely visible relational connections among people. Crowd behavior analysis can also be used for anomaly detection such as panic scenarios, dangerous situations, and illegal behaviors in public spaces. Hongxun Yao, Andrea Cavallaro, Thierry Bouwmans, Zhengyou Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2017 | Deep Feature Fusion for VHR Remote Sensing Scene ClassificationabstractThe rapid development of remote sensing technology allows us to get images with high and very high resolution (VHR). VHR imagery scene classification has become an important and challenging problem. In this paper, we introduce a framework for VHR scene understanding. First, the pretrained visual geometry group network (VGG-Net) model is proposed as deep feature extractors to extract informative features from the original VHR images. Second, we select the fully connected layers constructed by VGG-Net in which each layer is regarded as separated feature descriptors. And then we combine between them to construct final representation of the VHR image scenes. Third, discriminant correlation analysis (DCA) is adopted as feature fusion strategy to further refine the original features extracting from VGG-Net, which allows a more efficient fusion approach with small cost than the traditional feature fusion strategies. We apply our approach to three challenging data sets: 1) UC MERCED data set that contains 21 different areal scene categories with submeter resolution; 2) WHU-RS data set that contains 19 challenging scene categories with various resolutions; and 3) the Aerial Image data set that has a number of 10 000 images within 30 challenging scene categories with various resolutions. The experimental results demonstrate that our proposed method outperforms the state-of-the-art approaches. Using feature fusion technique achieves a higher accuracy than solely using the raw deep features. Moreover, the proposed method based on DCA fusion produces good informative features to describe the images scene with much lower dimension. Souleyman Chaib, Yanfeng Gu, Hongxun Yao |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2017 | Dancelets Mining for Video Recommendation Based on Dance StylesabstractDance is a unique and meaningful type of human expression, composed of abundant and various action elements. However, existing methods based on associated texts and spatial visual features have difficulty capturing the highly articulated motion patterns. To overcome this limitation, we propose to take advantage of the intrinsic motion information in dance videos to solve the video recommendation problem. We present a novel system that recommends dance videos based on a mid-level action representation, termed Dancelets. The Dancelets are used to bridge the semantic gap between video content and high-level concept, dance style, which plays a significant role in characterizing different types of dances. The proposed method executes automatic mining of dancelets with a concatenation of normalized cut clustering and linear discriminant analysis. This ensures that the discovered dancelets are both representative and discriminative. Additionally, to exploit the motion cues in videos, we employ motion boundaries as saliency priors to generate volumes of interest and extract C3D features to capture spatiotemporal information from the mid-level patches. Extensive experiments validated on our proposed large dance dataset, HIT Dances dataset, demonstrate the effectiveness of the proposed methods for dance style-based video recommendation. Tingting Han 0003, Hongxun Yao, Chenliang Xu, Xiaoshuai Sun, Yanhao Zhang 0001, Jason J. Corso |
IEEE Trans. Multim. | 2 |
| 2017 | Continuous Probability Distribution Prediction of Image Emotions via Multitask Shared Sparse RegressionabstractPrevious works on image emotion analysis mainly focused on predicting the dominant emotion category or the average dimension values of an image for affective image classification and regression. However, this is often insufficient in various real-world applications, as the emotions that are evoked in viewers by an image are highly subjective and different. In this paper, we propose to predict the continuous probability distribution of image emotions which are represented in dimensional valence-arousal space. We carried out large-scale statistical analysis on the constructed Image-Emotion-Social-Net dataset, on which we observed that the emotion distribution can be well-modeled by a Gaussian mixture model. This model is estimated by an expectation-maximization algorithm with specified initializations. Then, we extract commonly used emotion features at different levels for each image. Finally, we formalize the emotion distribution prediction task as a shared sparse regression (SSR) problem and extend it to multitask settings, named multitask shared sparse regression (MTSSR), to explore the latent information between different prediction tasks. SSR and MTSSR are optimized by iteratively reweighted least squares. Experiments are conducted on the Image-Emotion-Social-Net dataset with comparisons to three alternative baselines. The quantitative results demonstrate the superiority of the proposed method. Sicheng Zhao, Hongxun Yao, Yue Gao 0002, Rongrong Ji, Guiguang Ding |
IEEE Trans. Multim. | 2 |
| 2017 | A Biologically Inspired Appearance Model for Robust Visual TrackingabstractIn this paper, we propose a biologically inspired appearance model for robust visual tracking. Motivated in part by the success of the hierarchical organization of the primary visual cortex (area V1), we establish an architecture consisting of five layers: whitening, rectification, normalization, coding, and pooling. The first three layers stem from the models developed for object recognition. In this paper, our attention focuses on the coding and pooling layers. In particular, we use a discriminative sparse coding method in the coding layer along with spatial pyramid representation in the pooling layer, which makes it easier to distinguish the target to be tracked from its background in the presence of appearance variations. An extensive experimental study shows that the proposed method has higher tracking accuracy than several state-of-the-art trackers. Shengping Zhang, Xiangyuan Lan, Hongxun Yao, Huiyu Zhou 0001, Dacheng Tao, Xuelong Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2016 | Affective Computing and Applications of Image Emotion PerceptionsabstractImages can convey rich semantics and evoke strong emotions in viewers. The research of my PhD thesis focuses on image emotion computing (IEC), which aims to predict the emotion perceptions of given images. The development of IEC is greatly constrained by two main challenges: affective gap and subjective evaluation. Previous works mainly focused on finding features that can express emotions better to bridge the affective gap, such as elements-of-art based features and shape features. According to the emotion representation models, including categorical emotion states (CES) and dimensional emotion space (DES), three different tasks are traditionally performed on IEC: affective image classification, regression and retrieval. The state-of-the-art methods on the three above tasks are image-centric, focusing on the dominant emotions for the majority of viewers. For my PhD thesis, I plan to answer the following questions: (1) Compared to the low-level elements-of-art based features, can we find some higher level features that are more interpretable and have stronger link to emotions? (2) Are the emotions that are evoked in viewers by an image subjective and different? If they are, how can we tackle the user-centric emotion prediction? (3) For image-centric emotion computing, can we predict the emotion distribution instead of the dominant emotion category? Sicheng Zhao, Hongxun Yao |
AAAI | 2 |
| 2016 | User-Centric Affective Computing of Image Emotion PerceptionsabstractWe propose to predict the personalized emotion perceptions of images for each viewer. Different factors that may influence emotion perceptions, including visual content, social context, temporal evolution, and location influence are jointly investigated via the presented rolling multi-task hypergraph learning. For evaluation, we set up a large scale image emotion dataset from Flickr, named Image-Emotion-Social-Net, with over 1 million images and about 8,000 users. Experiments conducted on this dataset demonstrate the superiority of the proposed method, as compared to state-of-the-art. Sicheng Zhao, Hongxun Yao, Wenlong Xie, Xiaolei Jiang |
AAAI | 2 |
| 2016 | Hedged Deep TrackingabstractIn recent years, several methods have been developed to utilize hierarchical features learned from a deep convolutional neural network (CNN) for visual tracking. However, as features from a certain CNN layer characterize an object of interest from only one aspect or one level, the performance of such trackers trained with features from one layer (usually the second to last layer) can be further improved. In this paper, we propose a novel CNN based tracking framework, which takes full advantage of features from different CNN layers and uses an adaptive Hedge method to hedge several CNN based trackers into a single stronger one. Extensive experiments on a benchmark dataset of 100 challenging image sequences demonstrate the effectiveness of the proposed algorithm compared to several state-of-theart trackers. Yuankai Qi, Shengping Zhang, Hongxun Yao, Qingming Huang, Jongwoo Lim, Ming-Hsuan Yang 0001 |
CVPR | 4 |
| 2016 | Mining representative actions for actor identificationabstractPrevious works on actor identification mainly focused on static features based on face identification and costume detection, without considering the abundant dynamic information contained in videos. In this paper, we propose a novel method to mine representative actions of each actor, and show the remarkable power of such actions for actor identification task. Videos are firstly divided into shots and represented by BoW based on spatial-temporal features. Then we integrate the prototype theory with SVM to rank the shots and obtain the representative actions. Our method for actor identification combines representative actions with actors' appearance. We validate the method on episodes of the TV series "The Big Bang Theory". The experimental results show that the representative actions are consistent with human judgements and can greatly improve the matching performance as complementary to existing handcrafted static features for actor identification. Wenlong Xie, Hongxun Yao, Xiaoshuai Sun, Sicheng Zhao, Tingting Han 0003 |
ICASSP | 2 |
| 2016 | Crowd video retrieval via deep attribute-embedding graph rankingabstractSince the number of surveillance cameras in public areas increases very fast, massive crowd videos are captured and shared, which brings an urgent need to retrieve these videos efficiently and effectively. However, most recent research on crowd video mainly focused on crowd behavior understanding and abnormal detection. In this study, as the very first attempt, we propose a crowd video retrieval method via deep attribute-embedding graph ranking. Group profiling attributes are capable of reflecting rich crowd patterns in videos. To deeply embed the specific relationship and manifold structure of crowd patterns, we integrate graph ranking, optimized weights learning and deep metric transforming in a unified regularization framework for crowd video retrieval. To sufficiently explore the effects of multiple attributes and their complementation in crowds, we devise several scene-independent visual descriptors specifically for each crowd video. Interpretable descriptors are categorized into different levels and structures as group profiling attributes, according to semantic properties of crowd patterns. Extensive experiments conducted on CUHK crowd dataset demonstrate the effectiveness and superiority of the proposed approach. Yanhao Zhang 0001, Sicheng Zhao, Rongrong Ji, Xiusheng Lu, Hongxun Yao, Qingming Huang |
ICME | 6 |
| 2016 | A VHR scene classification method integrating sparse PCA and saliency computingabstractUnderstanding a scene provided by very high resolution (VHR) satellite imagery has become a more and more challenging problem. In this paper, we propose a new method for scene classification based on saliency computing of patches sampling from the VHR images. Sparse principal component analysis (sPCA) is then adopted to select the corresponding informative salient patches for image scene representation. The proposed technique for selecting informative salient patches is efficient and robust for scene understanding. We conduct experiments on the public UC Merced benchmark dataset, which contains 21 different areal categories with sub-meter resolution. Experimental results demonstrate the effectiveness of the proposed method, as compared with several state-of-the-art methods. Souleyman Chaib, Yanfeng Gu, Hongxun Yao, Sicheng Zhao |
IGARSS | 3 |
| 2016 | From Seed Discovery to Deep Reconstruction: Predicting Saliency in Crowd via Deep NetworksabstractAlthough saliency prediction in crowd has been recently recognized as an essential task for video analysis, it is not comprehensively explored yet. The challenges lie in that eye fixations in crowded scenes are inherently "distinct" and "multi-modal", which differs from those in regular scenes. To this end, the existing saliency prediction schemes typically rely on hand designed features with shallow learning paradigm, which neglect the underlying characteristics of crowded scenes. In this paper, we propose a saliency prediction model dedicated for crowd videos with two novelties: 1) Distinct units are discovered using deep representation learned by a Stacked Denoising Auto-Encoder (SDAE), considering perceptual properties of crowd saliency; 2) Contrast-based saliency is measured through deep reconstruction errors in the second SDAE trained on all units excluding distinct units. A unified model is integrated for online processing crowd saliency. Extensive evaluations on two crowd video benchmark datasets demonstrate that our approach can effectively explore crowd saliency mechanism in two-stage SDAEs and achieve significantly better results than state-of-the-art methods, with robustness to parameters. Yanhao Zhang 0001, Qingming Huang, Kuiyuan Yang, Jun Zhang 0017, Hongxun Yao |
ACM Multimedia | 6 |
| 2016 | Predicting Personalized Emotion Perceptions of Social ImagesabstractImages can convey rich semantics and induce various emotions to viewers. Most existing works on affective image analysis focused on predicting the dominant emotions for the majority of viewers. However, such dominant emotion is often insufficient in real-world applications, as the emotions that are induced by an image are highly subjective and different with respect to different viewers. In this paper, we propose to predict the personalized emotion perceptions of images for each individual viewer. Different types of factors that may affect personalized image emotion perceptions, including visual content, social context, temporal evolution, and location influence, are jointly investigated. Rolling multi-task hypergraph learning is presented to consistently combine these factors and a learning algorithm is designed for automatic optimization. For evaluation, we set up a large scale image emotion dataset from Flickr, named Image-Emotion-Social-Net, on both dimensional and categorical emotion representations with over 1 million images and about 8,000 users. Experiments conducted on this dataset demonstrate that the proposed method can achieve significant performance gains on personalized emotion classification, as compared to several state-of-the-art approaches. Sicheng Zhao, Hongxun Yao, Yue Gao 0002, Rongrong Ji, Wenlong Xie, Xiaolei Jiang, Tat-Seng Chua |
ACM Multimedia | 2 |
| 2016 | Exploring Discriminative Views for 3D Object Retrieval
Dong Wang 0030, Bin Wang 0032, Sicheng Zhao, Hongxun Yao, Hong Liu 0002 |
MMM (1) | 4 |
| 2016 | Unsupervised discovery of crowd activities by saliency-based clustering
Tingting Han 0003, Hongxun Yao, Xiaoshuai Sun, Sicheng Zhao, Yanhao Zhang 0001 |
Neurocomputing | 2 |
| 2016 | Auto-encoder based dimensionality reduction
Yasi Wang, Hongxun Yao, Sicheng Zhao |
Neurocomputing | 2 |
| 2016 | An Informative Feature Selection Method Based on Sparse PCA for VHR Scene ClassificationabstractUnderstanding the scenes provided by very high resolution satellite (VHR) imagery has become a critical task. In this letter, we propose a new informative feature selection method for VHR scene classification. First, scale-invariant feature transform and speeded up robust feature operators are used to extract local features from the original VHR images to construct a visual dictionary. A sparse principal component analysis (sPCA) is then adopted to learn a set of informative features from the visual dictionary for each category. Finally, the scenes are represented by sparse informative low-level features. We conducted experiments on the University of California at Merced data set containing 21 different areal scene categories with submeter resolution and the Sydney data set containing seven land-use categories with 0.5-m spatial resolution. The experimental results demonstrate that the proposed method outperforms the state-of-the-art methods even without saliency detection. Souleyman Chaib, Yanfeng Gu, Hongxun Yao |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2016 | Multi-modal microblog classification via multi-task learning
Sicheng Zhao, Hongxun Yao, Sendong Zhao, Xuesong Jiang, Xiaolei Jiang |
Multim. Tools Appl. | 2 |
| 2016 | Facial action unit recognition under incomplete data based on multi-label learning with missing labels
Baoyuan Wu, Bernard Ghanem, Yongping Zhao, Hongxun Yao |
Pattern Recognit. | 5 |
| 2015 | Learning a discriminative dictionary for facial expression recognitionabstractDictionary learning for sparse representation classifiers (SRC) has demonstrated great success for many classification problems, i.e., face recognition, object detection, etc. However, it has not enjoyed a similar reception in the facial expression recognition literature. In this paper, we applied dictionary learning methods to the task of facial expression recognition, which is then compared with SVM. In addition, we introduce a new dictionary learning method that incorporates side information, which is contained in the training data but not available in the testing phase. In particular, we introduce a new soft constraint derived from side information and combine it with the reconstruction error and the classification error to form a unified objective function. The optimal solution to the objective function is efficiently obtained using the K-SVD algorithm. Our algorithm learns the dictionary and an optimal linear classifier jointly. Experimental part demonstrates the effectiveness of the sparse representation classifiers for facial expression recognition problem, and dictionary learning with side information method achieves further improvement on low resolution facial expression recognition. Yongping Zhao, Hongxun Yao |
ACII | 3 |
| 2015 | Histograms of locally aggregated oriented gradientsabstractMotivated by the Vector of Locally Aggregated Descriptors (VLAD), we propose a new Histograms of Locally Aggregated Oriented Gradients descriptor (called HLAOG). In the Histograms of Oriented Gradients descriptor (HOG), the zero-order information of the gradients is captured. By contrast, in the HLAOG descriptor we accumulate the differences between gradient orientations and their nearest bin centers, which characterizes the distribution of the gradient orientations in regard to the bin centers. The HLAOG descriptor is demonstrated to be complementary to HOG in the experiments. Then, for setting the weights of the votes on different bins in a better way, we choose Gaussian function as the weighting method and present another new Gaussian Weighted Histograms of Oriented Gradients descriptor (called GWHOG) based on HOG. Evaluations on two public object recognition datasets (Caltech-101 and VOC2007) show that the combination of HOG and HLAOG outperforms HOG and the combination of HLAOG and GWHOG gets the best result. Xiusheng Lu, Shengping Zhang, Hongxun Yao, Xin Sun 0003, Yanhao Zhang 0001 |
ICIP | 3 |
| 2015 | Predicting discrete probability distribution of image emotionsabstractMost existing works on affective image classification tried to assign a dominant emotion category to an image. However, this is often insufficient, as the emotions that are evoked in viewers by an image are highly subjective and different. In this paper, we propose to predict the probability distribution of categorical image emotions. Firstly we extract commonly used features of different levels for each image. Then we formulize the emotion distribution prediction as a shared sparse leaning problem, which is optimized by iteratively reweighted least squares. Besides, we introduce three baseline algorithms. Experiments are carried out on a dataset of peer rated abstract paintings and the results demonstrate the superiority of our proposed method, as compared to some state-of-the-art approaches. Sicheng Zhao, Hongxun Yao, Xiaolei Jiang, Xiaoshuai Sun |
ICIP | 2 |
| 2015 | Distinctive action sketchabstractRecent developments in the field of image and video processing have led to a renewed interest in sketch correlated research. With that, there have emerged considerable solid evidence which revealed the significance of sketch to us. However, there have been few profound discussions on sketch based action analysis so far. In this paper, we present a framework of converting human actions to commendable sketches and discovering distinctive action sketches. The action sketches should satisfy three characteristics: sketchability, objective-ness and consistency. Primitive sketches are prepared according to the structured forests based fast edge detection. Meanwhile, we take advantage of Struck to accomplish adaptive object tracking in parallel. After that, we propose a method to ensure the spatio-temporal consistency between every sequential action sketches. On completion of previous stages, the process of distinctive mining of action sketches is carried out. The experimental results show that our approach has a promising potential. Ying Zheng 0009, Hongxun Yao, Xiaoshuai Sun, Sicheng Zhao |
ICIP | 2 |
| 2015 | Formation Period Matters: Towards Socially Consistent Group Detection via Dense Subgraph SeekingabstractGroup detection becomes an important task in crowd behavior surveillance. However, most existing methods ignore the formation persistency characteristics, which predict unreliable interactions when the crowd is realistic and complex. To address this issue, we propose a novel graph-based method to declare that the formation period really matters for detecting social groups in crowd. First, we develop a socially motivated representation by modeling the formation period probability in a Bayesian manner, which results in social and temporal consistency for group member interactions. A graph is then established using individuals as nodes and formation periods as edge weights to reflect pedestrian relationships. In this way, seeking of socially consistent groups is converted into an optimization problem which seeks dense subgraphs with maximum formation likelihood within the graph structure. We employ graph shift optimization to detect groups by finding all the dense subgraphs due to its robust performance. In the experimental results on public datasets, our proposed method clearly outperforms other related state-of-the-art methods. Yanhao Zhang 0001, Shengping Zhang, Hongxun Yao, Qingming Huang |
ICMR | 4 |
| 2015 | "Clustering of Dancelets": Towards Video Recommendation Based on Dance StylesabstractDance is a special and important type of action, composed of abundant and various action elements. However, the recommendation of dance videos on the web are still not well studied. It is hard to realize it in the way of traditional methods using associated texts or static features of video content. In this paper, we study the problem focusing on extraction and representation of action information in dances. We propose to recommend dance videos based on the automatically discovered ``Dance Styles'', which play a significant role in characterizing different types of dances. To bridge the semantic gap of video content and mid-level concept, style, we take advantage of a mid-level action representation method, and extract representative patches as ``Dancelets'', a sort of intermediation between videos and the concepts. Furthermore, we propose to employ Motion Boundaries as saliency priors and sparsely extract patches containing more representative information to generate a set of dancelet candidates. Dancelets are then discovered by Normalized-cut method, which is superior in grouping visually similar patterns into the same clusters. For the fast and effective recommendation, a random forest-based index is built, and the ranking results are derived according to the matching results in all the leaf notes. Extensive experiments validated on the web dance videos demonstrate the effectiveness of the proposed methods for dance style discovery and video recommendation based on styles. Tingting Han 0003, Hongxun Yao, Xiaoshuai Sun, Yanhao Zhang 0001, Sicheng Zhao, Xiusheng Lu, Yinghao Huang, Wenlong Xie |
ACM Multimedia | 2 |
| 2015 | Predicting Continuous Probability Distribution of Image Emotions in Valence-Arousal SpaceabstractPrevious works on image emotion analysis mainly focused on assigning a dominated emotion category or the average dimension values to an image for affective image classification and regression. However, this is often insufficient in many applications, as the emotions that are evoked in viewers by an image are highly subjective and different. In this paper, we propose to predict the continuous probability distribution of dimensional image emotions represented in valence-arousal space. By the statistical analysis on the constructed Image-Emotion-Social-Net dataset, we represent the emotion distribution as a Gaussian mixture model (GMM), which is estimated by the EM algorithm. Then we extract commonly used features of different levels for each image. Finally, we formulize the emotion distribution prediction as a multi-task shared sparse regression (MTSSR) problem, which is optimized by iteratively reweighted least squares. Besides, we introduce three baseline algorithms. Experiments conducted on the Image-Emotion-Social-Net dataset demonstrate the superiority of the proposed method, as compared to some state-of-the-art approaches. Sicheng Zhao, Hongxun Yao, Xiaolei Jiang |
ACM Multimedia | 2 |
| 2015 | Strategy for aesthetic photography recommendation via collaborative composition modelabstractIn this study, the authors propose a collaborative composition model for automatically recommending suitable positions and poses in the scene of photography taken by amateurs. By analysing aesthetic‐aware features, the authors' strategy jointly takes attention and geometry composition into account to learn the aesthetic manifestation knowledge of professional photographers. Firstly, aesthetic composition representation exploits the strength of visual saliency to explicitly encode the spatial correlation of the professional photos. Secondly, ℓ2 regularised least square is adopted to constrain the representation coefficients, which provides a fast solution in selecting aesthetic candidates collaboratively. In addition, a novel confidence measure scheme is further designed based on reconstruction errors and the reference photos are updated adaptively according to the composition rules. Both qualitative and quantitative evaluations show that the model performs well for the portrait photographing recommendation. Yanhao Zhang 0001, Qingming Huang, Sicheng Zhao, Xiusheng Lu, Xiaoshuai Sun, Hongxun Yao |
IET Comput. Vis. | 7 |
| 2015 | Strategy for dynamic 3D depth data matching towards robust action retrieval
Sicheng Zhao, Lujun Chen, Hongxun Yao, Yanhao Zhang 0001, Xiaoshuai Sun |
Neurocomputing | 3 |
| 2015 | Adaptive NormalHedge for robust visual tracking
Shengping Zhang, Huiyu Zhou 0001, Hongxun Yao, Yanhao Zhang 0001, Kuanquan Wang, Jun Zhang 0017 |
Signal Process. | 3 |
| 2015 | View-based 3D object retrieval via multi-modal graph learning
Sicheng Zhao, Hongxun Yao, Yanhao Zhang 0001, Yasi Wang, Shaohui Liu |
Signal Process. | 2 |
| 2015 | Social Attribute-Aware Force Model: Exploiting Richness of Interaction for Abnormal Crowd DetectionabstractInteractions among pedestrians usually play an important role in understanding crowd behavior. However, there are great challenges, such as occlusions, motion, and appearance variance, on accurate analysis of pedestrian interactions. In this paper, we introduce a novel social attribute-aware force model (SAFM) for detection of abnormal crowd events. The proposed model incorporates social characteristics of crowd behaviors to improve the description of interactive behaviors. To this end, we first efficiently estimate the scene scale in an unsupervised manner. Then, we introduce the concepts of social disorder and congestion attributes to characterize the interaction of social behaviors, and construct our crowd interaction model on the basis of social force by an online fusion strategy. These attributes encode social interaction characteristics and offer robustness against motion pattern variance. Abnormal event detection is finally performed based on the proposed SAFM. In addition, the attribute-aware interaction force indicates the possible locations of anomalous interactions. We validate our method on the publicly available data sets for abnormal detection, and the experimental results show promising performance compared with alternative and state-of-the-art methods. Yanhao Zhang 0001, Rongrong Ji, Hongxun Yao, Qingming Huang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2015 | Non-Rigid Object Contour Tracking via a Novel Supervised Level Set ModelabstractWe present a novel approach to non-rigid objects contour tracking in this paper based on a supervised level set model (SLSM). In contrast to most existing trackers that use bounding box to specify the tracked target, the proposed method extracts the accurate contours of the target as tracking output, which achieves better description of the non-rigid objects while reduces background pollution to the target model. Moreover, conventional level set models only emphasize the regional intensity consistency and consider no priors. Differently, the curve evolution of the proposed SLSM is object-oriented and supervised by the specific knowledge of the targets we want to track. Therefore, the SLSM can ensure a more accurate convergence to the exact targets in tracking applications. In particular, we firstly construct the appearance model for the target in an online boosting manner due to its strong discriminative power between the object and the background. Then, the learnt target model is incorporated to model the probabilities of the level set contour by a Bayesian manner, leading the curve converge to the candidate region with maximum likelihood of being the target. Finally, the accurate target region qualifies the samples fed to the boosting procedure as well as the target model prepared for the next time step. We firstly describe the proposed mechanism of two-phase SLSM for single target tracking, then give its generalized multi-phase version for dealing with multi-target tracking cases. Positive decrease rate is used to adjust the learning pace over time, enabling tracking to continue under partial and total occlusion. Experimental results on a number of challenging sequences validate the effectiveness of the proposed method. Xin Sun 0003, Hongxun Yao, Shengping Zhang |
IEEE Trans. Image Process. | 2 |
| 2015 | Learning Cross Space Mapping via DNN Using Large Scale Click-Through LogsabstractThe gap between low-level visual signals and high-level semantics has been progressively bridged by continuous development of deep neural network (DNN). With recent progress of DNN, almost all image classification tasks have achieved new records of accuracy. To extend the ability of DNN to image retrieval tasks, we proposed a unified DNN model for image-query similarity calculation by simultaneously modeling image and query in one network. The unified DNN is named the cross space mapping (CSM) model, which contains two parts, a convolutional part and a query-embedding part. The image and query are mapped to a common vector space via these two parts respectively, and image-query similarity is naturally defined as an inner product of their mappings in the space. To ensure good generalization ability of the DNN, we learn weights of the DNN from a large number of click-through logs which consists of 23 million clicked image-query pairs between 1 million images and 11.7 million queries. Both the qualitative results and quantitative results on an image retrieval evaluation task with 1000 queries demonstrate the superiority of the proposed method. Wei Yu 0004, Kuiyuan Yang, Yalong Bai, Hongxun Yao, Yong Rui |
IEEE Trans. Multim. | 4 |
| 2014 | DNN Flow: DNN Feature Pyramid based Image Matching
Wei Yu 0004, Kuiyuan Yang, Yalong Bai, Hongxun Yao, Yong Rui |
BMVC | 4 |
| 2014 | "Clustering by saliency" - Unsupervised discovery of crowd activitiesabstractIn this paper, we develop a novel unsupervised crowd activity discovery algorithm aiming to automatically explore latent action patterns among crowd activities and partition them into meaningful clusters. Inspired by computational model of human vision system, we present a spatiotemporal saliency-based representation to simulate visual attention mechanism and encode human-focused components in an activity stream. Combining with feature pooling, we could obtain a more compact and robust activity representation. Based on the affinity matrix of activities, N-cut is performed to generate clusters with meaningful activity patterns. We carry out experiments on our proposed HIT-BJUT dataset and another public UMN dataset. The experimental results demonstrate that the proposed unsupervised discovery method is capable of automatically mining meaningful activities from large-scale video data with mixed crowd activities. Tingting Han 0003, Hongxun Yao, Xiaoshuai Sun, Yanhao Zhang 0001 |
ICIP | 2 |
| 2014 | Structure-aware multi-object discovery for weakly supervised trackingabstractRecent progress on tracking has focused on designing robust statistical model or proposing effective appearance features to improve precision. This paper addresses another problem, namely the discovery and tracking of generic multi-object which have the similar appearance and motion pattern based on limited human annotations. We present a model-free tracking method that can automatically discover and track multi-object sharing the same spatial and motion structure, and update the structure during the tracking without prior acknowledge. The candidate objects are first selected by a SVM classifier trained on histogram-of-gradient (HOG) features. Then a segment algorithm is exploited to decide the suitable sizes of tracking boxes. The structure constrains are updated in a real-time manner according to the motion measure among the specified object and corresponding candidates. Experimental results reveal significant convenience and remarkable performance of our approach for the task of structure preserving multi-object discovery and tracking. Yuankai Qi, Hongxun Yao, Xiaoshuai Sun, Xin Sun 0003, Yanhao Zhang 0001, Qingming Huang |
ICIP | 2 |
| 2014 | Exploring covert attention for generic boosting of saliency modelsabstractCovert attention is an mental ability to attend to a stimulus without shifting ones gaze towards it. The covert attention mechanism allows small spatial displacement during saccadic eye-movements, which we believed to be responsible for the existence of a large number of imperfectly allocated eye fixations in current saliency benchmark datasets. Inspired by this new finding, we propose to use spatial pooling to integrate cover attention into the current saliency models. We test our pooling-based boosting strategy for 20 state-of-the-art attention models on two well acknowledged fixation datasets (YORK-120 & MIT-1003). The experimental results show that our method can stably improve the performance of over 95% of the tested models in the eye fixation prediction task. Xiaoshuai Sun, Hongxun Yao |
ICIP | 2 |
| 2014 | Exploring Principles-of-Art Features For Image Emotion RecognitionabstractEmotions can be evoked in humans by images. Most previous works on image emotion analysis mainly used the elements-of-art-based low-level visual features. However, these features are vulnerable and not invariant to the different arrangements of elements. In this paper, we investigate the concept of principles-of-art and its influence on image emotions. Principles-of-art-based emotion features (PAEF) are extracted to classify and score image emotions for understanding the relationship between artistic principles and emotions. PAEF are the unified combination of representation features derived from different principles, including balance, emphasis, harmony, variety, gradation, and movement. Experiments on the International Affective Picture System (IAPS), a set of artistic photography and a set of peer rated abstract paintings, demonstrate the superiority of PAEF for affective image classification and regression (with about 5% improvement on classification accuracy and 0.2 decrease in mean squared error), as compared to the state-of-the-art approaches. We then utilize PAEF to analyze the emotions of master paintings, with promising results. Sicheng Zhao, Yue Gao 0002, Xiaolei Jiang, Hongxun Yao, Tat-Seng Chua, Xiaoshuai Sun |
ACM Multimedia | 4 |
| 2014 | Affective Image Retrieval via Multi-Graph LearningabstractImages can convey rich emotions to viewers. Recent research on image emotion analysis mainly focused on affective image classification, trying to find features that can classify emotions better. We concentrate on affective image retrieval and investigate the performance of different features on different kinds of images in a multi-graph learning framework. Firstly, we extract commonly used features of different levels for each image. Generic features and features derived from elements-of-art are extracted as low-level features. Attributes and interpretable principles-of-art based features are viewed as mid-level features, while semantic concepts described by adjective noun pairs and facial expressions are extracted as high-level features. Secondly, we construct single graph for each kind of feature to test the retrieval performance. Finally, we combine the multiple graphs together in a regularization framework to learn the optimized weights of each graph to efficiently explore the complementation of different features. Extensive experiments are conducted on five datasets and the results demonstrate the effectiveness of the proposed method. Sicheng Zhao, Hongxun Yao, Yanhao Zhang 0001 |
ACM Multimedia | 2 |
| 2014 | Action recognition based on overcomplete independent components analysis
Shengping Zhang, Hongxun Yao, Xin Sun 0003, Kuanquan Wang, Jun Zhang 0017, Xiusheng Lu, Yanhao Zhang 0001 |
Inf. Sci. | 2 |
| 2014 | Preface: Internet multimedia computing and service
Shuqiang Jiang, Changsheng Xu, Yong Rui, Alberto Del Bimbo, Hongxun Yao |
Multim. Tools Appl. | 5 |
| 2014 | Where should I stand? Learning based human position recommendation for mobile photographing
Pengfei Xu 0001, Hongxun Yao, Rongrong Ji, Xianming Liu 0005, Xiaoshuai Sun |
Multim. Tools Appl. | 2 |
| 2014 | A refined particle filter based on determined level set model for robust contour tracking
Xin Sun 0003, Hongxun Yao |
Mach. Vis. Appl. | 2 |
| 2014 | Visual tracking via weakly supervised learning from multiple imperfect oracles
Bineng Zhong 0001, Hongxun Yao, Sheng Chen 0007, Rongrong Ji, Tat-Jun Chin, Hanzi Wang |
Pattern Recognit. | 2 |
| 2014 | Toward Statistical Modeling of Saccadic Eye-Movement and Visual SaliencyabstractIn this paper, we present a unified statistical framework for modeling both saccadic eye movements and visual saliency. By analyzing the statistical properties of human eye fixations on natural images, we found that human attention is sparsely distributed and usually deployed to locations with abundant structural information. This observations inspired us to model saccadic behavior and visual saliency based on super-Gaussian component (SGC) analysis. Our model sequentially obtains SGC using projection pursuit, and generates eye movements by selecting the location with maximum SGC response. Besides human saccadic behavior simulation, we also demonstrated our superior effectiveness and robustness over state-of-the-arts by carrying out dense experiments on synthetic patterns and human eye fixation benchmarks. Multiple key issues in saliency modeling research, such as individual differences, the effects of scale and blur, are explored in this paper. Based on extensive qualitative and quantitative experimental results, we show promising potentials of statistical approaches for human behavior research. Xiaoshuai Sun, Hongxun Yao, Rongrong Ji, Xianming Liu 0005 |
IEEE Trans. Image Process. | 2 |
| 2013 | Exploring Implicit Image Statistics for Visual Representativeness ModelingabstractIn this paper, we propose a computational model of visual representative ness by integrating cognitive theories of representative ness heuristics with computer vision and machine learning techniques. Unlike previous models that build their representative ness measure based on the visible data, our model takes the initial inputs as explicit positive reference and extend the measure by exploring the implic it negatives. Given a group of images that contains obvious visual concepts, we create a customized image ontology consisting of both positive and negative instances by mining the most related and confusable neighbors of the positive concept in ontological semantic knowledge bases. The representative ness of a new item is then determined by its likelihoods for both the positive and negative references. To ensure the effectiveness of probability inference as well as the cognitive plausibility, we discover the potential prototypes and treat them as an intermediate representation of semantic concepts. In the experiment, we evaluate the performance of representative ness models based on both human judgements and user-click logs of commercial image search engine. Experimental results on both Image Net and image sets of general concepts demonstrate the superior performance of our model against the state-of-the-arts. Xiaoshuai Sun, Xin-Jing Wang, Hongxun Yao, Lei Zhang 0001 |
CVPR | 3 |
| 2013 | A spatial-temporal constraint-based action recognition methodabstractIn this paper, we propose a spatial-temporal constraint-based action recognition method, in which two actions are compared by both the appearance features and the spatial-temporal structures. To represent the appearance information in videos, we utilize a random quantization method to obtain a more precise BoW-based representation. To calculate the similarity of two videos, we match the quantized interest point sets and map the matched pairs into the spatial-temporal offset space to compare the similarities of spatial-temporal structures. We leverage the KNN model to classify the actions. The experiment results on both KTH action dataset and YouTube action dataset demonstrate the effectiveness of our proposed method on action recognition. Tingting Han 0003, Hongxun Yao, Yanhao Zhang 0001, Pengfei Xu 0001 |
ICIP | 2 |
| 2013 | Night video enhancement using improved dark channel priorabstractVideos taken under low lighting condition usually have serious loss of visibility and contrast and are inconvenient for observation and analysis. To solve this problem, this paper presents a real-time night video enhancement approach. As observed that a pixel-wise inversion of a night video has quite similar appearance with the video acquired at foggy days, we use the similar idea of haze removal method to enhance the perceptual quality of the night videos. We present an improved dark channel prior model and integrate it with local smoothing and image Gaussian Pyramid operators. The experimental results demonstrate that the proposed approach can improve the perceptual quality of night videos in real-time in terms of not only enhancing details, but also effectively avoiding excessive enhancement phenomenon. Xuesong Jiang, Hongxun Yao, Shengping Zhang, Xiusheng Lu, Wei Zeng 0006 |
ICIP | 2 |
| 2013 | On dense sampling sizeabstractThis paper proposes a general method for size optimization in dense sampling to obtain a better representation of an image. Our method can be utilized to improve the performance of image classification and other tasks. We discuss the spatial consistency in global-scope restrained descriptors, by analyzing the appropriate sampling size. We apply the low rank method to solve the representative matrix of the descriptor sets at different scales, and obtain the optimized dense sampling size according to the lowest ranks of the representative matrices. Experimental results indicate that the proposed method gives an innovative and effective image representation, and it outperforms traditional dense sampling without size optimization. Hongxun Yao, Xiaoshuai Sun, Yanhao Zhang 0001 |
ICIP | 2 |
| 2013 | Real-time visual tracking using ℓ2 norm regularization based collaborative representationabstractRecently, sparse representation based visual tracking have been attracting increasing interests. Although reported desired performance, whether the sparse representation constrain is really useful is not clear. In addition, the high computation complexity also limits their usage in real-time applications. In this paper, we proposed a real-time visual tracking framework using ℓ2norm regularization based collaborative representation. Our framework represents any target candidate using a set of target templates and a set of background templates respectively, then combines their reconstruction errors to track the target accurately. By constraining ℓ2norm regularization on the representation coefficients, the coefficients can be solved analytically, which makes the proposed method run in real-time. The experimental results demonstrate that the proposed approach outperforms several state-of-the-art trackers. Xiusheng Lu, Hongxun Yao, Xin Sun 0003, Xuesong Jiang |
ICIP | 2 |
| 2013 | Non-rigid object tracking by adaptive data-driven kernelabstractWe derive an adaptive data-driven kernel in this paper to simultaneously address the kernel scale/orientation selection problem as well as the constant kernel shape in deformable object tracking applications. Level set technique is novelly introduced into the mean shift sample space to implement kernel evolution and update. Since the active contour model is designed to drive the kernel constantly to the direction that maximizes target likelihood, the kernel can adapt to target shape variation simultaneously with the mean shift iterations. Thus, it can give a better estimation bias to produce accurate shift of the mean and successfully avoid performance loss stemmed from pollution of the non-object regions hiding inside the kernel. Experimental results on a number of challenging sequences validate the effectiveness of the technique. Xin Sun 0003, Hongxun Yao, Shengping Zhang, Mingui Sun |
ICIP | 2 |
| 2013 | Sparse coding based motion attention for abnormal event detectionabstractIn this paper, we present a novel method based on sparsely coded motion attention for detecting abnormal events in crowded scenes. Unlike existing sparse coding based approaches, our model does not need to learn a dictionary and directly sparsely codes the motion features of the center patches with features of its surrounding patches. The sparse coding error is used to measure the motion attention intensity of the center patch. To reflect the crowd abnormal intensity, an online updated weighting scheme is designed to obtain the global activity intensity map. Two publicly available datasets-UMN dataset and UCSD Ped1 dataset are utilized to evaluate our approach in detecting global abnormal event and local abnormal event, respectively. The experiments show our method achieves the promising performance and is competitive with the state-of-the-art approaches. Shengping Zhang, Hongxun Yao |
ICIP | 3 |
| 2013 | Structured Textons for texture representationabstractIn this paper, we propose a novel texture descriptor, Structured Texton, to extract and characterize meaningful texture patterns in images. Structured Textons are constructed by grouping local extremum regions connected by the nesting relationship. To further improve the discriminative ability, high order texton words are generated from the Structured Textons, preserving both the appearance information and the spatial information. Finally, a semantic ranking criterion is proposed for selecting the discriminative high order texton words by means of finding informative patterns from images. The proposed Structured Texton is more discriminative than the single texton-based representation. Experimental results of texture classification and scene classification on public datasets demonstrate the effectiveness and discrimination of the proposed Structured Texton. Pengfei Xu 0001, Xianming Liu 0005, Hongxun Yao, Yanhao Zhang 0001, Shaopeng Tang |
ICIP | 3 |
| 2013 | The shortest warping path based multiple images alignmentabstractIn this paper, we propose a method to align multiple images of the same category. Images with large variations are aligned via a smooth transition formed by some intermediate images. These intermediate images are found by shortest warping path algorithm on a directed complete graph. Moreover, the common regions in the images are discovered to further improve alignment performance. The experimental results show that our method is effective to map and align images of the same category but with large variations of appearance, shape and view. Wei Yu 0004, Hongxun Yao, Kuiyuan Yang, Lei Zhang 0001 |
ICIP | 2 |
| 2013 | Beyond particle flow: Bag of Trajectory Graphs for dense crowd event recognitionabstractIn this paper, a novel crowd behavior representation, Bag of Trajectory Graphs (BoTG), is presented for dense crowd event recognition. To overcome huge loss of crowd structure and variability of motion in previous particle flow based methods, we design group-level representation beyond particle flow. From the observation that crowd particles are composed of atomic subgroups corresponding to informative behavior patterns, particle trajectories which simulate motion of individuals will be clustered to form groups at the first step. Then we connect nodes in each group as a trajectory graph and discover informative features to depict the graphs. A clip of crowd event can be further described by Bag of Trajectory Graphs (BoTG)-occurrences of behavior patterns, which provides critical clues for categorizing specific crowd event and detecting abnormality. The experimental results of abnormality detection and event recognition on public datasets demonstrate the effectiveness of our proposed BoTG on characterizing the group behaviors in dense crowd. Yanhao Zhang 0001, Hongxun Yao, Pengfei Xu 0001, Qingming Huang |
ICIP | 3 |
| 2013 | Flexible Presentation of Videos Based on Affective Content Analysis
Sicheng Zhao, Hongxun Yao, Xiaoshuai Sun, Xiaolei Jiang, Pengfei Xu 0001 |
MMM (1) | 2 |
| 2013 | Robust visual tracking based on online learning sparse representation
Shengping Zhang, Hongxun Yao, Huiyu Zhou 0001, Xin Sun 0003, Shaohui Liu |
Neurocomputing | 2 |
| 2013 | Video classification and recommendation based on affective analysis of viewers
Sicheng Zhao, Hongxun Yao, Xiaoshuai Sun |
Neurocomputing | 2 |
| 2013 | Visual attention modeling based on short-term environmental adaption
Xiaoshuai Sun, Hongxun Yao, Rongrong Ji |
J. Vis. Commun. Image Represent. | 2 |
| 2013 | Bidirectional-isomorphic manifold learning at image semantic understanding & representation
Xianming Liu 0005, Hongxun Yao, Rongrong Ji, Pengfei Xu 0001, Xiaoshuai Sun |
Multim. Tools Appl. | 2 |
| 2013 | Sparse coding based visual tracking: Review and experimental comparison
Shengping Zhang, Hongxun Yao, Xin Sun 0003, Xiusheng Lu |
Pattern Recognit. | 2 |
| 2013 | Weakly supervised codebook learning by iterative label propagation with graph quantization
Liujuan Cao, Rongrong Ji, Wei Liu 0005, Hongxun Yao, Qi Tian 0001 |
Signal Process. | 4 |
| 2013 | Learning from mobile contexts to minimize the mobile location search latency
Ling-Yu Duan, Rongrong Ji, Jie Chen 0006, Hongxun Yao, Tiejun Huang 0001, Wen Gao 0001 |
Signal Process. Image Commun. | 4 |
| 2013 | Learning to Distribute Vocabulary Indexing for Scalable Visual SearchabstractIn recent years, there is an ever-increasing research focus on Bag-of-Words based near duplicate visual search paradigm with inverted indexing. One fundamental yet unexploited challenge is how to maintain the large indexing structures within a single server subject to its memory constraint, which is extremely hard to scale up to millions or even billions of images. In this paper, we propose to parallelize the near duplicate visual search architecture to index millions of images over multiple servers, including the distribution of both visual vocabulary and the corresponding indexing structure. We optimize the distribution of vocabulary indexing from a machine learning perspective, which provides a “memory light” search paradigm that leverages the computational power across multiple servers to reduce the search latency. Especially, our solution addresses two essential issues: “What to distribute” and “How to distribute”. “What to distribute” is addressed by a “lossy” vocabulary Boosting, which discards both frequent and indiscriminating words prior to distribution. “How to distribute” is addressed by learning an optimal distribution function, which maximizes the uniformity of assigning the words of a given query to multiple servers. We validate the distributed vocabulary indexing scheme in a real world location search system over 10 million landmark images. Comparing to the state-of-the-art alternatives of single-server search,,and distributed search, our scheme has yielded a significant gain of about 200% speedup at comparable precision by distributing only 5% words. We also report excellent robustness even when partial servers crash. Rongrong Ji, Ling-Yu Duan, Jie Chen 0006, Lexing Xie, Hongxun Yao, Wen Gao 0001 |
IEEE Trans. Multim. | 5 |
| 2012 | The scale of edgesabstractAlthough the scale of isotropic visual elements such as blobs and interest points, e.g. SIFT[12], has been well studied and adopted in various applications, how to determine the scale of anisotropic elements such as edges is still an open problem. In this paper, we study the scale of edges, and try to answer two questions: 1) what is the scale of edges, and 2) how to calculate it. From the points of human cognition and physical interpretation, we illustrate the existence of the scale of edges and provide a quantitative definition. Then, an automatic edge scale selection approach is proposed. Finally, a cognitive experiment is conducted to validate the rationality of the detected scales. Moreover, the importance of identifying the scale of edges is also shown in applications such as boundary detection and hierarchical edge parsing. Xianming Liu 0005, Changhu Wang, Hongxun Yao, Lei Zhang 0001 |
CVPR | 3 |
| 2012 | What are we looking for: Towards statistical modeling of saccadic eye movements and visual saliencyabstractIn this paper, we present a unified statistical framework for modeling both saccadic eye movements and visual saliency. By analyzing the statistical properties of human eye fixations on natural images, we found that human attention is sparsely distributed and usually deployed to locations with abundant structural information. This new observations inspired us to model saccadic behavior and visual saliency based on Super Gaussian Component (SGC) analysis. The model sequentially obtains SGC using projection pursuit, and generates eye-movements by selecting the location with maximum SGC response. Beside human saccadic behavior simulation, we also demonstrated our superior effectiveness and robustness over state-of-the-arts by carrying out dense experiments on psychological patterns and human eye fixation benchmarks. These results also show promising potentials of statistical approaches for human behavior research. Xiaoshuai Sun, Hongxun Yao, Rongrong Ji |
CVPR | 2 |
| 2012 | Abnormal crowd behavior detection based on social attribute-aware force modelabstractIn this paper, a novel social attribute-aware force model is presented for abnormal crowd pattern detection in video sequences. We take social characteristics of crowd behaviors into account in order to improve the effectiveness of the simulation on the interaction behaviors of the crowd. A quick unsupervised method is proposed to estimate the scene scale. Both the social disorder attribute and congestion attribute are introduced to describe the realistic social behaviors by using statistical context feature. Through the semantic attribute-aware enhancement, we obtain an improved model on the basis of social force. We validate our method in public available datasets for abnormal detection, and the experimental results show promising performance compared with other state of the art methods. Yanhao Zhang 0001, Hongxun Yao, Qingming Huang |
ICIP | 3 |
| 2012 | Aesthetic composition represetation for portrait photographing recommendationabstractIn this paper, we present an intelligent portrait photographing framework for automatically recommending the suitable positions and poses in the scene of photography taken by amateurs. By analyzing aesthetic characteristics features, we propose a solution by constructing aesthetic composition representation which covers the attention composition and geometry composition to identify the underlying technique of professional photographer. First, we extract the attention composition feature of the professional photo by utilizing a visual saliency model. Then, a geometry composition feature is also presented to learn the spatial correlation. Finally, composition rules are applied to make appropriate pose and position. Experiments show our aesthetic composition representation performs well for portrait photographing recommendation. Yanhao Zhang 0001, Xiaoshuai Sun, Hongxun Yao, Qingming Huang |
ICIP | 3 |
| 2012 | Memorable basis: towards human-centralized sparse representationabstractPrevious studies of sparse representation in multimedia research focus on developing reliable and efficient dictionary learning algorithms. Despite the sparse prior, how to integrate other related perceptual factors of human being into dictionary learning process was seldom studied. In this paper, we investigate the influence of image memorability for human-centralized sparse representation. Based on the results of a photo memory game, we are able to quantitatively characterize an image's memorability which allows us to train sparse bases from the most memorable images instead of randomly selected natural images. We believed that such kind of basis is more consistent with neural networks in human brain and hence can better predict where human looks. To test our hypothesis, we choose human eye-fixation prediction problem for quantitative evaluation. The experimental results demonstrate the superior performance of our Memorable Basis compared to traditional sparse basis trained from unselected images. Xiaoshuai Sun, Hongxun Yao |
ACM Multimedia | 2 |
| 2012 | Action retrieval based on generalized dynamic depth data matchingabstractWith the great popularity and extensive application of Kinect, the Internet is sharing more and more depth data. To effectively use plenty of depth data would make great sense. In this paper, we propose a generalized dynamic depth data matching framework for action retrieval. Firstly we focus on single depth image matching utilizing both depth and shape feature. The depth feature used in our method is straightforward but proved to be very effective and robust for distinguishing various human actions. Then, we adopt shape context, which is widely used in shape matching, in order to strengthen the robustness of our matching strategy. Finally, we utilize Dynamic Time Warping to measure temporal similarity between two depth video sequences. Experiments based on a dataset of 17 classes of actions from 10 different individuals demonstrate the effectiveness and robustness of our proposed matching strategy. Lujun Chen, Hongxun Yao, Xiaoshuai Sun |
VCIP | 2 |
| 2012 | Location Discriminative Vocabulary Coding for Mobile Landmark Search
Rongrong Ji, Ling-Yu Duan, Jie Chen 0006, Hongxun Yao, Junsong Yuan 0001, Yong Rui, Wen Gao 0001 |
Int. J. Comput. Vis. | 4 |
| 2012 | Task-Dependent Visual-Codebook CompressionabstractA visual codebook serves as a fundamental component in many state-of-the-art computer vision systems. Most existing codebooks are built based on quantizing local feature descriptors extracted from training images. Subsequently, each image is represented as a high-dimensional bag-of-words histogram. Such highly redundant image description lacks efficiency in both storage and retrieval, in which only a few bins are nonzero and distributed sparsely. Furthermore, most existing codebooks are built based solely on the visual statistics of local descriptors, without considering the supervise labels coming from the subsequent recognition or classification tasks. In this paper, we propose a task-dependent codebook compression framework to handle the above two problems. First, we propose to learn a compression function to map an originally high-dimensional codebook into a compact codebook while maintaining its visual discriminability. This is achieved by a codeword sparse coding scheme with Lasso regression, which minimizes the descriptor distortions of training images after codebook compression. Second, we propose to adapt our codebook compression to the subsequent recognition or classification tasks. This is achieved by introducing a label constraint kernel (LCK) into our compression loss function. In particular, our LCK can model heterogeneous kinds of supervision, i.e., (partial) category labels, correlative semantic annotations, and image query logs. We validated our codebook compression in three computer vision tasks: 1) object recognition in PASCAL Visual Object Class 07; 2) near-duplicate image retrieval in UKBench; and 3) web image search in a collection of 0.5 million Flickr photographs. Our compressed codebook has shown superior performances over several state-of-the-art supervised and unsupervised codebooks. Rongrong Ji, Hongxun Yao, Wei Liu 0005, Xiaoshuai Sun, Qi Tian 0001 |
IEEE Trans. Image Process. | 2 |
| 2012 | Context-Aware Semi-Local Feature DetectorabstractHow can interest point detectors benefit from contextual cues? In this articles, we introduce a context-aware semi-local detector (CASL) framework to give a systematic answer with three contributions: (1) We integrate the context of interest points to recurrently refine their detections. (2) This integration boosts interest point detectors from the traditionally local scale to a semi-local scale to discover more discriminative salient regions. (3) Such context-aware structure further enables us to bring forward category learning (usually in the subsequent recognition phase) into interest point detection to locate category-aware, meaningful salient regions. Our CASL detector consists of two phases. The first phase accumulates multiscale spatial correlations of local features into a difference of contextual Gaussians (DoCG) field. DoCG quantizes detector context to highlight contextually salient regions at a semi-local scale, which also reveals visual attentions to a certain extent. The second phase locates contextual peaks by mean shift search over the DoCG field, which subsequently integrates contextual cues into feature description. This phase enables us to integrate category learning into mean shift search kernels. This learning-based CASL mechanism produces more category-aware features, which substantially benefits the subsequent visual categorization process. We conducted experiments in image search, object characterization, and feature detector repeatability evaluations, which reported superior discriminability and comparable repeatability to state-of-the-art works. Rongrong Ji, Hongxun Yao, Qi Tian 0001, Pengfei Xu 0001, Xiaoshuai Sun, Xianming Liu 0005 |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2012 | Robust Visual Tracking Using an Effective Appearance Model Based on Sparse CodingabstractIntelligent video surveillance is currently one of the most active research topics in computer vision, especially when facing the explosion of video data captured by a large number of surveillance cameras. As a key step of an intelligent surveillance system, robust visual tracking is very challenging for computer vision. However, it is a basic functionality of the human visual system (HVS). Psychophysical findings have shown that the receptive fields of simple cells in the visual cortex can be characterized as being spatially localized, oriented, and bandpass, and it forms a sparse, distributed representation of natural images. In this article, motivated by these findings, we propose an effective appearance model based on sparse coding and apply it in visual tracking. Specifically, we consider the responses of general basis functions extracted by independent component analysis on a large set of natural image patches as features and model the appearance of the tracked target as the probability distribution of these features. In order to make the tracker more robust to partial occlusion, camouflage environments, pose changes, and illumination changes, we further select features that are related to the target based on an entropy-gain criterion and ignore those that are not. The target is finally represented by the probability distribution of those related features. The target search is performed by minimizing the Matusita distance between the distributions of the target model and a candidate using Newton-style iterations. The experimental results validate that the proposed method is more robust and effective than three state-of-the-art methods. Shengping Zhang, Hongxun Yao, Xin Sun 0003, Shaohui Liu |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2011 | A novel supervised level set method for non-rigid object trackingabstractWe present a novel approach to non-rigid object tracking based on a supervised level set model (SLSM). In contrast with conventional level set models, which emphasize the intensity consistency only and consider no priors, the curve evolution of the proposed SLSM is object-oriented and supervised by the specific knowledge of the target we want to track. Therefore, the SLSM can ensure a more accurate convergence to the target in tracking applications. In particular, we firstly construct the appearance model for the target in an on-line boosting manner due to its strong discriminative power between objects and background. Then the probability of the contour is modeled by considering both the region and edge cues in a Bayesian manner, leading the curve converge to the candidate region with maximum likelihood of being the target. Finally, accurate target region qualifies the samples fed the boosting procedure as well as the target model prepared for the next time step. Positive decrease rate is used to adjust the learning pace over time, enabling tracking to continue under partial and total occlusion. Experimental results on a number of challenging sequences validate the effectiveness of the technique. Xin Sun 0003, Hongxun Yao, Shengping Zhang |
CVPR | 2 |
| 2011 | Sorting local descriptors for lowbit rate mobile visual searchabstractState-of-the-art mobile visual search systems put emphasis on developing compact visual descriptors, which enables low bit rate wireless transmission instead of delivering an entire query image. In this paper, we address the orderless nature of the transmission set of query descriptors . We propose to adapt the orders of local descriptors in transmission, which subsequently yields more consistent statistic distributions in each feature dimension towards more efficient residual coding based compression. Our scheme further enables lossy sorting by an adaptive quantization strategy within each feature dimension, which largely improves the compression rates of the residual coding in each dimension. We show that the performance degeneration of such lossy sorting is acceptable in our mobile landmark search applications. Our approach's effectiveness and efficiency is demonstrated via extensive experimental comparisons to state-of-the art works in both mobile visual descriptors and compact image signatures. Jie Chen 0006, Ling-Yu Duan, Rongrong Ji, Hongxun Yao, Wen Gao 0001 |
ICASSP | 4 |
| 2011 | A lowbit rate vocabulary coding scheme for mobile landmark searchabstractWe present a low bit rate vocabulary coding scheme in the context of mobile landmark search. Our scheme exploits location cues to boost a compact subset of visual vocabulary, which is discriminative for visual search and incurs low bit rate query for efficient upstream wireless transmission. To validate the coding scheme, we have developed mobile landmark search prototype systems within typical areas including Beijing, New York City, Lhasa, Singapore, and Florence. Our system maintains a single vocabulary in a mobile device, which can be efficiently adapted with the location information of city-scale mobile users. Thus multiple downloading of large vocabulary is completely avoided for normal city tourists. In landmark search domain, we have reported superior performance over the state-of-the art works in compact image descriptors or signatures. Rongrong Ji, Ling-Yu Duan, Jie Chen 0006, Hongxun Yao, Wen Gao 0001 |
ICASSP | 4 |
| 2011 | When codeword frequency meets geographical locationabstractWhen codeword frequency meets geographical location in landmark search applications, is it still discriminative for the search procedure. In this paper, we give a systematic investigation about how geographical location affects the effectiveness of codeword frequency. We explain why the standard IDF in the BoW models is less effective in location related search applications [11][12]. Consequently, we propose a “location discriminative codeword frequency” strategy to introduce the location context into the codeword discriminability measurement. This new codeword frequency is calculated in each geographical region, for which a spectral clustering scheme is proposed to partition the geographical map of each city into distinct regions. Extensive comparisons over the standard codeword frequency in state-of-the-art landmark search systems [1][1] demonstrates our approach's effectiveness. Rongrong Ji, Ling-Yu Duan, Jie Chen 0006, Hongxun Yao, Wen Gao 0001 |
ICASSP | 4 |
| 2011 | Affective Video Classification Based on Spatio-temporal Feature FusionabstractIn this paper, we propose a novel affective video classification method based on facial expression recognition by learning the spatio-temporal feature fusion of actors' and viewers' facial expressions. For spatial features, we integrate Haar-like features into compositional ones according to the features' correlation, and train a mid classifier during the period. Then this process is embedded into improved AdaBoost learning algorithm to obtain spatial features. And for temporal feature fusion, we adopt hidden dynamic conditional random fields (HDCRFs) based on HCRFs by introducing time dimension variable. Finally spatial features are embedded into HDCRFs to recognize facial expressions. Experiments on the well-known Cohn-Kanada database show that the proposed method has a promising recognition performance. And affective classification experimental results on our own videos show that most subjects are satisfied with the classification results. Sicheng Zhao, Hongxun Yao, Xiaoshuai Sun |
ICIG | 2 |
| 2011 | PKUBench: A context rich mobile visual search benchmarkabstractWhile there are ever growing focuses on mobile visual search in recent years, a comprehensive benchmark database with rich context information (such as GPS) for fair evaluation among different strategies is still missing. This paper introduces a PKUBench benchmark for the quantitative evaluations of mobile visual search with the support of GPS. It contains 13,179 images organized into 198 distinct landmark locations within the Peking University campus. Each location is captured with multiple shot sizes and viewing angles, using both digital cameras and phone cameras, each photo being tagged with rich contextual information in the mobile scenario. Moreover, this benchmark studies typical quality degeneration scenarios in mobile photographing, including variable resolutions, blurring, lighting changes, occlusions, as well as various viewing angles. Together with this benchmark, we provide the bag-of-visual-words search baselines involving contextual information refinement. Finally, distractor images are further introduced to evaluate the robustness of visual search methods in this database. Rongrong Ji, Ling-Yu Duan, Jie Chen 0006, Tiejun Huang 0001, Hongxun Yao, Wen Gao 0001 |
ICIP | 6 |
| 2011 | Learning the trip suggestion from landmark photos on the webabstractIn this paper, we introduce a novel touristic trip suggestion system to facilitate the traveling of mobile users in a given city. Given the current user location and his touristic destination, our system can suggest a shortest trip path that visits as many popular landmarks as possible. To this end, we collect geographical tagged photos from Flickr [1] and Panoramio [2] photo sharing websites. Then a geographical graph is constructed by modeling photos as vertices and their geographical and visual closenesses as connection strengths. In this graph, we mine a dominant subgraph by quantizing nearby and visually duplicated vertices, and then trimming unpopular subgraphs. Such dominant subgraph only retains the popular landmarks from the consensus of travelers in this city. In online suggestion, we map the current user location and the target location to the nearest vertices in this subgraph, based on which an optimal trip is suggested through a shortest path search. We have quantitatively validated our system in typical areas including Beijing and New York City, with quantitative comparisons to alternative approaches. Rongrong Ji, Ling-Yu Duan, Jie Chen 0006, Hongxun Yao, Tiejun Huang 0001, Wen Gao 0001 |
ICIP | 5 |
| 2011 | Sparse representation based visual element analysisabstractModern clothes are designed based on various visual elements of different fashion styles. Traditional vision-based clothes recommendation methods focused on searching clothes which are similar with user preferred samples in the aspects of colors and partial shape elements. In this paper, we propose a method of recommending clothes by mining visual elements of different fashion styles. Independent Component Analysis (ICA) is employed to extract sparse features, and then Term-Frequency (TF) analysis is applied to discover visual elements from these independent components. Finally, we test three ranking metrics for clothes recommendation including Euclidian distance of TFs, Cosine distance of TFs and Minimum TF. Experimental results based on web commercial images demonstrate the effectiveness of the proposed method. Hongxun Yao, Xiaoshuai Sun, Rongrong Ji, Xianming Liu 0005, Pengfei Xu 0001 |
ICIP | 2 |
| 2011 | Contour tracking via on-line discriminative appearance modeling based level setsabstractA novel level set method based on on-line discriminative appearance modeling (DAMLSM) is presented for contour tracking. In contrast with traditional level set models which emphasize the intensity consistent segmentation and consider no priors, the proposed DAMLSM takes the context of tracking into account and use a discriminative patch based target model to guide the curve evolution. By modeling both the region and edge cues in a Bayesian manner, the proposed level set method can lead an accurate convergence to the candidate region with maximum likelihood of being the target. Finally, we update the target model to adapt to the appearance variation, enabling tracking to continue under occlusion. Experiments confirm the robustness and reliability of our method. Xin Sun 0003, Hongxun Yao, Shengping Zhang |
ICIP | 2 |
| 2011 | Robust visual tracking via context objects computingabstractOcclusions are challenging issue for robust visual tracking. In this paper, motivated by the fact that a tracked object is usual- ly embedded into context that provides useful information for estimating the target, we propose a novel tracking algorithm named Tracking with Context Prediction (TCP). The context here includes the neighboring objects and specific parts of tar- get. The proposed method simultaneously track the target and context objects using the existing tracking methods. The positions of the context objects are used to predict the position of the target. Thus, the target can be stably tracked even when it is partially or fully occluded. By computing the probability of each prediction being target, our algorithm allows the drifting of context objects during tracking and do not require predictions from all context objects are correct. Experiments on challenging sequences show significant improvements especially in the case of occlusions and appearance changes. Zhongqian Sun, Hongxun Yao, Shengping Zhang, Xin Sun 0003 |
ICIP | 2 |
| 2011 | Stable Fast Rewiring Depends on the Activation of Skeleton Voxels
Sanming Song, Hongxun Yao |
ICONIP (1) | 2 |
| 2011 | Modular Scale-Free Function Subnetworks in Auditory Areas
Sanming Song, Hongxun Yao |
ICONIP (1) | 2 |
| 2011 | Learning Compact Visual Descriptor for Low Bit Rate Mobile Landmark Search
Rongrong Ji, Ling-Yu Duan, Jie Chen 0006, Hongxun Yao, Tiejun Huang 0001, Wen Gao 0001 |
IJCAI | 4 |
| 2011 | Probe the Potts States in the Minicolumn Dynamics
Sanming Song, Hongxun Yao |
ISNN (1) | 2 |
| 2011 | Towards low bit rate mobile visual search with multiple-channel codingabstractIn this paper, we propose a multiple-channel coding scheme to extract compact visual descriptors for low bit rate mobile visual search. Different from previous visual search scenarios that send the query image, we make use of the ever growing mobile computational capability to directly extract compact visual descriptors at the mobile end. Meanwhile, stepping forward from the state-of-the-art compact descriptor extractions, we exploit the rich contextual cues at the mobile end (such as GPS tags for mobile visual search and 2D barcodes or RFID tags for mobile product search), together with the visual statistics at the reference database, to learn multiple coding channels. Therefore, we describe the query with one of many forms of high-dimensional visual signature, which is subsequently mapped to one or more channels and compressed. The compression function within each channel is learnt based on a novel robust PCA scheme, with specific consideration to preserve the retrieval ranking capability of the original signature. We have deployed our scheme on both iPhone4 and HTC DESIRE 7 to search ten million landmark images in a low bit rate setting. Quantitative comparisons to the state-of-the-arts demonstrate our significant advantages in descriptor compactness (with orders of magnitudes improvement) and retrieval mAP in mobile landmark, product, and CD/book cover search. Rongrong Ji, Ling-Yu Duan, Jie Chen 0006, Hongxun Yao, Yong Rui, Shih-Fu Chang, Wen Gao 0001 |
ACM Multimedia | 4 |
| 2011 | Learning heterogeneous data for hierarchical web video classificationabstractWeb videos such as YouTube are hard to obtain sufficient precisely labeled training data and analyze due to the complex ontology. To deal with these problems, we present a hierarchical web video classification framework by learning heterogeneous web data, and construct a bottom-up semantic forest of video concepts by learning from meta-data. The main contributions are two-folds: firstly, analysis about middle-level concepts' distribution is taken based on data collected from web communities, and a concepts redistribution assumption is made to build effective transfer learning algorithm. Furthermore, an AdaBoost-Like transfer learning algorithm is proposed to transfer the knowledge learned from Flickr images to YouTube video domain and thus it facilitates video classification. Secondly, a group of hierarchical taxonomies named Semantic Forest are mined from YouTube and Flickr tags which reflect better user intention on the semantic level. A bottom-up semantic integration is also constructed with the help of semantic forest, in order to analyze video content hierarchically in a novel perspective. A group of experiments are performed on the dataset collected from Flickr and YouTube. Compared with state-of-the-arts, the proposed framework is more robust and tolerant to web noise. Xianming Liu 0005, Hongxun Yao, Rongrong Ji, Pengfei Xu 0001, Xiaoshuai Sun, Qi Tian 0001 |
ACM Multimedia | 2 |
| 2011 | Unsupervised fast anomaly detection in crowdsabstractIn this paper, we proposed a fast and robust unsupervised framework for anomaly detection and localization in crowed scenes. Our method avoids modeling the normal state of the crowds which is a very complex task due to the large within class variance of the normal target appearance and motion patterns. For each video frame, we extract the spatial temporal features of 3D blocks and generate the saliency map using a block-based center-surround difference operator. Then, motion vector matrix is obtained by adaptive rood pattern search block-matching algorithm and distance normalization. Attractive motion disorder descriptor is proposed to measure the global intensity of anomalies in the scene. Finally, we classify the frames into normal and anomalous ones by a binary classifier. In the experiments, we compared our method against several state-of-the-art approaches on UCSD dataset which is a widely used anomaly detection and localization benchmark. As the only unsupervised approach, our method outputs competitive results with near real-time processing speed Xiaoshuai Sun, Hongxun Yao, Rongrong Ji, Xianming Liu 0005, Pengfei Xu 0001 |
ACM Multimedia | 2 |
| 2011 | Video indexing and recommendation based on affective analysis of viewersabstractMost previous works on video indexing and recommendation were only based on the content of video itself, without considering the affective analysis of viewers, which is an efficient and important way to reflect viewers' attitudes, feelings and evaluations of videos. In this paper, we propose a novel method to index and recommend videos based on affective analysis, mainly on facial expression recognition of viewers. We first build a facial expression recognition classifier by embedding the process of building compositional Haar-like features into hidden conditional random fields (HCRFs). Then we extract viewers' facial expressions frame by frame through the videos, collected from the camera when viewers are watching videos, to obtain the affections of viewers. Finally, we draw the affective curve which tells the process of affection changes. Through the curve, we segment each video into affective sections, give the indexing result of the videos, and list recommendation points from views' aspect. Experiments on our collected database from the web show that the proposed method has a promising performance. Sicheng Zhao, Hongxun Yao, Xiaoshuai Sun, Pengfei Xu 0001, Xianming Liu 0005, Rongrong Ji |
ACM Multimedia | 2 |
| 2011 | Actor-independent action search using spatiotemporal vocabulary with appearance hashing
Rongrong Ji, Hongxun Yao, Xiaoshuai Sun |
Pattern Recognit. | 2 |
| 2011 | Mining flickr landmarks by modeling reconstruction sparsityabstractIn recent years, there have been ever-growing geographical tagged photos on the community Web sites such as Flickr. Discovering touristic landmarks from these photos can help us to make better sense of our visual world. In this article, we report our work on mining landmarks from geotagged Flickr photos for city scene summarization and touristic recommendations. We begin by exploring the geographical and visual statistics of the Web users' photographing manner, based on which we conduct landmark mining in two steps: First, we propose to partition each city into geographical regions based on spectral clustering over the geotags of Flickr photos. Second, in each landmark region, we present a representative photo mining scheme based on sparse representation. Our main idea is to regard the landmark mining problem as a process to find photos whose visual signatures can be reconstructed using other photos of this landmark region with a minimal coding length. This sparse reconstruction scheme offers a general perspective to mine the representative photos. Indeed, by simplifying the data correlation constraints in our scheme, several previous works in representative photo discovery and landmark mining can be derived. Finally, we introduce a Hyperlink-Induced Topic Search model to refine our landmark ranking, which incorporates the community knowledge to simulate the landmark ranking problem as a dynamic page ranking problem. We have deployed our proposed landmark mining framework on a city scene summarization and navigation system, which works on one million geotagged Flickr photos coming from twenty worldwide metropolises. We have also quantitatively compared our scheme with several state-of-the-art works. Rongrong Ji, Yue Gao 0002, Bineng Zhong 0001, Hongxun Yao, Qi Tian 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2010 | Towards semantic embedding in visual vocabularyabstractVisual vocabulary serves as a fundamental component in many computer vision tasks, such as object recognition, visual search, and scene modeling. While state-of-the-art approaches build visual vocabulary based solely on visual statistics of local image patches, the correlative image labels are left unexploited in generating visual words. In this work, we present a semantic embedding framework to integrate semantic information from Flickr labels for supervised vocabulary construction. Our main contribution is a Hidden Markov Random Field modeling to supervise feature space quantization, with specialized considerations to label correlations: Local visual features are modeled as an Observed Field, which follows visual metrics to partition feature space. Semantic labels are modeled as a Hidden Field, which imposes generative supervision to the Observed Field with WordNet-based correlation constraints as Gibbs distribution. By simplifying the Markov property in the Hidden Field, both unsupervised and supervised (label independent) vocabularies can be derived from our framework. We validate our performances in two challenging computer vision tasks with comparisons to state-of-the-arts: (1) Large-scale image search on a Flickr 60,000 database; (2) Object recognition on the PASCAL VOC database. Rongrong Ji, Hongxun Yao, Xiaoshuai Sun, Bineng Zhong 0001, Wen Gao 0001 |
CVPR | 2 |
| 2010 | Novel observation model for probabilistic object trackingabstractTreating visual object tracking as foreground and background classification problem has attracted much attention in the past decade. Most methods adopt mean shift or brute force search to perform object tracking on the generated probability map, which is obtained from the classification results; however, performing probabilistic object tracking on the probability map is almost unexplored. This paper proposes a novel observation model which is suitable to perform this task. The observation model considers both region and boundary cues on the probability map, and can be computed very efficiently by using the integral image data structure. Extensive experiments are carried out on several challenging image sequences, which include abrupt motion change, background clutter, partial occlusion, and significant appearance change. Quantitative experiments are further performed with several related trackers on a public benchmark dataset. The experimental results demonstrate the effectiveness of the proposed approach. Dawei Liang, Qingming Huang, Hongxun Yao, Shuqiang Jiang, Rongrong Ji, Wen Gao 0001 |
CVPR | 3 |
| 2010 | Visual tracking via weakly supervised learning from multiple imperfect oraclesabstractLong-term persistent tracking in ever-changing environments is a challenging task, which often requires addressing difficult object appearance update problems. To solve them, most top-performing methods rely on online learning-based algorithms. Unfortunately, one inherent problem of online learning-based trackers is drift, a gradual adaptation of the tracker to non-targets. To alleviate this problem, we consider visual tracking in a novel weakly supervised learning scenario where (possibly noisy) labels but no ground truth are provided by multiple imperfect oracles (i.e., trackers), some of which may be mediocre. A probabilistic approach is proposed to simultaneously infer the most likely object position and the accuracy of each tracker. Moreover, an online evaluation strategy of trackers and a heuristic training data selection scheme are adopted to make the inference more effective and fast. Consequently, the proposed method can avoid the pitfalls of purely single tracking approaches and get reliable labeled samples to incrementally update each tracker (if it is an appearance-adaptive tracker) to capture the appearance changes. Extensive comparing experiments on challenging video sequences demonstrate the robustness and effectiveness of the proposed method. Bineng Zhong 0001, Hongxun Yao, Sheng Chen 0007, Rongrong Ji, Xiao-Tong Yuan, Shaohui Liu, Wen Gao 0001 |
CVPR | 2 |
| 2010 | Exploring statistical properties for semantic annotation: sparse distributed and convergent assumptions for keywordsabstractDoes there exist a compact set of visual topics in form of keyword clusters capable to represent all images visual content within an acceptable error? In this paper, we answer this question by analyzing distribution laws for keywords from image descriptions and comparing with traditional techniques in NLP, thereby propose three assumptions: Sparse Distribution Attribute, Local Convergent Assumption and Global Convergent Conjecture. They are essential for keywords selection and image content understanding to overcome the semantic gap. Experiments are performed on a 60,000 web crawled images, and the correctness is validated by the performance. Xianming Liu 0005, Hongxun Yao, Rongrong Ji |
ICASSP | 2 |
| 2010 | SIGMA: Spatial Integrated Matching Association algorithm for logo detectionabstractIn this paper, we adopt the integration model of spatial feature correlations to order the indexing and matching features, and address the computational ineffectiveness and inefficiency of local features based logo detection methods. We propose a Spatial InteGrated Matching Association algorithm (SIGMA) for logo detection in natural scene that contains extremely variances in viewpoints, illuminations and occlusions. Our SIGMA algorithm consists of two phases: the Spatial InteGrated (SIG) phase and the Matching Association (MA) phase. The SIG phase integrates spatial correlations in feature representation, while the MA phase improves the matching performance by ordering an optimized matching sequence. We have collected a logo dataset containing 2,400 photos with 12 logo categories from Flickr, and experimental results demonstrate that the performance of proposed approach outperforms the state-of-the-art approaches on the dataset. Pengfei Xu 0001, Hongxun Yao, Rongrong Ji |
ICASSP | 2 |
| 2010 | Mining actor correlations with hierarchical concurrence parsingabstractMining actor correlations from TV series enables semantic level video understanding and facilitates users to conduct correlation-based query. In this paper, we introduce a graph-based actor correlations mining framework, which serves as the first attempt for effective actor association presentation and concurrence search. We leverage face detection and tracking to locate actors with 2D-PCA detector as pretreatment. To measure the actor association into a unified graph, we propose a context-based actor correlations hierarchical parsing approach, which considers video structure and hierarchical concurrence to refine actor association in our graph modeling. We not only can accomplish actor correlations mining, but also can acquire higher semantic information according to concurrence change. We present the actor correlation mining results in a graph-based interface to enable efficient users' navigation and search. Hongxun Yao, Rongrong Ji, Xiaoshuai Sun |
ICASSP | 2 |
| 2010 | Robust visual tracking using feature-based visual attentionabstractPsychophysical findings have shown that human vision system has an ability to improve target search by enhancing the representation of image components that are related to the searched target, which is the so-called feature-based visual attention. In this paper, motivated by these psychophysical findings, we propose a robust visual tracking algorithm by simulating such feature-based visual attention. Specially, we consider the general sparse basis functions extracted on a large set of natural image patches as features. We define that a feature is related to the target when succeeding activations of that feature cannot increase system's entropy. The target is finally represented by the probability distribution of those related features. The target search is performed by minimizing the Matusita distance measure between the distributions of the target model and candidate using Newton-style iterations. The experimental results verify that the proposed method is more robust and effective than widely used mean shift based methods. Shengping Zhang, Hongxun Yao, Shaohui Liu |
ICASSP | 2 |
| 2010 | Robust background modeling via standard variance featureabstractIn this paper, a novel standard variance feature is proposed for background modeling in dynamic scenes involving waving trees and ripples in water. The standard variance feature is the standard variance of a set of pixels' feature values, which captures mainly co-occurrence statistics of neighboring pixels in an image patch. The background modeling method based on standard variance feature includes two main components. First, we divide image into patches and represent each image patch as a standard variance feature. Then, assuming that standard variance feature fits a mixture of Gaussians distribution, we use mixture of Gaussians models to model it. Experimental results on several challenging video sequences demonstrate the effectiveness of our method. Bineng Zhong 0001, Hongxun Yao, Shaohui Liu |
ICASSP | 2 |
| 2010 | An Image Data Hiding Method Using Pixel-Based JND Model
Shaohui Liu, Feng Jiang 0001, Hongxun Yao, Debin Zhao |
ICIC (3) | 3 |
| 2010 | Visual saliency as sequential eye fixation probabilityabstractHuman vision system acquires essential information from the environment by sequentially sampling visual contents at important locations under the control of selective attention mechanism. We propose that bottom-up saliency is not based on global statistics but on information sampled at prior eye fixations. Our model calculates visual saliency using sequential eye fixation probability. However, the proposed model needs fixation priors, which are hard to simulate given current fixation data and experimental conditions. An approximation is proposed to generate a single saliency map by fusing all possible conditions of fixation prior. Our method outperforms all state-of-the-art models in predicting eye fixations, and shows reasonable response to various psychological patterns. Xiaoshuai Sun, Hongxun Yao, Rongrong Ji, Pengfei Xu 0001, Xianming Liu 0005, Shaohui Liu |
ICIP | 2 |
| 2010 | Saliency detection based on short-term sparse representationabstractRepresentation and measurement are two important issues for saliency models. Different with previous works that learnt sparse features from large scale natural statistics, we propose to learn features from short-term statistics of single images. For saliency measurement, we define background firing rate (BFR) for each sparse feature, and then we propose to use feature activation rate (FAR) to measure the bottom-up visual saliency. The proposed FAR measure is biological plausible and easy to compute, also with satisfied performance. Experiments on human eye fixations and psychological patterns demonstrate the effectiveness and robustness of our proposed method. Xiaoshuai Sun, Hongxun Yao, Rongrong Ji, Pengfei Xu 0001, Xianming Liu 0005, Shaohui Liu |
ICIP | 2 |
| 2010 | A steganography strategy based on equivalence partitions of hiding unitsabstractThis paper designs a novel hiding strategy based on an equivalence relation, which can remarkably enhance the quality of stego image without sacrificing the security and capacity of original steganography schemes. According to a constructed equivalence relation based on the capacity of hiding units, all hiding units can be partitioned into equivalence classes. Following that, the hiding procedure is performed in predefined order in equivalence classes as the traditional steganography scheme. Because of considering the relation between the length of message and capacity, the performance of the hiding method using proposed hiding strategy outperforms the original approaches when embedding the same length message. Experimental results indicate that the gain from the proposed strategy over existing hiding schemes can reach up to 4.0 dB. Shaohui Liu, Hongxun Yao, Shengping Zhang, Wen Gao 0001 |
ICME | 2 |
| 2010 | Localized Image Matte Evaluation by Gradient CorrelationabstractIn natural image matting, various kinds of algorithms have been recently proposed. Moreover, alpha matting results have also been generated for comparison and composition into new backgrounds. However, all these methods have to make an alpha matte comparison to the ground truth so that one can get the final pixel-wised evaluation of these results. Nevertheless, while the input datasets are just used for test and there are no ground truth mattes, it is not possible to perform comparisons and to generate the quantitative comparison results. In this paper we combine the two ideas above and propose a new pixel-wise alpha mattes evaluation method. This approach is based on using local windows to measure gradient correlation between image and the matte. An optimal image channel minimizing the image variance is also selected at each window in order to perform the correlation more correctly. Experimental result shows that, our system can generate precise evaluation result for each pixel of each matte without ground truth. Guilin Yao, Hongxun Yao, Qingming Huang |
ICPR | 2 |
| 2010 | A refined particle filter method for contour trackingabstractTraditional particle filter which uses simple geometric shapes for representation cannot track objects with complex shape accurately. In this paper, we propose a refined particle filter method for contour tracking based on a binary level set model. In contrast with other previous work, the computational efficiency is greatly improved due to the simple form of the level set function. In addition, we perform curve evolution in the update step to make good use of the observation at current time. Finally, we consider some appearance information as well as the energy function to measure the weight for particles, which can identify the target more accurately. Experiment results on several challenging video sequences have verified the proposed algorithm is efficient and effective in many complicated scenes. Xin Sun 0003, Hongxun Yao, Shengping Zhang |
VCIP | 2 |
| 2010 | 3D silhouette tracking with occlusion inferenceabstractIt is a challenging problem to robustly track moving objects from image sequences because of occlusions. Previous methods did not exploit depth information sufficiently. Based on multiple camera scenes, we propose a 3D silhouette tracking framework to resolve occlusions and recover the appearances in 3D space, which enhances tracking effectiveness. In the framework, 2D object silhouettes are initially gained by Snake. Then a Voxel Space Carving procedure is introduced to simultaneously generate the occlusion model and visual hull of objects. Next, we adopt Particle Filter to select the valuable parts of occlusion model and combine them with the initial object silhouettes to generate the updated visual hull. Finally, updated visual hull of the objects are re-projected to each view to obtain their final contours. The experiments under the public LAB and SCULPTURE datasets validate the feasibility and effectiveness of our framework. Hongxun Yao, Rongrong Ji, Tianqiang Liu, Debin Zhao |
VCIP | 2 |
| 2010 | A rotation and scale invariant texture description approachabstractThis paper presents a novel texture description approach, which is robust to variances in rotation, scale and illumination in images, to classify the texture of images. A limitation with traditional methods is that they are more or less sensitive to the mentioned changes in images. To overcome this problem, we propose a novel Local Haar Binary Pattern (LHBP) based framework to ensure invariance in global rotation, scale, and light change. Our method consists of two components: feature extraction and scale self-adaptive classification. The global rotation invariant LHBP histogram features are extracted against the variances of illumination and global rotation, and the scale self-adaptive strategy is used for optimizing the classification of different scale textures. Evaluation results on Outex and Brodatz databases illustrate the significant advantages of the proposed approach over existing algorithms. Pengfei Xu 0001, Hongxun Yao, Rongrong Ji, Xiaoshuai Sun, Xianming Liu 0005 |
VCIP | 2 |
| 2010 | Robust object tracking based on sparse representationabstractIn this paper, we propose a novel and robust object tracking algorithm based on sparse representation. Object tracking is formulated as a object recognition problem rather than a traditional search problem. All target candidates are considered as training samples and the target template is represented as a linear combination of all training samples. The combination coefficients are obtained by solving for the minimum l1-norm solution. The final tracking result is the target candidate associated with the non-zero coefficient. Experimental results on two challenging test sequences show that the proposed method is more effective than the widely used mean shift tracker. Shengping Zhang, Hongxun Yao, Xin Sun 0003, Shaohui Liu |
VCIP | 2 |
| 2010 | Robust object tracking combining color and scale invariant featuresabstractObject tracking plays a very important role in many computer vision applications. However its performance will significantly deteriorate due to some challenges in complex scene, such as pose and illumination changes, clustering background and so on. In this paper, we propose a robust object tracking algorithm which exploits both global color and local scale invariant (SIFT) features in a particle filter framework. Due to the expensive computation cost of SIFT features, the proposed tracker adopts a speed-up variation of SIFT, SURF, to extract local features. Specially, the proposed method first finds matching points between the target model and target candidate, than the weight of the corresponding particle based on scale invariant features is computed as the the proportion of matching points of that particle to matching points of all particles, finally the weight of the particle is obtained by combining weights of color and SURF features with a probabilistic way. The experimental results on a variety of challenging videos verify that the proposed method is robust to pose and illumination changes and is significantly superior to the standard particle filter tracker and the mean shift tracker. Shengping Zhang, Hongxun Yao, Peipei Gao |
VCIP | 2 |
| 2010 | Partial occlusion robust object tracking using an effective appearance modelabstractPartial occlusion is one of the most challenging difficulties for object tracking. In this paper, we present an approach to address this problem by using an effective appearance model which has two innovations. First, in contrast to widely used color histogram that models the appearance of an object using only color information, we assert that both color and texture are important cues for tracking, especially in the presence of complex background. We thus propose a novel local descriptor, named local color texture pattern (LCTP), to model the appearance of the object with color and texture information simultaneously. Second, global color histogram completely ignores the spatial layout information of an object and are sensitive to partial occlusion. In this work, we overcome this limitation based on a block-dividing way: 1) divide target into multiple blocks and then represent each block with LCTP histogram, 2) with a selectivity strategy, we select blocks that are not occluded and then combine similarities of those selected blocks to obtain final similarity measure. Experimental results demonstrate that the proposed method is more robust to partial occlusion than two state-of-the-art algorithms. Shengping Zhang, Hongxun Yao, Shaohui Liu |
VCIP | 2 |
| 2010 | Adaptive Sign Language Recognition With Exemplar Extraction and MAP/IVFSabstractSign language recognition systems suffer from the problem of signer dependence. In this letter, we propose a novel method that adapts the original model set to a specific signer with his/her small amount of training data. First, affinity propagation is used to extract the exemplars of signer independent hidden Markov models; then the adaptive training vocabulary can be automatically formed. Based on the collected sign gestures of the new vocabulary, the combination of maximum a posteriori and iterative vector field smoothing is utilized to generate signer-adapted models. Experimental results on six signers demonstrate that the proposed method can reduce the amount of the adaptation data and still can achieve high recognition performance. Yu Zhou 0015, Xilin Chen 0001, Debin Zhao, Hongxun Yao, Wen Gao 0001 |
IEEE Signal Process. Lett. | 4 |
| 2009 | Universal Steganalysis Based on Statistical Models Using Reorganization of Block-based DCT CoefficientsabstractThe goal of steganography is to hide information into media without disclosing the fact of existing communication. Currently, stganography such as least significant bit (LSB), quantization index modulation (QIM) and spread spectrum (SS), has become increasingly widespread. Steganalysis as a counterpart of stganography is to detect the presence of it. In this paper, we present a new universal steganalysis method based on statistical models of the imagepsilas discrete cosine transform (DCT) coefficients. In fact, the block-based DCT by proper reorganization of its coefficients can have similar characteristics to wavelet transforms. The presented universal steganalysis method utilizes these characteristics to build statistical models of the image and its prediction-error image. Features extracted from the re-organization DCT blocks of host images and theirs prediction-error images and features extracted from steg images and theirs prediction-error images are used to train the SVM classifier. In the testing, features from those potential images are inputted the trained-well classifier to determine where the potential images are stego images or not. The experiments have shown that the proposed method outperforms in general prior-arts of steganalysis methods based on wavelet transform domain. Shaohui Liu, Hongxun Yao, Debin Zhao |
IAS | 3 |
| 2009 | Local Spatial Co-occurrence for Background Subtraction via Adaptive Binned Kernel Estimation
Bineng Zhong 0001, Shaohui Liu, Hongxun Yao |
ACCV (3) | 3 |
| 2009 | Vocabulary hierarchy optimization for effective and transferable retrievalabstractScalable image retrieval systems usually involve hierarchical quantization of local image descriptors, which produces a visual vocabulary for inverted indexing of images. Although hierarchical quantization has the merit of retrieval efficiency, the resulting visual vocabulary representation usually faces two crucial problems: (1) hierarchical quantization errors and biases in the generation of “visual words”; (2) the model cannot adapt to database variance. In this paper, we describe an unsupervised optimization strategy in generating the hierarchy structure of visual vocabulary, which produces a more effective and adaptive retrieval model for large-scale search. We adopt a novel Density-based Metric Learning (DML) algorithm, which corrects word quantization bias without supervision in hierarchy optimization, based on which we present a hierarchical rejection chain for efficient online search based on the vocabulary hierarchy. We also discovered that by hierarchy optimization, efficient and effective transfer of a retrieval model across different databases is feasible. We deployed a large-scale image retrieval system using a vocabulary tree model to validate our advances. Experiments on UKBench and street-side urban scene databases demonstrated the effectiveness of our hierarchy optimization approach in comparison with state-of-the-art methods. Rongrong Ji, Xing Xie 0001, Hongxun Yao, Wei-Ying Ma |
CVPR | 3 |
| 2009 | Multl-resolution background subtraction for dynamic scenesabstractDynamic scenes (e.g. waving trees, ripples in water, illumination changes, camera jitters etc.) challenge many traditional background subtraction methods. In this paper, we present a novel background subtraction approach for dynamic scenes, in which the background is modeled in a multi-resolution framework. First, for each level of the pyramid, we run an independent mixture of Gaussians Models (GMM) that outputs a background subtraction map. Second, these background subtraction maps are combined via AND operator to finally get a more robust and accurate background subtraction map. This is a natural fusion because the original resolution and low resolution images have complementary strengths, which original resolution image contains rich information and low resolution image is insensitive to the noises and the small movement of dynamic scene. Experimental result shows that this real-time algorithm is able to detect moving objects accurately even in dynamic scenes. Bineng Zhong 0001, Shaohui Liu, Hongxun Yao, Baochang Zhang 0001 |
ICIP | 3 |
| 2009 | Neighboring Image Patches Embedding for background modelingabstractWe present a novel feature extraction framework, Neighboring Image Patches Embedding (NIPE), for robust and efficient background modeling. We divide image into patches and represent each image patch as a NIPE vector. Then, the background model of each image patch is constructed as a group of weighted adaptive NIPE vectors. The NIPE feature vector, whose components are similarities between current image patch and its neighbors, describes mainly the mutual relationship between neighboring patches. Since neighboring image patches tend to be similarly affected by environmental effects (e.g., dynamic background), the NIPE vectors are more robust in these conditions comparing with the conventional method. Experimental results demonstrate the efficiency and effectiveness of our proposed NIPE method. Bineng Zhong 0001, Hongxun Yao, Shaohui Liu |
ICIP | 2 |
| 2009 | Spatial-temporal nonparametric background subtraction in dynamic scenesabstractTraditional background subtraction methods model only temporal variation of each pixel. However, there is also spatial variation in real word due to dynamic background such as waving trees, spouting fountain and camera jitters, which causes the significant performance degradation of traditional methods. In this paper, a novel spatial-temporal nonparametric background subtraction approach (STNBS) is proposed to effectively handle dynamic background by modeling the spatial and temporal variations simultaneously. Specially, for each pixel in an image, we adaptively maintain a sample consisting of pixels observed in previous frames. At current frame, for a particular pixel, the proposed method estimates the probabilities of observing this pixel based on samples of its neighboring pixels. The pixel is labeled as background if one of these estimated probabilities is larger than a fixed threshold. All samples are adaptively updated over time. Experimental results on several challenging sequences show that the proposed method achieves the best performance than two state-of-the-art algorithms. Shengping Zhang, Hongxun Yao, Shaohui Liu |
ICME | 2 |
| 2009 | Mining city landmarks from blogs by graph modelingabstractRecent years have witnessed great prosperity in community-contributed multimedia. Discovering, extracting, and summarizing knowledge from these data enables us to make better sense of the world. In this paper, we report our work on mining famous city landmarks from blogs for personalized tourist suggestions. Our main contribution is a graph modeling framework to discover city landmarks by mining blog photo correlations with community supervision. This modeling fuses context, content, and community information in a style that simulates both static (PageRank) and dynamic (HITS) ranking models to highlight representative data from the consensus of blog users. Rongrong Ji, Xing Xie 0001, Hongxun Yao, Wei-Ying Ma |
ACM Multimedia | 3 |
| 2009 | Location sensitive indexing for image-based advertisingabstractThis paper introduces the architecture of our location sensitive indexing model which is used in a platform designed to deliver advertisements to users who primarily utilize images as queries instead of textual keywords. The indexing model facilitates an advertiser's ability to bid on images, such as billboards or logos, in order to obtain user feedback in judging image attractiveness. Additionally, the model enables automatic evaluation of advertisement popularity by mining users' query logs, which is critical for generating advertisement recommendations. The location sensitive architecture of this model enables effective and efficient functionality in large-scale scenarios. In the model's structure, our Location Sensitive Visual Indexing model (LSVI) incorporates location information that subdivides geographical regions for precise and localized image matching. By collecting feedback from mobile users, location-based mining can also help discover popular advertisements as well as their representative images. We have deployed our platform into a real-world advertising system in Beijing, China, which demonstrates effective results in comparative studies with both alternative and state-of-the-art approaches. Dechao Liu, Matthew R. Scott, Rongrong Ji, Hongxun Yao, Xing Xie 0001 |
ACM Multimedia | 5 |
| 2009 | What is a complete set of keywords for image description & annotation on the webabstractDoes there exist a compact set of keywords that can completely and effectively cover the image annotation problem by expanding from it? In this paper, we answer this question by presenting a complete set framework for image annotation, which is motivated by the existence of semantic ontology. To generate this set, we propose a cross model optimization strategy from both textual and visual information for topic decomposition, based on a so-called Bipartite LSA model, which minimize multimodal error energy functions in a probabilistic Latent Semantic Analysis model. To achieve complete set based annotation, we present a Gaussian-Kernel-Generative process based keyword generation procedure, which analogizes keyword annotation in a probabilistic generative manner. A group of experiments is performed on Washington University image database and 80,000 Flickr images with comparisons to the state-of-the-arts. Finally, potential advantages and future improvements of our framework are discussed outside the scope of topic modeling. Xianming Liu 0005, Hongxun Yao, Rongrong Ji, Pengfei Xu 0001, Xiaoshuai Sun |
ACM Multimedia | 2 |
| 2009 | Photo assessment based on computational visual attention modelabstractIt is difficult to be satisfied for automatic photo assessment using only low level visual features such as brightness, lighting, hue, contrast, color distribution and so on. Instead of using low level visual features, we present a novel computational visual attention model to assess photos. Firstly, a face-sensitive saliency map analysis is deployed to estimate attention distribution. Then, a Rate of Focused Attention (RFA) measurement is proposed to quantify photo quality. By integrating top-down supervision into the visual attention model, we further achieve personalized photo assessment to take user preference into quality evaluation, which can be extended into object or semantic oriented photo assessment scenarios. Experiments on personal photo albums with comparison to ground-truth user evaluations demonstrate the effeteness of the proposed method. Xiaoshuai Sun, Hongxun Yao, Rongrong Ji, Shaohui Liu |
ACM Multimedia | 2 |
| 2009 | Geometric and Algebraic Approaches of Planar Structure Recovery Based on Properties of Dual ConicabstractIn this article, approaches of planar Euclidean Structure recovery are proposed based on the geometric and algebraic characteristics of the dual conic. Algebraic relations between the matrix representation of the dual conic and the circular point enveloped are extensively analyzed, to attain a generalized framework for the problem. The dual conic has the decomposition with consistent geometric interpretation from which the circular point envelope can be solved. From the geometric viewpoint, based on the characteristics of the common elements of dual conics (their bitangents), a stratified approach is presented to establish the solution. Experiments on both synthetic and real images are conducted to demonstrate the robustness and accuracy of the presented approach. Liang Wang 0004, Hongxun Yao, Shaohui Liu |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2009 | Dynamic Background Subtraction Based on Local Dependency HistogramabstractTraditional background subtraction methods perform poorly when scenes contain dynamic backgrounds such as waving tree branches, spouting fountain, illumination changes, camera jitters, etc. In this paper, from the view of spatial context, we present a novel and effective dynamic background method with three contributions. First, we present a novel local dependency descriptor, called local dependency histogram (LDH), to effectively model the spatial dependencies between a pixel and its neighboring pixels. The spatial dependencies contain substantial evidence for differentiating dynamic background regions from moving objects of interest. Second, based on the proposed LDH, an effective approach to dynamic background subtraction is proposed, in which each pixel is modeled as a group of weighted LDHs. Labeling a pixel as foreground or background is done by comparing the LDH computed in current frame against its model LDHs. The model LDHs are adaptively updated by the current LDH. Finally, unlike traditional approaches using a fixed threshold to judge whether a pixel matches to its model, an adaptive thresholding technique is also proposed. Experimental results on a diverse set of dynamic scenes validate that the proposed method significantly outperforms traditional methods for dynamic background subtraction. Shengping Zhang, Hongxun Yao, Shaohui Liu |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2009 | Visual and textual fusion for semantically supervised region-based retrieval
Rongrong Ji, Hongxun Yao, Pengfei Xu 0001, Xiaoshuai Sun |
Multim. Syst. | 2 |
| 2009 | Synthetic data generation technique in Signer-independent sign language recognition
Feng Jiang 0001, Wen Gao 0001, Hongxun Yao, Debin Zhao, Xilin Chen 0001 |
Pattern Recognit. Lett. | 3 |
| 2009 | Contour-motion feature (CMF): A space-time approach for robust pedestrian detection
Yazhou Liu, Xilin Chen 0001, Hongxun Yao, Xinyi Cui, Wen Gao 0001 |
Pattern Recognit. Lett. | 3 |
| 2009 | Event Tactic Analysis Based on Broadcast Sports VideoabstractMost existing approaches on sports video analysis have concentrated on semantic event detection. Sports professionals, however, are more interested in tactic analysis to help improve their performance. In this paper, we propose a novel approach to extract tactic information from the attack events in broadcast soccer video and present the events in a tactic mode to the coaches and sports professionals. We extract the attack events with far-view shots using the analysis and alignment of web-casting text and broadcast video. For a detected event, two tactic representations, aggregate trajectory and play region sequence, are constructed based on multi-object trajectories and field locations in the event shots. Based on the multi-object trajectories tracked in the shot, a weighted graph is constructed via the analysis of temporal-spatial interaction among the players and the ball. Using the Viterbi algorithm, the aggregate trajectory is computed based on the weighted graph. The play region sequence is obtained using the identification of the active field locations in the event based on line detection and competition network. The interactive relationship of aggregate trajectory with the information of play region and the hypothesis testing for trajectory temporal-spatial distribution are employed to discover the tactic patterns in a hierarchical coarse-to-fine framework. Extensive experiments on FIFA World Cup 2006 show that the proposed approach is highly effective. Guangyu Zhu 0002, Changsheng Xu, Qingming Huang, Yong Rui, Shuqiang Jiang, Wen Gao 0001, Hongxun Yao |
IEEE Trans. Multim. | 7 |
| 2008 | Dynamic background modeling and subtraction using spatio-temporal local binary patternsabstractTraditional background modeling and subtraction methods have a strong assumption that the scenes are of static structures with limited perturbation. These methods will perform poorly in dynamic scenes. In this paper, we present a solution to this problem. We first extend the local binary patterns from spatial domain to spatio-temporal domain, and present a new online dynamic texture extraction operator, named spatio-temporal local binary patterns (STLBP). Then we present a novel and effective method for dynamic background modeling and subtraction using STLBP. In the proposed method, each pixel is modeled as a group of STLBP dynamic texture histograms which combine spatial texture and temporal motion information together. Compared with traditional methods, experimental results show that the proposed method adapts quickly to the changes of the dynamic background. It achieves accurate detection of moving objects and suppresses most of the false detections for dynamic changes of nature scenes. Shengping Zhang, Hongxun Yao, Shaohui Liu |
ICIP | 2 |
| 2008 | Vocabulary tree incremental indexing for scalable location recognitionabstractThis work aims at developing a scalable vision-based location recognition system where the backend database can be updated incrementally. Our proposed framework enables incremental indexing of vocabulary tree model, which efficiently includes new data into model refinement without re-generating entire model from overall dataset. An adaption trigger criterion is presented to lessen system computational cost, which is achieved by density-based relative entropy estimation between original dataset and newly coming data. Experiments on Seattle urban scene datasets with over 20K street-side images show the effectiveness of our work. Rongrong Ji, Xing Xie 0001, Hongxun Yao, Yongjian Wu 0001, Wei-Ying Ma |
ICME | 3 |
| 2008 | Directional correlation analysis of local Haar binary pattern for text detectionabstractTwo main restrictions exist in state-of-the-art text detection algorithms: 1. Illumination variance; 2. Text-background contrast variance. This paper presents a robust text characterization approach based on local Haar binary pattern (LHBP) to address these problems. Based on LHBP, a coarse-to-fine detection framework is presented to precisely locate text lines in scene images. Firstly, threshold-restricted local binary pattern is extracted from high-frequency coefficients of pyramid Haar wavelet. It preserves and uniforms inconsistent text-background contrasts while filtering gradual illumination variations. Subsequently, we propose a directional correlation analysis (DCA) approach to filter non-directional LHBP regions for locating candidate text regions. Finally, using LHBP histogram, an SVM-based post-classification is presented to refine detection results. Experimental results on ICDAR 03 demonstrate the effectiveness and robustness of our proposed method. Rongrong Ji, Pengfei Xu 0001, Hongxun Yao, Xiaoshuai Sun, Tianqiang Liu |
ICME | 3 |
| 2008 | Clustering-based subspace SVM ensemble for relevance feedback learningabstractThis paper presents a subspace SVM ensemble algorithm for adaptive relevance feedback (RF) learning. Our method deals with the case that user’s relevance feedback examples are usually insufficient and overlapped together in feature space, which decreases the learning effectiveness of RF classifiers. To enhance classification efficiency in such case, multiple SVMs are learned by clustering-based training set partition, each of which fits its cluster-specific sample distribution and gives labeling regressions to test samples that fall within this cluster. To adapt features to sample distribution within each cluster, AdaBoost feature selection is conducted onto pyramid Haar of H&I bands in HSI space. In AdaBoost, we evaluate the feature discriminative ability by an entropy-based uncertainty criterion, based on which an Eigen feature subspace is constructed in cluster-specific SVM training. Finally, regression results of multiple SVMs are probabilistic assembled to give the final labeling prediction for test image. We compare our cluster-based cascade SVMs (CSS) RF method in COREL 5,000 database with: 1. Single SVM; 2. Active Learning SVM [5]; 3. Bootstrap Sampling SVM [7]. The superior experimental results demonstrate the efficiency of our algorithm. Rongrong Ji, Hongxun Yao, Pengfei Xu 0001, Xianming Liu 0005 |
ICME | 2 |
| 2008 | Mahalanobis distance based Polynomial Segment Model for Chinese Sign Language RecognitonabstractSign Language Recognition (SLR) systems are mostly based on Hidden Markov Model (HMM) and have achieved excellent results. However, the assumption of frame independence in HMM makes it inconsistent with the characteristic of strong temporal correlation in sign language signals. Polynomial Segment Model (PSM) explicitly represents the temporal evolution of sign language features as a Gaussian process with time-varying parameters. In this paper PSM is first introduced to SLR framework to solve the temporal correlation problem. Considering the correlation among the coefficients of polynomial trajectorypsilas different orders, Mahalanobis distance is used as the classification criterion to evaluate the likelihood of test data. Experimental results show that our method outperform the conventional HMM methods by 6.81% in recognition accuracy. Yu Zhou 0015, Xilin Chen 0001, Debin Zhao, Hongxun Yao, Wen Gao 0001 |
ICME | 4 |
| 2008 | A covariance-based method for dynamic background subtractionabstractBackground subtraction in dynamic scenes is an important and challenging task. In this paper, we present a novel and effective method for dynamic background subtraction based on covariance matrix descriptor. The algorithm integrates two distinct levels: pixel level and region level. At the pixel level, spatial properties that are obtained from pixel coordinate values, and appearance properties, i.e., intensity, texture, gradient, etc, are used as features of each pixel. In the region level, the correlation of features extracted at the pixel level is represented by a covariance matrix that is calculated over a rectangle region around the pixel. Each pixel is modeled as a group of weighted adaptive covariance matrices. Experimental results on a diverse set of dynamic scenes show that the proposed method dramatically out-performs traditional methods for dynamic background subtraction. Shengping Zhang, Hongxun Yao, Shaohui Liu, Xilin Chen 0001, Wen Gao 0001 |
ICPR | 2 |
| 2008 | Hierarchical background subtraction using local pixel clusteringabstractWe propose a robust hierarchical background subtraction technique which takes the spatial relations of neighboring pixels in a local region into account to detect objects in difficult conditions. Our algorithm combines a per-pixel with a per-region background model in a hierarchical manner, which accentuates the advantages of each. This is a natural combination because the two models have complementary strengths. The per-pixel background model is achieved by mixture of Gaussians Models (GMM) with RGB feature. Although precisely describing background change in high resolution, it suffers from the sensitivity to quick variations in dynamic environment. To tolerate these quick variations, we further develop a novel GMM based per-region background model, which is updated by the cluster centers obtained from a k-means clustering of the pixels’ RGB feature in the region. Numerical and qualitative experimental results on challenging videos demonstrate the robustness of the proposed method. Bineng Zhong 0001, Hongxun Yao, Shiguang Shan, Xilin Chen 0001, Wen Gao 0001 |
ICPR | 2 |
| 2008 | Attention-driven action retrieval with DTW-based 3d descriptor matchingabstractFrom visual perception viewpoint, actions in videos can capture high-level semantics for video content understanding and retrieval. However, action-level video retrieval meets great challenges, due to the interferences from global motions or concurrent actions, and the difficulties in robust action describing and matching. This paper presents a content-based action retrieval framework to enable effective search of near-duplicated actions in large-scale video database. Firstly, we present an attention shift model to distill and partition human-concerned saliency actions from global motions and concurrent actions. Secondly, to characterize each saliency action, we extract 3D-SIFT descriptor within its spatial-temporal region, which is robust against rotation, scale, and view point variances. Finally, action similarity is measured using Dynamic Time Warping (DTW) distance to offer tolerance for action duration variance and partial motion missing. Search efficiency in large-scale dataset is achieved by hierarchical descriptor indexing and approximate nearest-neighbor search. In validation, we present a prototype system VILAR to facilitate action search within "Friends" soap operas with excellent accuracy, efficiency, and human perception revealing ability. Rongrong Ji, Xiaoshuai Sun, Hongxun Yao, Pengfei Xu 0001, Tianqiang Liu, Xianming Liu 0005 |
ACM Multimedia | 3 |
| 2008 | Shape from silhouettes based on a centripetal pentahedron model
Xin Liu 0047, Hongxun Yao, Xilin Chen 0001, Wen Gao 0001 |
Graph. Model. | 2 |
| 2008 | Effective and Automatic Calibration Using Concentric CirclesabstractIn this paper, we present an effective, flexible and completely automated camera calibration approach using only one pair of concentric circles. This approach utilizes the characteristics of concentric circles' tangent lines to locate the center of these circles, and finds the geometric constraints for calibration based on the orthogonality formed by a point on the circle and the two intersected points of the circle with the line through the center of the circle. The entire process requires no conic equation fitting and no metric measurement of the test pattern, which is very flexible to implement. Liang Wang 0004, Hongxun Yao, Heng-Da Cheng |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2007 | Combining Global and Local Classifiers for Lipreading
Shengping Zhang, Hongxun Yao, Yuqi Wan |
ACII | 2 |
| 2007 | Minimizing the Distortion Spatial Data Hiding Based on Equivalence Class
Shaohui Liu, Hongxun Yao, Wen Gao 0001, Dingguo Yang |
ICIC (1) | 2 |
| 2007 | Mean-Shift Blob Tracking with Adaptive Feature Selection and Scale AdaptationabstractWhen the appearances of the tracked object and surrounding background change during tracking, fixed feature space tends to cause tracking failure. To address this problem, we propose a method to embed adaptive feature selection into mean shift tracking framework. From a feature set, the most discriminative features are selected after ranking these features based on their Bayes error rates, which are estimated from object and background samples. For the selected features, a criterion is proposed to evaluate their stability for tracking and to guide feature reselection. The selected features are used to generate a weight image, in which mean shift is employed to locate the object. Moreover, a simple yet effective scale adaptation method is proposed to deal with object changing in size. Experiments on several video sequences show the effectiveness of the proposed method. Dawei Liang, Qingming Huang, Shuqiang Jiang, Hongxun Yao, Wen Gao 0001 |
ICIP (3) | 4 |
| 2007 | Trajectory based event tactics analysis in broadcast sports videoabstractMost of existing approaches on event detection in sports video are general audience oriented. The extracted events are then presented to the audience without further analysis. However, professionals, such as soccer coaches, are more interested in the tactics used in the events. In this paper, we present a novel approach to extract tactic information from the goal event in broadcast soccer video and present the goal event in a tactic mode to the coaches and sports professionals. We first extract goal events with far-view shots based on analysis and alignment of web-casting text and broadcast video. For a detected goal event, we employ a multi-object detection and tracking algorithm to obtain the players and ball trajectories in the shot. Compared with existing work, we proposed an effective tactic representation called aggregate trajectory which is constructed based on multiple trajectories using a novel analysis of temporal-spatial interaction among the players and the ball. The interactive relationship with play region information and hypothesis testing for trajectory temporal-spatial distribution are exploited to analyze the tactic patterns in a hierarchical coarse-to-fine framework. The experimental results on the data of FIFA World Cup 2006 are promising and demonstrate our approach is effective. Guangyu Zhu 0002, Qingming Huang, Changsheng Xu, Yong Rui, Shuqiang Jiang, Wen Gao 0001, Hongxun Yao |
ACM Multimedia | 7 |
| 2007 | Shape from silhouette outlines using an adaptive dandelion model
Xin Liu 0047, Hongxun Yao, Wen Gao 0001 |
Comput. Vis. Image Underst. | 2 |
| 2007 | Nonparametric background generation
Yazhou Liu, Hongxun Yao, Wen Gao 0001, Xilin Chen 0001, Debin Zhao |
J. Vis. Commun. Image Represent. | 2 |
| 2007 | Human Behavior Analysis for Highlight Ranking in Broadcast Racket Sports VideoabstractThe majority of existing work on sports video analysis concentrates on highlight extraction. Little work focuses on the important issue as how the extracted highlights should be organized. In this paper, we present a multimodal approach to organize the highlights extracted from racket sports video grounded on human behavior analysis using a nonlinear affective ranking model. Two research challenges of highlight ranking are addressed, namely affective feature extraction and ranking model construction. The basic principle of affective feature extraction in our work is to extract sensitive features which can stimulate user's emotion. Since the users pay most attention to player behavior and audience response in racket sport highlights, we extract affective features from player behavior including action and trajectory, and game-specific audio keywords. We propose a novel motion analysis method to recognize the player actions. We employ support vector regression to construct the nonlinear highlight ranking model from affective features. A new subjective evaluation criterion is proposed to guide the model construction. To evaluate the performance of the proposed approaches, we have tested them on more than ten-hour broadcast tennis and badminton videos. The experimental results demonstrate that our action recognition approach significantly outperforms the existing appearance-based method. Moreover, our user study shows that the affective highlight ranking approach is effective. Guangyu Zhu 0002, Qingming Huang, Changsheng Xu, Liyuan Xing, Wen Gao 0001, Hongxun Yao |
IEEE Trans. Multim. | 6 |
| 2006 | Visual Hull Embossment by Graph CutsabstractIn this paper, we present a visual hull embossment algorithm which works on the base of a visual hull to reconstruct a precise model from calibrated color images. The algorithm uses the geometry of the visual hull to compute voxel normals and visibilities. Then a photometric consistency field (PCF) is computed, which describes the photometric consistency score of every voxel. The visual hull is refined in a view dependent fashion. For each view, we compute a view dependent PCF (VDPCF) by resampling the PCF and then find an optimal depth image in VDPCF with graph cut based energy minimization, which takes into consideration data, smoothness and adjustment terms simultaneously. The depth images are used for carving voxels. Our algorithm can produce precise models with a moderate speed. Xin Liu 0047, Hongxun Yao, Xilin Chen 0001, Wen Gao 0001 |
ICIP | 2 |
| 2005 | The Bunch-Active Shape Model
Jingcai Fan, Hongxun Yao, Wen Gao 0001, Yazhou Liu, Xin Liu 0047 |
ACII | 2 |
| 2005 | An Information Acquiring Channel - Lip Movement
Xiaopeng Hong, Hongxun Yao, Qinghui Liu |
ACII | 2 |
| 2005 | Static Gesture Quantization and DCT Based Sign Language Generation
Feng Jiang 0001, Hongxun Yao, Guilin Yao, Wen Gao 0001 |
ACII | 3 |
| 2005 | An active volumetric model for 3D reconstructionabstractIn this paper, we present an active volumetric model (AVM) for 3D reconstruction from multiple calibrated images of a scene. The AVM is a physically motivated 3D deformable model which shrinks actively under the influence of multiple simulated forces towards the real scene by throwing away some of its voxels. It provides a computational framework to integrate several constraints in an intelligible way. In the current work, we use three forces derived respectively from the smooth constraint, the compulsory silhouette constraint, and the color consistency constraint. Based on the composition of the 3 forces, our algorithm can significantly restrain holes and floating voxels, which plague voxel coloring algorithms, and produce precise and smooth models. We test our algorithm by experiments based on both synthetic and real data. Xin Liu 0047, Hongxun Yao, Xilin Chen 0001, Wen Gao 0001 |
ICIP (2) | 2 |
| 2004 | Based on HMM and SVM multilayer architecture classifier for Chinese sign language recognition with large vocabularyabstractThis paper has put forward a new architecture classifier method for Chinese sign language recognition (CSLR) to improve the performance of recognition. It is a signer-independent method, to recognize Chinese sign language with large vocabulary using multilayer architecture classifier and making use of the advantages both of HMM (hidden Markov model) and SVM (support vector machines). Because HMM is good at dealing with sequential inputs, while SVM shows superior performance in classifying with good generalization properties especially for limited samples. Therefore, they can be combined to yield a better and effective multilayer architecture classifier. We apply SVMs to resolve the uncertainties of the remaining which are in confusable sets after the first-stage HMM-based recognizer. And the confusable sets would be updated dynamically according to the results of a recognition performance to optimize the discernment performance next time. Experimental results proved that it is an effective method for CSLR with large vocabulary keywords sign language recognition, HMM, SVM, multilayer architecture classifier. Jianjun Ye, Hongxun Yao, Feng Jiang 0001 |
ICIG | 2 |
| 2004 | Multilayer architecture in sign language recognition systemabstractUp to now analytical or statistical methods have been used in sign language recognition with large vocabulary. Analytical methods such as Dynamic Time Wrapping (DTW) or Euclidian distance have been used for isolated word recognition, but the performance is not satisfactory enough because it is easily interfered by noise. Statistical methods, especially hidden Markov Models are commonly used, for both continuous sign language and isolated words and with the expansion of vocabulary the processing time becomes increasingly unacceptable. Therefore, a multilayer architecture of sign language recognition for large vocabulary is proposed in this paper for the purpose of speeding up the recognition process. In this method the gesture sequence to be recognized is first located at a set of words that are easy to be confused (confusion set) through a global cursory search and then the gesture is recognized through a latter local search and the generation of confusion set is realized by DTW/ISODATA algorithm. Experiment results indicate that it is an effective algorithm for Chinese sign language recognition. Feng Jiang 0001, Hongxun Yao, Guilin Yao |
ICMI | 2 |
| 2003 | Neural network based steganalysis in still imagesabstractSeganalysis has recently attracted researchers' interests with the development of information hiding techniques. In this paper we propose a new method based neural network to get statistics features of images to identify the underlying hidden data. We first extract features of image embedded information, then input them into neural network to get output. And experiment results indicate this method is valid in steganalysis. This method will be used for Internet/network security, watermarking and so on. Shaohui Liu, Hongxun Yao, Wen Gao 0001 |
ICME | 2 |
| 2002 | Blind watermarking method based on DWT middle frequency pairabstractIn this paper a novel watermarking method based on the discrete wavelet transform (DWT) is proposed. Different from previous watermarking methods, the watermark is embedded into the middle frequency coefficients by quantizing the middle frequency coefficient pair, which is defined as a pair of coefficients that are at the same location in the LH and HL bands of the DWT coefficients. This watermarking technique can successfully resist severe transformation of gray levels while it is still robust to normal attacks such as blurring, sharpening and JPEG compression. In detection, it does not need the original image to recover the watermark. Hongxun Yao, Wen Gao 0001, Sanghyun Joo |
ICME (2) | 2 |
| 2001 | Face detection and location based on skin chrominance and lip chrominance transformation from color images
Hongxun Yao, Wen Gao 0001 |
Pattern Recognit. | 1 |
| 2000 | Towards robust lipreading
Wen Gao 0001, Jiyong Ma, Rui Wang 0016, Hongxun Yao |
INTERSPEECH | 4 |